An artificial intelligence-based full-text retrieval method and system for archives

Through artificial intelligence technology, the file is clustered with word segmentation and word embedding vectors, and the neighborhood radius adjustment is used to solve the accuracy of synonyms and synonyms in the full text search of archives, achieving a more efficient search effect.

CN117992606BActive Publication Date: 2025-07-08ZHENGZHOU RIXING ELECTRONICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410117225.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-08
Estimated Expiration
2044-01-26

AI Technical Summary

Technical Problem

The existing full-text search methods for archives are not accurate enough when dealing with synonyms, synonyms and alias, and it is easy to miss important information.

Method used

Artificial intelligence technology is used to segment the text in the archive, establish the correspondence between the word terms and the archive, cluster it through word embedding vectors, adjust the clustering results using different neighborhood radii, combine the inverted record table to obtain the search results, and sort and push the results.

Benefits of technology

It improves the accuracy and comprehensiveness of the full text of the archive, can better search files with similar semantics, and provide more accurate search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117992606B_ABST
    Figure CN117992606B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of artificial intelligence, and in particular, to a method and system for full-text retrieval of archives based on artificial intelligence. Specifically, a correspondence relationship between terms and archives is established to obtain an inverted record table; the retrieval terms input by the user and each term in the inverted record table are respectively encoded to obtain corresponding retrieval term embedding vectors and term embedding vectors; starting from the first neighborhood radius, the term embedding vectors are clustered according to the neighborhood radius, the cluster closest to the retrieval term embedding vector in the clustering result is calculated, and the full-text retrieval result of the archives is obtained according to the inverted record table and the cluster corresponding to each retrieval term. If the number of the retrieval results is less than the threshold, the full-text retrieval result corresponding to the next neighborhood radius is obtained, and so on, until the number of all retrieval results after duplicate removal is not less than the threshold; all the retrieval results after duplicate removal are sorted and then pushed to the user. The present invention improves the accuracy and comprehensiveness of full-text retrieval of archives.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Field

[0002] The present invention relates to the field of artificial intelligence, and in particular to a full-text retrieval method and system for archives based on artificial intelligence. Background Art

[0003] Archives refer to various historical records in different forms such as written words, charts, audio-visual materials, etc. directly formed by national institutions, social organizations, and individuals in the past and present in their political, economic, scientific, technological, cultural, religious, and other activities, which have preservation value for the country and society. The retrieval of archives has always been the focus of archive management. Especially in the current situation where the number of archives is huge and the types are diverse, an efficient and accurate retrieval method has become the key to ensuring the availability and accessibility of information. Full-text retrieval allows users to quickly locate relevant information by searching for specific words or phrases in a document, which is more efficient than traditional retrieval methods based on metadata or abstracts, especially for retrieving information hidden in an inconspicuous part of the archive. Inverted index is one of the most commonly used technologies in full-text retrieval. It first creates an index that associates each word with a list of documents containing that word. When searching, the algorithm directly looks up the list of documents containing the query word. However, the retrieval accuracy for synonyms, near-synonyms, aliases, etc. is not enough, and it is easy to miss important information. How to improve the accuracy and comprehensiveness of full-text retrieval of archives is an urgent problem to be solved in this field. Summary of the Invention

[0004] To solve the above problems, the present invention provides a full-text retrieval method for archives based on artificial intelligence, and the method includes the following steps:

[0005] Segment the text in the archives to obtain word items corresponding to each archive, establish the correspondence between the word items and the archives, and obtain an inverted record table; encode the retrieval word input by the user and each word item in the inverted record table respectively to obtain corresponding retrieval word embedding vectors and word item embedding vectors;

[0006] Starting from the first neighborhood radius, cluster the word item embedding vectors according to the neighborhood radius, calculate the cluster in the clustering result that is closest to the retrieval word embedding vector, obtain the full-text retrieval result of the archives according to the inverted record table and the cluster corresponding to each retrieval word. If the number of the retrieval results is less than the threshold, obtain the full-text retrieval result corresponding to the next neighborhood radius, and so on, until the number of all retrieval results after deduplication is not less than the threshold; sort all the deduplicated retrieval results and push them to the user; where the first neighborhood radius is the smallest, and the next neighborhood radius is greater than the previous neighborhood radius.

[0007] Preferably, the calculation of the cluster in the clustering result that is closest to the retrieval word embedding vector is specifically:

[0008] Calculate the average value of the term embedding vectors corresponding to all core objects in each cluster;

[0009] Obtain the distance between the retrieval term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the retrieval term embedding vector.

[0010] Preferably, calculating the cluster closest to the retrieval term embedding vector in the clustering result specifically includes:

[0011] Obtain the core objects and non-core objects in each cluster, add the core objects to the set, and if a non-core object is within the neighborhood radius of at least two core objects, add the non-core object to the set;

[0012] Calculate the average value of the term embedding vectors corresponding to all objects in the set to obtain the average value of the cluster;

[0013] Calculate the distance between the retrieval term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the retrieval term embedding vector.

[0014] Preferably, obtaining the full-text retrieval result of the file according to the inverted record table and the cluster corresponding to each retrieval term specifically includes:

[0015] Obtain all elements in the cluster corresponding to each retrieval term, find the file corresponding to the element from the inverted record table, and thus obtain the file set corresponding to each retrieval term;

[0016] If a file exists in the file sets corresponding to each retrieval term at the same time, take the file as the full-text retrieval result of the file.

[0017] Preferably, sorting all the deduplicated retrieval results and then pushing them to the user specifically includes:

[0018] Divide the retrieval results into at least one group according to the neighborhood radius corresponding to the retrieval results, and sort the groups in ascending order of the neighborhood radius;

[0019] For the elements in each group, calculate the term embedding vector closest to each retrieval term embedding vector, and thus obtain the corresponding relationship and distance between the retrieval term and the term, calculate the average value of the distances between all retrieval terms and the corresponding terms, and sort the elements within the group in ascending order of the average value;

[0020] Push the retrieval results to the user according to the group sorting and the sorting of the retrieval results within the group.

[0021] Preferably, sorting the elements within the group in ascending order of the average value specifically includes:

[0022] If the average values are the same, all the maximum distances between all the search terms and the corresponding terms in the elements with the same average value are deleted at the same time, the average value is recalculated, and the smaller one of the recalculated average values is arranged before the larger one; if the recalculated average values are still the same, continue to delete the maximum distances between the remaining search terms and the corresponding terms, recalculate the average value, and arrange the smaller one of the recalculated average values before the larger one; and so on until all the elements with the same average value are sorted.

[0023] In addition, the present invention also provides an archive full-text retrieval system based on artificial intelligence, and the system includes the following modules:

[0024] A word encoding module, configured to perform word segmentation on the text in the archive to obtain terms corresponding to each archive, establish the corresponding relationship between the terms and the archive, and obtain an inverted index table; encode the input search terms of the user and each term in the inverted index table respectively to obtain corresponding search term embedding vectors and term embedding vectors;

[0025] A retrieval module, configured to start from the first neighborhood radius, cluster the term embedding vectors according to the neighborhood radius, calculate the cluster in the clustering result that is closest to the search term embedding vector, obtain the archive full-text retrieval result according to the inverted index table and the cluster corresponding to each search term, if the number of the retrieval results is less than the threshold, obtain the archive full-text retrieval result corresponding to the next neighborhood radius, and so on until the number of all the retrieval results after duplicate removal is not less than the threshold; sort all the retrieval results after duplicate removal and push them to the user; wherein, the first neighborhood radius is the smallest, and the next neighborhood radius is greater than the previous neighborhood radius.

[0026] Preferably, the calculation of the cluster in the clustering result that is closest to the search term embedding vector is specifically:

[0027] Calculate the average value of the term embedding vectors corresponding to all the core objects in each cluster;

[0028] Obtain the distance between the search term embedding vector and the average value of each cluster, and use the cluster with the smallest distance as the cluster closest to the search term embedding vector.

[0029] Preferably, the calculation of the cluster in the clustering result that is closest to the search term embedding vector is specifically:

[0030] Obtain the core objects and non-core objects in each cluster, add the core objects to the set, and if the non-core objects are within the neighborhood radius of at least two core objects, add the non-core objects to the set;

[0031] Calculate the average value of the term embedding vectors corresponding to all the objects in the set to obtain the average value of the cluster;

[0032] Calculate the distances between the retrieval term embedding vectors and the averages of each cluster, and take the cluster with the minimum distance as the cluster closest to the retrieval term embedding vector.

[0033] Preferably, obtaining the full-text retrieval result of the file according to the inverted record table and the cluster corresponding to each retrieval term is specifically as follows:

[0034] Obtain all elements in the cluster corresponding to each retrieval term, find the file corresponding to the element from the inverted record table, and thus obtain the file set corresponding to each retrieval term;

[0035] If a file exists in the file sets corresponding to each retrieval term at the same time, then take the file as the full-text retrieval result of the file.

[0036] Preferably, pushing the deduplicated all retrieval results to the user after sorting is specifically as follows:

[0037] Divide the retrieval results into at least one group according to the neighborhood radius corresponding to the retrieval results, and sort the groups in ascending order of the neighborhood radius;

[0038] For the elements in each group, calculate the term embedding vectors closest to each retrieval term embedding vector, and thus obtain the correspondence and distance between the retrieval term and the term, calculate the average value of the distances between all retrieval terms and the corresponding terms, and sort the elements within the group in ascending order of the average value;

[0039] Push the retrieval results to the user according to the group sorting and the retrieval result sorting within the group.

[0040] Preferably, sorting the elements within the group in ascending order of the average value is specifically as follows:

[0041] If the average values are the same, simultaneously delete the maximum value of the distances between all retrieval terms and the corresponding terms among the elements with the same average value, recalculate the average value, and arrange the smaller recalculated average value before the larger one; if the recalculated average values are still the same, continue to delete the maximum value of the distances between the remaining retrieval terms and the corresponding terms, recalculate the average value, and arrange the smaller recalculated average value before the larger one; and so on until all the elements with the same average value are sorted.

[0042] In addition, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described above.

[0043] The present invention encodes the retrieval keywords and the terms after word segmentation in the file by using the word embedding vectors in artificial intelligence, and then clusters the term embedding vectors by setting different neighborhood radii. Since the clustering results are different for different neighborhood radii, the terms clustered with a smaller neighborhood radius are closer. If a preset number of files can be obtained through the clustering result with a smaller neighborhood radius, the retrieval result obtained from the clustering result with a smaller neighborhood radius is directly used; otherwise, the neighborhood radius is increased and the retrieval continues. By adopting the above method, the following effects can be achieved. On the one hand, the embedding vectors can retrieve files with similar semantics, and on the other hand, the retrieval result is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is a flowchart of Embodiment 1;

[0046] Figure 2 、 3 It is a graph of clustering results with different neighborhood radii;

[0047] Figure 4 、 5 It is the clusters corresponding to the same retrieval term under different neighborhood radii;

[0048] Figure 6 It is a structural diagram of Embodiment 2. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment 1, as Figure 1 shown, the present invention provides an artificial intelligence-based full-text retrieval method for archives, and the method includes the following steps:

[0052] S1, perform word segmentation on the text in the archives to obtain the term corresponding to each archive, establish the correspondence between the term and the archive, and obtain an inverted index table; encode the user input retrieval term and each term in the inverted index table respectively to obtain the corresponding retrieval term embedding vector and term embedding vector;

[0053] In the archive database, there are archives in different forms, such as pictures, texts, etc. Before performing word segmentation on the text in the archives, if there are pictures, first identify the text part in the pictures, and construct a text database for each archive, where the text database stores the text information of the archive, and the text information includes but is not limited to the text and / or the text content of the picture. The present invention can also be directly used for the retrieval of electronic archives. Then perform word segmentation on the text in the archives, and after preprocessing, obtain the term corresponding to each archive. Among them, the word segmentation method adopts Jieba segmentation, SnowNLP, LTP, etc. The preprocessing includes but is not limited to removing stop words such as "de" and "di", and can also include the replacement of synonyms, such as replacing the original school name of a certain university with the current school name, etc. After obtaining the terms, establish an inverted index table between the terms and the archives. Among them, the inverted index table records all the documents in which a certain word appears. For example, if the word or term is xx and its corresponding archives are 1 and 13, the record in the inverted index table is xx—{1, 15}.

[0054] Then encode the user input retrieval term and the terms in the inverted index table respectively to obtain the retrieval term embedding vector and the term embedding vector. In a more detailed embodiment, the term embedding vector is used as a new column in the inverted index table, and a record in the inverted index table is [502798]—xx—{1, 15} or xx—[502798]—{1, 15}, where [502798] represents the word embedding vector of xx. It should be understood that the above numbers are only used to explain the present invention and are not limited to the above content and length.

[0055] S2. Starting from the first neighborhood radius, cluster the term embedding vectors according to the neighborhood radius, calculate the cluster in the clustering result that is closest to the retrieval term embedding vector, and obtain the full-text retrieval result of the file according to the inverted record table and the cluster corresponding to each retrieval term. If the number of the retrieval results is less than the threshold, obtain the full-text retrieval result of the file corresponding to the next neighborhood radius, and so on, until the number of all retrieval results after deduplication is not less than the threshold; sort all the retrieval results after deduplication and push them to the user; wherein, the first neighborhood radius is the smallest, and the next neighborhood radius is greater than the previous neighborhood radius.

[0056] The encoding results of different terms are different, but the embedding vectors of terms with similar semantics are also relatively close. Based on this, the present invention uses different neighborhood radii (epsilon, eps) to cluster the term embedding vectors. If the neighborhood radius is small, the obtained result is more accurate. As the neighborhood radius increases, more terms with similar semantics are included in the clustering result, and the accuracy of the retrieval result will decrease, but the number of retrieval results will be larger. The present invention first clusters the term embedding vectors with a small neighborhood radius. For example, a small neighborhood radius can cluster the two synonyms "teacher" and "instructor" into one category. As the neighborhood radius increases, the clustering result will cluster synonyms and near-synonyms into one category. For example, "teacher", "instructor", and "professor" are clustered into one category. However, if the user's retrieval intention is a middle school teacher, obviously the clustering result with a smaller neighborhood radius is more accurate. The clustering results of the same term with different neighborhood radii are as Figure 2 、 3 shown, where esp1 is less than esp2. From Figure 2 、 3 it can be seen that the clustering results with different neighborhood radii are different. In a preferred embodiment, the DBSCAN algorithm is used for clustering.

[0057] After clustering, the clustering result, that is, multiple clusters, will be obtained. For example, 10 clusters are obtained, and the cluster closest to the retrieval keyword is found. If there are multiple retrieval keywords, each retrieval keyword will have a cluster closest to it, as Figure 4 、 5 shown. From Figure 4 、 5 it can be seen that the clusters corresponding to the same keyword are different. The so-called corresponding cluster refers to the cluster closest to the retrieval term embedding vector. Among them, there are various calculation methods for the distance, including but not limited to Euclidean distance and Manhattan distance.

[0058] In another embodiment, the calculation of the cluster in the clustering result that is closest to the retrieval term embedding vector is specifically as follows:

[0059] Calculate the average value of the term embedding vectors corresponding to all core objects in each cluster;

[0060] Among them, if the number of objects existing within the neighborhood radius of an object meets the requirements, it is called a core object. The density around the core object is higher. The objects in the cluster refer to the term embedding vectors. Since the term embedding vectors correspond to terms, the objects in the cluster also refer to terms, and the two express the same meaning. If the above requirements are not met, it is a non-core object. In this embodiment, by calculating the average value of the word embedding vectors corresponding to all core objects in each cluster, the influence of non-core objects can be removed. For example, due to inaccurate clustering results, some outliers are divided into the cluster.

[0061] Then calculate the distance between the retrieved term embedding vector and the above average value, and take the cluster with the smallest distance as the cluster closest to the retrieved term embedding vector.

[0062] In another embodiment, the calculation of the cluster closest to the retrieved term embedding vector in the clustering result is specifically as follows:

[0063] Obtain the core objects and non-core objects in each cluster, add the core objects to the set, and if the non-core object is within the neighborhood radius of at least two core objects, add the non-core object to the set; in this embodiment, the non-core objects are also filtered to reduce the influence of non-core objects. Specifically, if an object in the cluster is a core object, it is directly added to the set. If it is a non-core object, it is further determined whether it is within the neighborhood radius of at least two core objects. Here, the at least two core objects refer to at least two core objects in the same cluster. If it is within the neighborhood radius of at least two core objects, add the non-core object to the set. Then, calculate the average value of the term embedding vectors corresponding to all objects in the set to obtain the average value of the cluster; further calculate the distance between the retrieved term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the retrieved term embedding vector.

[0064] In an alternative way in the above embodiment, after obtaining the set, calculate the distance between the retrieved term embedding vector and each element in the set, and take the cluster with the smallest average value of the distances as the closest cluster.

[0065] After obtaining the cluster corresponding to each retrieved term, the specific method for obtaining the full-text retrieval result of the file according to the inverted index table and the cluster corresponding to each retrieved term is as follows:

[0066] Obtain all elements in the cluster corresponding to each retrieved term, find the files corresponding to the elements from the inverted index table, and thus obtain the file set corresponding to each retrieved term;

[0067] If a file exists in each file set corresponding to each search term, the file is used as the full-text search result of the file.

[0068] For example, there are two search terms. Search term 1 corresponds to cluster 1, and search term 2 corresponds to cluster 11. Since a cluster includes multiple term embedding vectors, and the term embedding vectors correspond to terms, through the above relationship, each search term corresponds to multiple terms, and each term corresponds to multiple files. This obtains the files corresponding to each search term. For example, search term 1 corresponds to files 1, 5, 7, and 10, and search term 2 corresponds to files 2, 5, 8, and 10. If the search terms entered by the user are only search term 1 and search term 2, then taking the intersection gives the full-text search result of the file as 5 and 10.

[0069] After obtaining the full-text search result of the file, it is necessary to push the result to the user. The higher the relevance, the more forward it should be sorted, so as to facilitate the user to quickly obtain the most relevant file search result. In one embodiment, the de-duplicated all search results are sorted and then pushed to the user, specifically:

[0070] The search results are divided into at least one group according to the neighborhood radius corresponding to the search results, and the groups are sorted in ascending order of the neighborhood radius;

[0071] First, the search results are divided into at least one group according to the neighborhood radius. The number of groups is the same as the number of neighborhood radii. There are multiple grouping methods. In a more specific embodiment, start from the group with a small neighborhood radius. If a file is retrieved through two neighborhoods, the file is placed in the group with a small neighborhood radius. For example, the search result with a neighborhood radius of 1 includes file 1, and the search result with a neighborhood radius of 2 also includes file 1. Then file 1 is placed in the group corresponding to the neighborhood radius of 1, where the neighborhood radius of 1 is less than the neighborhood radius of 2. Then the groups corresponding to the neighborhood radii are sorted in ascending order of the neighborhood radius.

[0072] For each element in the group, calculate the term embedding vector that is closest to each search term embedding vector, and then obtain the corresponding relationship and distance between the search term and the term. Calculate the average value of the distances between all search terms and the corresponding terms, and sort the elements within the group in ascending order of the average value;

[0073] After sorting the groups, it is necessary to further sort the elements within each group, calculate the distance between the search terms and each element within the group, obtain the term embedding vector closest to each search term, and then obtain the correspondence and distance between the search terms and the terms. Then, calculate the average value of the distances between all search terms and their corresponding terms. Suppose there are three elements in the group, namely archives, and the average distances between all search terms and each element within the group are 5, 18, and 9 respectively. Sorting the elements within the group in ascending order of the average values results in 1, 3, 2. Finally, push the search results to the user according to the group sorting and the sorting of the search results within the group. First, sort the groups, and then sort the elements within the group. For example, the elements of group 2 are always behind the elements of group 1, and the sorting of the elements within the group follows the sorting result within the group.

[0074] If the average values of the distances between all search terms and the archives are the same, the elements with the same average value within the group can be randomly sorted. In a more specific embodiment, the sorting of the elements within the group in ascending order of the average values is specifically as follows:

[0075] If the average values are the same, simultaneously delete the maximum value of the distances between all search terms and their corresponding terms among the elements with the same average value, recalculate the average value, and arrange the smaller recalculated average value before the larger one; if the recalculated average values are still the same, continue to delete the maximum value of the distances between the remaining search terms and their corresponding terms, recalculate the average value, and arrange the smaller recalculated average value before the larger one; and so on until all the elements with the same average value are sorted.

[0076] Embodiment 2, the present invention also provides an archive full-text retrieval system based on artificial intelligence, as Figure 6 shown, the system includes the following modules:

[0077] A word encoding module, which is used to segment the text in the archives to obtain the terms corresponding to each archive, establish the correspondence between the terms and the archives, and obtain an inverted record table; encode the user input search terms and each term in the inverted record table respectively to obtain the corresponding search term embedding vector and term embedding vector;

[0078] A retrieval module, which starts from the first neighborhood radius, clusters the term embedding vectors according to the neighborhood radius, calculates the cluster closest to the search term embedding vector in the clustering result, obtains the archive full-text retrieval result according to the inverted record table and the cluster corresponding to each search term. If the number of the retrieval results is less than the threshold, obtain the archive full-text retrieval result corresponding to the next neighborhood radius, and so on until the number of all the retrieval results after deduplication is not less than the threshold; sort all the deduplicated retrieval results and push them to the user; where the first neighborhood radius is the smallest, and the next neighborhood radius is greater than the previous one.

[0079] Preferably, calculating the cluster in the clustering result that is closest to the retrieval term embedding vector specifically involves:

[0080] Calculating the average value of the term embedding vectors corresponding to all core objects in each cluster;

[0081] Obtaining the distances between the retrieval term embedding vector and the average value of each cluster, and taking the cluster with the minimum distance as the cluster closest to the retrieval term embedding vector.

[0082] Preferably, calculating the cluster in the clustering result that is closest to the retrieval term embedding vector specifically involves:

[0083] Obtaining the core objects and non-core objects in each cluster, adding the core objects to a set, and if a non-core object is within the neighborhood radius of at least two core objects, adding the non-core object to the set;

[0084] Calculating the average value of the term embedding vectors corresponding to all objects in the set to obtain the average value of the cluster;

[0085] Calculating the distances between the retrieval term embedding vector and the average value of each cluster, and taking the cluster with the minimum distance as the cluster closest to the retrieval term embedding vector.

[0086] Preferably, obtaining the full-text retrieval result of the file according to the inverted record table and the cluster corresponding to each retrieval term specifically involves:

[0087] Obtaining all elements in the cluster corresponding to each retrieval term, finding the files corresponding to the elements from the inverted record table, and thus obtaining the file set corresponding to each retrieval term;

[0088] If a file exists in the file sets corresponding to each retrieval term at the same time, taking the file as the full-text retrieval result of the file.

[0089] Preferably, sorting all the retrieved results after duplicate removal and pushing them to the user specifically involves:

[0090] Dividing the retrieved results into at least one group according to the neighborhood radius corresponding to the retrieved results, and sorting the groups in ascending order of the neighborhood radius;

[0091] For the elements in each group, calculating the term embedding vector closest to each retrieval term embedding vector, and thus obtaining the corresponding relationship and distance between the retrieval term and the term, calculating the average value of the distances between all retrieval terms and the corresponding terms, and sorting the elements within the group in ascending order of the average value;

[0092] Pushing the retrieved results to the user according to the group sorting and the sorting of the retrieved results within the group.

[0093] Preferably, sorting the elements within the group in ascending order of the average value specifically involves:

[0094] If the average values are the same, simultaneously delete the maximum value of the distances between all search terms and the corresponding terms among the elements with the same average value, recalculate the average value, and arrange the element with the smaller recalculated average value before the one with the larger recalculated average value; if the recalculated average values are still the same, continue to delete the maximum value of the distances between the remaining search terms and the corresponding terms, recalculate the average value, and arrange the element with the smaller recalculated average value before the one with the larger recalculated average value; and so on, until all the elements with the same average value are sorted out.

[0095] Embodiment 3. The present invention further provides a computer program product, including a computer program which, when executed by a processor, implements the method as described in Embodiment 1.

[0096] Embodiment 4. The present invention further provides a computer-readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the method as described in Embodiment 1.

[0097] Embodiment 5. The present invention further provides a computing device, including a memory and a processor, on which a computer program is stored, and the computer program, when executed by a processor, implements the method as described in Embodiment 1.

[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Other embodiments can also be adopted; although the present invention has been described in detail with reference to the foregoing embodiments, those ordinary skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An artificial intelligence-based full-text retrieval method for archives, characterized in that, The method includes the following steps: Segment the text in the archives to obtain the terms corresponding to each archive, establish the correspondence between the terms and the archives, and obtain the inverted index table; Encode the search terms input by the user and each term in the inverted index table respectively to obtain the corresponding search term embedding vectors and term embedding vectors; Starting from the first neighborhood radius, cluster the term embedding vectors according to the neighborhood radius, calculate the cluster in the clustering result that is closest to the search term embedding vector, obtain the full-text search result of the archives according to the inverted index table and the cluster corresponding to each search term. If the number of the search results is less than the threshold, obtain the full-text search result of the archives corresponding to the next neighborhood radius, and so on, until the number of all search results after deduplication is not less than the threshold; Sort all the search results after deduplication and push them to the user; Among them, the first neighborhood radius is the smallest, and the next neighborhood radius is greater than the previous neighborhood radius; The calculation of the cluster in the clustering result that is closest to the search term embedding vector is specifically: Calculate the average value of the term embedding vectors corresponding to all core objects in each cluster; Obtain the distance between the search term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the search term embedding vector; The calculation of the cluster in the clustering result that is closest to the search term embedding vector is specifically: Obtain the core objects and non-core objects in each cluster, add the core objects to the set, and if the non-core object is within the neighborhood radius of at least two core objects, add the non-core object to the set; Calculate the average value of the term embedding vectors corresponding to all objects in the set to obtain the average value of the cluster; Calculate the distance between the search term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the search term embedding vector; The obtaining of the full-text search result of the archives according to the inverted index table and the cluster corresponding to each search term is specifically: Obtain all elements in the cluster corresponding to each search term, find the archives corresponding to the elements from the inverted index table, and thus obtain the archive set corresponding to each search term; If an archive exists in the archive sets corresponding to each search term at the same time, take the archive as the full-text search result of the archives; The sorting of all the search results after deduplication and pushing them to the user is specifically: Divide the search results into at least one group according to the neighborhood radius corresponding to the search results, and sort the groups in ascending order of the neighborhood radius; For the elements in each group, calculate the term embedding vector that is closest to each search term embedding vector, and thus obtain the correspondence and distance between the search term and the term, calculate the average value of the distances between all search terms and the corresponding terms, and sort the elements in the group in ascending order of the average value; Push the search results to the user according to the group sorting and the sorting of the search results within the group; The sorting of the elements in the group in ascending order of the average value is specifically: If the average values are the same, simultaneously delete the maximum distance between all search terms and the corresponding terms among the elements with the same average value, recalculate the average value, and arrange the elements with the smaller recalculated average value before those with the larger one; if the recalculated average values are still the same, continue to delete the maximum distance between the remaining search terms and the corresponding terms, recalculate the average value, and arrange the elements with the smaller recalculated average value before those with the larger one; and so on until all elements with the same average value are sorted.

2. An artificial intelligence-based full-text retrieval system for archives, characterized in that, The system includes the following modules: A word encoding module, which is used to segment the text in the archives to obtain the terms corresponding to each archive, establish the correspondence between the terms and the archives, and obtain an inverted index table; encode the input search terms of the user and each term in the inverted index table respectively to obtain the corresponding search term embedding vectors and term embedding vectors; A retrieval module, which starts from the first neighborhood radius, clusters the term embedding vectors according to the neighborhood radius, calculates the cluster in the clustering result that is closest to the search term embedding vector, obtains the full-text retrieval result of the archives according to the inverted index table and the cluster corresponding to each search term. If the number of the retrieval results is less than the threshold, obtain the full-text retrieval result of the archives corresponding to the next neighborhood radius, and so on until the number of the de-duplicated retrieval results is not less than the threshold; sort all the de-duplicated retrieval results and push them to the user; where the first neighborhood radius is the smallest, and the next neighborhood radius is greater than the previous one; The calculation of the cluster in the clustering result that is closest to the search term embedding vector is specifically: Calculate the average value of the term embedding vectors corresponding to all core objects in each cluster; Obtain the distance between the search term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the search term embedding vector; The calculation of the cluster in the clustering result that is closest to the search term embedding vector is specifically: Obtain the core objects and non-core objects in each cluster, add the core objects to the set, and if a non-core object is within the neighborhood radius of at least two core objects, add the non-core object to the set; Calculate the average value of the term embedding vectors corresponding to all objects in the set to obtain the average value of the cluster; Calculate the distance between the search term embedding vector and the average value of each cluster, and take the cluster with the smallest distance as the cluster closest to the search term embedding vector; The obtaining of the full-text retrieval result of the archives according to the inverted index table and the cluster corresponding to each search term is specifically: Obtain all elements in the cluster corresponding to each search term, find the archives corresponding to the elements from the inverted index table, and thus obtain the archive set corresponding to each search term; If an archive exists in the archive sets corresponding to each search term at the same time, take the archive as the full-text retrieval result of the archives; The sorting of all the de-duplicated retrieval results and pushing them to the user is specifically: Divide the retrieval results into at least one group according to the neighborhood radius corresponding to the retrieval results, and sort the groups in ascending order of the neighborhood radius; For each element in a group, calculate the term embedding vector that is closest to the retrieval term embedding vector, thereby obtaining the correspondence and distance between the retrieval term and the term. Calculate the average value of the distances between all retrieval terms and their corresponding terms, and sort the elements within the group in ascending order of the average value; Push the retrieval results to the user according to the group sorting and the sorting of the retrieval results within the group; The sorting of the elements within the group in ascending order of the average value is specifically as follows: If the average values are the same, simultaneously delete the maximum value of the distances between all retrieval terms and their corresponding terms among the elements with the same average value, recalculate the average value, and rank the element with the smaller recalculated average value before the one with the larger recalculated average value; if the recalculated average values are still the same, continue to delete the maximum value of the distances between the remaining retrieval terms and their corresponding terms, recalculate the average value, and rank the element with the smaller recalculated average value before the one with the larger recalculated average value; and so on until the sorting of all elements with the same average value is completed.

3. A computer program product comprising computer-executable instructions, characterized in that, The instruction, when executed, is used to implement the method described in claim 1.

Citation Information

Patent Citations

  • Cloud storage based power full text retrieval method and system

    CN102156711A

  • Label extracting method and device, apparatus and medium

    CN107861948A

  • Multi-model entropy weighted retrieval method and system based on Kmeans recall

    CN115309872A