A fast clustering retrieval method for large-scale business document data
By calculating the search matching degree, browsing confidence and search correlation degree of documents, clustering and false positive and false negative analysis, the problems of low search speed and efficiency of existing Chinese documents are solved, and more accurate document search results are achieved.
Patent Information
- Application Number
- CN202411716171.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-27
AI Technical Summary
When existing clustering methods process large-scale business document data, it is difficult to accurately identify semantic fuzzy and abstract features, resulting in false positive or false negative documents, which reduces the search speed and efficiency.
By obtaining the word vectors of the target search terms and historical search terms, calculating the similarity and browsing time, determining the search matching, browsing confidence and search correlation of the document, performing clustering, analyzing the possibility of false positives and false negatives, and adjusting the search category.
It improves the accuracy of document clustering, improves the speed and efficiency of document search, and makes the search results more accurate.
Smart Images

Figure CN119646216B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method for rapid clustering and retrieval of large-scale business document data. Background Art
[0002] Using clustering method to retrieve documents is a common method for processing large-scale data. Since the semantics of documents are extensible, when clustering documents through search terms, it is difficult to find clear, low-level semantic features to divide cluster boundaries due to the fuzzy and abstract characteristics of semantics. As a result, there will be false positive or false negative documents in the cluster, resulting in the erroneous retrieval of documents that are irrelevant to the search terms, and the failure to retrieve documents that are relevant to the search terms. As a result, the speed and efficiency of document retrieval are low during the process of retrieving documents based on the search terms, which cannot meet the requirements of accurate retrieval of documents, thus affecting the accuracy of the clustering results. Summary of the invention
[0003] In order to solve the above technical problems, a large-scale business document data fast clustering retrieval method is provided to solve the existing problems.
[0004] The solution to the technical problem of this application is to provide a method for rapid clustering retrieval of large-scale business document data, including the following steps:
[0005] Obtain the target search term of the current searcher; record other authority levels lower than the authority level of the current searcher as each subordinate level, and record the corresponding authority level of the current searcher and each subordinate level as each gradient level; obtain the historical search terms and document browsing time corresponding to each historical search record at each gradient level; respectively, form a search library for each level with all the documents that can be retrieved by the searchers at each gradient level;
[0006] According to the frequency of occurrence of the target search term in each document in the search library at any level, the search matching degree of each document in the search library at any level is determined; according to the similarity between the historical search term corresponding to each historical search record at each gradient level and the target search term, and the browsing time of each document in the search library at any level, the browsing confidence of each document in the search library at any level is determined; combined with the search matching degree, the search relevance of each document in the search library at any level is determined;
[0007] Clustering the search relevance of all documents in the search library at any level respectively, obtaining the search-related classes and search-irrelevant classes of the corresponding authority level of the current searcher, and the reference-related classes of any subordinate level;
[0008] Analyze the correlation between the documents in the search-related class, the search-irrelevant class and the reference-related class, and the difference between the corresponding authority level of the current searcher and any subordinate level, to determine the false positive possibility of each document in the search-related class and the false negative possibility of each document in the search-irrelevant class;
[0009] Based on the false positive possibility and the false negative possibility, the false positive documents within the retrieval related class and the false negative documents within the retrieval irrelevant class are determined; the false positive documents and the false negative documents are added and deleted to obtain adjusted retrieval related classes; based on the adjusted retrieval related classes, the retrieval results of the target search terms of the current searcher are displayed.
[0010] Preferably, the determining of the search matching degree of each document in the search library at any level includes:
[0011] Count the frequency of each keyword in the current searcher's target search term appearing in each document in the search library at any level;
[0012] The average frequency of all keywords in the target search term of the current searcher in each document in the search library at any level is used as the search matching degree of each document in the search library at any level.
[0013] Preferably, determining the browsing confidence of each document in the search library at any level includes:
[0014] Respectively obtain the word vector of the target search term and the word vector of the historical search term;
[0015] Calculating the similarity between the word vector of the historical search term and the word vector of the target search term;
[0016] Record all historical search records corresponding to the similarity degree greater than a preset relevant threshold at each gradient level as relevant search records;
[0017] Calculate the product of the browsing time of the book in any relevant search record at each gradient level and the absolute value of the similarity of the historical search term corresponding to the any relevant search record, and record it as the first product;
[0018] For all relevant search records at each gradient level, the normalized result of the mean of the first product in all relevant search records corresponding to each browsed document is used as the browsing confidence of each document browsed in the search library of any level, and the browsing confidence of the remaining documents in the search library of any level that have not been browsed in the relevant search records is assigned a preset value.
[0019] Preferably, the search relevance of each document in the search library at any level is the product of the search matching degree and the browsing confidence.
[0020] Preferably, the step of obtaining the search-related classes and search-irrelevant classes corresponding to the authority level of the current searcher includes:
[0021] Calculate the mean of the search relevance of all documents in the cluster corresponding to any level of search library, and record it as intra-cluster relevance;
[0022] The cluster with the largest correlation within the cluster in the level retrieval library corresponding to the authority level of the current search personnel is recorded as the retrieval-related class for the authority level of the current search personnel, and the remaining clusters in the level retrieval library corresponding to the authority level of the current search personnel are recorded as the retrieval-irrelevant classes for the authority level of the current search personnel.
[0023] Preferably, the method for obtaining the reference related class of any subordinate level is:
[0024] The cluster with the largest intra-cluster correlation in the level search library corresponding to any subordinate level is recorded as the reference related class of any subordinate level.
[0025] Preferably, determining the false positive probability of each document in the search-related category includes:
[0026] Calculate the difference between the authority level corresponding to the current searcher and any subordinate level, and record it as the level difference;
[0027] Calculating the similarity between each document in the search-related category and each document in any subordinate reference-related category, and recording it as a first similarity;
[0028] Selecting the maximum value of the first similarity between each document in the search-related class and all documents in any subordinate reference-related class, and recording the product of the maximum value of the first similarity and the level difference as the second product;
[0029] The normalized result of the mean of the second product of each document in the search-related class and all its subordinate levels is used as the false positive possibility of each document in the search-related class.
[0030] Preferably, the method for determining the false negative possibility of each document in the retrieved irrelevant category is:
[0031] Calculating the similarity between each document in the search irrelevant class and each document in any subordinate reference relevant class, recorded as the second similarity;
[0032] Selecting the maximum value of the second similarity between each document in the search irrelevant class and all documents in the reference relevant class at any subordinate level, and multiplying the maximum value of the second similarity by the level difference as the third product;
[0033] The normalized result of the mean of the third product of each document in the irrelevant class and all its subordinate levels is used as the false negative possibility of each document in the irrelevant class.
[0034] Preferably, the determining of the false positive documents in the search-related category and the false negative documents in the search-irrelevant category includes:
[0035] Record the document whose false positive probability in the search-related category is less than the preset first threshold as a false positive document;
[0036] The document corresponding to the false negative possibility in the irrelevant category of the retrieval being greater than the preset second threshold value is recorded as a false negative document.
[0037] Preferably, obtaining the adjusted retrieval-related class includes:
[0038] All false positive documents are deleted from the search-related class, all false negative documents in the search-irrelevant class are added to the search-related class, and all false negative documents are deleted from the search-irrelevant class to obtain an adjusted search-related class.
[0039] This application has at least the following beneficial effects:
[0040] The present application determines the search matching degree of each document in the search library at any level according to the frequency of the target search term in each document in the search library at any level. The beneficial effect is that the occurrence of the target search term in the document is taken into account to reflect the matching degree between the document and the target search term; the browsing confidence of each document in the search library at any level is determined according to the similarity between the historical search term corresponding to each historical search record at each gradient level and the target search term, and the browsing time of each document in the search library at any level; the search association of each document in the search library at any level is determined in combination with the search matching degree. The beneficial effect is that the correlation between the historical search terms and the target search terms in the historical search records is taken into account, so as to analyze the browsing time of the documents in the search records related to the target search terms, and further reflect the correlation between the documents and the target search terms; the search correlation of all the documents in the search library of any level is clustered respectively, and the search related classes and search irrelevant classes of the corresponding authority level of the current search personnel, as well as the reference related classes of any subordinate level are obtained; the correlation between the search related classes, the search irrelevant classes and the reference related classes, as well as the correlation between the current search personnel's corresponding authority level and the reference related classes, are analyzed. The beneficial effect of clustering the documents in each level of the search library is to divide the documents into two categories, and then analyze whether there are dissimilarities between the documents in the search-related category and the documents in the reference-related category of the subordinate level, so as to reflect whether the documents in the search-related category are false positive documents. Secondly, analyze whether there are similarities between the documents in the search-irrelevant category and the documents in the reference-related category of the subordinate level, so as to reflect whether the documents in the search-irrelevant category are false negative documents, thereby indicating whether there are documents irrelevant to the target search term that are clustered into In the retrieval-related class, documents related to the target search term are clustered into the retrieval-irrelevant class; based on the false positive possibility and the false negative possibility, the false positive documents in the retrieval-related class and the false negative documents in the retrieval-irrelevant class are determined; the false positive documents and the false negative documents are added and deleted to obtain an adjusted retrieval-related class, and based on the adjusted retrieval-related class, the retrieval results of the target search term of the current searcher are displayed, which has the beneficial effect of adjusting and correcting the documents in the retrieval-related class and the retrieval-irrelevant class, improving the accuracy of document clustering, improving the speed and efficiency of document retrieval, and making the retrieval results of documents more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The following is a further detailed description of a large-scale business document data rapid clustering retrieval method of the present application in conjunction with the accompanying drawings.
[0042] Figure 1A flowchart of the steps of a method for rapid clustering retrieval of large-scale business document data provided in an embodiment of the present application;
[0043] Figure 2 A flowchart of the steps of a method for obtaining the search relevance of each document in the search library at any level provided in an embodiment of the present application;
[0044] Figure 3 A flowchart of the steps of the method for obtaining the false positive possibility of each document in the retrieval-related category provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of this application more clear, the following is a further detailed description of a large-scale business document data fast clustering retrieval method proposed in this application in conjunction with the accompanying drawings and implementation examples. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.
[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0047] See also Figure 1 , which shows a flowchart of a method for rapid clustering retrieval of large-scale business document data provided by an embodiment of the present application, the method comprising the following steps:
[0048] Step 1: Obtain the target search terms of the current searcher, as well as the historical search terms and document browsing time corresponding to each historical search record at each gradient level.
[0049] To ensure the information security in the document database, access permissions are set for personnel at different levels. Therefore, when search personnel at different levels search for documents through the database, the scope of the documents retrieved may be different. Generally, search personnel with higher authority levels can view core documents. Core documents are often general and do not go into specific operational details, but provide a macro framework. The semantic information expressed in core documents is often vague and abstract. When clustering documents through search terms, it is difficult to find clear, low-level semantic features to divide the clustering boundaries, resulting in false positive or false negative documents in the clustering results.
[0050] When a searcher has a higher level of access rights, he or she is more likely to retrieve core documents, while when a searcher has a lower level of access rights, he or she is more likely to retrieve ordinary documents. Ordinary documents are extended from core documents. For example, in a task document issued once, a person with a high level of authority issues a core document corresponding to the task document. When the core document is issued to each level for execution and other specific operations, ordinary documents will be generated. Therefore, the semantics in ordinary documents are broader than those in core documents, so that the semantics in the task documents seen by searchers of different levels have a relationship of inclusion and being included. The documents retrieved by subordinates with a lower level of authority than the searcher must semantically contain the semantics in the documents retrieved by the superior searcher, that is, the documents retrieved by the subordinates are a subset or related extension of the content of the documents retrieved by the superior searcher.
[0051] Based on the above analysis, the target search term of the current searcher is obtained, and according to the log of the document database, the historical search terms and the browsing time of the document corresponding to each historical search record of the searcher at the corresponding authority level of the current searcher are obtained, as well as the historical search terms and the browsing time of the document corresponding to each historical search record of the searcher at other authority levels lower than the authority level of the current searcher, where one historical search record corresponds to one historical search term.
[0052] Other permission levels lower than the permission level of the current searcher are recorded as each subordinate level, and the corresponding permission level of the current searcher and each subordinate level are recorded as each gradient level; finally, the browsing time of the historical search terms and documents corresponding to each historical search record at each gradient level is obtained.
[0053] At this point, the target search term of the current searcher is obtained, and the historical search term and document browsing time corresponding to each historical search record at each gradient level are obtained.
[0054] Step 2: Determine the search matching degree of each document in the search library at any level based on the frequency of occurrence of the target search term in each document in the search library at any level; determine the browsing confidence of each document in the search library at any level based on the similarity between the historical search term corresponding to each historical search record at each gradient level and the target search term, as well as the browsing time of each document in the search library at any level; and determine the search relevance of each document in the search library at any level in combination with the search matching degree.
[0055] The higher the frequency of the target search term in a document, the greater the match between the corresponding document and the target search term, that is, the higher the relevance between the corresponding document and the target search term; secondly, the types of documents retrieved by searchers at the same level are similar. The corresponding documents can be analyzed through the historical search records of each gradient level, and the correlation between the target search term and the document can be determined to determine the search relevance.
[0056] Furthermore, the flowchart of the steps of the method for obtaining the search relevance of each document in the search library at any level provided in the embodiment of the present application is as follows: Figure 2 shown.
[0057] First, when searching, the search personnel search for multiple keywords. That is, a set of target search terms may contain multiple keywords. By analyzing the frequency of keywords in the target search terms in the documents, the search matching degree is determined, specifically:
[0058] All documents that can be retrieved by searchers at each level constitute the search library at each level;
[0059] Count the total number of words in each document in any level of search database;
[0060] Count the number of times each keyword in the current searcher's target search term appears in each document in the search library at any level;
[0061] The ratio of the number of times to the total number of words is recorded as frequency, and the average of the frequencies of all keywords in the target search term of the current searcher in each document of the search library at any level is taken as the search matching degree of each document in the search library at any level;
[0062] Preferably, in this embodiment, the calculation formula for the search matching degree of each document in the search library at any level is: Among them, D n,m is the search matching degree of the mth document in the nth level search library, n n,m,r N is the number of times the rth keyword in the current searcher's target search term appears in the mth document in the nth level search library. n,m is the total number of words in the mth document in the nth level search library, and R is the number of all keywords in the target search term of the current searcher; secondly, is the frequency.
[0063] It should be noted that, when the frequency of the keywords in the target search term appearing in the mth document in the nth level search library is higher, the obtained search matching degree is greater, which means that the corresponding document is more matched with the target search term.
[0064] Furthermore, when the historical search terms in the historical search records are more similar to the target search terms of the current searcher, and the browsing time of the corresponding document in the historical search records is longer, it means that the document has greater reference value; therefore, the similarity between the target search terms of the current searcher and the historical search terms is analyzed, specifically:
[0065] The word vector algorithm model is used to obtain the word vector of the target search term of the current searcher and the word vector of the historical search term corresponding to each historical search record at each gradient level;
[0066] Preferably, as an embodiment of the present application, a BERT (Bidirectional Encoder Representations from Transformers) algorithm is used to obtain word vectors, wherein the BERT algorithm is an existing well-known technology and is not elaborated in this embodiment. As other implementation methods, implementers can adopt other methods of the existing technology, for example, using GloVe or ELMo algorithms to obtain word vectors, etc. This embodiment does not impose any special restrictions on this.
[0067] Calculate the similarity between the word vector of the historical search term corresponding to each historical search record at each gradient level and the word vector of the target search term;
[0068] Preferably, in this embodiment, the cosine similarity between the word vector of the historical search term corresponding to each historical search record at each gradient level and the word vector of the target search term is calculated, wherein the calculation of cosine similarity is a well-known technology and will not be repeated here. As other implementation methods, the implementer may adopt other methods of the prior art, such as the Pearson correlation coefficient, etc., and this embodiment does not impose any special restrictions on this.
[0069] Record all historical search records corresponding to the similarity degree greater than a preset relevant threshold at each gradient level as relevant search records;
[0070] Preferably, in this embodiment, the preset relevant threshold value is 0.5. As other implementation modes, the implementer can set it according to the actual situation.
[0071] It should be noted that the greater the similarity, the greater the similarity between the historical search terms and the target search terms.
[0072] Furthermore, when the browsing time of the corresponding document in a relevant search record is longer, it means that the document is more referenceable and the matching degree between the document and the target search term is higher. When the similarity between the historical keywords corresponding to the relevant search record and the target search term is higher, the browsing time represented by the relevant search record is more credible. Therefore, the browsing time of the document corresponding to each relevant search record is analyzed to determine the browsing confidence, which is specifically:
[0073] Calculate the product of the browsing time of the book in any relevant search record at each gradient level and the absolute value of the similarity of the historical search term corresponding to the any relevant search record, and record it as the first product;
[0074] For all relevant search records at each gradient level, the normalized result of the mean of the first product in all relevant search records corresponding to each browsed document is used as the browsing confidence of each document browsed in the search library of any level, and the browsing confidence of the remaining documents in the search library of any level that have not been browsed in the relevant search records is assigned a preset value.
[0075] Preferably, in this embodiment, the preset value is 0.1. As for other implementation modes, the implementer can set it according to the actual situation.
[0076] Preferably, in this embodiment, the browsing confidence of each document browsed in the search library at any level is calculated as follows: Among them, γ n,m is the browsing confidence of the mth document browsed in the nth level search library, T n,m,q is the browsing time of the mth document browsed in the nth level search library in the qth related search record, Corr n,q is the similarity of the historical search terms corresponding to the qth related search record under the nth permission level, Q n,m is the number of all relevant search records corresponding to the mth document browsed in the nth level search library, || indicates the calculation of absolute value, norm{} is the normalization function, in this embodiment, the sigmoid function is used for normalization, as other implementation methods, the implementer can use other methods in the prior art, such as tanh function, etc., and this embodiment does not impose any special restrictions on this; secondly, T n,m,q ×|Corr n,q | is the first product.
[0077] It should be noted that, by using the similarity as a weight, the higher the similarity between the historical keywords corresponding to the relevant search records and the target search terms, the more credible the browsing time represented by the relevant search records, and the greater the browsing confidence obtained.
[0078] Further, based on the search matching degree and the browsing confidence, it is determined that:
[0079] The product of the search matching degree and the browsing confidence degree is used as the search relevance degree of each document in the search library at any level;
[0080] It should be noted that the greater the search relevance, the more relevant the corresponding document is to the target search term.
[0081] At this point, the search relevance of each document in the search library at any level is obtained.
[0082] Step 3, clustering the search relevance of all documents in the search library of any level, respectively, to obtain the search-related classes and search-irrelevant classes of the current search personnel's corresponding authority level, and the reference-related classes of any subordinate level; analyzing the correlation between the search-related classes, the search-irrelevant classes and the reference-related classes, respectively, and the difference between the current search personnel's corresponding authority level and any subordinate level, to determine the false positive possibility of each document in the search-related class and the false negative possibility of each document in the search-irrelevant class.
[0083] Furthermore, based on the search relevance, the documents are clustered, specifically:
[0084] Perform cluster analysis on the search relevance of all documents in each level of search library to obtain cluster clusters corresponding to each level of search library;
[0085] Preferably, in this embodiment, the k-means clustering algorithm is used for clustering, wherein the number of clusters is set to 2, and all documents in each level of retrieval library are divided into two categories; wherein the k-means clustering algorithm is a well-known technology and will not be described in detail here. As other implementation methods, implementers may adopt other methods of the prior art, such as DBSCAN clustering, etc., and this embodiment does not impose any special restrictions on this.
[0086] Calculate the mean of the search correlation of all documents in the cluster corresponding to each level of search library, and record it as the intra-cluster correlation;
[0087] The cluster with the largest correlation in the level search library corresponding to the authority level of the current searcher is recorded as the search-related class corresponding to the authority level of the current searcher, and the remaining clusters in the level search library corresponding to the authority level of the current searcher are recorded as the search-irrelevant class corresponding to the authority level of the current searcher;
[0088] Record the cluster with the largest correlation within the cluster in the level search library corresponding to any subordinate level lower than the corresponding authority level of the current searcher as the reference related class of any subordinate level;
[0089] It should be noted that the search-related class and the reference-related class represent documents related to the target search term, and the search-irrelevant class represents documents unrelated to the target search term.
[0090] Furthermore, due to the abstract limitations of the document semantics of the level retrieval library corresponding to the authority level of the search personnel, documents irrelevant to the target search terms may appear in the retrieval-related classes. These documents are called false positive documents. Documents related to the target search terms may appear in the retrieval-irrelevant classes. These documents are called false negative documents. Due to the possible existence of false positive documents and false negative documents, the clustering effect is not ideal.
[0091] Since ordinary documents are extensions and expansions based on the content of core documents, the semantic information of ordinary documents is more extensive. The reference-related class and the retrieval-related class have a relationship of inclusion and being included. When there is no document in the reference-related class that is similar to a document in the retrieval-related class, it means that the document is likely to be a false positive document; when there is a document in the reference-related class that is similar to a document in the retrieval-irrelevant class, it means that the document is likely to be a false negative document.
[0092] Furthermore, the flowchart of the steps of the method for obtaining the false positive possibility of each document in the search related category provided in the embodiment of the present application is as follows: Figure 3 shown.
[0093] Based on the above analysis, we analyze the similarity between the documents in the retrieval-related category and the documents in the reference-related category, remove the false positive documents in the retrieval-related category, analyze the similarity between the documents in the retrieval-irrelevant category and the documents in the reference-related category, and remove the false negative documents in the retrieval-irrelevant category. Specifically:
[0094] Calculate the difference between the level corresponding to the current search person and any subordinate level, and record it as the level difference;
[0095] Calculating the similarity between each document in the search-related category and each document in any subordinate reference-related category, and recording it as a first similarity;
[0096] Preferably, in this embodiment, the Jaccard similarity coefficient between each document in the search related class and each document in any subordinate reference related class is calculated, wherein the calculation of the Jaccard (Jaccard similarity coefficient) similarity coefficient is a well-known technology and will not be described in detail here.
[0097] Selecting the maximum value of the first similarity between each document in the search-related class and all documents in any subordinate reference-related class, and recording the product of the maximum value of the first similarity and the level difference as the second product;
[0098] taking the normalized result of the mean of the second product of each document in the search-related class and all its subordinate levels as the false positive possibility of each document in the search-related class;
[0099] Preferably, in this embodiment, the calculation method of the false positive probability of each document in the search related category is: Among them, PF m is the false positive probability of the mth document in the retrieval-related category, F m is the mth document in the retrieval related category, E h,g is the g-th instrument in the reference-related class of the h-th subordinate level, α h is the difference between the level corresponding to the current search person and the hth subordinate level, that is, the level difference; Jac() represents the calculation of the Jaccard similarity coefficient, max[] represents the maximum value function, H is the number of all subordinate levels under the authority level corresponding to the current search person, norm{} is the normalization function, in this embodiment, the sigmoid function is used for normalization, as other implementation methods, the implementer can use other methods of the prior art, such as the tanh function, etc., and this embodiment does not impose any special restrictions on this; secondly, max[Jac(F m ,E h,g )]×α h is the second product.
[0100] It should be noted that, when the subordinate level is lower, the scalability of the documents retrieved by the search personnel at the corresponding subordinate level is stronger. If the difference between each subordinate level and the level of the current search personnel is greater, the weight of the similarity is greater. If each document in the retrieval-related class is more similar to the document in the reference-related class, that is, the higher the similarity, the smaller the possibility that the corresponding document in the retrieval-related class is a false positive, the smaller the possibility of a false positive, the greater the possibility that the corresponding document in the retrieval-related class is related to the target search term. Conversely, the greater the possibility of a false positive, the greater the possibility that the corresponding document in the retrieval-related class is irrelevant to the target search term.
[0101] Furthermore, the possibility of false negatives is determined for documents in irrelevant categories, specifically:
[0102] Calculating the similarity between each document in the search irrelevant class and each document in any subordinate reference relevant class, recorded as the second similarity;
[0103] Preferably, in this embodiment, the Jaccard similarity coefficient between each document in the retrieval irrelevant class and each document in any subordinate reference relevant class is calculated, wherein the calculation of the Jaccard similarity coefficient is a well-known technique and will not be described in detail here.
[0104] Selecting the maximum value of the second similarity between each document in the search irrelevant class and all documents in the reference relevant class at any subordinate level, and multiplying the maximum value of the second similarity by the level difference as the third product;
[0105] Taking the normalized result of the mean of the third product of each document in the irrelevant class and all its subordinate levels as the false negative possibility of each document in the irrelevant class;
[0106] Preferably, in this embodiment, the calculation method of the false negative probability of each document in the irrelevant category is as follows: Among them, PT d is the false negative probability of the dth document in the irrelevant category of the retrieval, U d is the dth document in the irrelevant category of the retrieval, E h,g is the g-th instrument in the reference-related class of the h-th subordinate level, α h is the difference between the level corresponding to the current search person and the hth subordinate level, that is, the level difference; Jac() represents the calculation of the Jaccard similarity coefficient, max[] represents the maximum value function, H is the number of all subordinate levels under the authority level corresponding to the current search person, norm{} is the normalization function, in this embodiment, the sigmoid function is used for normalization, as other implementation methods, the implementer can use other methods of the prior art, such as the tanh function, etc., and this embodiment does not impose any special restrictions on this; secondly, max[Jac(U d ,E h,g )]×α h is the third product.
[0107] It should be noted that the more similar each document in the retrieval-irrelevant class is to the document in the reference-relevant class, that is, the higher the similarity, the greater the possibility that the corresponding document in the retrieval-irrelevant class is a false negative, and the greater the possibility of a false positive, indicating that the corresponding document in the retrieval-irrelevant class is more likely to be related to the target search term, reflecting that the corresponding documents in the retrieval-irrelevant class should be clustered in the retrieval-relevant class.
[0108] At this point, the false positive possibility of each document in the search-related class and the false negative possibility of each document in the search-irrelevant class are obtained.
[0109] Step 4, based on the false positive possibility and the false negative possibility, determine the false positive documents in the retrieval-related class and the false negative documents in the retrieval-irrelevant class; add and delete the false positive documents and the false negative documents to obtain adjusted retrieval-related classes, and based on the adjusted retrieval-related classes, display the retrieval results of the target search terms of the current searcher.
[0110] Further, based on the false positive possibility and the false negative possibility, the clustering result is adjusted, specifically:
[0111] Record the document whose false positive probability in the search-related category is less than the preset first threshold as a false positive document;
[0112] Recording the document whose false negative probability in the irrelevant category is greater than a preset second threshold as a false negative document;
[0113] Preferably, in this embodiment, the preset first threshold value is 0.3, and the preset second threshold value is 0.7. As other implementation modes, implementers can set them according to actual conditions.
[0114] Deleting all false positive documents from the search-related class, adding all false negative documents in the search-irrelevant class to the search-related class, and deleting all false negative documents from the search-irrelevant class, to obtain adjusted search-related class and search-irrelevant class;
[0115] All documents in the adjusted search-related categories are displayed on the search page as document search results of the target search term of the current searcher for the searcher to view.
[0116] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0117] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0118] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the present application. It should be pointed out that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application without departing from the content of the technical solution of the present application, shall fall within the protection scope of the technical solution of the present application.
Claims
1. A fast clustering retrieval method for large-scale business document data, characterized in that: The method comprises the following steps: Obtain the target search term of the current searcher; record other authority levels lower than the authority level of the current searcher as each subordinate level, and record the corresponding authority level of the current searcher and each subordinate level as each gradient level; obtain the historical search terms and document browsing time corresponding to each historical search record at each gradient level; respectively, form a search library for each level with all the documents that can be retrieved by the searchers at each gradient level; According to the frequency of occurrence of the target search term in each document in the search library at any level, the search matching degree of each document in the search library at any level is determined; according to the similarity between the historical search term corresponding to each historical search record at each gradient level and the target search term, and the browsing time of each document in the search library at any level, the browsing confidence of each document in the search library at any level is determined; combined with the search matching degree, the search relevance of each document in the search library at any level is determined; Clustering the search relevance of all documents in the search library at any level respectively, obtaining the search-related classes and search-irrelevant classes of the corresponding authority level of the current searcher, and the reference-related classes of any subordinate level; Analyze the correlation between the documents in the search-related class, the search-irrelevant class and the reference-related class, and the difference between the corresponding authority level of the current searcher and any subordinate level, to determine the false positive possibility of each document in the search-related class and the false negative possibility of each document in the search-irrelevant class; Based on the false positive possibility and the false negative possibility, the false positive documents within the retrieval related class and the false negative documents within the retrieval irrelevant class are determined; the false positive documents and the false negative documents are added and deleted to obtain adjusted retrieval related classes; based on the adjusted retrieval related classes, the retrieval results of the target search terms of the current searcher are displayed.
2. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: Determining the search matching degree of each document in the search library at any level includes: Count the frequency of each keyword in the current searcher's target search term appearing in each document in the search library at any level; The average frequency of all keywords in the target search term of the current searcher in each document in the search library at any level is used as the search matching degree of each document in the search library at any level.
3. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: Determining the browsing confidence of each document in the search library at any level includes: Respectively obtain the word vector of the target search term and the word vector of the historical search term; Calculating the similarity between the word vector of the historical search term and the word vector of the target search term; Record all historical search records corresponding to the similarity degree greater than a preset relevant threshold at each gradient level as relevant search records; Calculate the product of the browsing time of the book in any relevant search record at each gradient level and the absolute value of the similarity of the historical search term corresponding to the any relevant search record, and record it as the first product; For all relevant search records at each gradient level, the normalized result of the mean of the first product in all relevant search records corresponding to each browsed document is used as the browsing confidence of each document browsed in the search library of any level, and the browsing confidence of the remaining documents in the search library of any level that have not been browsed in the relevant search records is assigned a preset value.
4. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: The search relevance of each document in the search library at any level is the product of the search matching degree and the browsing confidence.
5. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: The acquisition of search-related classes and search-irrelevant classes corresponding to the authority level of the current searcher includes: Calculate the mean of the search relevance of all documents in the cluster corresponding to any level of search library, and record it as intra-cluster relevance; The cluster with the largest correlation within the cluster in the level retrieval library corresponding to the authority level of the current search personnel is recorded as the retrieval-related class for the authority level of the current search personnel, and the remaining clusters in the level retrieval library corresponding to the authority level of the current search personnel are recorded as the retrieval-irrelevant classes for the authority level of the current search personnel.
6. A large-scale business document data rapid clustering retrieval method as claimed in claim 5, characterized in that: The method for obtaining the reference related class of any subordinate level is: The cluster with the largest intra-cluster correlation in the level search library corresponding to any subordinate level is recorded as the reference related class of any subordinate level.
7. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: Determining the false positive probability of each document in the search-related category includes: Calculate the difference between the authority level corresponding to the current searcher and any subordinate level, and record it as the level difference; Calculating the similarity between each document in the search-related category and each document in any subordinate reference-related category, and recording it as a first similarity; Selecting the maximum value of the first similarity between each document in the search-related class and all documents in any subordinate reference-related class, and recording the product of the maximum value of the first similarity and the level difference as the second product; The normalized result of the mean of the second product of each document in the search-related class and all its subordinate levels is used as the false positive possibility of each document in the search-related class.
8. A large-scale business document data rapid clustering retrieval method as claimed in claim 7, characterized in that: The method for determining the false negative probability of each document in the irrelevant category is as follows: Calculating the similarity between each document in the search irrelevant class and each document in any subordinate reference relevant class, recorded as the second similarity; Selecting the maximum value of the second similarity between each document in the search irrelevant class and all documents in the reference relevant class at any subordinate level, and multiplying the maximum value of the second similarity by the level difference as the third product; The normalized result of the mean of the third product of each document in the irrelevant class and all its subordinate levels is used as the false negative possibility of each document in the irrelevant class.
9. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: The determining of the false positive documents in the search-related category and the false negative documents in the search-irrelevant category includes: Record the document whose false positive probability in the search-related category is less than the preset first threshold as a false positive document; The document corresponding to the false negative possibility in the irrelevant category of the retrieval being greater than the preset second threshold value is recorded as a false negative document.
10. A large-scale business document data rapid clustering retrieval method as claimed in claim 1, characterized in that: The step of obtaining the adjusted retrieval-related class includes: All false positive documents are deleted from the search-related class, all false negative documents in the search-irrelevant class are added to the search-related class, and all false negative documents are deleted from the search-irrelevant class to obtain an adjusted search-related class.
Citation Information
Patent Citations
Traditional Chinese medicine ancient book and literature retrieval system
CN115858739A
Literary work retrieval method and system based on semantic recognition
CN118170895A