Distributed aggregation retrieval method for law and regulation text database
Through pre-processing, inverted index optimization and distributed query processing of the regulatory text database, the problems of inaccurate and inefficient search in the existing technology are solved, and efficient and fast regulatory text retrieval is achieved.
Patent Information
- Application Number
- CN202510317882.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the retrieval of legal text databases has problems such as inaccurate keyword extraction, high redundancy of inverted indexes, low query efficiency and slow response speed, especially when facing a large number of query requests, it is difficult to achieve rapid response.
By preprocessing the legal text data, removing useless punctuation and blanks, building inverted indexes and merging similar items, eliminating low-frequency vocabulary, adopting distributed parallel query processing mode, allocating query requests to multiple data nodes for parallel processing, and summarizing and sorting the results.
It improves retrieval accuracy and efficiency, reduces index redundancy interference, enhances the system response speed and comprehensiveness of query results, and ensures that users obtain more accurate and valuable regulatory text information.
Smart Images

Figure CN120256566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aggregated retrieval, and particularly to a distributed aggregated retrieval method for a regulatory text database. Background Art
[0002] Aggregated retrieval technology is an important branch in the field of information retrieval, which refers to uniformly extracting information from multiple heterogeneous or distributed data sources, integrating, deduplicating, and sorting the retrieval results from different sources, and finally presenting them to users in a consistent manner.
[0003] In the prior art, retrieval is only relied on the basic indexing mode, resulting in inaccurate keyword extraction and making it difficult to improve the query accuracy and recall rate. At the same time, low-frequency words and similar items are not effectively processed during index construction, leading to an increase in the redundancy of the inverted index, thereby reducing the retrieval efficiency and response speed. The query processing process is mostly in a single-point or simple serial mode, lacking a distributed parallel mechanism and being difficult to achieve a fast response in the face of a large number of query requests. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the drawbacks existing in the prior art, and a distributed aggregated retrieval method for a regulatory text database is proposed.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. The distributed aggregated retrieval method for a regulatory text database includes the following steps:
[0006] Collect and integrate regulatory text data, preprocess the regulatory text data to remove useless punctuation and whitespace characters, and generate a preprocessed text set; perform word segmentation and word frequency statistics on the preprocessed text set to generate a text keyword set;
[0007] Based on the text keyword set, construct an inverted index to generate a preliminary inverted index; optimize the preliminary inverted index, merge similar items and eliminate low-frequency words to generate an optimized inverted index;
[0008] Based on the optimized inverted index, implement distributed query processing, distribute the query requests to multiple data nodes for parallel processing to generate distributed query results; summarize and sort the distributed query results of each data node to generate an aggregated query result;
[0009] Based on the aggregated query result, conduct result evaluation, calculate the retrieval accuracy and response time, and generate a performance evaluation result.
[0010] Preferably, the step of obtaining the preprocessed text set is as follows: extract all regulatory text data, perform format standardization on the text data, unify the encoding format of the text and the alignment of text lines, and at the same time check and clean the blank lines in the text data to obtain a formatted text data set;
[0011] Based on the formatted text data set, scan and remove all illegal characters and useless punctuation marks, including duplicate commas, quotation marks, and special symbols, and at the same time compress consecutive spaces in the text to obtain a preprocessed text set.
[0012] Preferably, the step of obtaining the text keyword set is as follows: perform word segmentation on the preprocessed text set, use Jieba or HanLP for word segmentation, identify the lexical boundaries in the Chinese text, and divide according to the part of speech to obtain a word-segmented text set;
[0013] Based on the word-segmented text set, by calculating the number of times each word appears in the text, use a hash table or dictionary data structure to store each word and its frequency of occurrence to obtain a word frequency statistics result set;
[0014] According to the word frequency statistics result set, by setting a frequency threshold, screen out high-frequency keywords to obtain a text keyword set.
[0015] Preferably, the step of obtaining the preliminary inverted index is as follows: extract each keyword from the text keyword set, establish an index, and the index records all the occurrence positions of each keyword in the document set to obtain a mapping table;
[0016] Based on the mapping table, calculate the inverse document frequency of each keyword;
[0017] Based on the inverse document frequency, construct an inverted index to generate a preliminary inverted index.
[0018] Preferably, the step of obtaining the optimized inverted index is as follows: based on the preliminary inverted index, scan all index entries, identify keywords that appear in multiple documents and point to the same document set, calculate the edit distance, and if the edit distance is lower than the set threshold, consider them as similar items and merge them to obtain a merged inverted index;
[0019] According to the merged inverted index, calculate the low-frequency value of each keyword;
[0020] Based on the low-frequency value, screen out keywords, and remove the screened keywords from the index table to generate an optimized inverted index.
[0021] Preferably, the steps for obtaining the distributed query result are as follows: extract query keywords from the optimized inverted index, parse the query request, identify the inverted index entries corresponding to the query keywords, determine all data nodes associated with the query, split the query request into multiple sub-query tasks according to the storage distribution and computing resources of the data nodes, and assign them to the data nodes to obtain a query task allocation table;
[0022] Based on the query task allocation table, calculate the computing load balance value of the query tasks;
[0023] Based on the computing load balance value, allocate computing resources to balance the task loads of each data node, perform parallel queries simultaneously, and merge the sub-query results returned by each data node to obtain the distributed query result.
[0024] Preferably, the steps for obtaining the aggregated query result are as follows: based on the distributed query result, check the sources of each query result, align the results returned by different data nodes for the same keyword according to the matching situation of the query keywords on the data nodes, merge duplicate items, and remove redundant content to obtain a preliminary aggregated query result;
[0025] Based on the preliminary aggregated query result, calculate the sorting priority of the query result;
[0026] Based on the sorting priority, sort the query results in descending order to generate the aggregated query result.
[0027] Preferably, the steps for obtaining the performance evaluation result are as follows: based on the aggregated query result, count the number of matching documents, compare with the number of matching documents expected by the query request, evaluate the retrieval accuracy, record the timestamp of query processing, calculate the query response time, and obtain the retrieval accuracy and response time data set;
[0028] Based on the retrieval accuracy and response time data set, evaluate the processing efficiency of the query result, analyze the fluctuation of different query tasks in terms of response time, and screen out abnormal query tasks to obtain the performance evaluation result.
[0029] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0030] Through data integration and preprocessing of regulatory texts, the present invention effectively removes useless punctuation marks and redundant whitespace characters, ensuring the purity of text data and improving the accuracy of subsequent word segmentation and word frequency statistics. Implementing the word segmentation and word frequency statistics process helps identify key feature words in the regulatory texts, thereby enhancing the precision in the subsequent inverted index construction stage. By initially constructing an inverted index and implementing operations to merge similar items and eliminate low-frequency words, the index scale and redundant interference are reduced, improving the index query efficiency and accuracy. With the help of a distributed parallel query processing mode, query requests are dynamically allocated to multiple nodes for concurrent processing, significantly reducing query latency and enhancing the overall system response speed. Aggregating and sorting the query results improves the comprehensiveness and sorting rationality of the query results, ensuring that users obtain more accurate and valuable regulatory text information. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] Please refer to Figure 1 , the present invention provides a technical solution, a distributed aggregation retrieval method for a regulatory text database, including the following steps:
[0034] Collect and integrate regulatory text data, preprocess the regulatory text data, remove useless punctuation marks and whitespace characters, and generate a preprocessed text set; perform word segmentation and word frequency statistics on the preprocessed text set to generate a text keyword set;
[0035] Based on the text keyword set, construct an inverted index to generate a preliminary inverted index; optimize the preliminary inverted index, merge similar items and eliminate low-frequency words to generate an optimized inverted index;
[0036] Based on the optimized inverted index, implement distributed query processing, allocate query requests to multiple data nodes for parallel processing to generate distributed query results; aggregate and sort the distributed query results of each data node to generate an aggregated query result;
[0037] Based on the aggregated query result, conduct result evaluation, calculate retrieval precision and response time, and generate a performance evaluation result.
[0038] The steps for obtaining the preprocessed text set are as follows: Extract all regulatory text data, perform format standardization on the text data, unify the encoding format of the text and the alignment of text lines, and at the same time check and clean the blank lines in the text data to obtain a formatted text data set;
[0039] Based on the formatted text data set, scan and remove all illegal characters and useless punctuation marks, including repeated commas, quotation marks, and special symbols, and at the same time compress consecutive spaces in the text to obtain the preprocessed text set.
[0040] Specifically, first, when reading the regulatory text data, identify the character encoding format of each file and count the types of line break characters in different files. If it is detected that there are multiple types of line break characters such as CRLF or LF coexisting, then perform a unified replacement on all the content. In terms of the encoding format, if it is found that some files use GBK or ANSI encoding, then batch convert them to UTF-8 format. During the conversion process, check whether there are extra bytes by detecting the BOM information of the file header and remove them. When aligning the text lines, the maximum display width of each line can be preset to 80 characters. If the length of a certain line exceeds this width, then split the extra characters to the next line. This value is determined by repeatedly testing the reading experience and combining the document display ratio. After performing the line alignment operation, check the number of visible characters in each line. If it only contains spaces or tab characters and the actual number of visible characters is less than or equal to 2, then judge that this line is a blank line. At this time, using 2 as the threshold is based on the statistical results of various format files and it is found that lines with less than or equal to 2 visible characters have no substantial text in many scenarios. If it is determined to be a blank line, then delete the content of this line. During this process, it is also necessary to split the redundant tab character sequences and detect whether there are repeated spaces after splitting. If it is detected that the length of consecutive spaces exceeds 3, then reduce it to 1 space. The setting of this 3 is selected based on the common typesetting situations in the previous sample documents. At the same time, perform additional verification on possible unresolvable symbols such as obsolete encoding characters or special embedded characters. The verification method is to compare the Unicode value of each character with the allowed range. If it is found that the Unicode value is less than 32 or greater than 126 and is not within the range of common Chinese, Japanese, and Korean characters, then mark this character as an unresolvable character and remove it. Repeat the above process until all text files are corrected to the same line alignment state.
[0041] First, check each file in the text dataset obtained in the previous step line by line for any possible illegal characters. The range of legal characters considers the basic Latin characters with Unicode values between 32 and 126, as well as Chinese characters with Unicode values between 0x4E00 and 0x9FA5 as normal characters. If some characters with ASCII values between 0 and 31 or exceeding 126 and not matching the common Chinese character range are detected, they are marked as illegal characters. This judgment range is determined by referring to common Chinese typesetting requirements and the definition of control characters in the ASCII encoding table. Then, continue to scan the punctuation marks in the document and delimit the set of common punctuation. Commas, periods, question marks, exclamation marks, semicolons, colons, etc. are recognized using their corresponding Unicode values. If a certain type of punctuation is found to repeat more than 2 times in the same position, the repeated part is deleted. The threshold of 2 times is selected based on the statistical frequency of symbol usage in multiple sample documents. In addition to commas, consecutive quotation marks and special symbols also need to be concerned. For characters determined to be special symbols, they are compared through a comparison table, which is constructed based on a list of special symbols obtained from previous statistics. If a symbol not included in the list and not within the legal character range is found, it is still judged that the character needs to be removed. After removing illegal characters and useless punctuation marks, it is also necessary to check whether there are consecutive repeated spaces in the text. If consecutive spaces with a length exceeding 2 are detected, they are compressed to 1 space. The number 2 is also determined by referring to the common space usage in sample documents. If there are extra spaces at the end of some lines, they are also deleted. Finally, text information without illegal characters is obtained uniformly.
[0042] The steps to obtain the set of text keywords are as follows: Perform word segmentation on the preprocessed text set, using Jieba Segmentation or HanLP for word segmentation, identify the lexical boundaries in the Chinese text, and divide according to part of speech to obtain the word-segmented text set;
[0043] Based on the word-segmented text set, by calculating the number of times each word appears in the text, use a hash table or dictionary data structure to store each word and its frequency of occurrence, and obtain the word frequency statistics result set;
[0044] According to the word frequency statistics result set, by setting a frequency threshold, filter out high-frequency keywords to obtain the set of text keywords.
[0045] Specifically, for the pre-processed text data, a corresponding word segmentation model needs to be established first and the specific word segmentation process needs to be performed. To this end, a batch of annotated corpus samples are first prepared to train the Chinese word segmentation model. These corpora contain common political document terms and frequently occurring daily vocabulary. Before training, the sample distribution is corrected by counting the occurrence intervals of each word in the corpus. To improve the annotation consistency, all the annotated data can be cross-compared multiple times. For example, after separating the word segmentation boundary and the part-of-speech annotation, each sentence is evaluated to check if there is any punctuation position deviation and corrected according to the comparison table. This comparison table consists of the positioning index and label information for each dictionary entry, and its source is obtained through manual screening and cataloging by domain experts when organizing political documents. Then, the maximum likelihood estimation method is used to calculate the probability distribution of the segmentation positions. During the process, a step size δ can be set to gradually update the segmentation probability. If the convergence still cannot be achieved when the number of iterations reaches the set upper limit (for example, set to 30 rounds), it is necessary to analyze whether the annotated data samples need to be expanded. After the model training is completed, the word segmentation operation is performed on each sentence of the text. Each line of characters is scanned and the possible candidate word segmentation cut-off points are matched. At the same time, non-Chinese characters and segments with mixed numbers and Chinese characters are divided into independent words, and entity categories such as suspected personal names and place names are separately identified according to the part-of-speech definitions in the annotated data. During the calculation process, the manually revised low-frequency vocabulary list can also be combined to proofread whether there are incorrect segmentations. For rare words that are difficult to classify, a secondary comparison is performed additionally. If a word is not within the known dictionary range and its occurrence frequency is less than 3 times, it can be temporarily classified as an out-of-vocabulary word and included in the additional list. The choice of 3 times is a balance value obtained from actual statistical samples. If it exceeds this value, it is considered to have a sufficient occurrence probability to be determined as a valid word. Finally, after confirming all the word boundaries, the results are output sentence by sentence as the segmented text data.
[0046] After obtaining the segmented text data, select all the identified words and record the number of times they appear in the text. To make the statistical results more intuitive, first map all the words to a hash table structure and perform hash hashing according to the Unicode or ASCII encoding of the words. When specifically executing, first read the words and calculate the hash value, then complete the insertion operation in the hash table and handle the conflict entries with the same hash value through chaining storage or open addressing. Then, traverse the text line by line and update the occurrence count of each word. If the same word is encountered, increment the corresponding count by 1. If it is a word that appears for the first time, create a count entry for the word in the hash table and set its initial value to 1. After the statistics are completed, check whether there are duplicate entries or homonym conflicts caused by polysemous words. If polysemous situations are found, distinguish them according to the pre-collected part-of-speech or context semantic information. After the distinction is completed, reallocate the counts and update the hash table again. If you want to further confirm the accuracy of the records, you can perform sampling comparison for different regions of the text. For example, randomly select several paragraphs and compare them with the identified word segmentation records to check for omissions or abnormal counts. Finally, summarize the count results of all words in the hash table to obtain a frequency distribution data structure.
[0047] After obtaining the word frequency statistics results, it is necessary to screen for high-frequency words based on specific criteria. The specific approach is to first calculate the average word frequency m1 of the entire text. For example, the total word frequency Fsum can be obtained by summing up the number of times all words appear, and then m1 can be obtained by dividing Fsum by the total number of valid words N. Next, a frequency threshold T can be set and its initial value can be set to m1 to distinguish high-frequency and non-high-frequency words. To select the optimal value of T, a decreasing or increasing scan can be performed according to the idea of Case 1. For example, set a step size δ and compare whether there are words with a frequency exceeding the current T in each round of calculation. If the number of words exceeding the current T is within a reasonable range, the adjustment can be stopped and this T can be used as the final threshold. If the range is not satisfied, continue to perform addition or subtraction operations on T according to a certain update rule, and re-statistics the number of words exceeding the threshold after each round. Through this method, a threshold range that can cover the main words and avoid too many low-frequency words from entering the list can be determined. At this time, it can be exemplified that if Fsum is 1200 and N is 300, then m1 is 4. If T is adjusted step by step from m1 with δ = 0.5, and the distribution of the number of high-frequency words in each round is detected until the target distribution is satisfied, then fix T as the final threshold and collect all the words whose number of occurrences exceeds this threshold as the high-frequency word set.
[0048] The steps to obtain the preliminary inverted index are as follows: Extract each keyword from the text keyword set, establish an index, and the index records all the occurrence positions of each keyword in the document set to obtain a mapping table;
[0049] Based on the mapping table, calculate the inverse document frequency of each keyword. The formula is:
[0050]
[0051] Among them, IDF i represents the inverse document frequency of keyword i, DF i is the number of documents in which keyword i appears, N is the total number of documents, and TF ij is the frequency of keyword i in document j;
[0052] Based on the inverse document frequency, an inverted index is constructed to generate a preliminary inverted index.
[0053] Specifically, based on the set of text keywords obtained previously, these keywords are read one by one and the corresponding document content is retrieved. During the process, an index mapping structure is set up internally to store the correspondence between each keyword and the specific document location. First, the order of appearance of the keyword is recorded by identifying its serial number position in the document and the start and end points of the character segment are marked. If the same keyword appears multiple times in the same document, these positions are marked in segments and arranged in an incrementing order. Then, the retrieval results of all documents are aggregated and the occurrence positions of the same keyword are merged together. Subsequently, after the retrieval is completed, a list of document distributions is formed for each keyword. If the keyword does not appear at all in a certain document, a null value identifier is reserved in the list to prevent subsequent mapping conflicts. Finally, all the keyword distribution records are traversed and sorted in the order of document numbers, and the sorted keyword position information is associated and recorded in the index mapping structure. When all the index records of the keywords are completed, the mapping result between the keywords and the document positions can be obtained.
[0054] The advantage of the formula is that by introducing the square factor of the overall word frequency mean on the basis of the traditional inverse document frequency calculation, it can consider the comprehensive frequency of the keyword appearing in each document while measuring the rarity of the keyword, thus taking into account the distribution of the keyword in different documents.
[0055] The steps to obtain N are as follows: This parameter represents the total number of documents in the text database. The specific value can be determined by querying the database that records the document list and obtaining the number of documents. If it is found during the operation that there are some duplicate document names, they should be deduplicated first and then counted. To illustrate the obtaining process more intuitively, it can be exemplified that there are 550 records in the actually retrieved regulation documents, among which 50 documents are duplicate entries. After removing the duplicates, the final value of N = 500 is obtained.
[0056] DF iThe obtaining steps of DF are as follows: This parameter represents the number of documents in which the keyword i appears, and it can be obtained by summarizing the distribution of the keyword in different documents through the retrieval index mapping structure. During the processing, it is necessary to query the document occurrence records of each keyword in turn. If the keyword appears at least once in a certain document, then this document is counted into DF i The statistical range of DF is as follows. Finally, the total number of documents that meet the occurrence conditions is accumulated to obtain DF i Taking a keyword "organizational discipline" as an example, after the retrieval is completed, it is found that this keyword appears in 120 out of 500 documents. Then DF i can be determined as 120
[0057] TF ij The obtaining steps of TF are as follows: This parameter represents the occurrence frequency of the keyword i in the document j, and it can be obtained by counting the occurrences of a single keyword in the specified document when establishing the index mapping. It is necessary to traverse the document j and count the matching times of the keyword i paragraph by paragraph, and then record this cumulative value as TF ij For example, if the query for "organizational discipline" is performed on a document with document number 8, and it is confirmed that the total number of occurrences after the retrieval is 15 times, then TF i8 can be set as 15
[0058] Calculation process:
[0059] The first step: First calculate where N = 500 and DF i = 120. Therefore
[0060] The second step: Take the natural logarithm of this result, ln(4.1667), and the value is approximately 1.4289;
[0061] The third step: Calculate where Dividing 700 by 500 gives 1.4, and then squaring 1.4 gives 1.96;
[0062] The fourth step: Multiply the results obtained in the first two steps, that is, 1.4289×1.96≈2.80. This value is the IDF i
[0063] This result shows that when the number of documents in which the keyword "organizational discipline" appears in 500 documents is 120 and the total number of overall occurrences is 700, the improved inverse document frequency operation can obtain IDF i = 2.80. If this value is relatively high, it indicates that the keyword distribution is more rare and discriminatory. If the value is relatively low, it indicates that the keyword is more common
[0064] Based on the inverse document frequency, first count the inverse document frequency data of the obtained keywords and compare them item by item. Subsequently, conduct a cross-analysis of the inverse document frequency of each keyword with the corresponding document occurrence records. If a keyword has retrieval results in most documents, mark the inverse document frequency value of the keyword and further identify its specific usage scenarios in different document paragraphs in the subsequent steps. Then, sort the inverse document frequencies of all keywords from largest to smallest to determine their degree of difference in the document set. During this period, also record the specific correspondence between keywords and documents and the coverage range of keywords at the paragraph level. Each time an addition or deletion operation is performed on the association between keywords and documents, it is necessary to re-induce according to the previously obtained inverse document frequency data to ensure that the current index order can reflect the actual differences in the document distribution of different keywords. If it is found that the inverse document frequencies of some keywords are relatively high, they are preferentially associated and marked with the document title or the first paragraph of the document. If the inverse document frequency is relatively low, it may be concentrated in paragraphs with general descriptions. In a similar way, continuously retrieve and align the specific line numbers and offset positions where keywords appear in all documents, generate an inverted index result covering all keywords, and use this inverted index as the basis for text retrieval after finishing the collation, and finally obtain a preliminary inverted index.
[0065] The steps to obtain the optimized inverted index are as follows: Based on the preliminary inverted index, scan all index entries, identify keywords that appear in multiple documents and point to the same document set, calculate the edit distance. If the edit distance is lower than the set threshold, consider them as similar items and merge them to obtain the merged inverted index;
[0066] According to the merged inverted index, calculate the low-frequency value of each keyword. The calculation formula is:
[0067]
[0068] where, L i is the low-frequency value of keyword i, M is the number of documents including keyword i, TF ij is the frequency of keyword i in document j, and P j is the number of paragraphs in document j;
[0069] Filter keywords based on the low-frequency value, remove the filtered keywords from the index table, and generate the optimized inverted index.
[0070] Specifically, based on the preliminary inverted index obtained previously, first read the keywords and their corresponding document sets and check each index item one by one. During the process, it is necessary to compare the string content of each keyword. In order to determine whether two keywords constitute similar items during the recognition process, it is necessary to first set the edit distance threshold and clarify the specific source of the threshold. For example, after analyzing 200 sample texts, the average length of common keywords is counted and combined with the number of characters of common spelling differences to determine that the threshold is 2. When the edit distance is less than 2, the two keywords can be regarded as the same type of keyword. To complete this step, the characters at the same position of the two keywords can be compared character by character and the total difference number can be calculated. If the difference number is less than 2, this pair of combinations will be merged into the same keyword entry. Subsequently, the remaining keyword groups are compared in turn and the above operations are repeated until the traversal is completed. After the similarity judgment of all keywords is completed, the document sets corresponding to each merged keyword are merged and the number of occurrences of the keyword in the merged set is counted. If a keyword only appears in a small number of documents, there is no need to merge and the original record is maintained. After the merger is completed, all the merged keyword entries are confirmed again. If there is still an edit distance lower than 2 between the merged string contents, another iterative merger is performed until there are no more items to be merged. Finally, all similar keywords are merged and the original inverted index is updated to obtain the merged inverted index.
[0071] The benefit of the formula is that by coupling the actual occurrence frequency of the keyword in each document with the number of document paragraphs and using the logarithmic function and the cube root operation method, it can highlight the particularity of high-frequency keywords appearing in a small number of documents when dealing with a large number of documents, and also balance the keywords that appear in multiple documents but with a low frequency, so as to more accurately identify low-frequency but still significant words in the screening.
[0072] The steps to obtain M are as follows: This parameter represents the number of documents including keyword i. Specifically, it can be obtained by retrieving keyword i in the merged inverted index and extracting all the document numbers where the keyword appears. Subsequently, these documents are counted one by one and duplicates are removed to obtain the final effective number of documents. To ensure statistical accuracy, it is necessary to ensure that the integration of similar keywords has been completed when merging keyword items in the previous operations, otherwise there will be duplicate counting. After determining M, it can be aligned with data such as TF ij and ensure that the document set participating in the operation is consistent with the keyword. For example, when the keyword "branch activity" corresponds to 32 documents in the merged inverted index, then M = 32.
[0073] TF ij The steps to obtain TF are as follows: This parameter represents the occurrence frequency of keyword i in document j. After the keywords are merged previously, it is necessary to read the number of occurrences according to the corresponding relationship between the keyword and the document in the inverted index and record this number of occurrences as TFij For subsequent statistics, when retrieving the keyword "branch activities", if a document with the number 10 is retrieved and it appears 12 times, then TF i10 = 12.
[0074] P j The steps to obtain P are as follows: This parameter represents the number of paragraphs in document j. By scanning the document text, the starting position of each paragraph can be used as a segmentation point, and paragraphs containing substantial text content are counted. During the process, if a blank paragraph is found, it is not included in the number of paragraphs. If a blank line between a separator and visible characters is detected, it is also treated as having no substantial content. For example, after actual collection, if document 10 is confirmed to have 45 effective paragraphs, then record P 10 = 45.
[0075] Calculation process:
[0076] The first step: Calculate the numerator
[0077] For example, the keyword "branch activities" appears in M = 3 documents, with document numbers 10, 12, and 15 respectively. Record TF i10 = 12, TF i12 = 5, TF i15 = 9 and the number of paragraphs are P 10 = 45, P 12 = 30, P 15 = 20 respectively, then
[0078]
[0079] where ln(1 + 45) = ln(46) ≈ 3.8286, ln(31) ≈ 3.4339, ln(21) ≈ 3.0445. Therefore, we get:
[0080] 12×3.8286 + 5×3.4339 + 9×3.0445 ≈ 45.9432 + 17.1695 + 27.4005 = 90.5132;
[0081] The second step: Divide this numerator by M = 3 to get:
[0082]
[0083] The third step: Take the cube root of the above result: From this, we can obtain L i = 3.1167;
[0084] The results show that when the frequency of occurrence of the keyword "branch activities" in 3 documents and the distribution of the number of document paragraphs are as above, the low-frequency value obtained by this formula is approximately 3.1167. If this value is relatively high, it often indicates that the keyword appears in a small number of documents and the number of paragraphs is relatively large. If it is found that this value exceeds a certain set threshold during subsequent retrieval of the entire database range, it may be necessary to further analyze whether this keyword needs to be retained or excluded in subsequent screening.
[0085] After obtaining the low-frequency values corresponding to each keyword, it is necessary to screen these values and remove relatively special keywords from the index table. First, it is necessary to sort according to the previously obtained low-frequency value distribution range and select a screening reference value to determine which keywords belong to the low-frequency range. The setting of the screening reference value often refers to the comprehensive consideration of the examination of the logarithmic value distribution curve and domain experience. For example, after statistics, it is found that the low-frequency values of most keywords are between 1.0 and 2.5, but when some keywords fall above 4.0, 2.5 is set as the preliminary threshold. When the low-frequency value is greater than 2.5, it is considered that the keyword appears frequently in local documents and the number of paragraphs cooperates to produce a relatively high calculation result, so it is listed as the key point to be screened. Subsequently, the keywords that meet this condition are marked and compared item by item again. For example, check its appearance in more documents and confirm whether it fully meets the screening requirements. If confirmed, it is removed from the index table. Finally, the index structure is updated and all the remaining keywords are renumbered or sorted again. At this time, the keyword screening process can be completed and an optimized inverted index can be generated.
[0086] The steps to obtain the distributed query results are as follows: Extract the query keywords from the optimized inverted index, parse the query request, identify the inverted index items corresponding to the query keywords, determine all the data nodes associated with the query, and according to the storage distribution and computing resources of the data nodes, split the query request into multiple sub-query tasks and assign them to the data nodes to obtain the query task allocation table;
[0087] Based on the query task allocation table, calculate the computing load balance value of the query task. The calculation formula is:
[0088]
[0089] where, F k is the computing load balance value of query task k, X is the number of data nodes associated with the query task, Q km is the computing request volume of query task k on data node m, S m is the number of pending query tasks on data node m currently, C k is the number of keywords that query task k needs to process;
[0090] Based on the calculated load balance value, computing resources are allocated to balance the task loads of each data node. Meanwhile, parallel queries are executed, and the sub-query results returned by each data node are merged to obtain the distributed query result.
[0091] Specifically, the optimized inverted index obtained previously is used to read the query keywords and record them item by item. Each query keyword needs to be disassembled to identify the corresponding inverted index items. First, all the documents and paragraph positions associated with the keyword are located in the inverted index, and this information is collected as preliminary identification information. Then, in combination with other constraint contents included in the query request, it is analyzed whether there is an additional keyword intersection requirement. For example, when there are multiple keywords in the retrieval instruction, all document entries that meet the matching conditions of all keywords need to be compared and screened one by one. To ensure that the resource distribution of data nodes can be fully utilized during the query process, the storage locations of data nodes need to be summarized first, and information such as the document sets already stored on each node, the current CPU occupancy rate, and memory occupancy rate of the node is recorded. Then, the document or paragraph content corresponding to the keyword is matched with the node location to determine which nodes the subsequent query calculations should be dispatched to. If the document entries corresponding to a keyword exist in only a very small number of nodes, this part of the calculation can be concentrated on these nodes to save the time of cross-node data transmission. If the document set pointed to by a keyword is distributed across multiple nodes, the computing resources of the nodes need to be additionally recorded to more reasonably allocate the load. During this process, the number of queries already queued on each node should also be counted to judge the load status. For example, a range of 0 to 50 for the number of queries to be processed is preset as a measurement standard. If the current number of query tasks to be processed on a node is above 50, it is considered that the load is relatively high, and it is necessary to consider dispatching the newly arrived query requests to other nodes with lighter loads. Then, all sub-query tasks in the query request are numbered, and the sub-query tasks are divided one by one according to the node allocation rules. Subsequently, a query task allocation table needs to be generated to summarize the relationship between these tasks and nodes. For example, the first sub-query of the query request is assigned to node A and node B, and the task amounts and the number of keywords to be processed by node A and node B at this time are recorded. When the allocation of all sub-queries is completed, a table containing information such as the sub-query task number, node number, number of keywords, and node load will be obtained. At this time, a final integrity check needs to be performed on it. For example, check whether the same sub-query task is repeatedly allocated or whether the assignment of a certain keyword is omitted in the allocation table. After confirming no conflicts, the query task allocation table can be submitted for subsequent processing.
[0092] The benefit of the formula is that it balances both the request volume of data nodes and the number of queued tasks on the nodes, and combines the scale of keywords that the query task itself needs to process, making the load assessment more targeted.
[0093] The steps to obtain X are as follows: This parameter represents the number of data nodes associated with the query task, which can be obtained by finding all assigned nodes for task k from the query task allocation table. To ensure statistical correctness, after the query request is split, it is necessary to remove duplicates from the node numbers one by one and exclude nodes without document data or nodes not participating in this task, thereby obtaining a unique node set. Finally, the number of elements in this node set is recorded as X. For example, if query task k1 involves three nodes with IDs 2, 5, and 9, then X = 3.
[0094] Q km The steps to obtain it are as follows: This parameter represents the calculation request volume of query task k on data node m, and it is necessary to count the specific operation or retrieval requirements according to the situation of query task and node allocation. In particular, it is necessary to quantify the actual processing processes such as document scanning, keyword comparison, and data result combination of this query on node m, and it is calculated by monitoring the number of node memory accesses and CPU instruction consumption time. For example, the total number of scanned lines and execution times of node m when executing this query can be accumulated as H 1m , and at the same time, the number of keyword matching times is accumulated as H 2m , and these summary values are converted into Q using a calculation formula similar to the following km : Q km =α1×H 1m +α2×H 2m , where α1 and α2 are set based on the resource consumption ratios of storage row scanning and keyword matching processing in the actual measurement environment respectively. Finally, Q km is recorded in the query task allocation table. For example, in a real measurement, when node 5 scans 30,000 lines and the number of matching operations is 500 times, combined with α1 = 0.001 and α2 = 0.02, Q k5 =0.001×30000 + 0.02×500 = 30 + 10 = 40.
[0095] S m The steps to obtain it are as follows: This parameter represents the number of query tasks to be processed by data node m currently, and it is necessary to count the total number of query requests that have been queued but not completed on node m during the task scheduling phase, which can usually be obtained through the queue management information at the node end.
[0096] C k The steps to obtain it are as follows: This parameter represents the number of keywords that query task k needs to process. First, all valid keywords are parsed from the query request, and then combined with the part that task k needs to process for summarization. Since a query request may contain multiple groups of keyword sets, it is necessary to clarify the number of keywords corresponding to each sub-query when splitting it into sub-queries, and then accumulate them to obtain C k , for example, if query k1 contains a keyword set with a size of 8, then C k =8.
[0097] Calculation process:
[0098] Step 1: First calculate the numerator
[0099] For example, query task k1 is associated with 3 data nodes m = 2, 5, 9 and the statistics are obtained for Q k1,2 = 50, Q k1,5 = 40, Q k1,9 = 35, and at the same time S2 = 10, S5 = 6, S9 = 15, then successively we can get Thus, the numerator is calculated as follows:
[0100] 50×3.3166 + 40×2.6458 + 35×4 = 165.83 + 105.83 + 140 = 411.66;
[0101] Step 2: Calculate the denominator
[0102] If X = 3 and C k = 8, then X + C k = 11,
[0103] Step 3: Divide the numerator by the denominator to get
[0104] This result shows that when query task k1 is assigned to 3 nodes and the request volume and queued tasks of each node are as above, the finally calculated computing load balance value is approximately 124.07. If this value is significantly higher than a certain benchmark value when comparing with other tasks in the follow-up, it can be judged that the load of this task is relatively large and more hardware resources need to be arranged further or priority adjustment is required. If it is lower than a certain range, the resources can be allocated to other tasks with higher load values first.
[0105] Based on the query task computing load balance value obtained previously, it is necessary to summarize the load sizes of each task on different data nodes and allocate computing resources with reference to the current available resource metrics of all nodes. To more comprehensively match the query requirements, it is necessary to first uniformly record the CPU idle ratio and memory free amount of the nodes as important data sources, and combine the execution times of the previously queued requests for each node to judge the bearable computing scale. For example, nodes with a CPU idle ratio between 20% and 50% can be defined as a moderate load range in the system. If the idle ratio of a certain node is lower than 20%, it is marked as having a high load, and new tasks need to be carefully redistributed. For example, when it is detected in a specific data center that the idle ratios of node 2 and node 5 are around 30% and the idle ratio of node 9 is about 10%, more subqueries can be preferentially scheduled to node 2 and node 5 to relieve the computing pressure on node 9. After comparing the load values of each task with the node resource situation one by one, an overall allocation plan covering all data nodes is finally generated, and then instructions are distributed to the corresponding nodes to start the parallel query process. During the process of the nodes executing the query, it is necessary to monitor the task processing progress of the nodes in real time and update the queuing information regularly. When the nodes complete the subqueries, they return data such as the matched documents, paragraphs, and keyword fields. Finally, the return information of all nodes is collected for merging and deduplication processing. The duplicate results from different nodes are uniformly summarized and then grouped and sorted in the combination mode of keywords. If it is detected that the same document is matched by multiple nodes at the same time, its record is counted for one merge. After all subquery results are processed, the integrated distributed query results will be obtained.
[0106] The steps for obtaining the aggregated query results are as follows: Based on the distributed query results, check the sources of each query result, align the results returned by different data nodes for the same keyword according to the matching situation of the query keywords on the data nodes, merge the duplicates, and remove the redundant content to obtain the preliminary aggregated query results;
[0107] Based on the preliminary aggregated query results, calculate the sorting priority of the query results. The calculation formula is:
[0108]
[0109] Among them, P r is the sorting priority of query result i, D i is the document matching degree corresponding to query result i, EM is the median matching degree of all query results, U i is the number of occurrences of query result i on different data nodes, and T i is the update time of the document corresponding to query result i;
[0110] Based on the sorting priority, sort the query results in descending order to generate the aggregated query results.
[0111] Specifically, based on the previously obtained distributed query results, it is necessary to scan and record the retrieval information returned by each data node one by one. First, collect the document numbers and paragraph positions corresponding to the keywords from the source of each retrieved data, and compare these positions sequentially. When it is detected that the same keyword appears with the same document number and the paragraph identifier also remains consistent in different data nodes, this duplicate information is merged, and the positions of the repeatedly appearing keywords are uniformly marked. In this way, the content that can be merged can be found during the comparison process. Subsequently, all duplicate items are removed at once. When removing, a comparison criterion can be set to define whether it is a duplicate item. For example, first read the keyword string, document number, and paragraph number and uniformly convert them into a fixed-length character sequence, and then perform character-by-character comparison. If the difference degree is 0, it is regarded as a duplicate and sorted out. There may be punctuation differences or case differences in the text content, and the text format needs to be unified in advance. For example, unify the English case to lowercase or uniformly convert full-width and half-width characters. These operations can determine the threshold according to experience or actual statistical results. If it is confirmed that the case form only appears in 0.5% of the entries in the actual data, the entries with inconsistent cases can be regarded as basically the same items for replacement. When all duplicate merges are completed, an integrated list of keyword-document pairs will be obtained. To further remove redundant content, it is also necessary to verify the appearance time or text fragment of the keywords again. For example, include the character intervals before and after the keyword in the record in the comparison. When the comparison results completely overlap, it indicates that the content is redundant and can be deleted. Before deletion, the paragraph range also needs to be checked to avoid misdeleting entries that are not in the same paragraph but are close in position. During the alignment stage, the document number lists of each node can be spliced in the order of node numbers, and the overlapping content is checked sequentially during splicing and the aforementioned merging operation is performed. Finally, the content of all nodes is aggregated to obtain a preliminary summary query result with duplicate entries removed.
[0112] The benefit of the formula is to combine the relative deviation of the query result corresponding document from the overall matching degree distribution, and incorporate the occurrence times and update time into a unified measure, so that the sorting can more comprehensively reflect the importance and recency of the results.
[0113] D i The acquisition steps of m are as follows: This parameter represents the document matching degree corresponding to query result i. It is necessary to first quantify the coincidence degree between the keyword and the document text during the retrieval process. The common approach is to comprehensively calculate based on multiple data such as the density and coverage of the keyword distribution in the document. Specifically, the total number of occurrences C of the keyword in the document can be counted m , and the total number of words W in the document is recorded i : Among them, β1 and β2 can be set by monitoring the influence degree of the correlation when keywords appear intensively, so as to balance the overall length of the document and the keyword clustering. For example, in actual tests, after referring to the keyword distribution of a batch of regulatory texts, β1 = 0.5 and β2 = 0.2, and the average matching degree under this configuration is the most balanced by summarizing about 300 documents. For a simple example, if the length W of the document m = 2000 and the number of keyword occurrences C m = 20, then After substituting the coefficients β1 and β2, D i = 0.5×0.01 + 0.2×0.2 = 0.005 + 0.04 = 0.045.
[0114] The acquisition step of EM is as follows: This parameter represents the median of the matching degrees of all query results, and it needs to be sorted after the matching degrees D i of all query results i are statistically counted to obtain the median. If the current result set size is N r , then all D i can be sorted in ascending order. If N r is odd, EM takes the value of the middle record in the sequence; if N r is even, EM takes the average of the two middle records. This is the standard median calculation process. For example, if there are 5 query result matching degrees of 0.03, 0.045, 0.06, 0.07, and 0.09 in sequence, then the third one is 0.06, and the median EM = 0.06.
[0115] U i The acquisition step of U is as follows: This parameter represents the number of occurrences of query result i on different data nodes. It is necessary to count the occurrence frequencies of the same query result on each node in the result summary table during the retrieval stage and assign the sum to U i . In actual operation, the document numbers, paragraph information, and keyword identifiers of the query results can be compared first. If they are all the same, it is regarded as the same query result. Subsequently, count how many nodes it is retrieved from, and this value is U i . For example, when the query result "The third paragraph of a certain document number 45 contains keyword A" appears in both node 2 and node 5, then U i = 2. If it is also reported by node 7, then U i = 3.
[0116] T iThe acquisition steps are as follows: This parameter represents the update time of the document corresponding to the query result i. It is necessary to read the latest update timestamp of the document in the database and convert it into a quantifiable value, such as subtracting the update time from a base date to get the number of days or hours, so as to bring it into the formula ln(1+T i ), for regulatory texts, the publication date or revision date in the document metadata can be obtained. This value can be set to the time difference corresponding to the record during system maintenance. For example, if the last revision time of the document is 30 days away from the current time, T can be set. i =30.
[0117] Calculation process:
[0118] Step 1: Calculate |D i -EM| and divide by
[0119] For example, when D i =0.045, EM = 0.06, we can get |0.045-0.06| = 0.015. If the query result appears on 3 nodes, then U i =3, Therefore, this part of the operation is
[0120] Step 2: Calculate ln(1+T i ).
[0121] If the latest revision time of the document is 30 days, then T i =30,ln(31)≈3.4339.
[0122] Step 3: Add up to get the sorting priority P r .
[0123] P r =0.0075+3.4339=3.4414;
[0124] The results show that for documents with a matching degree slightly lower than the median and retrieved from three nodes at the same time and updated 30 days ago, the sorting priority is 3.4414. r By sorting and arranging from large to small, you can see the position of this result in the entire list. If the value is high, it means that it is significantly different from the median or the document update time feature is more prominent. It may also be higher or lower under the same conditions. It all depends on the specific value range of the corresponding parameters and the calculation results in the formula.
[0125] Based on the sorting priority, all preliminary summary query results are read at once and sorted according to their respective P rTo rank the numerical values from high to low, first, a record area needs to be added to each result in the internal data structure to store the calculated P r value, and ensure that this value is consistent with the document number and paragraph information of the corresponding document of the query result. Then, sort all the records. To ensure the stability of the sorting process, the main sorting can be done according to the magnitude of the P r value. If there are multiple results with the same P r value, the document number or its order in the source data node can be used as the order for comparison. When sorting, it is also necessary to pay attention to whether the segment number or timestamp difference of the document needs to be additionally checked and further fine-tuned when the difference in P r values is within 0.05, otherwise, a large number of similar results may appear side by side. After sorting, a descending-order query result sequence will be obtained. Each record in this sequence carries information such as the matching degree, the number of occurrence nodes, and the update time. Subsequently, these information can be read and recorded in sequence to form the aggregated query result. If an incremental update is performed on new query results in subsequent operations, it is necessary to recalculate their P r values and insert them into the existing sorted list for merging, so as to obtain the overall aggregated query result including the latest retrieval results.
[0126] The steps to obtain the performance evaluation results are as follows: Based on the aggregated query result, count the number of matching documents, compare it with the number of matching documents expected by the query request, evaluate the retrieval accuracy, and record the timestamp of the query processing, calculate the query response time, and obtain the retrieval accuracy and response time data sets;
[0127] Based on the retrieval accuracy and response time data sets, evaluate the processing efficiency of the query results, analyze the fluctuation of the response time for different query tasks, and screen out abnormal query tasks to obtain the performance evaluation results.
[0128] Specifically, based on the aggregated query results, it is necessary to first read the corresponding document numbers from each query result and count the number of times each document appears item by item. Then, compare them one by one with the expected number of matching documents clearly recorded in the query request. For example, if the "number of documents to be matched" obtained through the query instruction is 30, and it is found that a total of 28 documents are retrieved in the actual statistics, the difference of 2 can be found. To further evaluate the retrieval accuracy, a known correct matching list can be introduced to compare which documents match the actual requirements. After the comparison, the number of accurately matched documents, mis-matched documents, and missed-matched documents can be counted, and quantitative indicators such as precision and recall can be calculated. If the thresholds of these indicators need to be set, the size of the document data involved in the statistical process can be referred to. For example, when the total number of documents is in the range of 100 to 500, the precision can be set at 80% as a baseline and adjusted flexibly according to the characteristics of the demand task. When the statistics are completed, the retrieval time needs to be recorded to generate the subsequent query processing timestamp. Here, two time nodes can be generated by the system clock at the start and end of the query respectively. For example, the query start time is s(t): t = 2025-03-08 10:15:20, and the query end time is s(t): t = 2025-03-08 10:15:25. The difference between the two gives the actual 5 seconds consumed for this query. Then, the precision and time data are associated and recorded. By comparing with the data of different queries, the distribution of retrieval precision under multiple tasks can be observed, and the precision and response time of this query can be calculated accordingly. For example, in a certain query, the precision is 85% and the response time is 5 seconds, while in another query, the precision is 82% but the time consumption reaches 8 seconds, which can be further analyzed. Finally, the precision and response time obtained from the above multiple queries are sorted out and summarized into a retrieval precision and response time dataset.
[0129] Based on the retrieval accuracy and response time dataset, a dedicated comparative analysis needs to be carried out to evaluate the processing efficiency of each query result. First, read the accuracy and response time corresponding to each query request in the above dataset and conduct grouped statistics on them. It can be divided according to multiple factors such as request type or the number of keywords. Compare whether the difference in response time for requests with the same number of keywords or similar document scales is significant. If the difference exceeds a certain empirical threshold, it can be marked to indicate an anomaly. For example, the error range can be set within 2 seconds as the normal interval. If the response time of some requests deviates from this interval, it is recorded as possibly abnormal. Subsequently, extract all the requests marked as abnormal and conduct targeted inspections. For example, combine the recorded CPU occupancy rate and network load to trace whether there are resource bottlenecks during the query process. The accuracy of this batch of abnormal queries can also be further investigated. For example, check whether there is an unexpectedly high error rate or missing document statistics during the keyword parsing stage. In this way, individual cases with large fluctuations in response time can be identified from a large number of query tasks. Finally, the confirmed abnormal entries are included in the screening results and summarized. Associate the information such as the query task numbers, the number of keywords, and the response time of these marked abnormal queries, so as to optimize or reallocate computing resources according to the specific situation during subsequent maintenance. After this screening, the performance evaluation results of this time can be formed.
[0130] The above is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A distributed aggregation retrieval method for a regulatory text database, characterized in that Including the following steps: Collect and integrate regulatory text data, preprocess the regulatory text data, remove useless punctuation and whitespace characters, and generate a preprocessed text set; Perform word segmentation and word frequency statistics on the preprocessed text set to generate a text keyword set; Based on the text keyword set, construct an inverted index to generate a preliminary inverted index; Optimize the preliminary inverted index, merge similar terms and remove low-frequency words to generate an optimized inverted index; Based on the optimized inverted index, implement distributed query processing, distribute the query request to multiple data nodes for parallel processing to generate a distributed query result; summarize and sort the distributed query results of each data node to generate an aggregated query result; Based on the aggregated query result, perform result evaluation, calculate retrieval accuracy and response time to generate a performance evaluation result.
2. The distributed aggregation retrieval method for the regulatory text database according to claim 1, wherein The steps for obtaining the preprocessed text set are as follows: extract all regulatory text data, perform format unification on the text data, unify the encoding format of the text and the alignment of text lines, and at the same time check and clean the blank lines in the text data to obtain a formatted text data set; Based on the formatted text data set, scan and remove all illegal characters and useless punctuation marks, including duplicate commas, quotation marks, and special symbols, and at the same time compress consecutive spaces in the text to obtain a preprocessed text set.
3. The distributed aggregation retrieval method for the regulatory text database according to claim 1, characterized in that The steps for obtaining the text keyword set are as follows: perform word segmentation on the preprocessed text set, use Jieba or HanLP for word segmentation, identify the lexical boundaries in the Chinese text, and divide according to part of speech to obtain a word-segmented text set; Based on the word-segmented text set, calculate the number of times each word appears in the text, and use a hash table or dictionary data structure to store each word and its occurrence frequency to obtain a word frequency statistics result set; According to the word frequency statistics result set, filter out high-frequency keywords by setting a frequency threshold to obtain a text keyword set.
4. The distributed aggregation retrieval method of the regulatory text database according to claim 1, wherein, The steps for obtaining the preliminary inverted index are as follows: extract each keyword from the text keyword set, establish an index, and the index records all the occurrence positions of each keyword in the document set to obtain a mapping table; Based on the mapping table, calculate the inverse document frequency of each keyword; Based on the inverse document frequency, construct an inverted index to generate a preliminary inverted index.
5. The distributed aggregation retrieval method for the regulatory text database according to claim 1, characterized in that, The steps for obtaining the optimized inverted index are as follows: based on the preliminary inverted index, scan all index entries, identify keywords that appear in multiple documents and point to the same document set, calculate the edit distance, and if the edit distance is lower than the set threshold, consider them as similar terms and merge them to obtain a merged inverted index; According to the merged inverted index, calculate the low-frequency value of each keyword; Based on the low-frequency value, filter keywords, remove the filtered keywords from the index table, and generate an optimized inverted index.
6. The distributed aggregation retrieval method of the regulatory text database according to claim 1, characterized in that The steps for obtaining the distributed query result are as follows: Extract the query keywords from the optimized inverted index, parse the query request, identify the inverted index entries corresponding to the query keywords, determine all the data nodes associated with the query, split the query request into multiple sub-query tasks according to the storage distribution and computing resources of the data nodes, and assign them to the data nodes to obtain a query task allocation table; Based on the query task allocation table, calculate the computing load balance value of the query tasks; Based on the computing load balance value, allocate computing resources to balance the task loads of each data node, perform parallel queries simultaneously, and merge the sub-query results returned by each data node to obtain the distributed query result.
7. The distributed aggregation retrieval method for the regulatory text database according to claim 1, characterized in that The steps for obtaining the aggregated query result are as follows: Based on the distributed query result, check the sources of each query result, align the results returned by different data nodes for the same keyword according to the matching situation of the query keywords on the data nodes, merge duplicate items, and remove redundant content to obtain a preliminary aggregated query result; Based on the preliminary aggregated query result, calculate the sorting priority of the query results; Based on the sorting priority, sort the query results in descending order to generate the aggregated query result.
8. The distributed aggregation retrieval method for the regulatory text database according to claim 1, wherein The steps for obtaining the performance evaluation result are as follows: Based on the aggregated query result, count the number of matching documents, compare it with the number of matching documents expected by the query request, evaluate the retrieval accuracy, record the timestamp of query processing, calculate the query response time, and obtain the retrieval accuracy and response time data set; Based on the retrieval accuracy and response time data set, evaluate the processing efficiency of the query results, analyze the fluctuation of the response time for different query tasks, and screen out abnormal query tasks to obtain the performance evaluation result.
Citation Information
Cited By
Method for aligning terms in foreign trade English file in real time
CN121093937A
Method for obtaining result text, electronic equipment and storage medium
CN121388149A
A method for obtaining a result text, an electronic device, and a storage medium
CN121388149B