An intelligent retrieval method and system for government affairs data based on the Internet of Things

By using inverted indexing and semantic expansion technologies in the government data retrieval system, intelligent search of government data has been solved, and the accuracy and comprehensiveness of the search results have been improved.

CN119961380BActive Publication Date: 2025-06-06SHANDONG ZHENGTU INFORMATION POLYTRON TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436409.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

Traditional government data retrieval algorithms based on precise keyword matching cannot automatically regard vocabulary with similar semantics as equivalent concepts, making it difficult to comprehensively mine government data files containing different expressions but with similar content semantics, affecting the accuracy and comprehensiveness of the search results.

Method used

An intelligent search method for government affairs data based on the Internet of Things is proposed. By obtaining a collection of government affairs data files, extracting keywords and establishing an inverted index, semantic expansion of the original vocabulary input by the user, determining the importance of the extended vocabulary, and filtering the search results based on the inverted index and importance.

Benefits of technology

It realizes intelligent matching of semantic similar vocabulary, improves the accuracy and comprehensiveness of government data retrieval, can quickly respond to keyword input and provide high-quality search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961380B_ABST
    Figure CN119961380B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing technology, and specifically to an intelligent retrieval method and system for government data based on the Internet of Things. The method includes: obtaining a set of government data files, extracting keywords, and establishing an inverted index for each keyword. If the user input is a keyword, the corresponding government data file is directly output based on the inverted index. If the user input is a non-keyword, semantic expansion is first performed to obtain multiple extended words, and the degree of interest of the extended words to historical users is determined based on the number of clicks on the extended words and the browsing time of historical users in the historical time period. The contextual relevance is determined by co-occurrence analysis of the extended words and the upper and lower words of the original words, and the importance of the extended words is determined based on the comprehensive interest level and contextual relevance. The target extended words are screened out according to the importance, and the retrieval results are determined according to the target extended words. This method improves the comprehensiveness and accuracy of retrieving government data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and more particularly to an intelligent retrieval method and system for government affairs data based on the Internet of Things. Background Art

[0002] Government data based on the Internet of Things refers to data related to government affairs collected through the Internet of Things technology, such as traffic flow, temperature, environmental index, etc. The government system generates corresponding government data files based on these data and stores them on the server side of the government system, such as traffic flow analysis reports, temperature statistical analysis, environmental monitoring analysis, etc.

[0003] Retrieving and analyzing these government data files is crucial to scientific government management and decision-making. Currently, government data retrieval mainly relies on precise keyword matching algorithms, that is, purely based on literal comparison. Specifically, when a user enters a keyword, the system will retrieve all government data files containing the keyword.

[0004] However, this retrieval method has serious defects in semantic understanding. Take the government data files in the fields of transportation and meteorology as an example. In terms of transportation, there are many synonyms, such as "congestion", "blockage" and "slow movement". If the user enters "congestion" for retrieval, using the precise keyword matching algorithm, those government data files that contain "blockage" or "slow movement" but actually describe the same traffic conditions will not be retrieved, resulting in the retrieval results missing a lot of relevant information. In the field of meteorology, "waterlogging" and "urban flooding" describe similar phenomena derived from meteorological disasters. If the user enters "urban flooding" for retrieval, using the traditional algorithm, those government data files that contain "waterlogging" but actually describe the same traffic conditions will not be retrieved.

[0005] It can be seen that the traditional retrieval algorithm based on precise keyword matching cannot automatically regard these semantically similar words as equivalent concepts, making it difficult to comprehensively mine government data files that contain different expressions but similar content semantics, affecting the accuracy and comprehensiveness of the retrieval results of government data files. Summary of the invention

[0006] In order to solve the problem that the traditional retrieval algorithm based on precise keyword matching leads to inaccurate and incomplete retrieval results of government data files, the present invention proposes an intelligent retrieval method and system for government data based on the Internet of Things.

[0007] On the one hand, the present invention provides an intelligent retrieval method for government affairs data based on the Internet of Things, comprising:

[0008] Obtain a set of government data files, extract keywords from each government data file, obtain a keyword set for the set of government data files, and establish an inverted index for each keyword;

[0009] If the original word input by the user is a keyword, the corresponding government data file is directly output as the search result based on the inverted index of the keyword;

[0010] If the original word entered by the user is not a keyword, perform the following operations to complete the search:

[0011] Perform semantic expansion on the original vocabulary to determine multiple expanded vocabulary corresponding to the original vocabulary;

[0012] According to the number of hits and browsing time of historical government data files corresponding to each extended word, the interest level of each extended word to historical users is determined; according to the number of co-occurrences of the hyponymy and hyponymy of each extended word and the original word in the government data files, the contextual relevance of each extended word to the original word is determined;

[0013] The importance of each extended word is determined according to the interest level of each extended word for historical users and the contextual relevance of each extended word to the original word. The extended words with importance greater than a preset importance threshold are taken as target extended words, and the retrieval results are determined according to the inverted index of all target extended words.

[0014] This technical solution first obtains a collection of government data files as the data source of the entire retrieval process, and extracts key information from a large number of unstructured government data files by extracting keywords. The inverted index of each keyword is established, which changes the traditional forward index search method from files to keywords to an efficient search mode from keywords to files, thereby improving the retrieval efficiency. When the original word entered by the user happens to be a keyword extracted by the government data retrieval system, the corresponding government data file can be directly output as a retrieval result based on the inverted index. This quick retrieval method designed for simple input situations gives full play to the advantages of the inverted index and can quickly respond to user requests. When the original word entered by the user is not a keyword, the technical solution first semantically expands the original word and mines multiple extended words related to it, thereby broadening the search scope and fully covering the user's potential search needs. Next, an in-depth assessment is conducted on the relevance of extended vocabulary to user needs from two different dimensions. By determining the degree of interest of historical users in each extended vocabulary and the contextual relevance of each extended vocabulary to the original vocabulary, the importance of the extended vocabulary can be judged comprehensively and accurately. Finally, extended vocabulary that is highly consistent with the actual needs of users is screened out based on importance. This processing method for non-keyword input can accurately select the most relevant parts to user needs from a large number of extended vocabulary, provide users with high-quality search results, and thereby improve the accuracy and comprehensiveness of intelligent search results for government data.

[0015] Furthermore, the government data file set is government data collected through Internet of Things devices and transmitted to a government data system. The government data system performs data analysis, and the generated government data file set is stored on the server side of the government data system.

[0016] Furthermore, the original vocabulary is semantically expanded to determine a plurality of expanded vocabulary corresponding to the original vocabulary as follows:

[0017] Obtain all synonyms and near synonyms of the original vocabulary as the original extended vocabulary set;

[0018] All words in the intersection of the original extended word set and the keyword set are taken as all extended words of the original words.

[0019] This technical solution uses the synonyms and antonyms of the original vocabulary through semantic expansion to establish a connection between the broad expressions that users may use and the keyword system within the government data system, thus bridging the gap between user input and system keywords and meeting users' diverse search needs. In addition, the final expanded vocabulary comes from the intersection of the original expanded vocabulary set and the keyword set. These expanded vocabulary are closely related to the keywords in the government data file, which means that the search results have a high correlation with the original vocabulary entered by the user in both semantics and the government data field, avoiding the interference of irrelevant information.

[0020] Furthermore, the interest level of each extended word for historical users is determined based on the following formula:

[0021] ;

[0022] In the formula, For any extended word, for Regarding the level of interest of historical users, is the normalization function, for The serial number of the corresponding historical government data file, for The number of corresponding historical government data files, for The corresponding The number of hits on historical government data files, is the maximum number of hits on the historical government data files corresponding to all extended terms, for The corresponding The historical user browsing time of historical government data files, is the maximum browsing time of the historical government data files corresponding to all extended vocabularies, for The corresponding The time decay index of a historical government data file.

[0023] This technical solution can accurately reflect the historical users' attention to each extended vocabulary by analyzing the number of clicks on the historical government data files corresponding to the extended vocabulary, the browsing time of historical users, the time sequence of the historical government data files, and the coverage of the extended vocabulary in the historical government data files.

[0024] Furthermore, the method for obtaining the number of hits and historical user browsing time of the historical government data file corresponding to each extended vocabulary is as follows:

[0025] Obtaining search logs within a historical time period, and obtaining government data files that have been searched within the historical time period based on the search logs, as well as the number of clicks and browsing time of each searched government data file;

[0026] For any extended word, the government data files that contain the extended word among the government data files that have been searched within the historical time period are used as the historical government data files corresponding to the extended word.

[0027] Furthermore, the contextual relevance of each expanded word to the original word is determined based on the following method:

[0028] If the vocabulary is expanded With the original vocabulary No. In a certain government data file, and of No. The co-occurrence count of the hyponyms is increased by 1, and the government data file is used as and No. Co-occurrence government data files of hyponyms and hyponyms; Calculate contextual relevance:

[0029] ;

[0030] In the formula, for and contextual relevance, for and The total number of co-occurrences of hyponyms and hyponyms of for exist and The total word frequency of all co-occurring hyponyms in the government data file, for The total number of hyponyms and hyponyms of for No. Hyponym With The total word frequency of all co-occurring hyponyms and hyponyms in the government data file.

[0031] This technical solution examines the co-occurrence of the extended vocabulary and the original vocabulary's hyponyms and hyponyms in the government data files, and deeply explores the contextual relevance between words from the semantic structure level. This method breaks through the simple surface matching of words, takes into account the mutual relationship of words in the semantic hierarchy system, and more comprehensively reflects the internal connection between the extended vocabulary and the original vocabulary's hyponyms and hyponyms in the actual government data context, thereby providing a more semantically relevant basis for retrieval.

[0032] Furthermore, one way to determine the importance of each expanded word is:

[0033] ;

[0034] In the formula, For extended words The weight of To expand vocabulary Regarding the level of interest of historical users, To expand vocabulary With the original vocabulary contextual relevance.

[0035] This technical solution can comprehensively and objectively evaluate the value of the extended vocabulary in the retrieval process by comprehensively considering the interest of historical users in the extended vocabulary and the contextual relevance of the extended vocabulary to the original vocabulary, thus avoiding the one-sidedness that may be caused by judging based on only a single factor. In addition, the formula adopts the form of harmonic mean, which can avoid a larger value overly covering the influence of another smaller value, so that these two factors can play a balanced and complementary role in determining the importance of the extended vocabulary, and more accurately reflect the true importance of the extended vocabulary.

[0036] Furthermore, another way to determine the importance of each extended word is:

[0037] ;

[0038] In the formula, For extended words The weight of To expand vocabulary Regarding the level of interest of historical users, To expand vocabulary With the original vocabulary contextual relevance, for and The joint entropy of .

[0039] The importance of each expanded word calculated by this technical solution can screen out government data files that are not only closely semantically related to the original word entered by the user, but also have been paid attention to by historical users, and the relationship between these two factors is relatively stable. This can effectively reduce the interference of irrelevant or low-relevant information, improve the accuracy of retrieval results, and enable users to obtain valuable information that meets their needs more quickly.

[0040] Furthermore, the method for determining the search results according to the inverted index of all target extended words is:

[0041] According to the inverted index of each target extended vocabulary, the corresponding government data file set is obtained as the corresponding candidate government data file set;

[0042] The candidate government data file sets corresponding to all target expansion words are arranged in descending order of importance of each target expansion word as the search result.

[0043] On the other hand, the present invention also provides an intelligent retrieval system for government data based on the Internet of Things, the intelligent retrieval system comprising a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement any step of the intelligent retrieval method.

[0044] The present invention has the following effects:

[0045] The present invention constructs an intelligent search strategy. For keyword input, according to the inverted index of keywords in the government data file set, the keyword input can be quickly responded to and the search results can be obtained. For non-keyword input, semantic expansion is used to mine related words, and the importance of the extended words is determined by comprehensive analysis from the time dimension and the structure dimension of the government data file. The government data files corresponding to the extended words with high importance are used as the search results, so that the search results are highly matched with the actual needs of users, and the accuracy and comprehensiveness of government data retrieval are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic flow chart of the method of the present invention;

[0047] Figure 2 It is a system structure block diagram of the present invention. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0049] The present invention provides an intelligent retrieval method for government affairs data based on the Internet of Things, such as Figure 1 As shown in, including:

[0050] S1: Extract the keywords of each government affairs data file and build an inverted index for each keyword.

[0051] Collect government affairs data through Internet of Things devices. The government affairs data files are the government affairs data collected by Internet of Things devices and transmitted to the government affairs data system. After data analysis by the government affairs system, a set of government affairs data files is generated. These government affairs data files are all unstructured data and are usually stored in the form of documents on the server side of the government affairs data system. For example, based on the traffic flow, temperature, environmental index, etc. of the city collected by the Internet of Things devices, the government affairs system generates corresponding traffic flow analysis reports, temperature statistics analysis, environmental monitoring analysis, etc.

[0052] First, extract the keywords of each government affairs data file, including:

[0053] Obtain all government affairs data files in the government affairs data system, and then use Chinese word segmentation tools for word segmentation. Common Chinese word segmentation tools include jieba, HanLP, etc. In this step, use jieba to segment each word in each government affairs data file, and then remove the stop words in each government affairs data file, such as words like "de", "di", "le" that have no actual analysis value.

[0054] Next, extract the keywords in each government affairs data file:

[0055] Obtain the corpus or dictionary of the field where each government affairs data file is located. For example, for a government affairs data file on traffic flow statistics, according to the theme "traffic flow statistics" of this government affairs data file, it can be determined that the field is the traffic field, then obtain the corpus or dictionary of the traffic field. Another example, for a government affairs data file on meteorological reports, according to the theme "meteorological reports" of this government affairs data file, it can be determined that the field is the meteorological field, then obtain the corpus or dictionary of the meteorological field.

[0056] For any word in any government affairs data file, if the word can be retrieved in the corpus or dictionary of the field where the government affairs data file is located, then the word is a keyword of this government affairs data file.

[0057] All keywords of each government affairs data file can be obtained according to the method of this step. All keywords (the same keywords are counted as one) of the government affairs data file set (government affairs data system) can be obtained based on the keywords in all government affairs data files.

[0058] Finally, build an inverted index for each keyword: the forward index first counts which keywords are contained in each government data file, and then builds an index for these keywords. The inverted index is the opposite. The inverted index first counts which government data files each keyword appears in, and then builds an index for the government data files containing the keyword. For example, if the keyword "congestion" appears in the 30th government data file and the 100th government data file, the inverted index of the keyword is .

[0059] By constructing an inverted index, we can quickly locate government data files containing specific keywords, greatly improving retrieval efficiency. After obtaining a collection of government data files, we extract keywords and construct an inverted index. This is an efficient way to organize data.

[0060] S2: In response to the judgment result of whether the original word input by the user is a keyword, different search strategies are executed.

[0061] When users retrieve government data through the government data system, there are usually two situations:

[0062] One is that the original word input by the user happens to be a keyword in the keyword set of the government data file set. At this time, the government data system can accurately judge the user's search intention. The user wants to obtain the government data file corresponding to the keyword corresponding to the original word input, and these government data files happen to exist in the government data system, then they will be directly returned as search results.

[0063] The other is that the original words input by the user are not keywords in the government data file collection. In this case, the user's search intention is not clear enough and further analysis is needed to determine the search results.

[0064] Therefore, this step needs to determine whether the original word entered by the user is a keyword in the government data file set. If the word entered by the user is a keyword, the government data system directly outputs the corresponding government data file as the search result based on the inverted index of the word entered by the user. This method is simple and direct and can quickly respond to user requests. If the word entered by the user is not a keyword, it will not directly return blank like the conventional search method, but execute steps S3-S6 to determine the final search result.

[0065] S3: semantically expand the original vocabulary input by the user to obtain multiple expanded vocabulary corresponding to the original vocabulary.

[0066] In actual search scenarios, users may not be able to accurately input the keywords defined by the system. Users often use synonyms, near synonyms or other related expressions of keywords to express their needs. Through semantic expansion, it can better match the user's natural language input, conform to the user's search habits, and improve the friendliness of the user-system interaction.

[0067] That is to say, when users search for government data through government websites, the original words they enter are often not accurate keywords, resulting in empty search results. This is usually because the user enters a synonym or near-synonym of the keyword. For example, when a user wants to search for government data files related to housing construction with provident funds, the keyword is "construction", but the user enters "construction", which will result in empty search results. For another example, when a user wants to check traffic conditions, the keyword is "blocking", but the user enters "congestion", which will result in empty search results.

[0068] Therefore, in order to further improve the accuracy of the retrieval results, it is necessary to semantically expand the original words input by the user and determine which government data file data to output as the retrieval results based on all the expanded words.

[0069] The method for semantically expanding the original vocabulary input by the user is:

[0070] Obtain all synonyms and near synonyms of the original vocabulary as the original extended vocabulary set;

[0071] The intersection of the original extended vocabulary set and the keyword set (all keywords in the government data file set) is calculated, and all the words in the intersection are used as all the extended words of the original vocabulary. Each of the obtained extended words is a keyword.

[0072] The operation of this step can mine the potential search intent behind the user input. By obtaining the synonyms and near synonyms of the original vocabulary and finding the intersection with the keyword set, it helps to find keywords that are related to the user's intent but not directly thought of by the user, so as to more comprehensively understand the user's needs and provide search results that are more in line with the user's expectations. It is an optimization of the search strategy. Traditional search methods may directly return blank results when users enter non-keywords, but semantic expansion avoids this situation, allowing the system to perform more in-depth analysis and processing when faced with imprecise input, improving the intelligence and adaptability of the search system.

[0073] S4: Evaluate the importance of each expanded vocabulary.

[0074] Although all the extended words of the original words input by the user can be obtained in step S3, it is not possible to directly output the corresponding government data files according to the inverted indexes corresponding to these extended words and return them as search results.

[0075] Because the user's search intent is often complex and subtle, the degree to which each extended vocabulary fits the user's real needs varies greatly in different scenarios. If the files corresponding to the inverted index of all extended vocabulary are directly returned without distinction, a large amount of irrelevant or secondary information will be flooded in the search results, which not only increases the cost of users to filter out valid information, but also is likely to cause users to miss the key data that really meets their needs. Therefore, in order to truly meet user needs and improve the accuracy and effectiveness of search results, it is particularly necessary to rigorously and comprehensively evaluate the importance of each extended vocabulary and filter it.

[0076] Therefore, this step first analyzes from the time dimension to obtain the number of clicks and browsing time of the historical government data files corresponding to the extended vocabulary within a period of time, which can intuitively reflect the historical users' interest in it. A high level of interest means that the extended vocabulary is more likely to meet the current user's search needs, and is thus used to determine the search results. From the file structure dimension of the government data file, if the extended vocabulary and the original vocabulary have many co-occurrences of superordinate and subordinate words, it indicates that it has a high correlation with the original vocabulary context, which is more helpful to accurately determine the search results. By evaluating the importance of extended vocabulary, the matching degree between the search results and the real needs of users can be improved, and the user's search experience can be optimized.

[0077] Specifically, the steps include:

[0078] S41: Determine the interest level of each extended word for historical users based on the number of clicks on the historical government data files corresponding to each extended word and the browsing time of historical users.

[0079] Set the length of the historical time period to 3 months (experience value);

[0080] The government data retrieval system obtains the retrieval log within 3 months before the user inputs the original word. The retrieval log contains all the government data files retrieved within this historical time period, as well as the number of clicks and historical user browsing time of each retrieved government data file.

[0081] Any expansion of the original word entered by the user , based on the search log, obtain the search results that contain the expanded vocabulary in the past three months The government data files were retrieved and the retrieved files containing the extended vocabulary Government data files as extended vocabulary The corresponding historical government data files record the number of these historical government data files. , and for the extended vocabulary The corresponding historical government data files are sorted, and the closer the time when the user inputs the original word is to the historical government data file, the smaller the sequence number.

[0082] Similarly, obtain the historical government data files corresponding to all extended words (in the historical time period). And among all extended words, if a certain extended word has the largest number and the largest scale of historical government data files, record the number as ;

[0083] By analyzing the logs, we can obtain the number of hits and historical user browsing time of each historical government data file corresponding to each extended vocabulary. And among all extended vocabulary, if a government data file corresponding to an extended vocabulary has the highest number of hits, we record the number of hits as , if a government data file corresponding to a certain extended vocabulary has the longest browsing time, it is recorded as .

[0084] Finally, calculate the interest level of each expanded word for historical users:

[0085]

[0086] In the formula, For any extended word, for Regarding the level of interest of historical users, It is an integral character. The larger the value, the more historical users have used the extended vocabulary in the historical period. The tendency to pay attention is higher. is the normalization function, To expand vocabulary The serial number of the corresponding government data file, for The number of corresponding historical government data files, which reflects the extended vocabulary The "richness" in historical government data files refers to the frequency with which the extended vocabulary appears in relevant government data files. The maximum number of government data files corresponding to all extended vocabularies, used for Normalize. To expand vocabulary The corresponding The number of clicks on a historical government data file. The more clicks, the more historical users are interested in the first Extended vocabulary involved in historical government data files More interested. is the maximum number of hits on the historical government data files corresponding to all extended terms. It is a benchmark value used to normalize the number of hits on the historical government data files corresponding to each extended term. , the number of clicks on government data files corresponding to different extended vocabulary can be unified into a relative scale, which is convenient for comparison and comprehensive calculation. To expand vocabulary The corresponding The historical user browsing time of a historical government data file can more deeply reflect the historical user's interest in the content of the government data file. The longer the browsing time, the more interested the user is in the extended vocabulary involved in the government data file. is the maximum browsing time of the historical government data files corresponding to all extended vocabularies, and Similarly, it is a benchmark value used to normalize browsing time. The browsing time of each historical government data file is divided by , so that the browsing time of historical government data files corresponding to different extended vocabulary are on the same relative scale, which is convenient for comprehensive calculation. To expand vocabulary The corresponding The time decay index of a historical government data file is used to evaluate the historical users’ interest in the extended vocabulary. The dynamic impact of time factors on the historical users’ interest should be fully considered. As time goes by, the historical users’ attention to the early government data files will usually decrease gradually. is a constant less than 1 (usually between 0 and 1), here we take , Indicates that the newer the government data file is ( The smaller it is), at this time, The larger the time decay exponential The closer it is to 1, the more recent historical government data files contribute to the historical users’ interest. On the contrary, the older government data files ( The larger the time decay exponent The smaller it is, the smaller its contribution to user interest is. The introduction of the time decay factor allows the formula to reflect the impact of time on user interest, which is more in line with the actual situation. In the government data scenario, users may pay more attention to newer events and other related content. The time decay index can more accurately calculate the actual interest of users in the extended vocabulary and highlight the importance of historical government data files closest to the user input time.

[0087] In this formula, It is the product of the normalized click volume and browsing time, taking into account the historical users' click behavior and browsing time on the historical government data files. If a historical government data file has a high number of clicks and a long browsing time, then this product will be large, which means that the historical users are more interested in the government data file.

[0088] In this formula, Is to calculate the expanded vocabulary first The relative ratio of the number of corresponding historical government data files to the maximum number , and then perform a square root operation on it. The purpose of this is to appropriately reduce the expansion vocabulary The corresponding number of historical government data files affects the results, because as the number of historical government data files increases, their marginal contribution to the interest of historical users decreases.

[0089] For example, for a certain extended vocabulary, when the number of corresponding historical government data files increases from 1 to 10, this number has a more obvious effect on the improvement of user interest; but when the number of files increases from 100 to 110, although the number of historical government data files increases by 10, the improvement in user interest is not so significant.

[0090] In this formula, the numerator is the historical user's use of the extended vocabulary The corresponding interest measures of all historical government data files are weighted and summed. The denominator is the accumulation of the interest measures of all extended words for historical users as a normalized total, so that the interest level of each extended word for historical users is between 0 and 1, which is convenient for comparison and analysis.

[0091] To summarize, this step first performs a comprehensive calculation on each historical government data file corresponding to each extended vocabulary based on the number of clicks, browsing time, and time decay index to obtain the contribution value of each historical government data file to the interest level of historical users, and then adds up these contribution values ​​and multiplies them by the normalized square root value of the number of historical government data files corresponding to the extended vocabulary.

[0092] This means that it not only takes into account the historical users' interest indicators in the historical government data files involving each extended vocabulary, but also takes into account the impact of the richness of each extended vocabulary in the historical government data files on the overall interest level. At the same time, it highlights the importance of newer historical government data files and historical government data files with high clicks and browsing time, thereby providing a scientific, comprehensive and effective method for evaluating the interest level of each extended vocabulary for historical users.

[0093] S42: Determine the contextual relevance of each extended word to the original word according to the number of co-occurrences of the superordinate and subordinate words of each extended word and the original word in the government data file.

[0094] Specifically, first obtain all the hypernyms and hyponyms of the original vocabulary, and select any hypernym or hyponym. If the vocabulary is expanded and At the same time in a certain government data file, it indicates that the extended vocabulary and Generates co-occurrence, expands vocabulary and The co-occurrence count of the extended vocabulary is increased by 1, and the co-occurrence count of the extended vocabulary is increased by 1. and The number of government data files is recorded as , that is, to expand vocabulary and The total number of co-occurrences of , and the expanded vocabulary will appear at the same time and All government data files as extended vocabulary and Co-occurrence government data files. Extended vocabulary and The more co-occurring government data files there are, the more the extended vocabulary The word frequency and The higher the frequency of the word, the more the vocabulary is expanded. and the original vocabulary The higher the contextual relevance.

[0095] In this way, a co-occurrence analysis is performed on each extended word and each hypernym or hyponym of the original word, and the extended words that co-occur with the hypernym or hyponym of the original word are obtained.

[0096] For example, the original word is "water pollution control", and its superordinate word is "environmental pollution control", and its subordinate words include "industrial wastewater pollution control", "domestic sewage pollution control", etc. The original word has an extended word "wastewater purification". In the relevant environmental protection government data files, "wastewater purification" often co-occurs with "water pollution control" and its superordinate and subordinate words. In the documents on industrial pollution prevention and control, it is mentioned that "industrial wastewater pollution control" requires effective "wastewater purification" means to achieve the goal of "water pollution control"; in the documents on urban environmental planning, the importance of "wastewater purification" technology in the process of "domestic sewage pollution control" will also be explained, which all belong to the category of "environmental pollution control".

[0097] From the perspective of co-occurrence, the close co-occurrence of "wastewater purification" and "water pollution control" and their superordinate and subordinate words shows that it is highly consistent with the user's intention to search for information related to "water pollution control". When the user enters the original word "water pollution control" for search, the extended word "wastewater purification" is more likely to provide search results that meet the user's needs, because it frequently appears in various contexts related to "water pollution control", reflecting the professional expression and practical application scenarios of the means of dealing with water pollution problems in the field of environmental protection.

[0098] In one embodiment, the contextual relevance of each expanded word to the original word is calculated based on the following formula:

[0099]

[0100] In this formula, for and contextual relevance, As a whole character, its value reflects the expanded vocabulary and the original vocabulary The closer the connection is in context, the larger the value is, the higher the contextual relevance between the two is.

[0101] for and The total number of co-occurrences of hyponyms and hyponyms of the word. The greater the total number, the more likely it is that the expansion word and the original vocabulary The more hyponyms and hyponyms of a word appear simultaneously in more government data file scenarios, the greater the possibility that there is some connection between them. for exist and The total word frequency of all co-occurring hyponyms in the government data file, As a whole character, it reflects the extended vocabulary In those cases where the original word The frequency of the occurrence of the hyponymy and hyponymy words in the government data files, the higher the total word frequency, the more the extended vocabulary The more "active" in these relevant government data files, the more expanded the vocabulary With the original vocabulary The more contextual relevance the word has, the higher the probability is. for The total number of hyponyms and hyponyms of the original vocabulary, which is a counting indicator used to measure the The number of related hyponyms. for No. Hyponym With The total word frequency of all co-occurring government data files of hyponyms and hypernyms reflects the frequency of occurrence of each specific hyponym and hypernym in these co-occurring government data files. is the ordinal number of the hyponym, ranging from 1 to , used to Each hyponym and hyponym of the word are numbered and distinguished.

[0102] In this formula, yes All the hyponyms of The average word frequency of the co-occurrence of these hyponyms and hypernyms in the government data files is taken into account by summing up the total word frequency of all hyponyms and hypernyms and then taking the average, taking into account the original words All hyponyms and hyponyms of the word as a whole are related to the extended word The higher the frequency of co-occurrence, the more likely it is that the vocabulary has been expanded. The stronger the contextual connection with the original word.

[0103] In this formula, Similar to the vector modulus calculation form, it is and The normalization process is performed so that the calculated The value should be within a reasonable range to avoid and The value is too large, resulting in Unreasonably increased.

[0104] In short, in government data files, the co-occurrence of words reflects their connection in context. The higher the co-occurrence frequency, the stronger the contextual relevance, which can more comprehensively evaluate the relationship between extended words and original words. Through co-occurrence analysis, it comprehensively considers multiple factors such as the number of co-occurring government data files, the word frequency of extended words, and the word frequency of the super- and subordinate words of the original words, and comprehensively reflects the contextual relevance of extended words in government data files with the original words.

[0105] S5: Determine the importance of each extended word according to the interest level of each extended word to historical users and the contextual relevance of each extended word to the original word.

[0106] In one embodiment, the importance of each expanded word is calculated based on the following relationship:

[0107]

[0108] In this formula, For extended words The importance of To expand vocabulary Regarding the level of interest of historical users, To expand vocabulary With the original vocabulary contextual relevance.

[0109] This formula uses the form of harmonic mean, which emphasizes the role of smaller values ​​when dealing with the combined impact of two variables on a certain result. In the current scenario, the interest level of each extended word to historical users and the contextual relevance of each extended word to the original word are both important, but if one of the values ​​is very small, it may mean that the relationship between the extended word and the original word in this respect is very weak. By using the harmonic mean, it is possible to avoid a larger relevance value masking the impact of another smaller value, thereby more comprehensively reflecting the combined effect of the two on the weight. For example, if an extended word has a strong contextual relevance to the original word, but the historical user has a very low interest in it, when using the harmonic mean to calculate, this lower interest will have a greater "pull-down" effect on the importance of the extended word, which avoids overestimating the importance of the extended word simply because of the strong contextual relevance, and can more objectively reflect its actual situation.

[0110] In this formula, the numerator This ensures that when both are large, the expansion word The importance of can be increased accordingly, and considering the interaction between the two, the denominator The influence of the two is balanced. and When one of the values ​​is too large, the denominator will also increase accordingly, for example, The value is very large. If we only look at the numerator, the importance of this expanded vocabulary may be over-inflated. , The impact will be limited. The effect can also be reflected, avoiding Calculation results of unilateral dominance importance.

[0111] In summary, this formula incorporates the interest of historical users in the extended vocabulary and the contextual relevance between the extended vocabulary and the original vocabulary into the same framework for calculation. These two factors reflect the importance of extended vocabulary in meeting user search needs from different perspectives. The interest of historical users in the extended vocabulary reflects the degree of attention paid by past users to the content related to the extended vocabulary, and reflects its popularity and attractiveness in actual usage scenarios. The contextual relevance between the extended vocabulary and the original vocabulary measures the closeness of the extended vocabulary to the original vocabulary entered by the user in context from a semantic level, ensuring the relevance of the extended vocabulary to the user's initial search intention. By considering these two factors at the same time, the value of extended vocabulary in the retrieval process can be comprehensively evaluated, rather than relying solely on the judgment of a single dimension.

[0112] In one embodiment, another way to determine the importance of each expansion word is:

[0113] The calculation formula for the importance of each extended word is constructed based on the concept of entropy. Entropy can measure the uncertainty or confusion of information. The importance of each extended word can be determined by calculating the joint entropy of the interest level of historical users and contextual relevance of each extended word.

[0114] First, suppose , ,according to and calculate and The joint entropy of

[0115]

[0116] In this formula, for and The joint entropy essentially measures the degree of uncertainty of the two variables as a whole. The smaller, the and The trend of change is more consistent. For example, and When both are higher or lower, and The value of becomes smaller, reflecting that they have better synergy and can better support each other in determining the importance of extended vocabulary.

[0117] Finally, calculate the importance of each expanded word:

[0118]

[0119] In this formula, For extended words The importance of For extended words Regarding the level of interest of historical users, For extended words and the original vocabulary contextual relevance.

[0120] In this formula, Weighed and The uncertainty of these two factors. When smaller, Larger, description and The relationship between them is relatively consistent, the uncertainty is low, and the higher the stability, the more stable these two factors can reflect the importance of the extended vocabulary more stably, and can further highlight the importance of the extended vocabulary. Therefore, using right On the contrary, when When larger, Smaller, description and The difference is large, the uncertainty is high, and the importance needs to be weakened. Therefore, using right Narrow it down and appropriately reduce the importance of the extended vocabulary.

[0121] The joint entropy allows the importance of the extended vocabulary to adapt to different contexts and user behavior patterns. In some cases, the historical user's interest level may be highly consistent with the contextual relevance of the original vocabulary. In this case, the joint entropy is small, and the weight calculation will tend to emphasize this consistency; in other cases, the historical user's interest level and contextual relevance may be relatively independent, and the joint entropy is large. The weight calculation will be adjusted accordingly to avoid over-emphasizing one aspect.

[0122] In short, this method can more accurately determine the importance of extended vocabulary because it takes into account the two key factors of historical user interest and contextual relevance. In the government data retrieval system, for the original vocabulary input by the user, the government data system can prioritize the display of government data files related to the extended vocabulary that has a high degree of interest to historical users and is closely related to the original vocabulary context according to the importance of each extended word, thereby improving the efficiency and accuracy of users in obtaining information.

[0123] S6: Determine the retrieval results according to the importance of each expanded word.

[0124] Arrange all extended words in descending order of importance, take the importance of the upper quartile (the importance corresponding to 25%) of the sorted importance as the importance threshold, and take the extended words whose importance is greater than the importance threshold as the target extended words. Quartile is an important indicator used in statistics to describe the distribution characteristics of data. The upper quartile can divide the data into four equal parts from high to low, and select the extended words with the highest importance (top 25%) to filter out the extended words with higher importance. Quartile can avoid the influence of extreme values ​​to a certain extent, while reflecting the higher level characteristics in the data set. In the current scenario, it helps to select the relatively most important part from many extended words as the target extended words.

[0125] According to the inverted index of each target extended vocabulary, the government data file set of the target extended vocabulary is obtained as the candidate government data file set of the target extended vocabulary; the candidate government data file sets corresponding to all target extended vocabulary are arranged in descending order according to the importance of each target extended vocabulary as the retrieval result.

[0126] For example, for the original words input by the user, the target expanded words are and , is more important than , The candidate government data file set includes 3 government data files. The candidate government data file set contains 1 government data file, and the final search result is:

[0127] Target extended vocabulary Candidate government data file 1;

[0128] Target extended vocabulary Candidate government data file 2;

[0129] Target extended vocabulary Candidate government data file 3;

[0130] Target extended vocabulary Candidate government data file 1.

[0131] Through this series of operations, the relevance of the government data files covered by the final search results to the actual needs of users has been improved. Since the government data files corresponding to important extended terms are displayed first, users can first access the information that is most likely to meet their needs when searching, which improves the relevance and quality of government data search results.

[0132] The present invention also provides an intelligent retrieval system for government data based on the Internet of Things, the intelligent retrieval system includes a memory and a processor, the memory stores a computer program, the processor executes the computer program to implement any step of the intelligent retrieval method, thereby realizing intelligent retrieval of government data.

[0133] Specifically, the intelligent retrieval system Figure 2 As shown, it includes an input module S100, a retrieval processing module S200, a data storage module S300 and an output module S400.

[0134] User: is the service object of the entire intelligent retrieval system, and initiates a retrieval request through the input module. The content input by the user can be keywords or non-keywords related to the government data that the user wants to query, and expects to obtain government data file retrieval results that meet the requirements from the system.

[0135] Input module S100: As the interactive interface between the user and the system, it is responsible for receiving the original vocabulary input by the user. Its main function is to preliminarily process and format the information input by the user so that the subsequent retrieval processing module can accurately identify and process it. For example, it performs encoding conversion on the input vocabulary, removes invalid characters, etc.

[0136] Retrieval processing module S200: It is the core processing unit of the entire retrieval system and undertakes multiple key tasks. Including:

[0137] Keyword processing: extract keywords of each government data file from the government data file set of the data storage module, construct a keyword set of the government data file set, and establish an inverted index for each keyword.

[0138] When the original word input by the user happens to be a keyword that has been indexed, the corresponding government data file can be quickly located and output as a search result based on the inverted index of the keyword. If the original word input by the user is a non-keyword, the search processing module will perform semantic expansion on it. The specific method is to first obtain all synonyms and near synonyms of the original word to form an original extended word set, and then take the intersection of the set and the keyword set to obtain all the extended words of the final original word. For each extended word, the interest level of each extended word for the historical user is calculated based on the number of clicks and historical user browsing time of the historical government data file obtained from the search log of the data storage module; at the same time, the contextual relevance of each extended word to the original word is determined by counting the number of co-occurrences of the upper and lower words of the extended word and the original word in the government data file. Finally, the importance of each extended word is determined based on these two factors, and the target extended words with an importance greater than the preset threshold are screened out. According to the inverted index of all target extended words, the corresponding government data file sets (i.e., candidate government data file sets) are obtained from the government data file set, and these file sets are arranged in descending order according to the importance of each target extended word to generate the final search result.

[0139] Data storage module S300: It contains two important parts. The first part is the collection of government data files: it stores government data collected through IoT devices. These data are analyzed by the government system to generate government data files and are stored on the server. These files are the basic data source for retrieval, and provide actual content support for keyword extraction, semantic expansion and the generation of retrieval results. The second part is the retrieval log: it records the information related to the retrieval operation in the historical time period, including the government data files that have been retrieved and the number of clicks and browsing time of each file. These log data play a key role in evaluating the interest of historical users in the extended vocabulary.

[0140] Output module S400: responsible for presenting the search results generated by the search processing module to the user. It will format and optimize the search results, such as displaying the file list according to certain typesetting rules and providing brief information of the file (such as file name, creation time, file size, etc.) so that the user can clearly and intuitively view and understand the search results.

[0141] The entire interaction process is:

[0142] The user inputs the original vocabulary that he wants to search into the system through the input module S100. After the input module S100 pre-processes the input content, it passes it to the search processing module S200. The search processing module S200 first determines whether the original vocabulary input by the user is a keyword. If it is a keyword, it directly obtains the corresponding government data file from the government data file set of the data storage module S300 based on the established inverted index, and then passes the result to the output module S400. If it is not a keyword, the search processing module S200 performs a semantic expansion operation to generate an extended vocabulary set of the original vocabulary. Then, using the search log and the government data file set in the data storage module S300, the interest level and context relevance of the extended vocabulary are calculated, its importance is evaluated, and the target extended vocabulary is screened out. Then, according to the inverted index of the target extended vocabulary, the candidate government data file set is obtained from the government data file set in the storage module S300, and it is sorted to generate the final search result, and finally the search result is passed to the output module S400. The output module S400 receives the search result, formats it and displays it to the user, completing a complete search interaction process.

[0143] During the entire process, the data storage module S300 provides the original data and historical retrieval information to the retrieval processing module S200. The retrieval processing module S200 performs complex retrieval and calculation operations based on this information, and finally the output module S400 presents the results to the user. The various modules work together to realize the intelligent retrieval function of government data based on the Internet of Things.

[0144] Although this specification has shown and described a number of embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will conceive of many modifications, changes and alternatives without departing from the ideas and spirit of the present invention. It should be understood that in the practice of the present invention, various alternatives to the embodiments of the present invention described herein may be employed.

Claims

1. An intelligent retrieval method for government affairs data based on the Internet of Things, characterized in that: include: Obtain a set of government data files, extract keywords from each government data file, obtain a keyword set for the set of government data files, and establish an inverted index for each keyword; If the original word input by the user is a keyword, the corresponding government data file is directly output as the search result based on the inverted index of the keyword; If the original word entered by the user is not a keyword, perform the following operations to complete the search: Perform semantic expansion on the original vocabulary to determine multiple expanded vocabulary corresponding to the original vocabulary; According to the number of hits and browsing time of historical government data files corresponding to each extended word, the interest level of each extended word to historical users is determined; according to the number of co-occurrences of the hyponymy and hyponymy of each extended word and the original word in the government data files, the contextual relevance of each extended word to the original word is determined; The importance of each extended word is determined according to the interest level of each extended word for historical users and the contextual relevance of each extended word to the original word. The extended words with importance greater than a preset importance threshold are taken as target extended words, and the retrieval results are determined according to the inverted index of all target extended words.

2. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: The government data file set is the government data collected by the Internet of Things devices and transmitted to the government data system. The government data system performs data analysis and generates a government data file set which is stored on the server side of the government data system.

3. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: The method of semantically expanding the original vocabulary and determining multiple expanded vocabulary corresponding to the original vocabulary is as follows: Obtain all synonyms and near synonyms of the original vocabulary as the original extended vocabulary set; All words in the intersection of the original extended word set and the keyword set are taken as all extended words of the original words.

4. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: The interest level of each extended word to historical users is determined based on the following formula: ; In the formula, For any extended word, for Regarding the level of interest of historical users, is the normalization function, for The serial number of the corresponding historical government data file, for The number of corresponding historical government data files, for The corresponding The number of hits on historical government data files, is the maximum number of hits on the historical government data files corresponding to all extended terms, for The corresponding The historical user browsing time of historical government data files, is the maximum browsing time of the historical government data files corresponding to all extended vocabularies, for The corresponding The time decay index of a historical government data file.

5. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: The method for obtaining the number of hits and historical user browsing time of the historical government data file corresponding to each extended vocabulary is as follows: Obtaining search logs within a historical time period, and obtaining government data files that have been searched within the historical time period based on the search logs, as well as the number of clicks and browsing time of each searched government data file; For any extended word, the government data files that contain the extended word among the government data files that have been searched within the historical time period are used as the historical government data files corresponding to the extended word.

6. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: The contextual relevance of each expanded word to the original word is determined based on the following method: If the vocabulary is expanded With the original vocabulary No. In a certain government data file, and of No. The co-occurrence count of the hyponyms is increased by 1, and the government data file is used as and No. The co-occurrence government data file of hyponyms and hyponyms; Calculate contextual relevance: ; In the formula, for and contextual relevance, for and The total number of co-occurrences of hyponyms and hyponyms of for exist and The total word frequency of all co-occurring hyponyms in the government data file, for The total number of hyponyms and hyponyms of for No. Hyponym With The total word frequency of all co-occurring hyponyms and hyponyms in the government data file.

7. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: One way to determine the importance of each expanded word is: ; In the formula, For extended words The weight of To expand vocabulary Regarding the level of interest of historical users, To expand vocabulary With the original vocabulary contextual relevance.

8. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: Another way to determine the importance of each expanded word is: ; In the formula, For extended words The weight of To expand vocabulary Regarding the level of interest of historical users, To expand vocabulary With the original vocabulary contextual relevance, for and The joint entropy of .

9. The intelligent retrieval method for government affairs data based on the Internet of Things according to claim 1 is characterized in that: The method for determining the search results according to the inverted index of all target expanded words is: According to the inverted index of each target extended vocabulary, the corresponding government data file set is obtained as the corresponding candidate government data file set; The candidate government data file sets corresponding to all target expansion words are arranged in descending order of importance of each target expansion word as the search result.

10. An intelligent retrieval system for government data based on the Internet of Things, characterized in that: The intelligent retrieval system includes a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the intelligent retrieval method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Semantic retrieval method and system for automatically extracting keywords

    CN118468881A

  • Digital archive intelligent retrieval method and system for storage device

    CN119202007A