Government affair intelligent search system based on semantic retrieval and use method thereof
Through the semantic retrieval-based government affairs intelligent search system, using the BERT model and a combination of multiple databases, the problems of the inability to accurately match user needs and the impact of data updates on search in the government affairs retrieval system have been solved, and the accuracy and efficiency of policy searches have been improved.
Patent Information
- Application Number
- CN202510928425.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
The existing government affairs retrieval system cannot accurately match user needs in policy retrieval, especially when the administrative division information provided by the user does not have relevant policies, it is unable to retrieve broader policy information, and the latest policy information cannot be retrieved during data updates, which increases the difficulty and time cost of retrieval.
A semantic retrieval-based intelligent government search system is adopted, sentence embedding is performed through the BERT pre-training model, vector and text retrieval is performed in combination with databases such as Milvus and ElasticSearch, policy metadata is stored in the metadata management database, and highly relevant keywords are recommended through the search hot word recommendation module. A data batch update mechanism is used to ensure that data updates do not affect searches.
The accuracy and efficiency of policy search results have been improved, and the search accuracy can be maintained during data updates. The recommended hot words are highly relevant to policies, allowing users to quickly find valuable policy information.
Smart Images

Figure CN120763318A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a government affairs intelligent search system based on semantic retrieval and a method for using the system, which belongs to the field of artificial intelligence data processing technology. Background Art
[0002] Efficient government policy retrieval is crucial for both government agencies and the public. Government agencies need to accurately locate relevant policies to make decisions and execute tasks, while the public also needs to query policies to understand relevant procedures. However, traditional government policy retrieval systems often have numerous limitations and struggle to meet the growing and diverse needs.
[0003] When existing government retrieval systems combine regional information to retrieve policies, users usually first provide location information, and the system then searches for policies based on this location information. However, this method has obvious shortcomings. If the administrative division information provided by the user does not have relevant policies, the system will not be able to retrieve the policies of the next higher level administrative division that applies to that first-level administrative division, resulting in users being unable to obtain a wider range of potentially applicable policy information. For high-frequency search content recommendations, the common hot word recommendation method is based on counting high-frequency words input by users. This method may cause some high-frequency words that are not related to policies to appear more frequently, resulting in the hot words recommended by the system not being strongly related to policies and unable to accurately match users' actual policy needs. This makes it difficult for users to quickly find valuable content in massive amounts of information, increasing the difficulty and time cost of retrieval. In addition, when updating policy data, the existing system may affect the policy search function, resulting in users being unable to retrieve the latest policy information during the data update period. Summary of the Invention
[0004] In view of the shortcomings of existing technologies, in order to solve technical problems such as the inability to accurately match users' actual policy needs in policy searches and the inability to search policy information when policy data is updated, a government affairs intelligent search system based on semantic retrieval and its usage method are proposed.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The government affairs intelligent search system based on semantic retrieval is characterized by including:
[0007] Sentence embedding module: The sentence embedding module embeds the query sentence, policy title, and full policy text through a pre-trained model and converts them into representations in a high-dimensional vector space;
[0008] Vector retrieval database: The vector retrieval database is used to store vector data of policy texts and search for relevant policies based on the embedded processing results of query statements;
[0009] Text retrieval database: The text retrieval database is used to store policy text data and search for relevant policies based on rule matching;
[0010] Metadata management database: the metadata management database is used to store policy metadata in text form;
[0011] Hot search word recommendation module: The hot search word recommendation module is used to statistically analyze the user's search statements and obtain high-frequency search keywords as hot search words.
[0012] Furthermore, the pre-training model is a BERT pre-training model, and the BERT pre-training model can simultaneously capture the contextual information of the text and achieve understanding of the two-way context of the text.
[0013] Furthermore, the vector retrieval database is any one of a Milvus database, a Qdrant database, a Faiss database, a Chroma database or a Weaviate database, which can realize the storage and management of vector data.
[0014] Furthermore, the text retrieval database is any one of the ElasticSearch database, Meilisearch database, Typesense database or Myscale database. The policy text is imported into the text retrieval database, so that the text data stored in the text retrieval database can reduce repeated calculations, improve response speed, and realize the rapid retrieval of data.
[0015] Furthermore, the metadata management database is implemented through a relational database, wherein the relational database is any one of a MySQL database, a PostgreSQL database, an Oracle Database database, an SQL Server database, and a MongoDB database. The policy metadata is imported into the metadata management database, and a cache mechanism is used to reduce duplicate connections and improve search efficiency.
[0016] Furthermore, policy metadata includes policy documents, policy URLs, policy regions, data batches, release times, and policy category information.
[0017] Furthermore, the step of statistically analyzing the user's search statements by the search hot word recommendation module includes:
[0018] 1) Segment the query entered by the user to obtain the search keywords;
[0019] 2) Analyze the relevance between search keywords and policies, perform relevance scoring, and obtain the policy relevance score of the search keywords;
[0020] 3) taking the search keywords with scores exceeding the threshold as candidate words;
[0021] 4) Perform word frequency statistics on the candidate words to obtain the number of times each candidate word appears, and select the N candidate words with the highest word frequency as the search hot words.
[0022] Preferably, in step 2) above, the method of analyzing the relevance between the search keywords and the policy and calculating the policy relevance score is:
[0023] ①Convert search keywords into vector form through sentence embedding module;
[0024] ② Send the keywords in vector form to the vector search database, analyze the similarity between the keywords in vector form and the policy topics in the vector database, and obtain the similarity scores between the keywords and the policy topics;
[0025] ③The highest similarity score is used as the policy relevance score of the search keyword.
[0026] A method for using a government affairs intelligent search system based on semantic retrieval is characterized by comprising the following steps:
[0027] (1) Policy ID generation: Generate a unique policy ID for each policy;
[0028] (2) Policy metadata is stored in the database and data is linked;
[0029] (3) User query processing: Process the query statement requested by the user to obtain the query keywords for subsequent text retrieval database queries;
[0030] (4) Vector database matching and similarity score calculation;
[0031] (5) Rule database matching and matching score calculation;
[0032] (6) The score results are weighted and the search results return the required policy metadata.
[0033] Furthermore, in the above step (2), when updating the policy metadata, a data batch-based data update mechanism is adopted, which includes the following steps:
[0034] S1. Confirm the data batch identifier and use the maximum data batch + 1 as the data batch for updating data;
[0035] S2. Write the updated data into the database and add a data batch identifier;
[0036] S3. Delete old batch data;
[0037] S4, when a user search request occurs during data updating, the policy of the largest data batch is taken as the search result.
[0038] The government affair intelligent search system based on semantic retrieval and the use method thereof have the following beneficial effects: the government affair intelligent search system based on semantic retrieval avoids the situation that the actual demand of a user for a policy cannot be accurately matched in a search process by combining semantic matching and rule matching, so that the search result of the policy is more accurate; the more extensive and applicable policy information can be searched by searching the policy first and then generating the region selection hierarchy according to the region information in the policy metadata; the search business is not affected when the policy metadata is updated by the data updating mechanism based on the data batch, so that the situation that the policy information cannot be searched when the policy metadata is updated is avoided, and the processing efficiency of the data is higher by replacing the modification operation with the adding and deleting operations; the search hot words recommended by the system can be strongly related to the policy by screening the search keywords in the semantic approximate matching mode. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 The flowchart of the use method of the government affair intelligent search system based on semantic retrieval is shown, and the diagram is taken as the abstract drawing. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0041] The government affair intelligent search system based on semantic retrieval comprises:
[0042] The sentence embedding module: wherein the sentence embedding module processes the embedding of the query sentence, the policy title and the policy full-text by a pre-training model, and converts the representation in a high-dimensional vector space;
[0043] Specifically, the pre-training model is a BERT pre-training model, and the BERT pre-training model can capture the context information of the text at the same time, so as to realize the understanding of the bidirectional context of the text.
[0044] The BERT pre-training model is used to convert the related text content of the query sentence, the policy title and the policy full-text into vector data and store the vector data.
[0045] Specifically, the BERT pre-trained model completes the specific process of text embedding: first, the input query statement, policy title and full policy text are preprocessed, including word segmentation, adding special tags and constructing input embeddings; then it is encoded through a multi-layer Transformer encoder, each layer of the encoder contains a self-attention mechanism and a feedforward neural network; finally, the embedding is output, and the hidden state after processing by the last layer of encoder is output to obtain a high-dimensional vector representation of the input content.
[0046] Vector retrieval database: The vector retrieval database is used to store vector data of policy texts and search for relevant policies based on the embedded processing results of query statements;
[0047] Specifically, the vector retrieval database can be any of the following: Milvus, Qdrant, Faiss, Chroma, or Weaviate databases. Different vector retrieval databases can be selected based on actual functional requirements to enable storage and management of vector data. The vector data converted by the BERT pre-trained model is stored in the vector retrieval database, generating relevant search documents when users query.
[0048] Text retrieval database: The text retrieval database is used to store policy text data and search for relevant policies based on rule matching;
[0049] Specifically, the text retrieval database is any one of the ElasticSearch database, Meilisearch database, Typesense database or Myscale database. Different text retrieval databases can be selected according to actual functional requirements, and the policy text can be imported into the text retrieval database to reduce repeated calculations of the text data stored in the text retrieval database, improve the response speed, and realize the rapid retrieval of data.
[0050] By combining sentence embedding modules, vector retrieval databases, and text retrieval databases, a hybrid search of "semantics + keywords" is achieved, making policy search results more accurate.
[0051] Metadata management database: The metadata management database is used to store policy metadata in text form;
[0052] Specifically, the metadata management database is implemented through a relational database, wherein the relational database is any one of a MySQL database, a PostgreSQL database, an Oracle Database database, an SQL Server database, and a MongoDB database. The policy metadata is imported into the metadata management database, and a cache mechanism is used to reduce repeated connections and improve search efficiency.
[0053] Among them, the policy metadata includes policy files, URLs of policies, policy regions, data batches, release times and policy category information, so that the search can obtain more applicable policy information.
[0054] The search hot word recommendation module is used to statistically analyze the search statements of users to obtain high-frequency search keywords as search hot words.
[0055] Specifically, the search hot word recommendation module statistically analyzes the search statements of users, and the steps include:
[0056] 1) The query statement input by the user during searching is processed by the BERT pre-training model for word segmentation to obtain search keywords;
[0057] 2) The relevance of the search keywords to the policies is analyzed, and a relevance score is calculated to obtain the policy relevance score of the search keywords;
[0058] Specifically, the way to analyze the relevance of the search keywords to the policies and calculate the policy relevance score is:
[0059] ① The search keywords are converted into vector form by the sentence embedding module;
[0060] ② The vector form of the keywords is sent to the vector retrieval database, and the similarity between the vector form of the keywords and each policy topic in the vector database is analyzed to obtain the similarity score of the keywords and each policy topic;
[0061] ③ The highest similarity score is taken as the policy relevance score of the search keywords.
[0062] 3) The search keywords with a score exceeding a threshold value are taken as candidate words;
[0063] Specifically, in different use scenarios, different thresholds are set according to different vector similarity calculation methods, and the threshold value is not limited; common vector similarity calculation methods include but are not limited to cosine similarity and inner product.
[0064] In the implementation process, the cosine similarity is used to calculate the vector similarity, and the calculation result of the cosine similarity is taken as the relevance score; the calculation result of the cosine similarity is between -1 and 1, and the closer to 1 indicates the higher similarity. The threshold value of the cosine similarity is usually set to 0.7 or 0.8, indicating a higher similarity. The specific threshold value is 0.7, and if the relevance score of the search keywords is greater than 0.7, the search keywords are taken as candidate words.
[0065] 4) The word frequency of the candidate words is counted to obtain the number of occurrences of each candidate word, and the N candidate words with the highest word frequency are taken as search hot words.
[0066] A method for using a government affairs intelligent search system based on semantic retrieval is characterized by comprising the following steps:
[0067] (1) Policy ID generation: Generate a unique policy ID for each policy;
[0068] (2) Policy metadata is stored in the database and data association is performed;
[0069] Specifically, preprocessing is first performed to convert the policy title text into vector form through the sentence embedding module, and then the preprocessed vector data and the policy title text are added to the vector retrieval database; then, the policy title text is imported into the text retrieval database; finally, the policy metadata is imported into the metadata management database, and the data association between the vector retrieval database, text retrieval database and metadata management database is realized through the generated unique policy ID.
[0070] (3) User query processing: Process the query statement requested by the user to obtain the query keywords for subsequent text retrieval database queries;
[0071] Specifically, the query statement requested by the user is converted into a vector form through the sentence embedding module to obtain a query vector, which is used for subsequent vector retrieval database queries; the user query statement is segmented to obtain query keywords, which are used for subsequent text retrieval database queries.
[0072] (4) Vector database matching and similarity score calculation;
[0073] Specifically, the query vector is sent to the vector retrieval database, and the similarity between the query vector and each policy title in the vector retrieval database is calculated to obtain a similarity score between the query vector and each policy title in the vector retrieval database.
[0074] (5) Rule database matching and matching score calculation;
[0075] Specifically, the query keyword is sent to the text retrieval database, and a rule matching query is performed with the policy title column in the text retrieval database through rule matching to obtain the matching score between the query keyword and each policy in the text retrieval database.
[0076] (6) The score results are weighted and the search results return the required policy metadata.
[0077] Specifically, the similarity score of each policy in the vector retrieval database and the score of the rule matching in the text retrieval database are weighted, and the weighted scores are sorted from high to low. The top K policies are taken as the matching results. According to the corresponding policy ID, the corresponding policy is queried from the metadata management database, and the required policy metadata is returned;
[0078] The K value can be designed according to actual display requirements. If a search result display page wants to display several search results, different K values are designed. Usually, the K value is greater than 1 and less than infinity.
[0079] If multiple matching results have the same policy title, the corresponding policy metadata is queried from the metadata management database based on the corresponding policy ID. Based on the regional information in the policy metadata, a regional selection hierarchy is generated. After prompting the user to select a region, the policy metadata required for the corresponding region is returned.
[0080] In addition, in the process of searching using the semantic retrieval-based government affairs intelligent search system of the present invention, the policy metadata in the metadata management database needs to be updated regularly to meet more accurate search requirements. Therefore, the policy metadata in the above step (2) needs to be updated regularly. When updating, a data update mechanism based on data batches is adopted to ensure that the search service is not affected when the data is updated. By replacing the modification operation with the addition and deletion operation, the data processing efficiency is improved. The specific steps include the following:
[0081] S1. Confirm the data batch identifier and use the maximum data batch + 1 as the data batch for updating data;
[0082] Specifically, when updating policy data, first query the data batch field of the policy data in the metadata database, and use the maximum data batch + 1 as the data batch of the updated data;
[0083] S2. Write the updated data into the database and add a data batch identifier;
[0084] Specifically, the updated policy data is subjected to policy storage operations, preprocessed, and written into the vector retrieval database, text retrieval database, and metadata management database. When writing into the metadata management database, the data batch field is the maximum data batch obtained in step S1 + 1.
[0085] S3. Delete old batch data;
[0086] Specifically, the policy metadata of the smaller data batch is deleted in the metadata management database, and the policy is synchronously deleted in the vector retrieval database and the text retrieval database according to the policy ID of the deleted policy metadata.
[0087] S4. When a user search request occurs during data update, the policy with the largest data batch is used as the retrieval result.
[0088] Specifically, when a user search request occurs during data update, the search results may include different data batch versions of a certain policy. In this regard, the data batches in the metadata of each policy are checked, and the policy with the largest data batch is output as the search result.
[0089] The above embodiments are merely examples for clarity of explanation and are not intended to limit the embodiments. Those skilled in the art will appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all embodiments here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. The government affairs intelligent search system based on semantic retrieval is characterized by: include: Sentence embedding module: The sentence embedding module embeds the query sentence, policy title, and full policy text through a pre-trained model and converts them into representations in a high-dimensional vector space; Vector retrieval database: The vector retrieval database is used to store vector data of policy texts and search for relevant policies based on the embedded processing results of query statements; Text retrieval database: The text retrieval database is used to store policy text data and search for relevant policies based on rule matching; Metadata management database: the metadata management database is used to store policy metadata in text form; Hot search word recommendation module: The hot search word recommendation module is used to statistically analyze the user's search statements and obtain high-frequency search keywords as hot search words.
2. The semantic retrieval-based government affairs intelligent search system according to claim 1 is characterized in that: The pre-trained model is the BERT pre-trained model, and the BERT pre-trained model can simultaneously capture the contextual information of the text and achieve understanding of the two-way context of the text.
3. The semantic retrieval-based government affairs intelligent search system according to claim 1 is characterized in that: The vector retrieval database is any one of the Milvus database, the Qdrant database, the Faiss database, the Chroma database or the Weaviate database, and can realize the storage and management of vector data.
4. The semantic retrieval-based government affairs intelligent search system according to claim 1 is characterized in that: The text retrieval database is any one of the ElasticSearch database, Meilisearch database, Typesense database or Myscale database. The policy text is imported into the text retrieval database to reduce repeated calculations of the text data stored in the text retrieval database, improve the response speed, and realize the rapid retrieval of data.
5. The semantic retrieval-based government affairs intelligent search system according to claim 1 is characterized in that: The metadata management database is implemented through a relational database, where the relational database is any one of MySQL database, PostgreSQL database, Oracle Database database, SQL Server database and MongoDB database. The policy metadata is imported into the metadata management database, and the cache mechanism is used to reduce duplicate connections and improve search efficiency.
6. The semantic retrieval-based government affairs intelligent search system according to claim 5 is characterized in that: Policy metadata includes policy document, policy URL, policy region, data batch, release time and policy category information.
7. The semantic retrieval-based government affairs intelligent search system according to claim 1 is characterized in that: The steps of the search hot word recommendation module to statistically analyze user search statements include: 1) Segment the query entered by the user to obtain the search keywords; 2) Analyze the relevance between search keywords and policies, perform relevance scoring, and obtain the policy relevance score of the search keywords; 3) taking the search keywords with scores exceeding the threshold as candidate words; 4) Perform word frequency statistics on the candidate words to obtain the number of times each candidate word appears, and select the N candidate words with the highest word frequency as the search hot words.
8. The semantic retrieval-based government affairs intelligent search system according to claim 7 is characterized in that: In step 2) above, the method for analyzing the relevance between the search keywords and the policy and calculating the policy relevance score is as follows: ①Convert search keywords into vector form through sentence embedding module; ② Send the keywords in vector form to the vector search database, analyze the similarity between the keywords in vector form and the policy topics in the vector database, and obtain the similarity scores between the keywords and the policy topics; ③The highest similarity score is used as the policy relevance score of the search keyword.
9. The method for using the government affairs intelligent search system based on semantic retrieval is characterized in that: The following steps are involved: (1) Policy ID generation: Generate a unique policy ID for each policy; (2) Policy metadata is stored in the database and data is linked; (3) User query processing: Process the query statement requested by the user to obtain the query keywords for subsequent text retrieval database queries; (4) Vector database matching and similarity score calculation; (5) Rule database matching and matching score calculation; (6) The score results are weighted and the search results return the required policy metadata.
10. The method for using the semantic retrieval-based government affairs intelligent search system according to claim 9, characterized in that: In step (2), when updating the policy metadata, a data batch-based data update mechanism is adopted, which includes the following steps: S1. Confirm the data batch identifier and use the maximum data batch + 1 as the data batch for updating data; S2. Write the updated data into the database and add a data batch identifier; S3. Delete old batch data; S4. When a user search request occurs during data update, the policy with the largest data batch is used as the retrieval result.
Citation Information
Patent Citations
Enterprise policy matching method based on label similarity
CN112380318A
Data sub-library processing method, device and platform based on Domino platform
CN113377876A
Distributed data synchronization method and system
CN113901141A
Intelligent retrieval recommendation method and system
CN116186381A
Multi-knowledge granularity text retrieval method and device for RAG
CN119415623A