Government affair industry intelligent information retrieval and pushing system and method based on large model

Through the modularly designed intelligent information retrieval system of the government affairs industry, combined with large language models and search enhancement generation technology, the problem that traditional government affairs information retrieval system is difficult to deal with complex semantics and personalized needs is solved, and efficient and accurate information retrieval and personalized push is achieved, which improves the user experience of government affairs information services.

CN120336504APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510415677.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional government information retrieval systems are difficult to deal with complex semantics and personalized needs, and cannot meet the diverse information needs of users.

Method used

Modular design is adopted, combining natural language processing, large language model and search enhancement generation technology, and receiving user queries through intelligent semantic understanding and query analysis modules, using pre-trained large language models for semantic analysis and intent recognition, optimizing query logic with context perception mechanism, and quickly searching related document fragments from the government cloud through semantic-driven efficient search modules, generating enhanced intelligent content, supporting multi-modal output, and achieving accurate push through personalized push and feedback optimization modules.

Benefits of technology

It realizes the full process optimization from semantic understanding to personalized push, provides efficient and accurate information retrieval and push, supports a variety of content forms, ensures the accuracy and credibility of generated content, and dynamically optimizes system performance based on user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336504A_ABST
    Figure CN120336504A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and machine learning, in particular to a government affair industry intelligent information retrieval and pushing system and method based on a large model, and the system comprises an intelligent semantic understanding and query analysis module, a semantic-driven efficient retrieval module, a generation-enhanced intelligent content generation module, and a personalized pushing and feedback optimization module. The method has the beneficial effects that a natural language query request input by a user is received through the intelligent semantic understanding and query analysis module, semantic analysis and intention recognition are performed by utilizing a pre-trained large language model, key information is extracted, and query logic is optimized through a context sensing mechanism. Then, a semantic-driven efficient retrieval module quickly retrieves document fragments most relevant to user query from mass data of government affair cloud, precise matching is achieved through semantic vectorization and an efficient vector retrieval technology, and a retrieval strategy is dynamically optimized in combination with user feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and machine learning, and particularly to an intelligent information retrieval and push system and method for the government affairs industry based on large models. Background Art

[0002] With the advancement of digital government affairs, a large amount of data such as policy documents and regulatory announcements needs to be efficiently managed and accurately pushed to meet the diverse information needs of users. Traditional retrieval systems mostly rely on keyword matching and are difficult to handle complex semantics and personalized requirements. Through modular design, this patent combines natural language processing (NLP), large language models (LLM), retrieval augmented generation (RAG) technology, and reinforcement learning to achieve full-process optimization from semantic understanding to personalized push.

[0003] Currently, designing an intelligent information retrieval and push system based on large models can meet the personalized needs of users for government services, and can provide accurate information push based on the user's historical behavior and preference information, providing effective support for the government affairs industry to serve the public. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent information retrieval and push system and method for the government affairs industry based on large models to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An intelligent information retrieval and push system for the government affairs industry based on large models, including an intelligent semantic understanding and query parsing module, which is used for:

[0006] Receiving natural language query requests submitted by users through various input methods;

[0007] Encoding the queries input by users using a pre-trained large language model, extracting semantic features, understanding the true intentions of the user's queries, including the topics, objectives, and possible context information of the queries, and extracting key information in the queries based on the semantic understanding results, such as entities, time, and actions;

[0008] Preprocessing the user's queries, including word segmentation, part-of-speech tagging, and entity recognition, and converting the natural language queries into structured query information;

[0009] Providing a query optimization function, and performing correlation analysis on the user's consecutive queries through a context awareness mechanism to optimize the query logic.

[0010] Preferably, it further includes a semantically driven efficient retrieval module, which is used for:

[0011] Extract structured and unstructured text data from the government cloud, perform data cleaning, word segmentation, part-of-speech tagging and other natural language processing operations, use a pre-trained large language model to encode the processed text data to generate high-dimensional semantic vectors, and store the generated semantic vectors in an efficient vector retrieval system;

[0012] Receive the user query semantic representation passed by the intelligent semantic understanding and query parsing module, encode the query to generate a query vector, use the vector retrieval algorithm to quickly retrieve the document vector most similar to the query vector in the semantic index library, and initially screen out the top N most relevant document fragments according to the similarity score;

[0013] Use machine learning algorithms to evaluate the relevance of the initially screened retrieval results, optimize the result relevance by combining the context information of the user query and historical behavior data, and sort the retrieval results according to the relevance score;

[0014] Record the user's interaction behavior with the retrieval results as feedback data, and dynamically adjust the retrieval strategy and model parameters according to the user feedback and system operation data.

[0015] Preferably, it further includes a module for generating enhanced intelligent content, which is used for:

[0016] Receive the relevant document fragments retrieved by the semantic-driven efficient retrieval module, use the generation part of the RAG technology, inject the retrieved fragments as external knowledge into the generation process, encode and decode the document fragments through a large language model, and achieve semantic fusion and content integration by combining the attention mechanism and context awareness technology to generate a natural and fluent answer;

[0017] Dynamically adjust the style and format of the generated content according to the user's query intention and context information, and generate a coherent and personalized answer by combining the user's historical interaction records;

[0018] Support the generation of various content forms such as text, tables, and charts, and ensure the diversity and richness of the generated content through multi-modal fusion technology;

[0019] Introduce a verification mechanism during the generation process to ensure the accuracy and credibility of the generated content, and continuously optimize the parameters of the generation model by combining user feedback;

[0020] Provide explanatory support for the generated content, indicating the source or key basis of the cited document fragments.

[0021] Preferably, it further includes a personalized push and feedback optimization module, which is used for:

[0022] Build a user portrait system, analyze the user's historical queries, browsing behaviors and feedback through machine learning algorithms, collect the user's basic information, query records, click behaviors, and stay time data, generate a user preference model, and identify the user's interest points and behavior patterns using clustering analysis and association rule mining;

[0023] Based on the user portrait and preference model, combined with the content generated by the enhanced intelligent content generation module, achieve personalized information push, and dynamically adjust the type and priority of the pushed content according to the user's interest fields, query frequencies and historical feedback;

[0024] Provide multi-channel push interfaces, support multiple push methods such as web pages, mobile terminals, and government affairs APPs, and combine with the user management system of the government affairs cloud to achieve accurate information push and user notification functions;

[0025] Collect and analyze the user's real-time feedback, including click-through rate, satisfaction score, complaints, and use reinforcement learning technology to dynamically adjust the retrieval and generation strategies according to the user's feedback, and continuously optimize the accuracy of the system and the user experience.

[0026] Preferably, the pre-trained large language model is a model based on the Transformer architecture, such as BERT, SentenceTransformers or a model custom-trained based on this architecture, and the efficient vector retrieval system is FAISS or Milvus.

[0027] A method for a government affairs industry intelligent information retrieval and push system based on a large model, comprising the following steps:

[0028] Receive a natural language query request submitted by the user through multiple input methods;

[0029] Use the pre-trained large language model to encode the query input by the user, extract semantic features, understand the real intention of the user's query, including the query topic, target and possible context information, and extract key information in the query based on the semantic understanding result, such as entities, time and actions;

[0030] Preprocess the user's query, including word segmentation, part-of-speech tagging, entity recognition, and convert the natural language query into structured query information;

[0031] Provide a query optimization function, perform correlation analysis on the user's consecutive queries through a context awareness mechanism, and optimize the query logic.

[0032] Preferably, it further comprises the following steps:

[0033] Extract structured and unstructured text data from the government cloud, perform data cleaning, word segmentation, part-of-speech tagging, and other natural language processing operations. Use a pre-trained large language model to encode the processed text data to generate high-dimensional semantic vectors, and store the generated semantic vectors in an efficient vector retrieval system;

[0034] Receive the user's query semantic representation after semantic understanding and query parsing, encode the query to generate a query vector, use a vector retrieval algorithm to quickly retrieve the document vectors most similar to the query vector in the semantic index library, and preliminarily screen out the top N most relevant document fragments according to the similarity score;

[0035] Use machine learning algorithms to evaluate the relevance of the preliminarily screened retrieval results, optimize the result relevance by combining the context information and historical behavior data of the user's query, and sort the retrieval results according to the relevance score;

[0036] Record the user's interaction behavior with the retrieval results as feedback data, and dynamically adjust the retrieval strategy and model parameters according to the user feedback and system operation data.

[0037] Preferably, the following steps are further included:

[0038] Receive the relevant document fragments retrieved by the semantic-driven efficient retrieval module, use the generation part of the RAG technology, inject the retrieved fragments as external knowledge into the generation process, encode and decode the document fragments through a large language model, and realize semantic fusion and content integration by combining the attention mechanism and context awareness technology to generate a natural and fluent answer;

[0039] Dynamically adjust the style and format of the generated content according to the user's query intention and context information, and generate a coherent and personalized answer by combining the user's historical interaction records;

[0040] Support the generation of various content forms such as text, tables, and charts, and ensure the diversity and richness of the generated content through multimodal fusion technology;

[0041] Introduce a verification mechanism during the generation process to ensure the accuracy and credibility of the generated content, avoid information deviation or errors by comparing the retrieved document fragments and the generated content, and continuously optimize the parameters of the generation model by combining user feedback;

[0042] Provide explanatory support for the generated content, indicating the source or key basis of the cited document fragments.

[0043] Preferably, the following steps are further included:

[0044] Build a user portrait system, analyze the user's historical queries, browsing behaviors, and feedback through machine learning algorithms, collect the user's basic information, query records, click behaviors, and stay time data, generate a user preference model, and identify the user's interest points and behavior patterns using clustering analysis and association rule mining;

[0045] Based on the user portrait and preference model, combined with the content generated by the generation-enhanced intelligent content generation module, achieve personalized information push, and dynamically adjust the type and priority of the pushed content according to the user's interest fields, query frequencies, and historical feedback;

[0046] Provide multi-channel push interfaces, support various push methods such as web pages, mobile devices, and government affairs APPs, and combine with the user management system of the government affairs cloud to achieve accurate information push and user notification functions.

[0047] Preferably, it further includes the following steps:

[0048] Collect and analyze the user's real-time feedback, including click-through rate, satisfaction score, and complaints;

[0049] Use reinforcement learning technology to dynamically adjust the retrieval and generation strategies according to the user's feedback. For example, optimize the accuracy and relevance of the generated content through a reward mechanism, and adjust the push frequency and content type according to the user's negative feedback to continuously optimize the accuracy of the system and the user experience.

[0050] Compared with the prior art, the beneficial effects of the present invention are:

[0051] The intelligent information retrieval and push system and method for the government affairs industry based on a large model proposed by the present invention receive a natural language query request input by the user through the intelligent semantic understanding and query parsing module, use a pre-trained large language model for semantic parsing and intent recognition, extract key information, and optimize the query logic through a context awareness mechanism. Subsequently, the semantic-driven efficient retrieval module quickly retrieves the most relevant document fragments to the user's query from the massive data in the government affairs cloud, achieves accurate matching through semantic vectorization and efficient vector retrieval technology, and dynamically optimizes the retrieval strategy in combination with the user's feedback. The generation-enhanced intelligent content generation module receives the retrieved document fragments, uses the RAG technology and the large language model to generate high-quality content that meets the user's needs, supports various multi-modal output forms such as text, tables, and charts, and optimizes the accuracy and credibility of the generated content through a verification mechanism and user feedback. Finally, the personalized push and feedback optimization module achieves accurate information push based on the user portrait and preference analysis, and dynamically adjusts the push strategy according to the user's real-time feedback through the multi-channel push interface and reinforcement learning technology to continuously optimize the system performance and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1Flow chart of intelligent semantic understanding and query parsing for the present invention;

[0053] Figure 2 Flow chart of efficient retrieval driven by semantics for the present invention;

[0054] Figure 3 Flow chart of generating enhanced intelligent content generation for the present invention;

[0055] Figure 4 Flow chart of personalized push and feedback optimization for the present invention;

[0056] Figure 5 Flow chart of the method for the present invention. Detailed implementation manners

[0057] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages clearer, the following further elaborates on the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are some, but not all, embodiments of the present invention, and are only used to explain the embodiments of the present invention, rather than to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0058] Embodiment 1. Please refer to Figures 1 to 4 , the present invention provides a technical solution: an intelligent information retrieval and push system for the government affairs industry based on a large model, including:

[0059] Module 1: Intelligent semantic understanding and query parsing module

[0060] The intelligent semantic understanding and query parsing module is the front-end entry of the entire information retrieval and push system, ensuring that the system can understand the true intentions of users and lay a solid foundation for subsequent information retrieval and generation modules.

[0061] Receiving user input: The module first receives natural language query requests submitted by users through various input methods (such as text boxes, voice input, etc.). These requests may be vague, colloquial, and may even contain spelling mistakes or ungrammatical situations.

[0062] Semantic understanding and intention recognition: Using a pre-trained large language model (such as a model based on the Transformer architecture), encode the query input by the user and extract its semantic features. Through the context awareness ability of the model, understand the true intention of the user's query, including the topic, target, and possible context information of the query. For example, when the user inputs "I want to query the recent changes in social security policies", the system needs to recognize that the user's intention is to query the latest changes in social security policies.

[0063] Key information extraction: Based on the results of semantic understanding, further extract the key information in the query, such as entities (e.g., "social security policy"), time (e.g., "recently"), and actions (e.g., "query"). These key information will serve as important inputs for the subsequent retrieval and generation modules.

[0064] Query preprocessing: Preprocess the user query, including natural language processing techniques such as word segmentation, part-of-speech tagging, and entity recognition. Word segmentation divides the continuous text input by the user into independent lexical units; part-of-speech tagging labels the part of speech of each word (e.g., noun, verb, etc.); entity recognition identifies the key entities in the query (e.g., person names, place names, organization names, etc.). These preprocessing steps help transform the natural language query into structured query information for subsequent module processing.

[0065] Query optimization and context association: Provide query optimization functions, and through the context awareness mechanism, conduct correlation analysis on the user's continuous queries. For example, the user may first query "social security policy" and then ask "how to handle social security transfer". The system needs to identify the relevance of the two queries and optimize the query logic according to the context information to improve the query efficiency and accuracy.

[0066] Module 2: Semantic-driven efficient retrieval module

[0067] Module 2 quickly retrieves the most relevant document fragments or data records from the massive data in the government cloud for the user query.

[0068] Construction and optimization of the semantic index library: 1) Data preprocessing: Extract structured and unstructured text data from the government cloud, including policy documents, regulations, announcements, etc. Clean the data to remove noise information such as HTML tags and irrelevant symbols, and perform natural language processing operations such as word segmentation and part-of-speech tagging. 2) Semantic vectorization: Use pre-trained large language models (such as BERT, Sentence Transformers) to encode the preprocessed text data to generate high-dimensional semantic vectors. These vectors can capture the semantic information of the text and provide a basis for subsequent vector retrieval. 3) Index storage: Store the generated semantic vectors in an efficient vector retrieval system, such as FAISS or Milvus. These systems support the rapid retrieval of large-scale vectors.

[0069] Vector retrieval and semantic matching: 1) User query vectorization: Receive the semantic representation of the user query passed by Module 1, and use the same language model as the data preprocessing to encode the query to generate a query vector. 2) Retrieval execution: Use a vector retrieval algorithm (such as FAISS) to quickly retrieve the document vectors most similar to the query vector in the semantic index library. 3) Initial result screening: According to the similarity score, initially screen the top N most relevant document fragments to ensure the diversity and coverage of the retrieval results.

[0070] Retrieval Result Sorting and Optimization: 1) Relevance Assessment: Use machine learning algorithms (such as neural network - based ranking models) to assess the relevance of the initially screened retrieval results. Combine the context information of the user's query and historical behavior data to further optimize the relevance of the results. 2) Result Sorting: Sort the retrieval results according to the relevance scores to ensure that the document fragments that best meet the user's needs are ranked at the front. For consecutive queries, prioritize the results that are more relevant to the previous query to enhance the coherence and accuracy of the retrieval.

[0071] Feedback and Dynamic Optimization: 1) User Feedback Collection: Record the user's interaction behaviors with the retrieval results (such as clicks, dwell time, etc.) as feedback data. 2) Dynamic Adjustment: Dynamically adjust the retrieval strategies and model parameters based on user feedback and system operation data to continuously optimize the retrieval effect.

[0072] Module Three: Generation - Enhanced Intelligent Content Generation Module

[0073] Module Three efficiently generates high - quality content that meets the user's needs and supports multi - modal output to meet the diverse needs of the government cloud.

[0074] Semantic Fusion and Content Integration: Module Three receives the relevant document fragments retrieved by Module Two and uses the generation part (Generation) in the RAG technology to inject the retrieved fragments as external knowledge into the generation process. Encode and decode the document fragments through a large - language model (LLM), and combine attention mechanisms and context - awareness technologies to achieve semantic fusion and content integration, generating natural and fluent answers.

[0075] Dynamic Content Generation: Dynamically adjust the style and format of the generated content according to the user's query intent and context information. For example, in the government affairs consultation scenario, generate formal or colloquial answers according to the formality of the user's question. At the same time, combine the user's historical interaction records to generate coherent and personalized answers.

[0076] Multi - modal Generation Ability: Support the generation of various content forms such as text, tables, and charts. For example, for policy interpretation questions, generate detailed text answers; for data - related questions, generate intuitive tables or charts. Through multi - modal fusion technologies such as cross - modal attention mechanisms, ensure the diversity and richness of the generated content.

[0077] Optimization and Verification of Generated Content: During the generation process, introduce a verification mechanism to ensure the accuracy and credibility of the generated content. By comparing the retrieved document fragments and the generated content, avoid information deviation or errors. At the same time, combine user feedback to continuously optimize the parameters of the generation model and improve the generation quality.

[0078] Enhanced interpretability: Provide interpretive support for the generated content, such as indicating the source of the cited document fragments or key bases in the answer. This design of interpretability helps to enhance users' trust in the generated content, especially in the government affairs field, ensuring the transparency and authority of information.

[0079] Module Four: Personalized Push and Feedback Optimization Module

[0080] Module Four realizes precise personalized information push and dynamically optimizes the system performance according to user feedback, improving the user experience and the overall efficiency of the system.

[0081] User profile construction and preference analysis: Build a user profile system to analyze users' historical queries, browsing behaviors, and feedback through machine learning algorithms. Collect data such as users' basic information, query records, click behaviors, and stay times to generate a user preference model. Use clustering analysis and association rule mining to identify users' interest points and behavior patterns, providing data support for personalized push.

[0082] Personalized push strategy: Based on the user profile and preference model, combined with the content generated by Module Three, realize personalized information push. Dynamically adjust the type and priority of the push content according to users' interest areas, query frequencies, and historical feedback. For example, for users who often query social security policies, give priority to pushing the latest policy interpretations related to them.

[0083] Multi-channel push function: Provide multi-channel push interfaces, supporting multiple push methods such as web pages, mobile terminals, and government affairs APPs. Combine with the user management system of the government affairs cloud to realize precise information push and user notification functions.

[0084] Real-time feedback and dynamic optimization: Collect and analyze users' real-time feedback, including click-through rates, satisfaction scores, complaints, etc. Use reinforcement learning technology to dynamically adjust the retrieval and generation strategies according to users' feedback. For example, optimize the accuracy and relevance of the generated content through a reward mechanism, and adjust the push frequency and content type according to users' negative feedback. Continuously optimize the accuracy of the system and the user experience.

[0085] Embodiment 2, based on Embodiment 1, proposes a method for an intelligent information retrieval and push system in the government affairs industry based on a large model, including the following steps: receiving a natural language query request submitted by a user through various input methods; using a pre-trained large language model to encode the query input by the user, extract semantic features, understand the true intention of the user's query, including the query topic, target, and possible context information, and extract key information in the query based on the semantic understanding result, such as entities, time, and actions; preprocessing the user's query, including word segmentation, part-of-speech tagging, and entity recognition, to convert the natural language query into structured query information; providing a query optimization function to perform correlation analysis on the user's consecutive queries through a context-aware mechanism and optimize the query logic.

[0086] It also includes the following steps: extracting structured and unstructured text data from the government affairs cloud, performing natural language processing operations such as data cleaning, word segmentation, and part-of-speech tagging, using a pre-trained large language model to encode the processed text data to generate high-dimensional semantic vectors, and storing the generated semantic vectors in an efficient vector retrieval system; receiving the user's query semantic representation after semantic understanding and query parsing, encoding the query to generate a query vector, using a vector retrieval algorithm to quickly retrieve the document vectors most similar to the query vector in the semantic index library, and initially screening out the top N most relevant document fragments according to the similarity score; using a machine learning algorithm to evaluate the relevance of the initially screened retrieval results, optimizing the result relevance by combining the context information of the user's query and historical behavior data, and sorting the retrieval results according to the relevance score; recording the user's interaction behavior with the retrieval results as feedback data, and dynamically adjusting the retrieval strategy and model parameters according to the user feedback and system operation data.

[0087] It also includes the following steps: receiving the relevant document fragments retrieved by the semantic-driven efficient retrieval module, using the generation part of the RAG technology, injecting the retrieved fragments as external knowledge into the generation process, encoding and decoding the document fragments through a large language model, and realizing semantic fusion and content integration by combining the attention mechanism and context-aware technology to generate a natural and fluent answer; dynamically adjusting the style and format of the generated content according to the intention and context information of the user's query, and generating a coherent and personalized answer by combining the user's historical interaction records; supporting the generation of various content forms such as text, tables, and charts, and ensuring the diversity and richness of the generated content through multi-modal fusion technology; introducing a verification mechanism during the generation process to ensure the accuracy and credibility of the generated content, avoiding information deviation or errors by comparing the retrieved document fragments and the generated content, and continuously optimizing the parameters of the generation model according to the user feedback; providing explanatory support for the generated content, indicating the source or key basis of the cited document fragments.

[0088] It also includes the following steps: constructing a user portrait system, analyzing the user's historical queries, browsing behaviors and feedback through machine learning algorithms, collecting the user's basic information, query records, click behaviors, and stay time data, generating a user preference model, and identifying the user's interest points and behavior patterns by using clustering analysis and association rule mining; based on the user portrait and preference model, combining with the content generated by the enhanced intelligent content generation module to achieve personalized information push, and dynamically adjusting the type and priority of the pushed content according to the user's interest fields, query frequencies and historical feedbacks; providing multi-channel push interfaces, supporting multiple push methods such as web pages, mobile terminals, and government affairs APPs, and combining with the user management system of the government affairs cloud to achieve accurate information push and user notification functions.

[0089] It also includes the following steps: collecting and analyzing the user's real-time feedback, including click-through rate, satisfaction score, and complaints; using reinforcement learning technology to dynamically adjust the retrieval and generation strategies according to the user's feedback, for example, optimizing the accuracy and relevance of the generated content through a reward mechanism, adjusting the push frequency and content type according to the user's negative feedback, and continuously optimizing the accuracy of the system and the user experience.

[0090] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirits of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent information retrieval and push system for the government affairs industry based on large models, characterized in that: It includes an intelligent semantic understanding and query parsing module, which is used for: Receiving natural language query requests submitted by users through various input methods; Encoding the queries input by users using a pre-trained large language model, extracting semantic features, understanding the true intentions of users' queries, including the topics, objectives, and possible context information of the queries, and extracting key information in the queries based on the semantic understanding results, such as entities, time, and actions; Preprocessing the users' queries, including word segmentation, part-of-speech tagging, and entity recognition, and converting natural language queries into structured query information; Providing a query optimization function, and performing correlation analysis on the consecutive queries of users through a context awareness mechanism to optimize the query logic.

2. The intelligent information retrieval and push system for the government affairs industry based on the large model according to claim 1, wherein: It also includes a semantically driven efficient retrieval module, which is used for: Extracting structured and unstructured text data from the government cloud, performing natural language processing operations such as data cleaning, word segmentation, and part-of-speech tagging, encoding the processed text data using a pre-trained large language model to generate high-dimensional semantic vectors, and storing the generated semantic vectors in an efficient vector retrieval system; Receiving the semantic representation of the users' queries transmitted by the intelligent semantic understanding and query parsing module, encoding the queries to generate query vectors, using vector retrieval algorithms to quickly retrieve the most similar document vectors to the query vectors in the semantic index library, and initially screening out the top N most relevant document fragments according to the similarity scores; Using machine learning algorithms to evaluate the relevance of the initially screened retrieval results, optimizing the result relevance by combining the context information of the users' queries and historical behavior data, and sorting the retrieval results according to the relevance scores; Recording the interaction behaviors of users on the retrieval results as feedback data, and dynamically adjusting the retrieval strategies and model parameters according to the users' feedback and system operation data.

3. The intelligent information retrieval and push system for the government affairs industry based on a large model according to claim 2, characterized in that: It also includes a generation-enhanced intelligent content generation module, which is used for: Receiving the relevant document fragments retrieved by the semantically driven efficient retrieval module, using the generation part in the RAG technology, injecting the retrieved fragments as external knowledge into the generation process, encoding and decoding the document fragments through a large language model, and realizing semantic fusion and content integration by combining the attention mechanism and context awareness technology to generate natural and fluent answers; Dynamically adjusting the style and format of the generated content according to the intentions and context information of the users' queries, and generating coherent and personalized answers by combining the historical interaction records of the users; Supporting the generation of various content forms such as text, tables, and charts, and ensuring the diversity and richness of the generated content through multi-modal fusion technology; Introducing a verification mechanism during the generation process to ensure the accuracy and credibility of the generated content, and continuously optimizing the parameters of the generation model by combining the users' feedback; Providing explanatory support for the generated content, and indicating the sources or key bases of the cited document fragments.

4. The intelligent information retrieval and push system for the government affairs industry based on a large model according to claim 3, wherein: It also includes a personalized push and feedback optimization module, which is used for: Constructing a user portrait system, analyzing the historical queries, browsing behaviors, and feedback of users through machine learning algorithms, collecting the basic information, query records, click behaviors, and residence time data of users, generating a user preference model, and identifying the interest points and behavior patterns of users by using cluster analysis and association rule mining; Based on the user profile and preference model, combined with the content generated by the generation-enhanced intelligent content generation module, personalized information push is realized, and the type and priority of the push content are dynamically adjusted according to the user's interest areas, query frequencies, and historical feedback; Provide multi-channel push interfaces, support various push methods such as web pages, mobile terminals, and government affairs APPs, and combine with the user management system of the government affairs cloud to achieve accurate information push and user notification functions; Collect and analyze the user's real-time feedback, including click-through rate, satisfaction score, complaints, and use reinforcement learning technology to dynamically adjust the retrieval and generation strategies according to the user's feedback, and continuously optimize the accuracy of the system and the user experience.

5. The intelligent information retrieval and push system for the government affairs industry based on a large model according to claim 4, wherein: The pre-trained large language model is a model based on the Transformer architecture, such as BERT, SentenceTransformers, or a model custom-trained based on this architecture, and the efficient vector retrieval system is FAISS or Milvus.

6. A method for the intelligent information retrieval and push system in the government affairs industry based on the large model according to claim 5, characterized in that: Include the following steps: Receive the natural language query request submitted by the user through various input methods; Use the pre-trained large language model to encode the user's input query, extract semantic features, understand the true intention of the user's query, including the query topic, target, and possible context information, and extract key information in the query based on the semantic understanding result, such as entities, time, and actions; Preprocess the user's query, including word segmentation, part-of-speech tagging, entity recognition, and convert the natural language query into structured query information; Provide a query optimization function, and perform correlation analysis on the user's consecutive queries through a context-aware mechanism to optimize the query logic.

7. A method according to claim 6, characterized in that: Also include the following steps: Extract structured and unstructured text data from the government affairs cloud, perform natural language processing operations such as data cleaning, word segmentation, and part-of-speech tagging, use the pre-trained large language model to encode the processed text data to generate high-dimensional semantic vectors, and store the generated semantic vectors in an efficient vector retrieval system; Receive the user query semantic representation after semantic understanding and query parsing, encode the query to generate a query vector, use the vector retrieval algorithm to quickly retrieve the document vector most similar to the query vector in the semantic index library, and initially screen out the top N most relevant document fragments according to the similarity score; Use machine learning algorithms to evaluate the relevance of the initially screened retrieval results, optimize the result relevance in combination with the context information of the user's query and historical behavior data, and sort the retrieval results according to the relevance score; Record the user's interaction behavior with the retrieval results as feedback data, and dynamically adjust the retrieval strategy and model parameters according to the user feedback and system operation data.

8. A method according to claim 6, characterized in that: Also include the following steps: Receive the relevant document fragments retrieved by the semantic-driven efficient retrieval module, use the generation part of the RAG technology, inject the retrieved fragments as external knowledge into the generation process, encode and decode the document fragments through the large language model, and combine the attention mechanism and context-aware technology to achieve semantic fusion and content integration, and generate a natural and fluent answer; Dynamically adjust the style and format of the generated content according to the user's query intention and context information, and generate a coherent and personalized answer by combining the user's historical interaction records; Support the generation of various content forms such as text, tables, and charts, and ensure the diversity and richness of the generated content through multi-modal fusion technology; Introduce a verification mechanism during the generation process to ensure the accuracy and credibility of the generated content. Avoid information deviation or errors by comparing the retrieved document fragments with the generated content, and continuously optimize the parameters of the generation model in combination with user feedback; Provide explanatory support for the generated content, indicating the source of the cited document fragments or key bases.

9. A method according to claim 6, characterized in that: It also includes the following steps: Build a user portrait system, analyze the user's historical queries, browsing behaviors, and feedback through machine learning algorithms, collect the user's basic information, query records, click behaviors, and dwell time data, generate a user preference model, and identify the user's interest points and behavior patterns using clustering analysis and association rule mining; Based on the user portrait and preference model, combine the content generated by the generation-enhanced intelligent content generation module to achieve personalized information push, and dynamically adjust the type and priority of the push content according to the user's interest areas, query frequencies, and historical feedback; Provide multi-channel push interfaces, support various push methods such as web pages, mobile terminals, and government affairs APPs, and realize accurate information push and user notification functions in combination with the user management system of the government affairs cloud.

10. A method according to claim 6, characterized in that: It also includes the following steps: Collect and analyze the user's real-time feedback, including click-through rate, satisfaction rating, and complaints; Use reinforcement learning technology to dynamically adjust the retrieval and generation strategies according to the user's feedback. For example, optimize the accuracy and relevance of the generated content through a reward mechanism, adjust the push frequency and content type according to the user's negative feedback, and continuously optimize the accuracy of the system and the user experience.

Citation Information

Patent Citations

  • Power data security policy large model question-answering system and method based on relation pooling

    CN119646160A

  • Enterprise knowledge base query method based on large language model

    CN119719345A

Cited By

  • Server retrieval result generation method and device, computer equipment and medium

    CN120541277A

  • Contract retrieval enhancement optimization method, equipment and medium

    CN120596647A

  • A contract search enhancement optimization method, device, and medium

    CN120596647B

  • Smart city service method and system based on Internet of Things large model, and medium

    CN120750978A

  • Government affair data processing method and device based on large model, equipment and storage medium

    CN120873069A