Machine learning intelligent question and answer analysis method and question and answer system based on government affair service
Through multi-channel data collection and multi-level semantic analysis technology, a government service question-and-answer combination matrix and hierarchical theme relationship data are constructed, which solves the problem of insufficient understanding of government service question-and-answer in the existing technology, and realizes the automation of intelligent sorting of multiple answers and feedback, significantly improving the user experience.
Patent Information
- Application Number
- CN202510685604.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing question-and-answer analysis method based on government services does not have enough depth of understanding and feedback accuracy when dealing with complex and diverse government service content, and cannot automatically generate targeted multi-answer intelligent sorting and feedback.
By deploying a multi-channel government service data collection engine, historical question data, answer data and knowledge base data are collected, and two-way word segmentation analysis, multi-scale word segmentation window analysis, semantic trunk triple analysis and LDA theme model are used to build a government service question-and-answer combination matrix and hierarchical theme relationship data, and an intelligent question-and-answer model is established to achieve intelligent sorting of multiple answers.
It improves the semantic understanding ability of government service data, enhances the depth of processing of complex natural languages by the Q&A system, realizes the automation of intelligent sorting and feedback of multiple answers, and significantly improves user experience and response efficiency.
Smart Images

Figure CN120198086A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a machine learning intelligent question and answer analysis method and a question and answer system based on government services. Background Art
[0002] With the comprehensive digital transformation of government services, more and more government service needs are shifting from offline to online. Government affairs are relatively dependent on manual work, which is not only labor-intensive but also highly repetitive. With the development of artificial intelligence technology, intelligent question and answer has been widely used in intelligent customer service question and answer systems for government services. However, the existing government service question and answer analysis methods usually have diverse and complex government service data sources, including historical question and answer data, government service knowledge base data, etc. When processing content with strong logic and multi-level associations in government services, a complete solution has not yet been formed, resulting in insufficient understanding depth and feedback accuracy of the question and answer model. In the context of diversified user questions, it is impossible to automatically generate targeted multi-answer intelligent sorting and feedback. Summary of the invention
[0003] Based on this, the present invention provides a machine learning intelligent question and answer analysis method and question and answer system based on government services to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a machine learning intelligent question-answering analysis method based on government services includes the following steps: Step S1: deploying a multi-channel government service data collection engine, and collecting government service data based on the multi-channel government service data collection engine to obtain government service data, wherein the government service data includes historical government service question data, historical government service answer data, and government service knowledge base data; Step S2: Perform bidirectional word segmentation and parsing processing on the government service data to generate government service word segmentation and parsing data; perform government service semantic trunk triple analysis based on the government service word segmentation and parsing data to generate government service semantic trunk triple data; perform government service question-answer association combination matrix design on the government service semantic trunk triple data to generate a government service question-answer combination matrix; Step S3: performing a topic distribution probability analysis of the government service question and answer combination matrix, and generating government service question and answer combination topic distribution probability data; performing a hierarchical topic relationship analysis of the question and answer combination based on the government service question and answer combination topic distribution probability data, and generating question and answer combination hierarchical topic relationship data; Step S4: Establish an intelligent feedback mapping relationship for government service Q&A through the hierarchical topic relationship data of Q&A combinations to obtain a government service intelligent Q&A model; receive instant government service user question data; transmit the instant government service user question data to the government service intelligent Q&A model for intelligent feedback processing of multiple answers to government services to obtain intelligent sorting data of multiple answers to government services; transmit the intelligent sorting data of multiple answers to government services to the terminal to execute the intelligent Q&A feedback operation of government services.
[0005] Further, step S1 includes the following steps: Step S11: Obtain the government service authorization API interface; Step S12: Perform interface source address identification processing on the government service authorization API interface to obtain an identified government service authorization API interface; Step S13: Design a government service data verification script; Step S14: Deploy a multi-channel government service data collection engine according to the identified government service authorization API interface and the government service data verification script; Step S15: Collect government service data based on the multi-channel government service data collection engine to obtain government service data.
[0006] Further, the government service data verification script in step S13 is used to perform government service text compliance verification and screening operations and government service text redundancy screening operations.
[0007] Further, step S2 includes the following steps: Step S21: Perform bidirectional word segmentation parsing processing on government service data to generate government service word segmentation parsing data; Step S22: Perform multi-scale word segmentation window analysis based on the government service word segmentation parsing data to generate multi-scale word segmentation window data; Step S23: Use the multi-scale word segmentation window data to perform window word segmentation part-of-speech identification analysis on the government service word segmentation parsing data to generate government service word segmentation part-of-speech data; Step S24: Perform government service part-of-speech entity relationship analysis based on the government service word segmentation part-of-speech data to generate government service part-of-speech entity relationship data; Step S25: Perform government service semantic backbone triple analysis based on the government service part-of-speech entity relationship data to generate government service semantic backbone triple data; Step S26: Perform government service Q&A association analysis based on the government service data to generate government service Q&A association data; Step S27: Design a combination matrix for government service Q&A association for the government service semantic backbone triple data through the government service Q&A association data to generate a government service Q&A combination matrix.
[0008] Further, step S25 includes the following steps: Step S251: Establish a syntactic dependency relationship tree for government services based on the preset BERT algorithm and part-of-speech data of government service word segmentation, and generate a government service syntactic dependency tree model; Step S252: Perform tree model dependency type optimization processing on the government service syntactic dependency tree model to generate an optimized government service syntactic dependency tree model; Step S253: Use the optimized government service syntactic dependency tree model to perform syntactic dependency backbone feature analysis on government service part-of-speech entity relationship data, and generate government service syntactic dependency backbone feature data; Step S254: Perform government service semantic backbone triple analysis based on the government service syntactic dependency backbone feature data to generate government service semantic backbone triple data.
[0009] Further, step S3 includes the following steps: Step S31: Use the LDA topic model to perform topic distribution probability analysis on the government service question-and-answer combination matrix, and generate government service question-and-answer combination topic distribution probability data; Step S32: Perform question-and-answer combination topic identification based on the government service question-and-answer combination topic distribution probability data to generate question-and-answer combination topic data; Step S33: Perform intra-cluster subset topic distribution probability identification processing on question-and-answer combination clustering based on the government service question-and-answer combination matrix and government service question-and-answer combination topic distribution probability data, and generate question-and-answer combination intra-cluster subset topic distribution probability data; Step S34: Perform intra-cluster sub-topic analysis identification processing on question-and-answer combination intra-cluster based on the question-and-answer combination intra-cluster subset topic distribution probability data, and generate question-and-answer combination intra-cluster sub-topic data; Step S35: Perform hierarchical topic relationship analysis on question-and-answer combinations based on the question-and-answer combination topic data and question-and-answer combination intra-cluster sub-topic data, and generate question-and-answer combination hierarchical topic relationship data.
[0010] Further, step S33 includes the following steps: Step S331: Perform question-and-answer combination similarity analysis on the government service question-and-answer combination matrix to generate government service question-and-answer combination similarity data; perform government service question-and-answer combination clustering cluster analysis based on the government service question-and-answer combination similarity data to generate government service question-and-answer combination clustering cluster data; Step S332: Map the government service question-and-answer combination topic distribution probability data to the government service question-and-answer combination clustering cluster data for intra-cluster subset topic distribution probability identification processing to generate question-and-answer combination intra-cluster subset topic distribution probability data.
[0011] Further, step S34 includes the following steps: Step S341: Analyze the subset topic feature of the question-and-answer combination cluster for the subset topic distribution probability data of the question-and-answer combination cluster to generate subset topic feature data of the question-and-answer combination cluster; Step S342: Reconstruct the subset topic of the question-and-answer combination cluster according to the subset topic feature data of the question-and-answer combination cluster to generate subset topic reconstruction data of the question-and-answer combination cluster; Step S343: Identify the sub-topics within the question-and-answer combination cluster according to the subset topic reconstruction data of the question-and-answer combination cluster to generate sub-topic data within the question-and-answer combination cluster.
[0012] Further, step S4 includes the following steps: Step S41: Establish an intelligent feedback mapping relationship for government service question-and-answer through the question-and-answer combination hierarchical topic relationship data to obtain a government service intelligent question-and-answer model; Step S42: Receive instant government service user question data; Step S43: Transmit the instant government service user question data to the government service intelligent question-and-answer model for multi-answer intelligent feedback processing of government services, including: matching the user question data with the question-and-answer combination hierarchical topic relationship data through the government service intelligent question-and-answer model to generate multi-matching answer data for government services; analyzing the confidence of the multi-matching answers for government services according to the multi-matching answer data for government services to generate confidence data for the multi-matching answers for government services; using the confidence data for the multi-matching answers for government services as a multi-answer sorting mechanism for government services, and performing multi-answer intelligent sorting and confidence output on the multi-matching answer data for government services through the multi-answer sorting mechanism for government services to obtain multi-answer intelligent sorting data for government services; Step S44: Transmit the multi-answer intelligent sorting data for government services to the terminal to execute the government service intelligent question-and-answer feedback operation.
[0013] This specification provides a machine learning intelligent question-and-answer system based on government services for performing the machine learning intelligent question-and-answer analysis method based on government services as described above. The machine learning intelligent question-and-answer system based on government services includes: A data collection module for deploying a multi-channel government service data collection engine and collecting government service data based on the multi-channel government service data collection engine to obtain government service data, where the government service data includes historical government service question data, historical government service answer data, and government service knowledge base data; The Q&A combination standardization module is used to perform bidirectional word segmentation and parsing processing on government service data to generate government service word segmentation and parsing data; perform government service semantic backbone triple analysis based on the government service word segmentation and parsing data to generate government service semantic backbone triple data; design a combination matrix for government service Q&A association for the government service semantic backbone triple data to generate a government service Q&A combination matrix; The Q&A combination theme identification module is used to perform an analysis of the theme distribution probability of the Q&A combination on the government service Q&A combination matrix to generate government service Q&A combination theme distribution probability data; perform an analysis of the hierarchical theme relationship of the Q&A combination based on the government service Q&A combination theme distribution probability data to generate Q&A combination hierarchical theme relationship data; The intelligent Q&A feedback module is used to establish an intelligent feedback mapping relationship for government service Q&A through the Q&A combination hierarchical theme relationship data to obtain a government service intelligent Q&A model; receive instant government service user question data; transmit the instant government service user question data to the government service intelligent Q&A model for multi-answer intelligent feedback processing of government services to obtain government service multi-answer intelligent sorting data; transmit the government service multi-answer intelligent sorting data to the terminal to execute the government service intelligent Q&A feedback operation.
[0014] The beneficial effects of this application are as follows. The deployment of the multi-channel government service data collection engine and the government service data collection play a significant role in improving data quality and the accuracy of the question-and-answer system. By obtaining the government service authorization API interface and performing interface source address identification processing, it is possible to ensure the legal and secure acquisition of government service data from official channels, ensuring the reliability of data sources and avoiding interference from external untrusted data on system performance. The designed government service data verification script executes compliance verification screening operations and text redundancy screening operations, further enhancing the accuracy of data processing. Text compliance verification can effectively eliminate question and answer content that does not conform to government service specifications, while text redundancy screening can remove redundant data and duplicate content, thus ensuring the quality and uniqueness of the data. Through this layer of screening, the system can focus more on high-quality data, improving the efficiency and accuracy of subsequent processing. The multi-channel data collection engine that combines the government service authorization API interface identification and the verification script deployment can automatically and efficiently collect government service-related data from multiple channels. Through bidirectional word segmentation parsing and semantic backbone triple analysis techniques, the semantic understanding ability of government service data is significantly improved, enabling the question-and-answer system to better handle complex and diverse natural language questions. Bidirectional word segmentation parsing can solve the problem that traditional unidirectional word segmentation cannot accurately capture context information by analyzing the context relationship of words from both left and right directions. By generating high-quality word segmentation parsing data, the system can more accurately understand the keywords and important information in the user's question, improving the accuracy of semantic analysis. Based on the generated word segmentation parsing data, multi-scale word segmentation window analysis helps the system capture semantic information at different levels, further enhancing the understanding ability of diverse language structures. Multi-scale word segmentation window analysis solves the ambiguity problem in government service texts by performing word combination and semantic analysis at different scales, improving the accuracy of the system in processing long sentences, complex sentences, and polysemous words. Based on the part-of-speech data of word segmentation, part-of-speech identification analysis further refines the depth of semantic understanding. By adding grammatical identifiers (such as nouns, verbs, adjectives, etc.) to each word, the system can distinguish different types of information and correctly understand the overall structure of the sentence, which helps improve the accuracy of subsequent semantic analysis. During the syntactic dependency relationship tree model and semantic backbone triple analysis process, the system deeply mines the core grammar and semantic structure in the text through the BERT algorithm and the optimized syntactic dependency tree model, and then extracts the key information in government service questions and answers. This grammar tree structure can not only solve the grammar dependency problem but also effectively identify important components such as the subject, predicate, and object in the question and answer, thus providing more accurate semantic support for subsequent question and answer analysis and matching. Through the LDA topic model, the topic distribution probability analysis of the government service question-and-answer combination matrix is carried out, and then the potential topic structure in government service questions and answers is revealed.The application of the LDA (Latent Dirichlet Allocation) topic model can effectively mine the distributions of different topics from a large amount of Q&A data, generate the probability data of the combined topic distribution of government service Q&A, automatically identify and distinguish different categories of government service questions, enabling the system to accurately identify and classify various government service needs. Through further analysis and processing of the topic distribution probability data, the system can assign clear topic identifiers to each Q&A combination, making subsequent Q&A processing clearer and more standardized. Each Q&A combination has a corresponding topic label, which helps the system quickly find relevant reference information when processing user questions, avoiding inefficient irrelevant retrievals and duplicate processing, thus improving the response speed and accuracy. The analysis of Q&A combination similarity and clustering cluster analysis further enhances the system's processing ability in the face of complex questions. Similarity analysis helps the system understand the relationships between different government service questions, and clustering cluster analysis groups Q&A combinations with similar topics into one category, effectively reducing the computational burden and improving the retrieval efficiency. By mapping the topic distribution probability to the clustering cluster, the system can identify the topic distribution probability for subsets within each cluster, thereby improving the system's accurate response ability in multiple similar question scenarios. The feature analysis and reconstruction of the topic distribution probability data for subsets within the cluster generate more refined sub-topic data for Q&A combinations within the cluster. This refined processing ensures that each Q&A combination can be further analyzed and responded to based on its sub-topics, thus achieving more detailed and diverse government service responses. By establishing an intelligent feedback mapping relationship through the hierarchical topic relationship data of Q&A combinations, the system can accurately match the user's immediate question with relevant information in the government service knowledge base, ensuring that highly relevant and accurate answers can be provided. Through this mapping, the system not only improves the matching degree of user questions but also reduces incorrect answers caused by semantic understanding deviations. Receiving immediate government service user question data and transmitting it to the government service intelligent Q&A model for processing can ensure that each user question receives a timely and accurate response. In the stage of intelligent feedback processing of multiple answers for government services, by analyzing the confidence levels of multiple matching answers, the system can, based on the credibility of different answers, preferentially recommend the most relevant and correct answers. In this stage, by generating confidence data for multiple matching answers in government services, the accuracy and robustness of the intelligent Q&A system are further enhanced. The application of the multi-answer sorting mechanism for government services enables the system to intelligently sort multiple answers according to the confidence level and ensure that users can obtain the most reliable answers through the confidence output. The introduction of this mechanism greatly improves the reliability and user experience of the Q&A system because users can obtain the answers that best meet their needs instead of a bunch of redundant and uncertain options. Transmitting the intelligent sorting data of multiple answers for government services to the terminal ensures that the terminal can successfully execute the intelligent Q&A feedback operation for government services.In this way, users can not only obtain high-quality answers, but also accelerate the feedback of question and answer through the intelligent sorting mechanism, significantly improving the response efficiency of government services.
[0015] Therefore, through the deployment of a multi-channel government service data collection engine, the machine learning intelligent question and answer analysis method based on government services of the present invention enables the system to efficiently integrate historical question data, answer data, and government service knowledge base data, and uses data verification scripts to screen the collected data to ensure the accuracy and compliance of the data. By adopting bidirectional word segmentation parsing and multi-scale word segmentation window technology, the system can deeply mine the semantic features of government service data, and further optimize the semantic parsing process through a syntactic dependency tree model based on the BERT algorithm, and through semantic backbone triple analysis technology, effectively improve the ability to understand government service semantics. In particular, by optimizing with a syntactic dependency tree model based on BERT, the logical associations and hierarchical structures in government service texts can be accurately captured, significantly enhancing the depth of processing complex natural languages by the question and answer model, and solving the deficiencies in semantic parsing and logical reasoning of the prior art. Through the design of a question and answer combination matrix and the analysis of topic distribution probability based on the LDA topic model, the construction of hierarchical topic relationships and intelligent feedback mapping mechanisms for question and answer combinations is realized, and diverse and highly accurate intelligent feedback can be provided according to the user's immediate question content, further improving the rationality of answer priorities and significantly enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the step flow of a machine learning intelligent question and answer analysis method based on government services of the present invention; Figure 2 is Figure 1 a detailed implementation step flow schematic diagram of step S2 in Figure 3 is Figure 1 a detailed implementation step flow schematic diagram of step S3 in The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0018] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0019] It should be understood that although the terms "first", "second", etc. are used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0020] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a machine learning intelligent question and answer analysis method based on government affairs services. In the embodiments of the present invention, please refer to Figure 1 shown in the following figure, which is a schematic flowchart of the steps of a machine learning intelligent question and answer analysis method based on government affairs services according to the present invention. The machine learning intelligent question and answer analysis method based on government affairs services includes the following steps: Step S1: Deploy a multi-channel government affairs service data collection engine, and collect government affairs service data based on the multi-channel government affairs service data collection engine to obtain government affairs service data, where the government affairs service data includes historical government affairs service question data, historical government affairs service answer data, and government affairs service knowledge base data; In the embodiments of the present invention, in the implementation of the government service system, the corresponding API interface permissions are obtained by accessing the authorization management platform of the government service provider. In the specific implementation process, it is necessary to clarify the API interface directory list of the government service data sources. For example, "user basic information query interface", "approval process progress interface", and "government affairs interpretation content interface". During the implementation process, first, an interface authorization application needs to be submitted to the government service provider, including the description of the interface usage, the access frequency requirements, and the content of the security authentication protocol. After passing the review, the government service provider generates the corresponding API interface key (such as OAuth 2.0 key) and returns an interface authorization file containing the API call document. For the obtained government service API interfaces, it is necessary to perform source address identification processing to achieve accurate calls in a multi-interface environment. The source address (i.e., the request URL) of each API interface is parsed into a unique identifier. Each time the interface is called, the system will dynamically obtain the corresponding source address according to the identifier to ensure the accuracy and consistency of the call. For the collected government service data, a verification script needs to be designed to ensure the integrity and validity of the data. The core logic of the verification script includes two parts: text compliance verification and redundant data screening. Text compliance verification refers to checking the collected data content according to the standards in the government service field. Based on the aforementioned identified API interface information and verification script, a multi-channel government service data collection engine is constructed. In the implementation, the open-source framework Scrapy is used to build a distributed crawling module, and Celery is combined to achieve task scheduling. First, the interface identifier and its corresponding URL are configured as the input parameters of the collection task. Secondly, the verification script is used as the post-collection processing module to filter out non-compliant data, retain valid data, and store the valid data in a structured format (such as JSON or CSV) in a local or cloud database. The deployed data collection engine is started, and different API interfaces are automatically called by setting the scheduling frequency and execution period. For example, for the "approval process progress interface", the call frequency is set to once per hour, and a request is initiated through the HTTP GET method using the scheduling script, and the returned data is stored locally. After each collection task is completed, the task ID, execution time, return status code, and the number of collected data entries are recorded through the log. If the collected data contains an error status code or is a null value, the exception alert module is immediately triggered to send an email or SMS to notify the administrator to troubleshoot the problem. Through this step, a structured government service data set is finally obtained, providing high-quality input for subsequent Q&A analysis.
[0021] Step S2: Perform bidirectional word segmentation and parsing processing on government service data to generate government service word segmentation and parsing data; perform government service semantic backbone triple analysis based on the government service word segmentation and parsing data to generate government service semantic backbone triple data; design a combined matrix for government service question-answer association for the government service semantic backbone triple data to generate a government service question-answer combined matrix; In the embodiments of the present invention, the text of government service data is first parsed through bidirectional word segmentation technology. Open-source word segmentation tools such as Jieba or HanLP are used to first perform Chinese word segmentation, splitting the input text (such as "The implementation of government affairs has been released to the public") into a sequence of words. Then, bidirectional word segmentation processing is carried out, that is, word recognition is performed simultaneously from left to right and from right to left to ensure that the semantic connections between words can be captured during the processing. In this process, for specific named entities, such as "the implementation of government affairs", proper noun recognition is carried out to ensure the accuracy of the word segmentation results, and the word segmentation results are converted into a sequence of words and stored as structured data. The multi-scale word segmentation window analysis aims to capture semantic hierarchical relationships through lexical aggregation analysis at different scales (window sizes). In implementation, using the word segmentation parsed data, different-sized analysis windows are first set, and typical window sizes are 1 word, 2 words, 3 words, etc. Through the method of sliding windows, window word sequences of different scales are generated. Using the multi-scale word segmentation window data, window word segmentation part-of-speech tagging analysis is carried out on the government service word segmentation parsed data. Based on the multi-scale word segmentation window data, a part-of-speech tagging tool (such as Stanza or HanLP) is used to tag the part of speech for each word. In the specific operation process, by inputting the word segmentation results, the tool is used for part-of-speech tagging to determine the grammatical role of each word (such as noun, verb, adjective, etc.). This process includes: parsing the word segmentation results; using the existing part-of-speech tagging model to tag each word; recording the tagging results in the word segmentation parsed data and generating part-of-speech tagging data. The tagging data includes the part-of-speech label, part-of-speech category (such as verb, noun) and position identifier of each word, which is convenient for subsequent analysis. According to the government service word segmentation part-of-speech data, government service part-of-speech entity relationship analysis is carried out. Through the already tagged part-of-speech data, the entity relationships between words are identified and analyzed. Using the dependency parsing method (such as Spacy or StanfordNLP), syntactic dependency analysis is carried out on the word segmentation part-of-speech data, and a relationship graph of all relevant words is constructed. Each entity and its relationship are stored as a triple data of "entity 1 - relationship - entity 2" (such as "government affairs - implementation - situation"). According to the part-of-speech entity relationship data, semantic backbone triple analysis is carried out to extract the core semantic information in the text. Applying NLP tools (such as BERT or RoBERTa) for in-depth semantic understanding, through further processing of the entity relationship data, the key triples in the sentence are identified, and the semantic structure of the triples is optimized through a deep learning model to ensure the accuracy of the analysis results. The generated triple data is stored as a structured database table for subsequent query and analysis. Association analysis is carried out on the questions and answers in the government service data. Through statistical methods or machine learning techniques, the correlation between questions and answers is identified. By analyzing the historical question and answer data, effective association rules can be extracted from the text.Classify and cluster historical Q&A data, and then use pattern recognition methods to generate association information based on historical Q&A and corresponding similar questions and similar answers, and store and mark it in the form of data association. Design a combined matrix for government service Q&A association of government service semantic backbone triple data through government service Q&A association data to generate a government service Q&A combined matrix. Extract the relationship data between users' historical questions and historical answers from the government service Q&A association data. These data usually include the keywords, intentions of users' questions and the answer content provided by the system. By analyzing the semantic backbone triple data, extract representative semantic backbone information, such as the relationship between the theme, verb and object. Then, based on the above government service Q&A association data and semantic backbone triple data, conduct a combined matrix design. Specifically, when designing, each row represents the key triple of a Q&A scenario, and each column represents different Q&A themes or intentions. By comparing the semantic similarity and relevance between different Q&A scenarios, the system will generate a Q&A combined matrix, where each element represents the association strength or matching degree between different Q&A scenarios. Through the design of this combined matrix, the system can identify which Q&A contents are relevant, thereby effectively improving the accuracy and intelligence level of Q&A matching.
[0022] Step S3: Analyze the topic distribution probability of Q&A combinations in the government service Q&A combined matrix to generate government service Q&A combined topic distribution probability data; conduct a hierarchical topic relationship analysis of Q&A combinations based on the government service Q&A combined topic distribution probability data to generate Q&A combination hierarchical topic relationship data; In the embodiments of the present invention, the Latent Dirichlet Allocation (LDA) topic model is used to analyze the topic distribution probability of the government service question-and-answer combination matrix. Taking the government service question-and-answer combination matrix as input data, the text in the matrix is first preprocessed, including removing stop words, word segmentation, and word frequency statistics. Then, the preprocessed text data is vectorized into a bag-of-words model and input into the LDA model for topic modeling analysis. Through iterative calculations, the LDA model generates a set of topic distribution probability data for each question-and-answer combination, and the output results are stored in matrix form, where the rows represent the question-and-answer combinations, the columns represent the topics, and the cell values are the probability values of the topics. According to the topic distribution probability data generated by the LDA model, topic identification processing is performed on each question-and-answer combination, and the dominant topic of each question-and-answer combination is determined based on the topic distribution probability. The basis for determining the dominant topic is the one with the largest probability value. The generated question-and-answer combination topic data is stored in a structured table form, and each record includes the question-and-answer combination, the dominant topic number, and the corresponding keyword set. Based on the question-and-answer combination matrix and the topic distribution probability data, question-and-answer combination clustering and intra-cluster subset topic distribution probability analysis are carried out. The hierarchical clustering algorithm is used to perform clustering analysis on the question-and-answer combinations. The input data is the question-and-answer combination matrix and its topic distribution probability. The cosine similarity is used to calculate the similarity between question-and-answer combinations, and a clustering tree is constructed based on the similarity and a threshold is set for cluster division. For each clustering cluster, the intra-cluster subset topic distribution probability is further calculated for the question-and-answer combinations within it. The output intra-cluster subset topic distribution probability data includes the cluster number, the topic probability distribution, and the dominant topic. Further analysis is performed on the intra-cluster subset topic distribution probability data of the question-and-answer combination clusters to identify the sub-topics within each clustering cluster. First, the sub-topics with significant differences are extracted based on the intra-cluster subset topic distribution probability data, and the sub-topics are identified by calculating the distribution differences between topics (such as KL divergence). For subsets with significant topic differences, keyword extraction methods (such as TF-IDF) are used to extract keywords to identify the semantic content of the sub-topics. The output intra-cluster sub-topic data is stored in a structured form, including the cluster number, the sub-topic number, and the corresponding keyword set. Based on the question-and-answer combination topic data and the intra-cluster sub-topic data, the hierarchical topic relationship of the question-and-answer combinations is constructed. A hierarchical structure is constructed according to the subordinate relationship between the topics and sub-topics, using a tree-like data structure, where each node represents a topic or a sub-topic, and the edge represents the subordinate relationship. Through relationship analysis, the association path between each question-and-answer combination and its corresponding topic is determined, and the generated hierarchical topic relationship data is stored in a tree-like structure or a nested table form, and each record includes the topic hierarchical path and the question-and-answer combination association information.
[0023] Step S4: Establish an intelligent feedback mapping relationship for government service Q&A through the Q&A combination hierarchical topic relationship data to obtain a government service intelligent Q&A model; receive instant government service user question data; transmit the instant government service user question data to the government service intelligent Q&A model for intelligent feedback processing of multiple answers for government services to obtain government service multiple answer intelligent sorting data; transmit the government service multiple answer intelligent sorting data to the terminal to execute the government service intelligent Q&A feedback operation.
[0024] In the embodiment of the present invention, the Q&A combination hierarchical topic relationship data is used as input, and a graph neural network (GNN) model is adopted to construct an intelligent feedback mapping relationship. First, the Q&A combination hierarchical topic relationship data is mapped into a graph structure, where nodes represent topics, edges represent the hierarchical relationship between topics, and the edge weights are determined by the topic relationship strength. The GNN model is used to extract features and model relationships for this graph structure. During training, a supervised learning method is adopted, the objective function is optimized as cross-entropy loss, the learning rate is set to 0.001, and a government service intelligent Q&A model is output. This model can receive user questions and efficiently map them to corresponding answers through the hierarchical topic relationship. The government service question data input by the user is received in real time through the API interface, and the input format is a JSON object. The system uses regular expressions to verify the format of the input data to ensure the integrity of fields and the validity of characters. After passing the verification, the question text content is extracted and normalized, including removing extra spaces, converting to lowercase letters, and removing stop words. The processed question text is stored in vector form, for example, the question is represented as a vector through TF-IDF. The instant question data is matched with the Q&A combination hierarchical topic relationship data in the intelligent Q&A model. The sentence embedding method based on BERT is adopted to vectorize the user question and the hierarchical topic question, and the cosine similarity is calculated. Based on the similarity analysis results, matching association is performed to generate government service multiple matching answer data. The similarity scores of the multiple matching answer data are normalized, for example, the similarity is converted into a confidence distribution using the Softmax function. The generated multiple matching answer confidence data includes the matching answers and their corresponding confidences. A sorting mechanism is constructed based on the confidence data. A descending sorting strategy is adopted to preferentially output answers with high confidence. Confidence annotations are added to the sorted answers to generate government service multiple answer intelligent sorting data. After receiving by the user terminal, the display module is called to generate a user-friendly feedback result, for example, the answer content and confidence are displayed in the form of a card.
[0025] Further, step S1 includes the following steps: Step S11: Obtain the government service authorization API interface; Step S12: Perform interface source address identification processing on the government service authorization API interface to obtain the identified government service authorization API interface; Step S13: Design a government service data verification script; Step S14: Deploy a multi-channel government service data collection engine according to the identified government service authorization API interface and the government service data verification script; Step S15: Collect government service data based on the multi-channel government service data collection engine to obtain government service data.
[0026] In the embodiment of the present invention, in the implementation of the government service system, the corresponding API interface permissions are obtained by accessing the authorization management platform of the government service provider. In the specific implementation process, it is necessary to clarify the API interface directory list of the government service data source. For example, the "user basic information query interface", the "approval process progress interface", and the "government affairs interpretation content interface". In the implementation process, first, an interface authorization application needs to be submitted to the government service provider, including the purpose description of the application interface, the access frequency requirements, and the content of the security authentication protocol. After passing the review, the government service provider generates the corresponding API interface key (such as the OAuth 2.0 key) and returns an interface authorization file containing the API call document. Taking the interface "government affairs interpretation content interface" as an example, the authorization file includes the HTTP request URL of the interface, the authentication fields required in the request header (such as Authorization: Bearer <token>(0), the supported HTTP methods (such as GET or POST) and the necessary parameter fields (such as policy_id and request_time). For the obtained government service API interfaces, it is necessary to perform identification processing on the source addresses to achieve accurate calls in a multi-interface environment. Specifically, when implementing, the source address (i.e., the request URL) of each API interface is parsed into a unique identifier. For example, for the URL of the "Government Affairs Interpretation Content Interface", the identifier GOV_API_POLICY_DETAILS is generated and stored in the local database table api_interface_mapping. The recorded fields in the table include the interface name (such as "Government Affairs Interpretation Content Interface"), the interface URL, the interface identifier (such as GOV_API_POLICY_DETAILS), and the authentication key. Each time an interface is called, the system will dynamically obtain the corresponding source address from the api_interface_mapping table according to the identifier to ensure the accuracy and consistency of the call. For the collected government service data, it is necessary to design a verification script to ensure the integrity and validity of the data. The core logic of the verification script includes two parts: text compliance verification and redundant data screening. Text compliance verification refers to checking the collected data content according to the standards in the government service field. For example, in the approval progress data, it is necessary to verify whether the field status contains predefined values (such as "Accepted", "Approval in Progress", "Completed"); in the time field apply_date, it is necessary to ensure that the date format is YYYY-MM-DD. Redundant data screening is based on MD5 hash value calculation to compare whether the newly collected data is repeated with the historical data. If the hash value of a piece of data already exists in the historical data table, it is marked as redundant data and filtered. The verification script is implemented in Python, combined with regular expressions for content matching, and SQL stored procedures are used to complete redundant detection. Based on the aforementioned identified API interface information and verification script, a multi-channel government service data collection engine is constructed. In the implementation, the open-source framework Scrapy is used to build a distributed crawling module, and Celery is combined to achieve task scheduling. First, the interface identifier and its corresponding URL are configured as the input parameters of the collection task. For example, the task configuration file api_tasks.json records: {"task_id":"1","api_id":"GOV_API_POLICY_DETAILS","frequency":"hourly"}. Secondly, the verification script is used as the post-collection processing module to filter out non-compliant data, retain valid data, and store the valid data in a structured format (such as JSON or CSV) in the local or cloud database.For example, the collected government affairs interpretation data is stored in the table policy_details according to the field names policy_id, title, content, and publish_date respectively to ensure data quality. Start the deployed data collection engine, and make automated calls to different API interfaces by setting the scheduling frequency and execution cycle. For example, for the "approval process progress interface", set the call frequency to once per hour, use the scheduling script to initiate a request via the HTTP GET method, and store the returned data locally. After each data collection task is completed, record the task ID, execution time, return status code (such as 200 or 500), and the number of collected data entries through the log. If the collected data contains an error status code or is a null value, immediately trigger the exception alert module to send an email or SMS to notify the administrator to troubleshoot the problem. Through this step, a structured collection of government affairs service data is finally obtained, providing high-quality input for subsequent Q&A analysis.
[0027] Further, the government affairs service data verification script described in step S13 is used to perform the government affairs service text compliance verification and screening operation and the government affairs service text redundancy screening operation.
[0028] In the embodiment of the present invention, the government affairs service text compliance verification and screening operation checks the legality of the data format and field content according to predefined rules. Taking the data of the "government affairs interpretation content interface" as an example, the compliance verification script checks the fields in the returned government affairs interpretation data. The script first verifies whether the field policy_id is of string type and meets a specific length range (for example, the length is 8 to 12 characters). Then, it verifies whether the title field is a non-empty string and does not exceed 100 characters; it verifies whether the content field contains valid content of the government affairs interpretation and the length is between 500 and 5000 characters. Finally, the script verifies whether the publish_date field conforms to the date format of YYYY-MM-DD. The purpose of the government affairs service text redundancy screening operation is to remove duplicate data to improve data quality and avoid unnecessary duplicate storage. This process calculates the hash value of the text content and compares whether the newly collected data is the same as the data already stored in the database. If the hash values are the same, it is regarded as redundant data. Use the MD5 algorithm to generate the hash value of the data content. For example, use the hashlib library to calculate the MD5 value of the field content. The data already stored in the database will be compared with the hash value of the new data. If there is already data with the same hash value, the new data will not be stored.
[0029] Further, as an embodiment of the present invention, refer to Figure 2 shown, for Figure 1 the detailed step flow diagram of step S2 in Step S21: Perform bidirectional word segmentation and parsing on government service data to generate government service word segmentation and parsing data; In the embodiment of the present invention, the text of government service data is first parsed through bidirectional word segmentation technology. Open-source word segmentation tools such as Jieba or HanLP are used. First, Chinese word segmentation is performed to split the input text (such as "The government affairs execution situation has been released to the public") into a sequence of words. Then, bidirectional word segmentation processing is carried out, that is, word recognition is performed simultaneously from left to right and from right to left to ensure that the semantic connections between words can be captured during the processing. During this process, for specific named entities, such as "government affairs execution situation", proper noun recognition is performed to ensure the accuracy of the word segmentation results. Finally, the word segmentation results are converted into a sequence of words (such as "government affairs", "execution", "situation", "has", "to", "the public", "released") and stored as structured data, such as JSON format, recording the start position and end position of each word. The generated word segmentation and parsing data serves as the basis for subsequent steps.
[0030] Step S22: Perform multi-scale word segmentation window analysis based on the government service word segmentation and parsing data to generate multi-scale word segmentation window data; In the embodiment of the present invention, multi-scale word segmentation window analysis aims to capture semantic hierarchical relationships through lexical aggregation analysis at different scales (window sizes). In implementation, using the word segmentation and parsing data, different-sized analysis windows are first set, and typical window sizes are 1 word, 2 words, 3 words, etc. Through the method of sliding windows, window word sequences of different scales are generated. For example, for the word segmentation result "The government affairs execution situation has been released to the public", window data of scales 1, 2, and 3 are generated respectively: 1-word window: ["government affairs", "execution", "situation", "has", "to", "the public", "released"]; 2-word window: ["government affairs execution", "execution situation", "situation has", "has to", "to the public", "the public released"]; 3-word window: ["government affairs execution situation", "execution situation has", "situation has to", "has to the public", "to the public released"]. Through this analysis, the potential semantics of words in different contexts can be understood more deeply. The generated window data is stored in a structured format with multiple fields for subsequent processing.
[0031] Step S23: Use the multi-scale word segmentation window data to perform window word segmentation part-of-speech identification analysis on the government service word segmentation and parsing data to generate government service word segmentation part-of-speech data; In the embodiments of the present invention, when performing part-of-speech tagging analysis, based on multi-scale word segmentation window data, a part-of-speech tagging tool (such as Stanza or HanLP) is used to tag the part of speech for each word. In the specific operation process, by inputting the word segmentation result, the tool is used for part-of-speech tagging to determine the grammatical role of each word (such as noun, verb, adjective, etc.). For example, for the window "government affairs execution situation", the part-of-speech tagging result is: "government affairs (noun), execution (verb), situation (noun)". This process includes: parsing the word segmentation result; using the existing part-of-speech tagging model to tag each word; recording the tagging result in the word segmentation parsing data, and generating part-of-speech tagging data. The tagging data includes the part-of-speech label, part-of-speech category (such as verb, noun) and position identifier of each word, which is convenient for subsequent analysis.
[0032] Step S24: Perform government affairs service part-of-speech entity relationship analysis based on the government affairs service word segmentation part-of-speech data to generate government affairs service part-of-speech entity relationship data; In the embodiments of the present invention, when performing part-of-speech entity relationship analysis, through the already labeled part-of-speech data, the entity relationship between words is identified and analyzed, and the dependency syntax analysis method (such as Spacy or StanfordNLP) is used to perform syntactic dependency analysis on the word segmentation part-of-speech data. For example, for the sentence "The government affairs execution situation has been released to the public", first, "government affairs" and "execution" are identified as the subject-predicate relationship, and the relationship graph is constructed: "government affairs (subject) to execution (predicate)". Through dependency relationship analysis, the relationship between "government affairs" and "execution" is generated, and combined with other part-of-speech tags, the relationship graph of all related words is constructed. Finally, each entity and its relationship are stored as triple data in the form of "entity 1-relationship-entity 2" (such as "government affairs-execution-situation").
[0033] Step S25: Perform government affairs service semantic backbone triple analysis based on the government affairs service part-of-speech entity relationship data to generate government affairs service semantic backbone triple data; In the embodiments of the present invention, according to the part-of-speech entity relationship data, semantic backbone triple analysis is performed to extract the core semantic information in the text. In the specific implementation, NLP tools (such as BERT or RoBERTa) are applied for in-depth semantic understanding. By further processing the entity relationship data, the key triples in the sentence are identified. For example, for the sentence "The government affairs execution situation has been released to the public", the analysis process can identify that "government affairs" is the subject, "execution" is the action, and "situation" is the object, and finally generate the triple data "government affairs-execution-situation". In addition, the semantic structure of the triple is optimized through a deep learning model to ensure the accuracy of the analysis result. The generated triple data is stored in a structured database table for subsequent query and analysis.
[0034] Step S26: Perform association analysis on government service Q&A based on government service data to generate government service Q&A association data; In the embodiment of the present invention, the Q&A in government service data is subjected to association analysis, and the correlation between questions and answers is identified through statistical methods or machine learning techniques. By analyzing historical Q&A data, effective association rules can be extracted from the text. For example, the question "What is government affairs interpretation?" is associated with the answer "Government affairs interpretation refers to the detailed explanation of government affairs content". Classify and cluster historical Q&A data, and then use pattern recognition methods to generate association information based on historical answers and corresponding similar questions and similar answers, and store and mark them in a data association manner.
[0035] Step S27: Design a combined matrix for government service Q&A association on the government service semantic backbone triple data through the government service Q&A association data to generate a government service Q&A combined matrix.
[0036] In the embodiment of the present invention, a combined matrix for government service Q&A association is designed on the government service semantic backbone triple data through the government service Q&A association data to generate a government service Q&A combined matrix. Specifically, when implementing, first extract the relationship data between the user's historical questions and historical answers from the government service Q&A association data. These data usually include the keywords, intentions of the user's questions, and the answer content provided by the system. Then, by analyzing the semantic backbone triple data, extract representative semantic backbone information, such as the relationship between the theme, verb, and object. Then, based on the above government service Q&A association data and semantic backbone triple data, perform combined matrix design. Specifically, when designing, each row represents the key triple of a Q&A scenario, and each column represents different Q&A topics or intentions. By comparing the semantic similarity and correlation between different Q&A scenarios, the system will generate a Q&A combined matrix, where each element represents the association strength or matching degree between different Q&A scenarios. Through the design of this combined matrix, the system can identify which Q&A contents are relevant, thereby effectively improving the accuracy and intelligence level of Q&A matching.
[0037] Further, step S25 includes the following steps: Step S251: Establish a government service syntactic dependency tree based on the preset BERT algorithm and government service word segmentation part-of-speech data to generate a government service syntactic dependency tree model; Step S252: Perform optimization processing on the dependency types of the tree model of the government service syntactic dependency tree model to generate an optimized government service syntactic dependency tree model; Step S253: Perform government service syntactic dependency trunk feature analysis on the government service part-of-speech entity relationship data using the optimized government service syntactic dependency tree model to generate government service syntactic dependency trunk feature data; Step S254: Perform government service semantic trunk triple analysis based on the government service syntactic dependency trunk feature data to generate government service semantic trunk triple data.
[0038] In the embodiments of the present invention, based on the preset BERT algorithm (Bidirectional Encoder Representations from Transformers) and the part-of-speech data of government service word segmentation, a syntactic dependency relationship tree of government services is established. The part-of-speech data of government service word segmentation is input into the BERT model. The BERT model adopts a bidirectional Transformer structure and can capture the deep semantics of words through context information. After encoding each word, using a dependency parsing algorithm (such as the dependency parser in StanfordNLP), combined with the part-of-speech tags of each word, a dependency relationship tree is constructed. The grammatical structure of each sentence extracts the dependency relationship between words through the BERT model and represents it as a tree structure. For example, in the sentence "The implementation of government affairs has been released to the public", the subject-predicate relationship is "government affairs (subject) to implementation (verb)", thus generating a complete syntactic dependency relationship tree model. The nodes of the tree represent words, and the edges represent the dependency relationships between words. The generated dependency relationship tree model is saved as structured data, including the parent node, child node, and relationship type of each word. Further optimize the dependency types of the syntactic dependency tree model. According to the characteristics of government service data, define and optimize the types and weights of dependency relationships. The optimization process includes: First, analyze whether each dependency type in the dependency tree (such as "subject-predicate relationship", "verb-object relationship") conforms to the structural characteristics of government service data. For example, interrogative sentences in government affairs texts have specific dependency types, such as the "question-action-answer" structure. Then, use a rule-based method to adjust the dependency types in the tree, especially for complex syntactic structures, by adjusting the dependency types to better reflect semantic relationships. For example, set the dependency type of the fixed phrase "government affairs interpretation" to the "fixed collocation" type, rather than the simple "verb-object relationship". At the same time, by re-weighting the dependency paths of the words in the tree, key information nodes such as "government affairs" or "implementation" have higher weights in the tree model. The optimized dependency relationship tree model is output in a standardized format (such as JSON or XML) to ensure that the grammatical structure information can be accurately obtained and utilized in subsequent processing. Use the optimized syntactic dependency tree model to perform syntactic dependency backbone feature analysis on the part-of-speech entity relationship data of government services. Identify the backbone features in the syntactic dependency tree, that is, the core grammatical structure part of the sentence, usually including the subject-predicate-object relationship and key information points. By analyzing the optimized syntactic dependency tree, extract the backbone feature data. For example, for the sentence "The implementation of government affairs has been released to the public", the backbone relationship in the tree model is "government affairs - implementation - situation". When analyzing, select the nodes and edges most closely related to the syntactic structure and use them as the backbone features to extract. For complex government service problems with nested clauses or modifying components, the optimized tree model can help accurately identify the backbone part.For example, in the sentence "The implementation status of government affairs has been released to the public", the main features are "government affairs - implementation - status", while "has been" is an additional component. The generated syntactic dependency main feature data is stored in the form of triples, and each triple includes two entities and the relationship between them. The semantic main triple analysis is carried out using the syntactic dependency main feature data of government affairs services. By extracting the main part of the syntactic dependency relationship tree (i.e., the most important grammatical relationship), it is transformed into structured triple data. Natural language processing techniques (such as BERT, Word2Vec, or rule - based semantic analysis methods) are used to perform semantic analysis on the extracted main features to ensure the integrity and accuracy of semantic information. For example, for the sentence "The implementation status of government affairs has been released to the public", the obtained triple is: "government affairs - implementation - status". For a sentence containing multiple verbs or predicates, such as "The progress of government affairs implementation has been released on the website", context analysis is required to ensure the generation of correct triples, such as: "government affairs - implementation - progress" and "progress - release - website". On this basis, the optimization algorithm further confirms the correctness and semantic level of the triples through semantic relevance analysis, and sorts or annotates them to improve the response quality of the question - answering system. The generated semantic main triple data serves as the basis for subsequent question - answering analysis, enabling the system to more accurately understand user questions and provide correct answers.
[0039] Further, as an embodiment of the present invention, refer to Figure 3 shown in Figure 1 the detailed step - by - step schematic diagram of step S3. In this embodiment, step S3 includes the following steps: Step S31: Use the LDA topic model to analyze the topic distribution probability of the question - answering combination of the government affairs service question - answering combination matrix, and generate the government affairs service question - answering combination topic distribution probability data; In the embodiments of the present invention, the Latent Dirichlet Allocation (LDA) topic model is used to analyze the topic distribution probability of the government service question-answer combination matrix. Taking the government service question-answer combination matrix as input data, the text in the matrix is first preprocessed, including stop word removal, word segmentation, and word frequency statistics. For example, a preset automated text processing engine is used to perform stop word removal on the text in the matrix, removing general vocabulary such as "how" and "handle", and retaining domain keywords. The Jieba word segmentation tool is used to perform bidirectional word segmentation on questions and answers such as "how to apply for a residence permit" and "what materials are required for handling a residence permit" to generate a word list. For example, "residence permit handling conditions" is segmented into ["residence permit", "handle", "conditions"], and the word frequencies are counted (e.g., "residence permit" appears 300 times and "handle" appears 250 times). Then, the preprocessed text data is vectorized into a bag-of-words model and input into the LDA model for topic modeling analysis. The LDA model generates a set of topic distribution probability data for each question-answer combination through iterative calculations. For example, for the question-answer combinations "how to apply for a residence permit" and "residence permit handling conditions", the LDA model of the Scikit-learn library is used, setting the number of topics K = 3 (such as "residence permit application", "household registration transfer", "certification materials"), the number of iterations is 500 times, and the learning rate is 0.01. After inputting the bag-of-words model matrix, the model generates a topic distribution probability vector for each question-answer combination. For example: the topic distribution of the question-answer combination "how to apply for a residence permit" is: topic 1 (residence permit application) 0.7, topic 2 (household registration transfer) 0.2, topic 3 (certification materials) 0.1; the topic distribution of the question-answer combination "what materials are required for household registration transfer" is: topic 1 (residence permit application) 0.3, topic 2 (household registration transfer) 0.6, topic 3 (certification materials) 0.1. The output result is stored as a 500*3 matrix, with each row corresponding to a question-answer combination and each column corresponding to the probability value of a topic. The output result is stored in matrix form, with rows representing question-answer combinations, columns representing topics, and cell values being the probability values of topics.
[0040] Step S32: Perform topic identification on the government service question-answer combination according to the topic distribution probability data of the government service question-answer combination to generate government service question-answer combination topic data; In the embodiments of the present invention, according to the topic distribution probability data generated by the LDA model, topic identification processing is performed on each question-answer combination. First, the dominant topic of each question-answer combination is determined according to the topic distribution probability. The determination basis for the dominant topic is the one with the largest probability value, and the calculation formula is: , It is the probability value of the i-th Q&A combination on the k-th topic. For example, the topic distribution of the Q&A combination "Conditions for obtaining a residence permit" is [0.7, 0.2, 0.1], and the dominant topic is Topic 1 (Application for residence permit); the topic distribution of the Q&A combination "Household registration migration process" is [0.2, 0.6, 0.2], and the dominant topic is Topic 2 (Household registration migration). For each topic, the top 5 words with the highest probability are extracted as keywords. The TF-IDF algorithm is used to calculate the relevance between words and topics, and the formula is: Among them, is the word frequency of word w in topic k, is the inverse document frequency of word w. For example, the keywords of Topic 1 (Application for residence permit) are "Residence permit" (TF-IDF = 0.8), "Application" (TF-IDF = 0.7), "Conditions" (TF-IDF = 0.6), etc. By calculating the keyword distribution of each topic, the representative keywords of each topic are further extracted. For example, the keywords of Topic 1 are "Residence permit", "Application", "Conditions". The generated Q&A combination topic data is stored in a structured table form, and each record contains the Q&A combination, the dominant topic number, and the corresponding keyword set.
[0041] Step S33: Based on the government service Q&A combination matrix and the government service Q&A combination topic distribution probability data, perform processing on the topic distribution probability identification of the subset within the cluster for Q&A combination clustering, and generate Q&A combination cluster subset topic distribution probability data; In the embodiment of the present invention, based on the Q&A combination matrix and the topic distribution probability data, Q&A combination clustering and the analysis of the topic distribution probability of the subset within the cluster are performed. The hierarchical clustering algorithm is used to perform clustering analysis on the Q&A combinations. The input data is the Q&A combination matrix and its topic distribution probability. The cosine similarity is used to calculate the similarity between Q&A combinations, and the formula is: Among them, and are the topic distribution vectors of the i-th and j-th Q&A combinations. The similarity threshold is set to 0.6, and 500 Q&A are divided into 2 clusters through agglomerative hierarchical clustering: Cluster C1 (300, the dominant topic is "Application for residence permit"), Cluster C2 (200, the dominant topic is "Household registration migration"). For the 300 Q&A combinations within Cluster C1, calculate the mean vector of the topic distribution: , the dominant theme is Theme 1 (residence permit application, 0.65), and the secondary theme is Theme 2 (household registration transfer, 0.2). Similarly, the mean vector of cluster C2 is [0.2, 0.6, 0.2], and the dominant theme is Theme 2 (household registration transfer, 0.6). A clustering tree is constructed based on similarity and a threshold is set for cluster division. For each clustering cluster, the subset topic distribution probability of the Q&A combinations within it is further calculated. The output subset topic distribution probability data within the cluster includes the cluster number, the topic probability distribution, and the dominant theme.
[0042] Step S34: Perform Q&A combination intra-cluster sub-topic analysis and identification processing based on the subset topic distribution probability data within the Q&A combination cluster to generate Q&A combination intra-cluster sub-topic data; In the embodiment of the present invention, the subset topic distribution probability data within the Q&A combination cluster is further analyzed to identify the sub-topics within each clustering cluster. First, the sub-topics with significant differences are extracted based on the subset topic distribution probability data within the cluster. The sub-topics are identified by calculating the distribution difference between topics (such as KL divergence). The formula is: , when the divergence is greater than a certain threshold, it is determined that the difference between Theme 1 and Theme 2 is significant and they need to be divided into different sub-topics. For the subsets with significant theme differences, a theme feature extraction method (such as TF-IDF) is used to extract keywords to identify the semantic content of the sub-topics. For example, for the Q&A text within cluster C1, the TextRank algorithm is used to extract sub-topic keywords. Taking the Q&A "What materials are required for applying for a residence permit" as an example, after word segmentation, a word sequence ["residence permit", "application", "submit", "materials"] is generated. Through iterative calculation (damping coefficient = 0.85, number of iterations = 10), the keyword weights are obtained: "materials" (1.3), "submit" (1.1), "residence permit" (1.0). The final output intra-cluster sub-topic data is stored in a structured form, including the cluster number, the sub-topic number, and its keyword set.
[0043] Step S35: Perform hierarchical theme relationship analysis of the Q&A combination based on the Q&A combination theme data and the Q&A combination intra-cluster sub-topic data to generate Q&A combination hierarchical theme relationship data.
[0044] In the embodiments of the present invention, a hierarchical topic relationship of a Q&A combination is constructed based on the Q&A combination topic data and the sub-topic data within a cluster. In a specific implementation, first, a hierarchical structure is constructed according to the subordinate relationship between the topic and the sub-topic. For example, for the leading topic "Residence Permit Application", its subordinate sub-topics include "Application Conditions", "Application Process", etc. When constructing the hierarchical relationship, a tree-like data structure is adopted, where each node represents a topic or a sub-topic, and the edge represents the subordinate relationship. Through relationship analysis, the association path between each Q&A combination and its corresponding topic is determined. For example, the path of "How to apply for a residence permit" is from "Residence Permit Application" to "Application Conditions". The finally generated hierarchical topic relationship data is stored in the form of a tree-like structure or a nested table, and each record contains the topic hierarchical path and the Q&A combination association information.
[0045] Further, step S33 includes the following steps: Step S331: Perform Q&A combination similarity analysis on the government service Q&A combination matrix to generate government service Q&A combination similarity data; perform government service Q&A combination clustering cluster analysis based on the government service Q&A combination similarity data to generate government service Q&A combination clustering cluster data; Step S332: Map the government service Q&A combination topic distribution probability data to the government service Q&A combination clustering cluster data for in-cluster subset topic distribution probability identification processing to generate Q&A combination in-cluster subset topic distribution probability data.
[0046] In the embodiments of the present invention, when performing Q&A combination similarity analysis on the government service Q&A combination matrix, the Q&A combination matrix is standardized to ensure that the data in different dimensions of the matrix have the same dimension. Then, the cosine similarity or Pearson correlation coefficient is used to calculate the similarity between each two Q&A combinations in the matrix. The cosine similarity is used to calculate the semantic correlation between Q&A combinations, and the formula is: Taking Q&A combination A "How to apply for a passport" (vector 1, 0.3, 0.1) and Q&A combination B "What materials are required for passport application" (vector 0.9, 0.4, 0.2) as an example, the calculation shows It shows that the semantics of the two are highly similar. Finally, the similarity values between all Q&A combinations are stored as a symmetric matrix to generate the similarity data of government service Q&A combinations. This data is used for subsequent cluster analysis. Based on the generated Q&A combination similarity data, hierarchical clustering (Hierarchical Clustering) or K-means clustering algorithm is used to divide the Q&A combinations into clusters. For hierarchical clustering, a bottom-up merging method is used. According to the similarity values in the similarity matrix, similar Q&A combinations are gradually merged until the preset cluster number target is reached. Using the agglomerative hierarchical clustering algorithm, with the cosine similarity matrix as the input, the merging threshold is set to 0.85. Initially, each Q&A is an independent cluster, and the clusters with similarity greater than the threshold are gradually merged. For example: Q&A A (0.95) is merged with Q&A B to form cluster C1; Q&A C "Renewal process of Hong Kong and Macau Pass" (vector 0.2, 0.8, 0.1) and Q&A D "How to renew the Hong Kong and Macau Pass" (vector 0.3, 0.7, 0.1) are merged into cluster C2 with a similarity of 0.92. Finally, cluster data is generated, including cluster numbers, member lists, and inter-cluster distances. Using the topic distribution probability data of government service Q&A combinations, a mapping operation is performed with the Q&A combination cluster data. For the Q&A combinations in each cluster, their corresponding topic distribution probabilities are extracted, and the topic distribution probabilities of all Q&A combinations within the cluster are weighted and averaged to generate the overall topic distribution probability within the cluster. For example, the topic distribution probabilities of each Q&A in cluster C1 are weighted and averaged, and the weight is the similarity ranking of the Q&A within the cluster (the higher the ranking, the higher the weight, range 0.1 - 1.0). The calculation formula is: . Among them, is the weight of the i-th Q&A (such as the weight of Q&A A is 1.0, and the weight of Q&A B is 0.9), is the probability value of the i-th Q&A in topic k. Suppose the average topic distribution of cluster C1 is: μ = [0.75 (topic 1), 0.2 (topic 2), 0.05 (topic 3). For cluster 1, it contains Q&A combinations A and B, and the topic distribution probabilities are topic 1 (0.7), topic 2 (0.3) and topic 1 (0.6), topic 2 (0.4) respectively. According to the overall topic distribution probability within the cluster, the subset topic distribution characteristics are further analyzed. By setting a topic contribution threshold (such as 0.5), a subset of Q&A combinations with significant topic contributions within the cluster is screened out. For example, for cluster 1, if the distribution probability of topic 1 exceeds 0.5, then this topic is identified as the dominant topic, and all Q&A combinations with a topic 1 proportion higher than 0.5 are classified into the topic 1 subset. At the same time, the topic distribution probabilities within the subset are statistically analyzed to generate subset-level topic distribution probability data. The finally generated Q&A combination cluster subset topic distribution probability data includes cluster numbers, subset numbers, topic numbers, and their probability distributions. For example, cluster 1 contains a topic 1 subset, and its topic distribution probability is topic 1 (0.8), topic 2 (0.2).
[0047] Further, step S34 includes the following steps: Step S341: Analyze the subset topic feature within the Q&A combination cluster for the subset topic distribution probability data of the Q&A combination cluster to generate subset topic feature data within the Q&A combination cluster; Step S342: Perform subset topic reconstruction processing within the Q&A combination cluster based on the subset topic feature data within the Q&A combination cluster to generate subset topic reconstruction data within the Q&A combination cluster; Step S343: Perform sub-topic identification processing within the Q&A combination cluster based on the subset topic reconstruction data within the Q&A combination cluster to generate sub-topic data within the Q&A combination cluster.
[0048] In the embodiment of the present invention, when analyzing the subset topic feature within the Q&A combination cluster for the subset topic distribution probability data of the Q&A combination cluster, the subset topic distribution probability data of the Q&A combination cluster is input, including the topic number and its distribution probability of each subset. For example, for subset 1, the topic distribution probability is topic 1 (0.6), topic 2 (0.3), and topic 3 (0.1). Statistical analysis is performed on the subset topic distribution probability data of each subset, including the calculation of the importance weight of the topic and the evaluation of the collaborative relationship between topics. The importance weight is obtained through normalization processing, and the collaborative relationship is calculated through a correlation matrix. The mutual information (MI) is used to calculate the relevance between topic 1 and topic 2, and the formula is: , where x = 1 indicates that the Q&A contains topic 1, and y = 1 indicates that it contains topic 2. It is statistically obtained that: there are 30 Q&As that contain both topic 1 and 2, p(1,1) = 30 / 120 = 0.25; there are 60 Q&As that only contain topic 1, p(1,0 = 60 / 120 = 0.5; there are 24 Q&As that only contain topic 2, p(0,1) = 24 / 120 = 0.2; and there are 6 Q&As that contain neither, p(0,0) = 6 / 120 = 0.05. It is calculated that It is -0.34, and the mutual information is negative, indicating that the correlation between Topic 1 and 2 is weak and needs to be processed independently. The subset topic feature data within the generated Q&A combination cluster includes structured information such as subset number, dominant topic, secondary topic, and topic collaboration relationship. Input the subset topic feature data within the Q&A combination cluster and perform feature analysis on the dominant topic and secondary topic of each subset. For example, the dominant topic of subset 1 is Topic 1, and the secondary topic is Topic 2. Optimize and reconstruct the Q&A combinations within the subset according to the weight distribution of the dominant topic and secondary topic. The reconstruction operations include reallocating the topic weights of the Q&A combinations and adjusting the topic distribution. Set the dominant topic weight increase coefficient α = 1.3 and the secondary topic decay coefficient β = 0.7. Then, the changed Topic 1 is 0.78, Topic 2 is 0.21, and Topic 3 remains unchanged. After normalization again, the focus on the dominant topic is enhanced. Introduce a feature optimization model, such as the Kullback-Leibler divergence, during the reconstruction process to calculate the change in the topic distribution before and after reconstruction and adjust the optimization strategy until the optimal balance state is reached. Finally, output the reconstructed topic distribution data. The subset topic reconstruction data within the generated Q&A combination cluster includes subset number, reconstructed topic distribution probability, and adjusted weights. For example, the reconstruction data of subset 1 is Topic 1 (0.72), Topic 2 (0.19), and Topic 3 (0.09). Input the subset topic reconstruction data within the Q&A combination cluster. For each subset, extract the reconstructed topic distribution data and identify the dominant topic according to the preset topic identification threshold (such as 0.6). Identify the dominant topic after reconstruction according to the topic distribution probability in the reconstruction data. For example, for subset 1, the weight of Topic 1 is 0.72, which exceeds the threshold of 0.6. Therefore, identify Topic 1 as the dominant sub-topic of this subset. In addition, calculate the feature contribution of the secondary sub-topic. For example, the weight of Topic 2 is 0.19 and is identified as the secondary sub-topic. If the weight of a newly added topic is less than the secondary sub-topic threshold (such as 0.1), it is marked as the background topic and excluded from subsequent analysis. The sub-topic data within the generated Q&A combination cluster includes subset number, dominant sub-topic number, secondary sub-topic number, and background topic number. For example, the sub-topic data of subset 1 is dominant sub-topic 1, secondary sub-topic 2, and background topic 3.
[0049] Further, step S4 includes the following steps: Step S41: Establish an intelligent feedback mapping relationship for government service Q&A through the Q&A combination hierarchical topic relationship data to obtain a government service intelligent Q&A model; Step S42: Receive instant government service user question data; Step S43: Transmit the instant government service user question data to the government service intelligent Q&A model for intelligent feedback processing of multiple answers for government services, including: performing matching processing of the user question data and the Q&A combination level theme relationship data on the instant government service user question data and the Q&A combination level theme relationship data through the government service intelligent Q&A model to generate government service multiple matching answer data; performing confidence analysis on the government service multiple matching answer data to generate government service multiple matching answer confidence data; using the government service multiple matching answer confidence data as a government service multiple answer sorting mechanism, and performing intelligent sorting and confidence output on the government service multiple matching answer data through the government service multiple answer sorting mechanism to obtain government service multiple answer intelligent sorting data; Step S44: Transmit the government service multiple answer intelligent sorting data to the terminal to execute the government service intelligent Q&A feedback operation.
[0050] In the embodiment of the present invention, the Q&A combination level theme relationship data is used as input, and a graph neural network (GNN) model is used to construct an intelligent feedback mapping relationship. First, the Q&A combination level theme relationship data is mapped into a graph structure, where nodes represent themes, edges represent the hierarchical relationships between themes, and the edge weights are determined by the theme relationship strength (for example, the edge weight between theme A and theme B is 0.8).
[0051] Feature extraction and relationship modeling of the graph structure are performed through the GNN model. The GraphSAGE algorithm is used to construct an intelligent Q&A model. The input node features are TF-IDF vectors of topic keywords (e.g., the vector of T1 is [0.8, 0.3, 0.1]), and the edge features are weight values. The training process is as follows: Initialize the dimension of the node embedding vector to 128; Use two layers of aggregators (Aggregation Layer), the first layer aggregates the features of neighbor nodes, and the second layer generates the final embedding; The objective function is the cross-entropy loss, the optimizer is Adam, the learning rate is 0.001, and the number of training epochs is 200; Input the labeled Q&A pairs (e.g., the question "What materials are needed for applying for a residence permit" corresponds to the answer path T1→T1-1), and the model learns the mapping relationship between the question vector and the topic nodes. After training, an intelligent Q&A model for government services is output. This model can receive the user's question vector and output the association probability values with each topic node. For example, when inputting the question "Application conditions for a residence permit", the model outputs the association probability with T1-1 as 0.95 and the probability with T1-2 as 0.1. Receive the government service question data input by the user in real time through the API interface, and the input format is a JSON object. The system uses regular expressions to perform format verification on the input data to ensure the integrity of fields and the validity of characters. After passing the verification, extract the content of the question text and perform normalization processing, including removing extra spaces, converting to lowercase letters, and removing stop words. The processed question text is stored in vector form, for example, representing the question as a vector through TF-IDF. The instant question data is matched with the hierarchical topic relationship data of the Q&A combination in the intelligent Q&A model. Adopt the sentence embedding method based on BERT to vectorize the user's question and the hierarchical topic question, and calculate the cosine similarity. For example, the similarity between the user's question vector and the vector of topic A is 0.98, and the similarity with the vector of topic B is 0.76. The topics with similarity higher than the preset threshold (e.g., 0.7) are used as the preliminary matching results to generate multi-matching answer data for government services. Normalize the similarity scores of the multi-matching answer data, for example, use the Softmax function to convert the similarity into a confidence distribution. The matching result is that the confidence of topic A is 0.85 and the confidence of topic B is 0.15. The generated multi-matching answer confidence data contains the matching answers and their corresponding confidences, for example, {"Topic A": 0.85, "Topic B": 0.15}. Build a sorting mechanism based on the confidence data. Adopt the descending sorting strategy and give priority to outputting the answers with high confidence. For example, the sorting result is that topic A is greater than topic B. Add confidence annotations to the sorted answers to generate multi-answer intelligent sorting data for government services. Transmit the multi-answer intelligent sorting data to the user terminal in the standard JSON format through the HTTP protocol. The output data includes the sorting result and the confidence. After the user terminal receives it, it calls the display module to generate a user-friendly feedback result, for example, displaying the answer content and the confidence in the form of a card.
[0052] This specification provides a machine learning intelligent Q&A system based on government services, which is used to execute the machine learning intelligent Q&A analysis method based on government services as described above. The machine learning intelligent Q&A system based on government services includes: A data collection module, which is used to deploy a multi-channel government service data collection engine, collect government service data based on the multi-channel government service data collection engine to obtain government service data, where the government service data includes historical government service question data, historical government service answer data, and government service knowledge base data; A Q&A combination standardization module, which is used to perform bidirectional word segmentation and parsing processing on government service data to generate government service word segmentation and parsing data; perform government service semantic backbone triple analysis based on the government service word segmentation and parsing data to generate government service semantic backbone triple data; design a combination matrix for government service Q&A association for the government service semantic backbone triple data to generate a government service Q&A combination matrix; A Q&A combination theme identification module, which is used to perform an analysis of the theme distribution probability of Q&A combinations on the government service Q&A combination matrix to generate government service Q&A combination theme distribution probability data; perform an analysis of the hierarchical theme relationship of Q&A combinations based on the government service Q&A combination theme distribution probability data to generate Q&A combination hierarchical theme relationship data; An intelligent Q&A feedback module, which is used to establish an intelligent feedback mapping relationship for government service Q&A through the Q&A combination hierarchical theme relationship data to obtain a government service intelligent Q&A model; receive instant government service user question data; transmit the instant government service user question data to the government service intelligent Q&A model for multi-answer intelligent feedback processing of government services to obtain government service multi-answer intelligent sorting data; transmit the government service multi-answer intelligent sorting data to the terminal to execute the government service intelligent Q&A feedback operation.
[0053] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be included in the present invention.
[0054] The above description is only a specific implementation manner of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.< / token>
Claims
1. A machine learning intelligent question answering analysis method based on government affairs services, characterized in that It includes the following steps: Step S1: Deploy a multi-channel government service data collection engine, and collect government service data based on the multi-channel government service data collection engine to obtain government service data, where the government service data includes historical government service question data, historical government service answer data, and government service knowledge base data; Step S2: Perform bidirectional word segmentation and parsing processing on the government service data to generate government service word segmentation and parsing data; perform government service semantic backbone triple analysis based on the government service word segmentation and parsing data to generate government service semantic backbone triple data; Design a combination matrix for government service question and answer association for the government service semantic backbone triple data to generate a government service question and answer combination matrix; Step S3: Perform a topic distribution probability analysis of the question and answer combinations on the government service question and answer combination matrix to generate government service question and answer combination topic distribution probability data; perform a hierarchical topic relationship analysis of the question and answer combinations based on the government service question and answer combination topic distribution probability data to generate question and answer combination hierarchical topic relationship data; Step S4: Establish an intelligent feedback mapping relationship for government service questions and answers through the question and answer combination hierarchical topic relationship data to obtain a government service intelligent question and answer model; Receive instant government service user question data; Transmit the instant government service user question data to the government service intelligent question and answer model for intelligent feedback processing of multiple answers to government services to obtain government service multiple answer intelligent sorting data; Transmit the government service multiple answer intelligent sorting data to the terminal to execute the government service intelligent question and answer feedback operation.
2. The machine learning intelligent question answering analysis method based on government affairs services according to claim 1, wherein Step S1 includes the following steps: Step S11: Obtain a government service authorization API interface; Step S12: Perform interface source address identification processing on the government service authorization API interface to obtain an identified government service authorization API interface; Step S13: Design a government service data verification script; Step S14: Deploy a multi-channel government service data collection engine according to the identified government service authorization API interface and the government service data verification script; Step S15: Collect government service data based on the multi-channel government service data collection engine to obtain government service data.
3. The machine learning intelligent question and answer analysis method based on government affairs services according to claim 2, wherein, The government service data verification script in Step S13 is used to perform government service text compliance verification and screening operations and government service text redundancy screening operations.
4. The machine learning intelligent question-answering analysis method based on government affairs services according to claim 1, wherein Step S2 includes the following steps: Step S21: Perform bidirectional word segmentation and parsing processing on the government service data to generate government service word segmentation and parsing data; Step S22: Perform multi-scale word segmentation window analysis based on the government service word segmentation and parsing data to generate multi-scale word segmentation window data; Step S23: Use the multi-scale word segmentation window data to perform window word segmentation part-of-speech identification analysis on the government service word segmentation and parsing data to generate government service word segmentation part-of-speech data; Step S24: Perform government service part-of-speech entity relationship analysis based on the government service word segmentation part-of-speech data to generate government service part-of-speech entity relationship data; Step S25: Perform government service semantic backbone triple analysis based on the government service part-of-speech entity relationship data to generate government service semantic backbone triple data; Step S26: Conduct government service question-answer association analysis based on government service data to generate government service question-answer association data; Step S27: Design a combined matrix for government service question-answer association for the government service semantic backbone triple data through the government service question-answer association data to generate a government service question-answer combined matrix.
5. The machine learning intelligent question and answer analysis method based on government affairs services according to claim 4, wherein Step S25 includes the following steps: Step S251: Establish a government service syntactic dependency tree based on the preset BERT algorithm and government service word segmentation part-of-speech data to generate a government service syntactic dependency tree model; Step S252: Perform tree model dependency type optimization processing on the government service syntactic dependency tree model to generate an optimized government service syntactic dependency tree model; Step S253: Use the optimized government service syntactic dependency tree model to conduct government service syntactic dependency backbone feature analysis on the government service part-of-speech entity relationship data to generate government service syntactic dependency backbone feature data; Step S254: Conduct government service semantic backbone triple analysis based on the government service syntactic dependency backbone feature data to generate government service semantic backbone triple data.
6. The machine learning intelligent question and answer analysis method based on government affairs services according to claim 1, characterized in that Step S3 includes the following steps: Step S31: Use the LDA topic model to conduct topic distribution probability analysis of question-answer combinations for the government service question-answer combined matrix to generate government service question-answer combined topic distribution probability data; Step S32: Conduct question-answer combination topic identification based on the government service question-answer combined topic distribution probability data to generate question-answer combination topic data; Step S33: Perform intra-cluster subset topic distribution probability identification processing for question-answer combination clustering based on the government service question-answer combined matrix and the government service question-answer combined topic distribution probability data to generate intra-cluster subset topic distribution probability data for question-answer combinations; Step S34: Conduct intra-cluster sub-topic analysis and identification processing based on the intra-cluster subset topic distribution probability data for question-answer combinations to generate intra-cluster sub-topic data for question-answer combinations; Step S35: Conduct hierarchical topic relationship analysis for question-answer combinations based on the question-answer combination topic data and the intra-cluster sub-topic data for question-answer combinations to generate hierarchical topic relationship data for question-answer combinations.
7. The machine learning intelligent question and answer analysis method based on government affairs services according to claim 6, wherein Step S33 includes the following steps: Step S331: Conduct question-answer combination similarity analysis for the government service question-answer combined matrix to generate government service question-answer combined similarity data; conduct government service question-answer combination clustering cluster analysis based on the government service question-answer combined similarity data to generate government service question-answer combination clustering cluster data; Step S332: Map the government service question-answer combined topic distribution probability data to the government service question-answer combination clustering cluster data for intra-cluster subset topic distribution probability identification processing to generate intra-cluster subset topic distribution probability data for question-answer combinations.
8. The machine learning intelligent question and answer analysis method based on government affairs services according to claim 6, characterized in that, Step S34 includes the following steps: Step S341: Conduct intra-cluster subset topic feature analysis for the intra-cluster subset topic distribution probability data for question-answer combinations to generate intra-cluster subset topic feature data for question-answer combinations; Step S342: Conduct intra-cluster subset topic reconstruction processing based on the intra-cluster subset topic feature data for question-answer combinations to generate intra-cluster subset topic reconstruction data for question-answer combinations; Step S343: Process the identification of sub - topics within the Q&A combination cluster based on reconstructing data for sub - topics within the Q&A combination cluster subset theme, and generate sub - topic data within the Q&A combination cluster.
9. The machine learning intelligent question-answering analysis method based on government affairs services according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Establish an intelligent feedback mapping relationship for government service Q&A through Q&A combination hierarchical theme relationship data to obtain a government service intelligent Q&A model; Step S42: Receive instant government service user question data; Step S43: Transmit the instant government service user question data to the government service intelligent Q&A model for intelligent feedback processing of multiple answers for government services, including: performing a matching process between the instant government service user question data and the Q&A combination hierarchical theme relationship data through the government service intelligent Q&A model to generate government service multiple - matching answer data; analyzing the confidence level of the government service multiple - matching answer data to generate government service multiple - matching answer confidence level data; using the government service multiple - matching answer confidence level data as a government service multiple - answer sorting mechanism, and performing intelligent sorting and confidence level output of the government service multiple - matching answer data through the government service multiple - answer sorting mechanism to obtain government service multiple - answer intelligent sorting data; Step S44: Transmit the government service multiple - answer intelligent sorting data to the terminal to execute the government service intelligent Q&A feedback operation.
10. A machine learning intelligent question answering system based on government affairs services, characterized in that, For executing the machine - learning - based intelligent Q&A analysis method for government services as described in claim 1, the machine - learning - based intelligent Q&A system for government services includes: A data collection module, used to deploy a multi - channel government service data collection engine, and collect government service data based on the multi - channel government service data collection engine to obtain government service data, where the government service data includes historical government service question data, historical government service answer data, and government service knowledge base data; A Q&A combination standardization module, used to perform two - way word - segmentation parsing processing on government service data to generate government service word - segmentation parsing data; perform government service semantic backbone triple analysis based on the government service word - segmentation parsing data to generate government service semantic backbone triple data; design a combination matrix for government service Q&A association for the government service semantic backbone triple data to generate a government service Q&A combination matrix; A Q&A combination theme identification module, used to analyze the theme distribution probability of Q&A combinations in the government service Q&A combination matrix to generate government service Q&A combination theme distribution probability data; perform hierarchical theme relationship analysis of Q&A combinations based on the government service Q&A combination theme distribution probability data to generate Q&A combination hierarchical theme relationship data; An intelligent Q&A feedback module is used to establish an intelligent feedback mapping relationship for government service Q&A through the Q&A combination hierarchical topic relationship data, so as to obtain a government service intelligent Q&A model; receive instant government service user question data; transmit the instant government service user question data to the government service intelligent Q&A model for intelligent feedback processing of multiple answers to government services, so as to obtain intelligent sorting data of multiple answers to government services; and transmit the intelligent sorting data of multiple answers to government services to the terminal to execute the intelligent Q&A feedback operation of government services.
Citation Information
Patent Citations
A method and a device for solving a demand configuration problem of a servitization information system
CN109783127A
Online medical community question and answer text clustering method based on multi-feature fusion
CN114969283A
Intelligent question and answer method based on service item elements
CN116737887A
Knowledge graph question and answer method and device, equipment and storage medium
CN117290478A
Telephone customer service processing method and system based on personalized robot
CN118433311A
Cited By
Effect matrix construction method based on topic clustering and two rounds of questions and answers
CN121256033A
A method for constructing an efficacy matrix based on topic clustering and two-round question answering
CN121256033B
Government affair intelligent question-answering method, government affair intelligent question-answering system, medium and product
CN121303150A
Government affair intelligent question and answer method, government affair intelligent question and answer system, medium and product
CN121303150B