A Fine-tuning Method, Device, Terminal Device and Medium for a Government Affairs Q&A System
By constructing a question-and-answer data set containing noise data and logical reasoning, and fine-tuning the large language model, the illusion problem when the search results in the government affairs question-and-answer system is solved, the accuracy and robustness of the model are improved, and the effectiveness of government affairs question-and-answer is ensured.
Patent Information
- Application Number
- CN202510121461.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The existing government Q&A system has shortcomings in accuracy and relevance, especially when large language models are prone to hallucinations when the search results are not correlated, providing misleading information.
By constructing the original Q&A dataset, adding noise data and logical reasoning processes, generating a third Q&A dataset, and using Lora fine-tuning method to train a large language model to improve the model's understanding of the logical relationship between policy content and user problems.
It significantly reduces the phenomenon of large language models having hallucinations when the search results are not correlated, reduces the risk of overfitting, and improves the robustness of the model in actual government affairs scenarios and the accuracy of the answers.
Smart Images

Figure CN120067682B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to natural language processing technology, and in particular to a fine-tuning method, device, terminal device and medium for a government affairs Q&A system. Background Art
[0002] With the rapid development of information technology, government affairs Q&A systems have gradually become an important tool for communication between the government and the public. Such systems aim to provide the public with answers to relevant policies, laws and regulations quickly and accurately in an automated manner. However, existing government affairs Q&A systems still face many challenges in practical applications, especially in terms of accuracy and relevance.
[0003] At present, many government affairs Q&A systems adopt the Retrieval-Augmented Generation (RAG) technology framework based on retrieval and generation. This framework retrieves relevant laws, regulations and policy clauses from a pre-constructed knowledge base, and then combines the questions raised by users, and a large language model generates answers. This method can theoretically provide relatively rich and accurate answers. However, in practical applications, there are still some technical defects.
[0004] First of all, the RAG framework depends on the accuracy of the retrieval system. Existing retrieval technologies are difficult to ensure that the specific context of the user's question can be fully matched each time, resulting in the retrieved policy clauses may not be completely relevant to the user's question. For example, when a user asks about the process of appealing against a specific building violation, the system may retrieve the regulations on real estate transaction management that are not directly related to it. Although these regulations may contain some seemingly relevant content, they are not specific regulations for the user's question. Secondly, when the large language model generates answers, it is easily misled by the retrieval results, resulting in the so-called "hallucination" phenomenon, that is, the model may generate inaccurate or even wrong answers based on incompletely matched retrieval content. For example, in the case of appealing against an illegal sunroom, the model may wrongly quote an irrelevant complaint process, misleading the user. The existence of these defects causes the government affairs Q&A system to be unable to fully meet the needs of users when providing services, and may even provide misleading information.
[0005] Therefore, there is an urgent need for a fine-tuning method, device, terminal device and medium for a government affairs Q&A system. Summary of the Invention
[0006] The present invention provides a fine-tuning method, device, terminal device and medium for a government affairs Q&A system to solve the above problems existing in the prior art.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A fine-tuning method for a government affairs Q&A system, comprising:
[0009] S101: Construct the original Q&A dataset D1 based on the questions of users and the answers of staff in the scenario;
[0010] S102: Add noise data to the original Q&A dataset D1 to obtain the second Q&A dataset D2;
[0011] S103: Add logical reasoning processes to the second Q&A dataset D2 to generate the third Q&A dataset D3;
[0012] S104: Construct a fine-tuned large language model according to the third Q&A dataset D3.
[0013] Among them, before step S102 includes:
[0014] Scrape relevant texts from government websites. The relevant texts include relevant laws, regulations, ordinances and policies, and segment the scraped relevant texts into text blocks with a length of L to obtain a policy library {Chunk(m,n)}, where Chunk(m,n) represents the nth text block of the mth policy.
[0015] Among them, step S101 includes:
[0016] S1011: Collect the questions of users and the answers of staff in the real scenario to obtain the original Q&A data;
[0017] S1012: Conduct detailed annotation and classification on each user question Question and staff answer response in the original Q&A data. The annotation content includes:
[0018] Policy(i,k) represents the kth policy name cited by the staff for the ith question; Content(i,k) represents the relevant content in the kth policy cited by the staff; Answer(i) represents the direct answer given by the staff to the ith question;
[0019] S1013: Through the annotation and classification of the original Q&A data, form a standardized original Q&A dataset D1 = {Question(i), response(i)}.
[0020] Among them, step S102 includes:
[0021] S1021: Preset a threshold P0 and set the maximum number K of noise data to be added;
[0022] S1022: Generate a random number between 0 and 1 for each question Question(i). If the random number is greater than the preset threshold P0, do not modify the Q&A data; if the random number is less than or equal to P0, perform the following steps:
[0023] Generate a random integer K0 between 0 and K, where K0 represents the number of text blocks to be added, and repeat this step K times;
[0024] S1023: Randomly select a text block from the policy library {Chunk(m,n)}, record the corresponding policy name of the text block, and then add the text block to {Policy(i,k), Content(i,k)} of the original answer Response(i).
[0025] Among them, the steps of S103 include:
[0026] S1031: Extract M questions Question(m), as well as the corresponding policy names Policy(m,k) and policy contents Content(m,k) from the second Q&A dataset D2, and write the logical reasoning process to obtain the reasoning result logic(i);
[0027] S1032: When writing the logical reasoning process, for each question question(i) and each policy policy(i,k), content(i,k) in the second Q&A dataset D2, generate logic(i) according to the following two situations:
[0028] Situation 1, if policy(i,k) and content(i,k) are noise data, then logic(i) = "The policy content has nothing to do with the user's question";
[0029] Situation 2, if policy(i,k) and content(i,k) are not noise data, then generate using a large language model. The generation of the large language model takes the constructed prompt as input to obtain the output of the large language model, and logic(i) = the output of the large language model;
[0030] Among them, the constructed prompt includes:
[0031] The known user question Question(i);
[0032] The known policy Policy(i,k), Content(i,k);
[0033] Please combine the policy and content to perform logical reasoning and output the result;
[0034] Example:
[0035] User question question(1), given policies policy(1,k), content(1,k), output logic(1),
[0036] User question question(2), given policies policy(2,k), content(2,k), output logic(2), ......
[0038] User question question(M), given policies policy(M,k), content(M,k), output: logic(M);
[0039] S1033: Based on the inference result logic(i), generate the third Q&A dataset D3, where D3 = {Question(i), response(i)}, and response(i) = {{policy(i,k), content(i,k)}, logic(i), answer(i)}.
[0040] Among them, the S104 step includes:
[0041] Construct a fine-tuning dataset according to the third Q&A dataset D3, and the specific form is:
[0042] Input = {Question(i), {Policy(i,k), Content(i,k)}}
[0043] Output = {Logic(i), Answer(i)};
[0044] Use the Lora fine-tuning method to train the large language model. Through K epochs of iterative training, adjust the parameters of the model to optimize the model's understanding of the logical relationship between user questions and policy content, and finally obtain the fine-tuned large language model.
[0045] Among them, after the S104 step includes:
[0046] Use the fine-tuned large language model to answer new government questions. The fine-tuned large language model generates the corresponding logical reasoning Logic(i) and the final answer Answer(i) according to the input Question(i) and {Policy(i,k), Content(i,k)}.
[0047] Among them, a fine-tuning device for the government affairs Q&A system includes:
[0048] An original Q&A dataset construction unit for constructing an original Q&A dataset D1 based on the questions of users and the answers of staff in a scenario;
[0049] A second Q&A dataset construction unit for adding noise data to the original Q&A dataset D1 to obtain a second Q&A dataset D2;
[0050] A third Q&A dataset construction unit for adding a logical reasoning process to the second Q&A dataset D2 to generate a third Q&A dataset D3;
[0051] A large language model set construction unit for constructing a fine-tuned large language model according to the third Q&A dataset D3.
[0052] Wherein, an electronic device includes: at least one processor and a memory, wherein:
[0053] The memory is used to store computer execution instructions;
[0054] At least one processor is used to execute the computer execution instructions stored in the memory, so that at least one processor executes a fine-tuning method of a government affairs Q&A system.
[0055] Wherein, a computer storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement a fine-tuning method of a government affairs Q&A system.
[0056] Compared with the prior art, the present invention has the following advantages:
[0057] A fine-tuning method of a government affairs Q&A system includes: constructing an original Q&A dataset D1 based on the questions of users and the answers of staff in a scenario; adding noise data to the original Q&A dataset D1 to obtain a second Q&A dataset D2; adding a logical reasoning process to the second Q&A dataset D2 to generate a third Q&A dataset D3; constructing a fine-tuned large language model according to the third Q&A dataset D3. It significantly reduces the phenomenon of the large language model hallucinating when the retrieval results are irrelevant. Incorporating logical reasoning into the training data effectively reduces the risk of overfitting in the training process of the large language model.
[0058] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention.
[0059] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Brief Description of the Drawings
[0060] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation on the present invention. In the accompanying drawings:
[0061] Figure 1 It is a flowchart of a fine-tuning method for a government affairs Q&A system in an embodiment of the present invention;
[0062] Figure 2 It is a flowchart of constructing an original Q&A dataset D1 in an embodiment of the present invention;
[0063] Figure 3 It is a structural diagram of a fine-tuning device for a government affairs Q&A system in an embodiment of the present invention. Detailed implementation manners
[0064] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0065] An embodiment of the present invention provides a fine-tuning method for a government affairs Q&A system, including:
[0066] S101: Based on the questions of users and the answers of staff in the scenario, construct an original Q&A dataset D1;
[0067] S102: Add noise data to the original Q&A dataset D1 to obtain a second Q&A dataset D2;
[0068] S103: Add a logical reasoning process to the second Q&A dataset D2 to generate a third Q&A dataset D3;
[0069] S104: Construct a fine-tuned large language model according to the third Q&A dataset D3.
[0070] The working principle of the above technical solution is as follows: Collect user questions and answers in the real scenario: Question(i): Collect the i-th question of each user, ensuring that the question is an inquiry about certain specific government affairs or policies. Response(i): Collect the answer of the staff to the i-th question. The answer of the staff includes three parts:
[0071] Policy(i,k): The k-th policy name cited by the staff in the answer.
[0072] Content(i,k): The content related to the question in the k-th policy cited.
[0073] Answer(i): The specific answer given by the staff.
[0074] Adding noise data: Add irrelevant policy content to the original Q&A dataset D1 through the following steps:
[0075] Generate a random number for each question Question(i) and determine whether to add noise data. If the random number is less than the preset threshold P0, add noise data.
[0076] Generate a random integer K0, representing the number of noise text blocks to be added.
[0077] Randomly select K0 text blocks from the policy library and add them to the corresponding answer.
[0078] For each question and its related policy content in the second Q&A dataset D2, generate a detailed logical reasoning process (logic) through inference. Use large language models such as GPT-4 to generate logical reasoning content. This step is very important because it can prevent the model from overfitting and only learning surface features. Through inference, the model can learn how to make reasonable inferences based on the relationship between policy content and questions.
[0079] Specific steps: Label samples: Extract M samples (questions and corresponding policies) and write the corresponding logical reasoning processes to form the Sample set. Generate the reasoning process: For each question-policy combination in the second Q&A dataset D2, use a large language model (such as GPT-4) to generate the reasoning process. The reasoning can include the understanding of policies, the judgment of relevance, etc. If the policy is not relevant to the question, output "The policy content has nothing to do with the user's question".
[0080] Question: "My child's household registration is in Chongqing. What conditions need to be met to enroll in a local school nearby?"
[0081] Reasoning process:
[0082] According to the content of Policy 1: "School-age children and adolescents with household registration in this city shall enroll in schools nearby without taking an entrance examination at their place of household registration."
[0083] Information provided by the user: "The child's household registration is in Chongqing and meets the school-age conditions."
[0084] Reasoning result: Combining this information, it can be inferred that the child meets the conditions and can enroll in a local school nearby.
[0085] Finally, obtain the third Q&A dataset D3, whose structure is:
[0086] D3 = {Question(i), response(i)}
[0087] Among them, response(i) includes: policy(i,k), content(i,k), logic(i) (reasoning process), and answer(i) (final answer).
[0088] Using the third Q&A dataset D3, a fine-tuning dataset is constructed, and the parameters of the large language model are adjusted through the LoRA fine-tuning technique to enable it to better understand and process policy-related Q&A tasks.
[0089] The fine-tuning process optimizes the performance of the model by iteratively updating the model parameters over K epochs.
[0090] Input and output format:
[0091] Input: {Question(i), {policy(i,k), content(i,k)}}
[0092] Output: {logic(i), answer(i)}.
[0093] The beneficial effects of the above technical solution are as follows: significantly reducing the phenomenon of hallucinations generated by the large language model when the retrieval results are irrelevant. Incorporating logical reasoning into the training data effectively reduces the risk of overfitting during the training process of the large language model. Through the introduction of noisy data and the logical reasoning process, the model is prevented from overfitting to surface features, enhancing its robustness in actual government affairs scenarios. The fine-tuned model can give more accurate and interpretable answers, especially performing excellently in complex policy questions.
[0094] In another embodiment, before step S102, it includes:
[0095] Relevant texts are crawled from government websites. The relevant texts include relevant laws, regulations, ordinances, and policies, and the crawled relevant texts are segmented into text blocks of length L to obtain a policy library {Chunk(m,n)}, where Chunk(m,n) represents the nth text block of the mth policy.
[0096] The working principle of the above technical solution is as follows: Data source selection: Select official government websites or relevant laws and regulations databases as data sources. These websites provide documents such as laws, regulations, and policies issued by national or local governments. Use web crawling tools (such as the requests library in Python, or more complex tools such as Scrapy) to download web page data from the specified website. Extract useful texts from the web pages through an HTML parser (such as BeautifulSoup). These texts contain the full text of legal documents, document summaries, or specific clauses. During parsing, irrelevant content such as advertisements, menus, and pictures on the web pages needs to be filtered out, and only the main text related to laws and regulations is retained.
[0097] Extract valid laws, regulations, or policy provisions from the crawled web pages. Usually, legal documents have a complex structure, including titles, articles, attachments, etc. When crawling, ensure that the main text part is extracted and redundant tags, annotations, etc. are removed. Clean up extra characters, such as web page encoding garbled characters and special symbols, through regular expressions or other text processing methods (such as the re module in Python) to ensure unified text format.
[0098] Split the text into text blocks of length L and determine the maximum length L of each text block. The setting of L is usually based on actual application requirements. For example, L can be 200 words, 500 words, etc. This process is to split long legal documents into manageable small pieces for subsequent storage, query, and analysis. Text splitting: Split according to the set length L. In specific operations, it will be split according to sentences, paragraphs, or natural delimiters. Usually, try to avoid splitting in the middle of a sentence to ensure the semantic integrity of the text. Boundary handling: If the length of a certain paragraph or part of the text exceeds L, the system will automatically split it into multiple text blocks to ensure that each block does not exceed the maximum length L.
[0099] Policy library definition: Organize the crawled laws, regulations, ordinances, and policy documents by blocks to form a policy library. Each policy document will be split into multiple text blocks, and a unique identifier will be assigned to each text block. Storage format: The storage format of the policy library can be a database, JSON file, CSV file, etc. Each policy text block will contain the following information:
[0100] Chunk(m,n): Among them, m represents the mth policy, and n represents the nth text block of that policy. Each Chunk(m,n) identifier uniquely corresponds to a text block, and specific text content can be quickly queried through this identifier.
[0101] The beneficial effects of the above technical solution are as follows: Splitting large legal and regulatory documents into multiple small text blocks can make retrieval and query more efficient. After block-level splitting of policy texts, the system can more easily perform intelligent analysis and processing, such as natural language processing (NLP) tasks (e.g., entity recognition, sentiment analysis, etc.). The independence and small size of each text block help reduce the computational burden on the model and improve processing efficiency. By storing each policy text in small chunks, the memory consumption during single-document processing and storage can be effectively reduced. In addition, the small text block storage format also facilitates subsequent version control, updating, or merging operations. The structure of small text blocks makes the update and revision of policy documents more convenient. When modifying a specific article, only the corresponding text block needs to be updated, rather than reorganizing the entire document, improving the maintainability of the document. The split policy library can not only support traditional full-text retrieval but also perform more fine-grained analysis through the block-level structure. Users can perform multi-dimensional analysis operations such as statistical analysis, content association, and article comparison at the block level. Small text blocks are convenient for machine learning models to process. For legal texts, splitting into smaller blocks helps perform more accurate automated tasks such as classification and prediction, e.g., automatic classification and named entity recognition (NER) through deep learning algorithms.
[0102] In another embodiment, step S101 includes:
[0103] S1011: Collect user questions and staff answers in real scenarios to obtain original Q&A data;
[0104] S1012: Perform detailed annotation and classification on each user question Question and staff answer response in the original Q&A data. The annotation content includes:
[0105] Policy(i,k) represents the kth policy name cited by the staff for the ith question; Content(i,k) represents the relevant content in the kth policy cited by the staff; Answer(i) represents the direct answer given by the staff to the ith question;
[0106] S1013: Through the annotation and classification of the original Q&A data, form a standardized original Q&A data set D1 = {Question(i), response(i)}.
[0107] The working principle of the above technical solution is: collect questions raised by users in actual scenarios and answers given by staff to obtain original question and answer data. Data source selection: First, choose a suitable scenario or channel, such as a customer service platform, a government consultation window, a legal consultation platform, etc., where users often ask staff questions related to policies and regulations. Data collection: Collect data on user questions and staff answers through conversation records, chat records or manual questionnaires. In actual operation, question and answer pairs can be recorded using automated tools or manual methods. Question formatting: Perform preliminary formatting and cleaning of the collected raw data to ensure the neatness of the data. For example, remove irrelevant questions, simplify repeated questions, etc.
[0108] Each user question (Question) and staff response (Response) in the raw Q&A data is carefully annotated and categorized to provide structured information for subsequent dataset construction. Annotating Question(i): The user's i-th question (Question(i)) is the first part of the question-answer pair and the subject of the staff response. Questions typically involve specific areas (such as law, administration, and policy) that require understanding and categorization.
[0109] Annotating Response(i): The staff member's response (Response(i)) consists of three parts: Policy(i,k): The name of the kth policy cited by the staff member for the i-th question (Policy(i,k)). For example, if a user asks "How to apply for a subsidy," the staff member will cite policies or regulations related to subsidies. Annotation involves extracting the specific policy document name cited in the staff member's response. Content(i,k): The relevant content from the kth policy cited in the staff member's response (Content(i,k)). For example, the staff member might cite a specific article or paragraph in the policy regarding subsidy application conditions. Annotation involves extracting and storing this specific content. Answer(i): The staff member's direct response to the i-th question (Answer(i)). This is typically a concise response provided by the staff member based on the question content and relevant policies. Annotation Tools and Methods: Annotation can be performed using manual annotation or semi-automated tools (such as named entity recognition and text classification in NLP). Manual annotation is typically performed by a team of experts, while semi-automated methods leverage existing semantic understanding models to help accelerate the annotation process.
[0110] By labeling and classifying the original question-answering data, a standardized original question-answering dataset D1 is formed, which is convenient for subsequent analysis, model training or other applications. Dataset formatting: Based on the labeled data, a standardized question-answering dataset is constructed. This dataset consists of two main parts:
[0111] Question(i): The i-th question of the user.
[0112] Response(i): The answer of the staff to the i-th question. Response(i) is a composite structure, including Policy(i,k) (the name of the policy cited by the staff), Content(i,k) (the relevant content cited in the policy), and Answer(i) (the direct answer of the staff).
[0113] Data storage: The dataset D1 can be stored in common data formats such as JSON, CSV, or database tables. Each record should contain the following fields:
[0114] Question(i): The text of the user's question.
[0115] Response(i): A dictionary or nested structure containing Policy(i,k), Content(i,k), and Answer(i).
[0116] Data verification: Ensure that the annotation information in the dataset D1 is accurate and error-free, and avoid mislabeling or omission.
[0117] The beneficial effects of the above technical solutions are as follows: Through annotation and classification, the original Q&A data is converted into a structured format. Each Q&A pair is decomposed into multiple analyzable parts (question, policy, policy content, answer), which is not only convenient for storage but also supports subsequent precise queries and analysis. By annotating Policy(i,k) and Content(i,k), the specific policies and their content cited by the staff in the answer are clearly marked. This annotation not only facilitates subsequent policy tracing but also helps answerers quickly retrieve relevant regulations, providing precise support for policy implementation and dissemination. The annotated Q&A dataset D1 can be used for the training of machine learning models, and then an automated Q&A system can be realized. By constructing a high-quality training dataset, the system can automatically match relevant policies when users ask questions and give accurate answers. Based on the standardized dataset, the answers of the staff are refined into specific policy references and article content. In this way, the answers are not only more authoritative but also more in line with policy specifications, improving the accuracy and integrity of the responses. After structuring a large amount of Q&A data, multi-dimensional statistics and analysis can be carried out. For example, it can be analyzed which questions are the most common, which policies are cited the most, and which fields of policies are most concerned by users. This provides data support for decision-making and policy adjustment. Through this dataset, a more intelligent Q&A system can be trained. The system can not only answer users' specific questions but also automatically quote relevant policy documents according to the context of the questions, improving the accuracy and efficiency of automatic Q&A.
[0118] In another embodiment, step S102 includes:
[0119] S1021: Preset a threshold P0 and set the maximum number K of noise data to be added.
[0120] S1022: Generate a random number between 0 and 1 for each question Question(i). If the random number is greater than the preset threshold P0, do not modify this Q&A data. If the random number is less than or equal to P0, perform the following steps:
[0121] Generate a random integer K0 between 0 and K, where K0 represents the number of text blocks to be added, and repeat this step K times.
[0122] S1023: Randomly select a text block from the policy library {Chunk(m,n)}, record the policy name corresponding to this text block, and then add this text block to {Policy(i,k),Content(i,k)} of the original answer Response(i).
[0123] The working principle of the above technical solution is as follows: Preset a threshold P0, which indicates whether to modify a piece of Q&A data each time it is processed. The range of P0 is between 0 and 1. When the generated random number is less than or equal to P0, noise data will be added to this Q&A data, otherwise no modification is made.
[0124] For each Q&A entry (Question(i), Response(i)) in the dataset D1, perform the following steps: Generate a random number: Generate a random number between 0 and 1 for the current Q&A entry. If the random number is greater than P0, do not modify this data and directly skip this data. Add noise: If the random number is less than or equal to P0, then noise needs to be added to this Q&A data. At this time, generate a random integer K0 between 0 and K, which represents the number of text blocks to be added. The range of K0 is from 0 to K, and K is the preset maximum number of noise blocks. Add text blocks: Repeat the following operation K0 times:
[0125] Randomly select a text block from the policy library {Chunk(m,n)}. The text blocks in the policy library are predefined text segments, and each text block corresponds to a policy name. Record the policy name corresponding to this text block and add it to the original answer Response(i). Specifically, add the {Policy(i,k),Content(i,k)} information of the selected text block to the answer of Response(i).
[0126] Updated Q&A data: For the modified Q&A data, the original answer part will be supplemented with new text blocks, making the final answer Response(i) more verbose and containing more policy information. The generated dataset D2 is a Q&A dataset with added noise based on D1.
[0127] Suppose we have a question and answer dataset D1, and one of the data is as follows:
[0128] Question: "What are China's science and technology innovation policies?"
[0129] Answer: "The Chinese government promotes science and technology innovation and has formulated a number of supportive policies with the goal of improving the independent innovation ability of science and technology."
[0130] Now, assume that the set threshold P0 is 0.5 and the maximum number of noise blocks K is 3. We process the data item by item:
[0131] For this question, a random number 0.3 is generated. Obviously, 0.3 <= 0.5, so noise needs to be added.
[0132] Generate a random number K0 from 0 to 3. Assume K0 is 2, which means we will add 2 text blocks.
[0133] In this way, our original Q&A data becomes more complex by randomly adding noise (policy blocks), contains more information, and at the same time enhances the model's understanding of policy-related content.
[0134] The beneficial effects of the above technical solution are as follows: By randomly adding text blocks, the Q&A dataset D2 is more diverse than the original dataset D1. This can help enhance the robustness of the model because the model can learn various different answering methods and information structures. For machine learning models, especially in natural language processing tasks, adding noise helps improve the generalization ability of the model. It enables the model to better handle the noisy data that may be encountered in actual applications and prevents the model from simply memorizing the training data. In the real world, Q&A systems may produce noisy answers due to reasons such as incomplete data and missing information. By adding noise to the data, the model can better simulate and adapt to the situations in real applications. By adding policy blocks, it can ensure that the model can handle more complex policy information when answering questions, thereby enhancing its understanding and answering ability of the policy background.
[0135] In another embodiment, step S103 includes:
[0136] S1031: Extract M questions Question(m), corresponding policy names Policy(m,k), and policy contents Content(m,k) from the second Q&A dataset D2, and write the logical reasoning process to obtain the reasoning result logic(i);
[0137] S1032: When writing the logical reasoning process, for each question question(i) and each policy policy(i,k), content(i,k) in the second Q&A dataset D2, generate logic(i) according to the following two cases:
[0138] Case 1: If policy(i,k) and content(i,k) are noise data, then logic(i) = "The policy content has nothing to do with the user's question";
[0139] Case 2: If policy(i,k) and content(i,k) are not noise data, then use the large language model to generate. The generation of the large language model takes the constructed prompt as input to obtain the output of the large language model, and logic(i) = the output of the large language model;
[0140] Among them, the constructed prompt includes:
[0141] Known user question Question(i);
[0142] Known policy Policy(i,k), Content(i,k);
[0143] Please combine the policy and content to conduct logical reasoning and output the result;
[0144] Example:
[0145] User question question(1), known policy policy(1,k), content(1,k), output logic(1),
[0146] User question question(2), known policy policy(2,k), content(2,k), output logic(2), ......
[0148] User question question(M), known policy policy(M,k), content(M,k), output: logic(M);
[0149] S1033: Based on the inference result logic(i), generate a third question-answering dataset D3, where D3 = {Question(i), response(i)}, response(i) = {{policy(i,k), content(i,k)}, logic(i), answer(i)}.
[0150] The working principle of the above technical solution is:
[0151] Step 1: Extract M questions and related policy data from the second question-answering dataset D2
[0152] Extract M user questions Question(m) from dataset D2, along with the policy name Policy(m,k) and policy content Content(m,k) corresponding to each question. Each question, policy, and content here constitutes a complete data unit.
[0153] Assume that the following three questions are selected from dataset D2:
[0154] Question 1: My child is 7 years old and wants to enroll in primary school. How can he enroll?
[0155] Policy 1: Chongqing Municipal Compulsory Education Student Registration Management Methods
[0156] Content 1: School-age children and teenagers with local household registration can go to school nearby without taking any entrance exams; new students for compulsory education in urban areas can go to school nearby according to the principle of "three matchings", that is, the household registration, property ownership certificate and actual residence of school-age children are consistent with those of their parents.
[0157] Question 2: I work in Yubei District, Chongqing. How can I transfer my child to a school here?
[0158] Policy 2: Chongqing Yubei District 2023 Compulsory Education School Enrollment Plan
[0159] Content 2: "All eligible children and teenagers who were registered in our district before August 31, 2022, must log in to the registration system to fill in their household registration information. After verification and confirmation of the authenticity of the information, the District Education Commission will send the student information to the corresponding school for eligible students, and the school will directly issue the admission notice."
[0160] Question 3: I live in Area A of Huanhu Yaju. All the children around me have received notification letters, why hasn’t my child received one yet?
[0161] Policy 3: Chongqing Yubei District 2023 Compulsory Education School Enrollment Plan
[0162] Content 3: "Schools enroll students based on factors such as household registration information, age, and place of residence of eligible children, in accordance with the principle of school zoning."
[0163] Step 2: Logical reasoning based on policy and content
[0164] For each question Question(i) and each related policy Policy(i,k) and policy content Content(i,k), perform logical reasoning to ensure that the reasoning results can clearly explain the relationship between the policy and the question.
[0165] Logical reasoning examples:
[0166] Question 1: My child is 7 years old and wants to enroll in primary school. How can he enroll?
[0167] Policy 1: School-age children and teenagers with local household registration can go to school nearby without taking any entrance exams; new students for compulsory education in urban areas can go to school nearby according to the principle of "three matchings", that is, the household registration, property ownership certificate and actual residence of school-age children are consistent with those of their parents.
[0168] Logical reasoning: Based on policy content 1, "school-age children with local household registration will attend school near their registered residence." Combined with the fact that the child will soon be seven years old, this indicates that the child meets the school enrollment age requirement. Furthermore, the policy clearly states that enrollment requires compliance with the "three matching" principle: household registration, property ownership certificate, and actual place of residence. If parents provide these supporting documents, the child should be able to attend school nearby.
[0169] Question 2: I work in Yubei District, Chongqing. How can I transfer my child to a school here?
[0170] Policy 2: All eligible children and teenagers who were registered in our district before August 31, 2022, must log in to the registration system to fill in their household registration information. After verification and verification of the information, the District Education Commission will send the student information to the corresponding school, and the school will directly issue an admission notice.
[0171] Logical reasoning: According to Policy 2, household registration information must first be submitted in the registration system and verified to be accurate. If approved, the district education committee will send the student's information to the designated school, which will then directly issue the admission letter. Therefore, parents need to confirm that they have logged into the system and submitted their household registration information.
[0172] Question 3: I live in Area A of Huanhu Yaju. All the children around me have received notification letters, why hasn’t my child received one yet?
[0173] Policy 3: Schools enroll students based on factors such as the household registration information, age, and place of residence of children of school age, in accordance with the principle of school zoning.
[0174] Logical reasoning: First, you need to confirm whether the relevant household registration information and proof of residence have been submitted as required by the policy. If so, you also need to confirm whether the child meets the age requirement. If other eligible children have already received a notification, it may be due to incomplete systems or review processes. Parents should contact the school or district education committee to inquire about the review progress and admission status.
[0175] Step 3: Generate logical reasoning dataset D3
[0176] Based on the logical reasoning results, combined with questions, policies and content, the third question-answering dataset D3 is generated. Each data unit contains:
[0177] Question: The specific question asked by the user.
[0178] Policy and content: The name and content of the policy related to the issue.
[0179] Logical reasoning: the logical reasoning process based on policy content and issues.
[0180] Answer: The final answer is a detailed explanation based on logical reasoning and policies.
[0181] Assume that we have completed logical reasoning for each question and generated a D3 dataset. The D3 data format is as follows:
[0182] D3=[
[0183] {
[0184] "Question":"My child is 7 years old and wants to enroll in primary school. How can I enroll him?"
[0185] "response":{
[0186] "policy":"Chongqing Municipal Compulsory Education Student Registration Management Methods",
[0187] "content":"Children and teenagers of school age with local household registration are exempt from entrance exams and enrolled in schools nearest their registered residence. New students for compulsory education in urban areas are assigned to schools nearest their registered residence, implementing the "three matching" principle, meaning that the school-age child's household registration, property ownership certificate, and actual residence must be consistent with their parents'."
[0188] "logic":"According to policy 1, "school-age children with local household registration will attend school near their registered residence." Combined with the fact that "the child will soon be 7 years old," we can conclude that the child meets the school enrollment age requirement. Furthermore, the policy clearly states that enrollment requires compliance with the "three matching" principle: household registration, property ownership certificate, and actual place of residence. If parents provide these supporting documents, the child should be able to attend school nearby."
[0189] Parents need to provide household registration, property ownership certificate, and relevant materials that are consistent with their actual place of residence. In accordance with the policy, their children can attend school near their registered place of residence.
[0190] }
[0191] },
[0192] {
[0193] "Question":"I work in Yubei District, Chongqing. How can I transfer my child to a school here?",
[0194] "response":{
[0195] "policy":"2023 Compulsory Education School Enrollment Plan in Yubei District, Chongqing",
[0196] "content":"All eligible children and teenagers who were registered in our district before August 31, 2022, must log in to the registration system to fill in their household registration information. After verification and confirmation of the authenticity of the information, the District Education Commission will send the student information to the corresponding school for eligible students, who will then directly issue admission letters."
[0197] "logic":"Parents need to log in to the registration system to fill in their household registration information. After ensuring that the information is accurate, the school will send an admission letter. If this step has not been completed, parents should log in to the system immediately to fill in the information.",
[0198] Please log in to the registration system and fill in the relevant information. After confirming that the information has been reviewed and approved, wait for the admission notice.
[0199] }
[0200] },
[0201] {
[0202] "Question":"I live in Huanhu Yaju District A. All the children around me have received notifications, but why hasn't my child received one yet?"
[0203] "response":{
[0204] "policy":"2023 Compulsory Education School Enrollment Plan in Yubei District, Chongqing",
[0205] "content":"Schools recruit students based on factors such as household registration information, age, and place of residence of eligible children, and in accordance with the principle of school zoning.",
[0206] "logic": "Parents need to confirm whether all the required household register and residence certificates have been submitted. If they have been submitted but the notice has not been received, they should contact the school to confirm whether there is an incomplete review.",
[0207] "answer": "Please ensure that all the required supporting materials have been submitted and contact the school to confirm the review progress.",
[0208] }
[0209] }
[0211] Step 4: Model Fine-tuning and Application
[0212] Fine-tune the large language model: Use the Q&A dataset D3 containing logical reasoning to fine-tune the large language model. Through training, the model can better understand the logical relationship between policy content and user questions, and avoid overfitting and refusal to answer phenomena.
[0213] Generate question answers: In actual applications, when a new user question is received, the model can provide more accurate and reasonable answers based on its internal logical reasoning.
[0214] The beneficial effects of the above technical solutions are as follows: Without a logical reasoning process, directly providing questions and policies to the large language model for fine-tuning may cause the model to learn shallow features, resulting in serious overfitting. The introduction of logical reasoning can help the model understand the deep logical relationship between policies and questions, thereby improving the generalization ability of the model; By introducing logical reasoning steps, the model can make judgments based on a more rigorous reasoning process, rather than just relying on word-level matching, which can significantly improve the accuracy of the model's answers to complex questions; Through logical reasoning, the model can clearly identify which policies and contents are relevant to the user's question and which are noise data, thereby preventing the model from wrongly connecting irrelevant policy content to the user's question and improving the quality and reliability of the answer; Through the sample data generated based on logical reasoning, the diversity of the model training data can be increased, which helps to improve the model's ability to handle different types of questions, especially when facing complex or ambiguous questions.
[0215] In another embodiment, the S104 step includes:
[0216] Construct a fine-tuning dataset according to the third Q&A dataset D3, and the specific form is:
[0217] Input = {Question(i), {Policy(i,k), Content(i,k)}}
[0218] Output = {Logic(i), Answer(i)};
[0219] Use the LoRA fine-tuning method to train the large language model. Through K epochs of iterative training, adjust the model's parameters to optimize the model's understanding of the logical relationship between user questions and policy content, and finally obtain a fine-tuned large language model.
[0220] Among them, obtaining a fine-tuned large language model includes:
[0221] Through K epochs of iterative training, gradually update the model's weights, optimize its response ability to input data, and gradually adjust the parameters to improve the model's understanding of the logical relationship between user questions and policy content;
[0222] Within each epoch, based on the feedback of the current training data, adjust the model's specific parameters and gradually reduce the loss function value to improve the accuracy and reasoning ability of the large language model on specific tasks;
[0223] According to the training results, at the end of each training cycle, evaluate the model performance to check whether the model meets the preset optimization criteria;
[0224] Based on the evaluation results, if the model performance does not meet the expectations, continue to perform iterative training, adjust the training set or optimize the parameter settings until a fully fine-tuned large language model is obtained.
[0225] The working principle of the above technical solution is: construct a fine-tuned instruction dataset according to the third Q&A dataset (D3). Each piece of data in the dataset contains the user's question (Question(i)), relevant policies and content (Policy(i,k), Content(i,k)), and the model's target output: logical reasoning result (Logic(i)) and the final answer (Answer(i)).
[0226] Data format: Input: {Question(i), {Policy(i,k), Content(i,k)}}, indicating that the relevant policy information is included in the question. Output: {Logic(i), Answer(i)}, indicating that the model outputs the reasoning logic and the final answer based on the policy content and the question.
[0227] Suppose the D3 dataset contains the following questions:
[0228] Question(i): "How to apply for an education subsidy?"
[0229] Policy(i,1): "Education subsidy policy"
[0230] Content(i,1): "Students who meet the criteria can apply for educational subsidies online."
[0231] For this problem, the target output of the model should be:
[0232] Logic(i): "Students who meet the criteria can apply, and the application needs to be done online."
[0233] Answer(i): "You can apply for educational subsidies through the official website as long as you meet the relevant criteria."
[0234] LoRA Fine-tuning Method: LoRA (Low-Rank Adaptation) is a fine-tuning method for large language models. It reduces the computational cost and memory requirements during fine-tuning by decomposing the weights of the pre-trained model into low-rank matrices. During the LoRA fine-tuning process, the original parameters of the model are not directly modified. Instead, appropriate low-rank matrices are added to adjust the model's response.
[0235] During the fine-tuning process, we will iteratively train the model for K epochs. Each time during training, the model will be adjusted according to the input questions and relevant policy content, combined with logical reasoning rules and target answers. Prepare the fine-tuning dataset: Based on the third Q&A dataset D3, construct the input-output dataset. Each piece of data includes: Input: questions and relevant policy content (e.g., "How to apply for educational subsidies?" and "Educational subsidy policy: Eligible students can apply."), Output: reasoning logic and answers (e.g., "Students who meet the criteria can apply, and the application needs to be done online." and "You can apply for educational subsidies through the official website."). Initialize the large language model: Select a pre-trained large language model (such as GPT-3 or GPT-4) and load its original parameters. Apply the LoRA method: Embed low-rank matrices into the model through the LoRA fine-tuning method, so that we can optimize the model's performance on specific tasks with a small number of adjustments. LoRA focuses on changing some key parameters in the model to make it more suitable for tasks in specific domains (e.g., reasoning based on policy content).
[0236] Iteratively train the model using K epochs (i.e., K complete traversals of the dataset). In each training, the model adjusts its weights by comparing its predicted output (based on the input questions and policy content) with the true target output (logic and answers). As the number of training times increases, the model will gradually optimize its understanding of the logical relationship between user questions and policy content. After training for K epochs, the parameters of the model will be optimized, and finally, a large language model that can better answer user questions involving policy content will be obtained.
[0237] The beneficial effects of the above technical solution are as follows: Through the adjustment of the low-rank matrix, the LoRA method avoids the comprehensive modification of all parameters of the model, thus significantly reducing the computing resources and memory required in the fine-tuning process. This makes the fine-tuning of large language models more efficient, especially in resource-constrained situations. The fine-tuned model can better understand the logical relationship between user questions and policy content. Especially when dealing with questions involving policies and regulations, it can provide more accurate and logical answers. The model can infer reasonable answers based on the existing policy content, thereby effectively improving the user experience. The LoRA method reduces the amount of parameter updates in the fine-tuning process, greatly reducing the computing resources and storage requirements for training, which is particularly important for dealing with large-scale language models. In addition, due to the high efficiency of LoRA fine-tuning, the training time of the model is also significantly shortened. Through LoRA fine-tuning, the model can not only adapt to the needs of specific domains (such as policy analysis and question answering), but also retain its performance in other general tasks. In this way, the fine-tuned model can handle customized questions and continue to perform well in other tasks.
[0238] In another embodiment, after step S104, it includes:
[0239] Use the fine-tuned large language model to answer new government affairs questions. The fine-tuned large language model generates the corresponding logical reasoning Logic(i) and the final answer Answer(i) according to the input Question(i) and {Policy(i,k), Content(i,k)}.
[0240] The working principle of the above technical solution is as follows: The input data includes: Question(i): The government affairs question raised by the user. For example, "How to apply for enterprise subsidies?" Policy(i,k) and Content(i,k): Relevant policy content and detailed descriptions. For example, enterprise subsidy policies and related application requirements.
[0241] Logic(i): The reasoning process of the model based on the question and policy content, used to clarify the key elements of the question. Answer(i): The final generated answer, combining logical reasoning and policy content.
[0242] The model receives the question (Question(i)) and the relevant policy content (Policy(i,k) and Content(i,k)), and starts to process. The model first conducts logical reasoning (Logic(i)) by analyzing the question and policy content. For example, if the question mentions "enterprise subsidies", the model will identify key information such as application eligibility and application procedures involved in the policy content.
[0243] Suppose the user asks: "How to apply for enterprise subsidies?"
[0244] Question (i): "How to apply for enterprise subsidies?"
[0245] Policy (i,k): "Enterprise subsidy policy"
[0246] Content (i,k): "Eligible enterprises can apply for enterprise subsidies online. Financial statements and tax registration certificates need to be provided."
[0247] Logic (i): Based on the logical relationship between the question and the policy content, the model infers that the user needs to apply through the online platform and prepare two materials: financial statements and tax registration certificates.
[0248] Generate Answer: Based on the logical reasoning, the model will generate the final answer (Answer (i)) in combination with the existing policy content. Here, the model will comprehensively consider all known policy elements and construct a clear and concise answer that meets the actual requirements.
[0249] Answer (i): "You can apply for enterprise subsidies through the official website. When applying, you need to provide financial statements and tax registration certificates. If you meet the relevant conditions, your application will be approved."
[0250] The large language model fine-tuned by LoRA can automatically generate a clear and compliant reasoning process and the final answer based on the logical relationship between the user's question and the policy content. Through multiple rounds of training, the model has a deeper understanding of the policy background and can provide more accurate reasoning and answers.
[0251] The beneficial effects of the above technical solution are as follows: The fine-tuned large language model can deeply understand and analyze the policy content and the rules behind it, and can accurately extract the policy elements crucial to the user's question. Whether it is about the application process, eligibility requirements, or other policy-related details, the model can quickly provide corresponding answers. By fine-tuning on a large number of government affairs questions and policy content, the model enhances its logical reasoning ability and can automatically infer the most reasonable answer based on the context of the question and the provided policy content. This means that the model not only relies on direct factual information but can also reason based on the context to obtain an answer that is more in line with the actual situation.
[0252] In another embodiment, a fine-tuning device for a government affairs Q&A system includes:
[0253] An original Q&A dataset construction unit, configured to construct an original Q&A dataset D1 based on the questions of users and the answers of staff in the scenario;
[0254] The second Q&A dataset construction unit is used to add noisy data to the original Q&A dataset D1 to obtain the second Q&A dataset D2;
[0255] The third Q&A dataset construction unit is used to add logical reasoning processes to the second Q&A dataset D2 to generate the third Q&A dataset D3;
[0256] The large language model set construction unit is used to construct a fine-tuned large language model based on the third Q&A dataset D3.
[0257] The working principle of the above technical solution is as follows: First, construct the original Q&A dataset D1, which mainly consists of user questions and staff answers in the government affairs Q&A system. The original dataset includes questions and standard answers in real scenarios.
[0258] Specific operation steps: Collect scenario data: Collect various questions raised by users in the government affairs Q&A system. For example, users may ask, "How to apply for subsistence allowances?" Extract questions and answers: Staff answer questions based on relevant government service regulations and policies. For example, the staff answer, "You can apply for subsistence allowances through the local civil affairs department and need to provide materials such as ID cards and income certificates." Organize Q&A data: Organize each pair of Q&A into data entries as follows:
[0259] Question: How to apply for subsistence allowances? Answer: You can apply for subsistence allowances through the local civil affairs department and need to provide materials such as ID cards and income certificates.
[0260] Based on the original Q&A dataset D1, add noisy data to obtain the second Q&A dataset D2. The noisy data can be errors, ambiguities, or non-standard expressions in user questions, or non-standard answers that may exist in staff answers.
[0261] Specific operation steps: Noise introduction: Randomly introduce some common noises in the questions and answers in the D1 dataset. For example: Question noise: The user's input may not be clear enough, such as "How to apply for subsistence allowances?" or "What to do about applying for subsistence allowances?" Answer noise: The staff's answer may not be complete or accurate enough, such as "Just go to the local department and bring all the materials." Generate the D2 dataset: Use these Q&A pairs with noise as new data points to finally obtain a Q&A dataset D2 containing noise. For example: Question: How to apply for subsistence allowances? Answer: Just go to the local department and bring all the materials. Or: Question: What to do about applying for subsistence allowances? Answer: You can go to the civil affairs department and bring your ID card and other relevant materials.
[0262] In the second question-and-answer dataset D2, add a logical reasoning process to generate the third question-and-answer dataset D3. Here, the "logical reasoning process" refers to adding background information, reasoning steps, policy explanations, etc. based on the given questions and answers to make the question-and-answer process clearer and more reasonable.
[0263] Specific operation steps: Introduce the reasoning process: For each question and answer in D2, add reasoning steps. For example, assume a question is "How to apply for subsistence allowances?" The staff may explain in detail the policy background, procedures, and required materials. Adding the reasoning process can help users better understand the relevant policies and procedures. Build a complete reasoning chain: In some cases, it may be necessary to break down complex questions and explain each step. For example: Question: How to apply for subsistence allowances? Answer: First, you need to confirm whether you meet the application conditions for subsistence allowances. According to national regulations, applying for subsistence allowances requires meeting a certain income standard. Second, you need to bring your ID card, family income certificate, and other relevant materials to the local civil affairs department to submit an application. After submitting the application, the civil affairs department will review your materials and decide whether to approve your application for subsistence allowances based on the results.
[0264] Generate the D3 dataset: Combine the noise data and the reasoning process to finally form the D3 dataset. The reasoning process will make the answers more complete and easier to understand. For example: Question: How to apply for subsistence allowances? Answer: To apply for subsistence allowances, you need to first confirm whether you meet the conditions. Income and family status will be used as the review criteria. If you meet the conditions, you need to prepare materials such as ID cards and family income certificates and go to the local civil affairs department to submit an application. After submission, the relevant department will review your materials and issue subsistence allowances after confirmation.
[0265] Based on the third question-and-answer dataset D3, build a fine-tuned large language model. This model will, on the basis of the original pre-trained large language model, use the D3 dataset for targeted training to make it more suitable for government affairs question-and-answer tasks.
[0266] Specific operation steps: Select a pre-trained model: First, select a pre-trained large language model suitable for natural language understanding and generation (such as GPT, BERT, etc.). Data fine-tuning: Use the D3 dataset to fine-tune the pre-trained model so that the model can generate high-quality answers related to government affairs. Training process: By continuously inputting question-and-answer pairs, train the model to gradually adjust the generated answers to better meet the needs of government affairs questions. Testing and optimization: After completing the fine-tuning, conduct tests to check the performance of the model in actual scenarios. If the answers generated by the model meet the government affairs standards and are logical, it can be further deployed. Generated results: The fine-tuned large language model can understand the user's questions and provide detailed answers that comply with policy regulations. For example: Question: How to apply for subsistence allowances? Answer: According to your specific situation, the conditions for applying for subsistence allowances include family income, the employment situation of family members, etc. You can carry relevant supporting materials and go to the local civil affairs department to apply.
[0267] The beneficial effects of the above technical solutions are as follows: By constructing the original question-and-answer dataset D1 and introducing noise and the reasoning process, the model can understand various expressions and question complexities, enabling it to give more accurate and complete answers when facing user questions. The second question-and-answer dataset D2 after adding noise data can train the model to still generate correct answers when facing various non-standard, ambiguous, or incorrect expressions. This makes the model more robust and adaptable. For relatively ambiguous questions such as "How to handle the application for subsistence allowances?", through learning from the noise data, the model can understand and give reasonable answers without being misled. By adding the reasoning process to the third question-and-answer dataset D3, the model's answers are no longer simple factual descriptions but include clear reasoning and background explanations. This enables users to better understand government affairs policies and the logic behind them. Through the fine-tuned large language model, the government affairs question-and-answer system can automatically and efficiently respond to a large number of user consultations, reducing the pressure on artificial customer service and improving the efficiency and quality of government affairs services. Users can quickly obtain information on policies, procedures, applications, etc. through the intelligent question-and-answer system, enhancing the accessibility and transparency of government services.
[0268] In another embodiment, an electronic device includes: at least one processor and a memory, where:
[0269] The memory is used to store computer execution instructions;
[0270] At least one processor is used to execute the computer execution instructions stored in the memory.
[0271] The working principle of the above technical solution is as follows: A memory (such as a memory) is used to store the instructions and data executed by a computer. In a government affairs Q&A system, these instructions are usually program codes (such as algorithms, data query logics, etc.), and the data may include a Q&A database, information input by users, etc. A common instruction in the government affairs Q&A system is "query relevant answers to the user's question". This instruction is stored in the memory, indicating that the system provides the most relevant answer to the user by comparing the questions and answer pairs in the database.
[0272] When the computer starts up, the processor (CPU) reads the instructions from the memory and starts to execute. The role of the processor is to decode the instructions and perform corresponding calculations or operations. The processor usually fetches an instruction from the memory and then performs specific tasks (such as arithmetic operations, data operations, etc.) according to the content of the instruction. Suppose the user inputs a question in the government affairs Q&A system: "Where can I apply for a social security card?" The processor fetches the instruction "query user questions" from the memory; the processor recognizes that this is a query operation and then starts to retrieve relevant answers from the database in the memory (for example: "A social security card can be applied for through the local social security bureau website."); finally, the processor returns the query result to the user.
[0273] During the execution process, the processor needs to perform a large amount of data interaction with the memory. For example, the processor of the government affairs Q&A system will access the knowledge base stored in the memory, the historical data input by users, and various rule and logic libraries. The user inputs the question: "How to apply for unemployment benefits?" The processor first loads the rules and steps related to "application for unemployment benefits" from the memory. Then, according to the specific query content, the processor decides what materials the user needs to submit and where to handle it, etc. Finally, the processor feeds back the query result to the user.
[0274] After the processor executes the instructions in the memory, the processor feeds back the result to the user through an output device. This process usually includes the presentation in the form of text, graphics, or other forms. For example: The user inputs the question: "How to handle the house transfer procedures?" After steps such as querying and calculating, the processor will display a series of detailed steps, required materials, and handling locations on the interface of the government affairs Q&A system.
[0275] The beneficial effects of the above technical solution are as follows: Through the cooperation of the memory and the processor, the government affairs Q&A system can quickly respond to a large number of user questions in a short time and provide instant answers. The computing speed of the processor and the data access speed of the memory can greatly improve the working efficiency of the system.
[0276] In another embodiment, a computer storage medium stores computer-executable instructions.
[0277] The working principle of the above technical solution is as follows: A computer storage medium refers to the hardware component used by a computer to store data and instructions. Common storage media include hard disk drives (HDDs), solid-state drives (SSDs), random access memory (RAM), optical discs, flash memories, etc. In a government affairs Q&A system, the computer storage medium stores various data of the system (such as Q&A databases, user history records) and program codes (such as logical instructions for processing query requests).
[0278] The program code of the government affairs Q&A system is first loaded and stored in the storage medium. The storage medium is usually used to save the execution instructions and data files of the system. When the system starts, the operating system loads the program code into the memory so that the processor can execute it.
[0279] Suppose there is a module in the government affairs Q&A system that specifically processes user query requests, such as "steps to apply for a business license". The execution instructions of the system include:
[0280] Retrieve relevant information from the database;
[0281] Parse the user's natural language question;
[0282] Return the query result.
[0283] These instructions will be written into a program during system development and stored in the hard disk or solid-state drive (SSD).
[0284] When a user submits a query request, the computer processes the request by executing the instructions stored in the storage medium. First, the program instructions in the storage medium are loaded into the computer's random access memory (RAM), and the processor (CPU) reads and executes the instructions from the memory. The processor will perform tasks such as calculations and data operations according to the content of the instructions.
[0285] Suppose the user queries the question: "How to apply for a business license?" The program instructions stored in the storage medium instruct the system to first read the relevant information about "business license application" in the database. The system parses the user's question through a natural language processing (NLP) module and converts it into a machine-readable format. Then, the processor executes the instructions, retrieves the relevant steps from the database, and displays the answer to the user.
[0286] During the execution of instructions, the computer needs to perform frequent data interactions. The storage medium stores not only program instructions but also data files, such as Q&A databases, user query records, etc. The processor extracts the required data from the storage medium, processes it, and returns the query result. For example: The user asks, "How to apply for a social security card?" The program instructions of the system tell the processor how to extract the information about applying for a social security card from the database. The data file in the storage medium will contain the steps required for the application, the materials needed, etc. The processor formats this information into an answer and displays it to the user.
[0287] After the processor finishes executing the instructions, the result will be fed back to the user through a display device (such as a computer screen). The content of these feedbacks is calculated and sorted out through the program code and data in the storage medium. For example: After the system processes the user's question, it returns the answer: "A social security card can be applied for through the official website. The materials required include ID card, household register, etc." This content is generated by combining and calculating the program instructions and data files stored in the storage medium.
[0288] The beneficial effects of the above technical solution are as follows: The storage medium provides an efficient environment for storing computer-executed instructions. The separate storage of program instructions and data enables the computer to efficiently call instructions and execute tasks during operation. The computer storage medium supports the storage of a large amount of data. The government affairs Q&A system can store a large amount of policy information, regulations, common questions, etc., and perform quick queries and updates as needed. The scalability of the storage medium enables the system to continuously increase the storage capacity according to requirements, thereby supporting more service contents.
[0289] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A fine-tuning method for a government affairs Q&A system, characterized in that, Including: S101: Based on the questions from users and the answers from staff in the scenario, construct the original Q&A dataset D1; S102: Add noise data to the original Q&A dataset D1 to obtain the second Q&A dataset D2; S103: Add logical reasoning processes to the second Q&A dataset D2 to generate the third Q&A dataset D3; S104: Construct a fine-tuned large language model based on the third Q&A dataset D3; The steps of S101 include: S1011: Collect the questions from users and the answers from staff in the real scenario to obtain the original Q&A data; S1012: Conduct detailed annotation and classification on each user question Question and staff answer response in the original Q&A data. The annotation content includes: Policy(i,k) represents the k-th policy name cited by the staff for the i-th question; Content(i,k) represents the relevant content in the k-th policy cited by the staff; Answer(i) represents the direct answer given by the staff to the i-th question; S1013: Through the annotation and classification of the original Q&A data, form the standardized original Q&A dataset D1 = {Question(i), response(i)}; The steps of S103 include: S1031: Extract M questions Question(m), as well as the corresponding policy names Policy(m,k) and policy contents Content(m,k) from the second Q&A dataset D2, and write the logical reasoning process to obtain the reasoning result logic(i); S1032: When writing the logical reasoning process, for each question question(i) and each policy policy(i,k), content(i,k) in the second Q&A dataset D2, generate logic(i) according to the following two situations: Situation 1, if policy(i,k) and content(i,k) are noise data, then logic(i) = "The policy content has nothing to do with the user's question"; Situation 2, if policy(i,k) and content(i,k) are not noise data, then generate using the large language model. The generation of the large language model takes the constructed prompt as the input, obtains the output of the large language model, and logic(i) = the output of the large language model; Among them, the constructed prompt includes: Known user question Question(i); Known policies Policy(i,k), Content(i,k); Please combine the policies and content to conduct logical reasoning and output the result; S1033: Based on the reasoning result logic(i), generate the third Q&A dataset D3, where D3 = {Question(i), response(i)}, and response(i) = {{policy(i,k), content(i,k)}, logic(i), answer(i)}.
2. A fine-tuning method for a government affairs Q&A system according to claim 1, characterized in that, Before the steps of S102, Scrape relevant texts from government affairs websites, where the relevant texts include relevant laws, regulations, ordinances, and policies, and split the scraped relevant texts into text chunks of length L to obtain a policy library {Chunk(m,n)}, where Chunk(m,n) represents the nth text chunk of the mth policy.
3. A fine-tuning method for a government affairs Q&A system according to claim 1, characterized in that, Step S102 includes: S1021: Preset a threshold P0 and set the maximum number K of noise data to be added; S1022: Generate a random number between 0 and 1 for each question Question(i). If the random number is greater than the preset threshold P0, do not modify the Q&A data of this item. If the random number is less than or equal to P0, then perform the following steps: Generate a random integer K0 between 0 and K, where K0 represents the number of text chunks to be added, and repeat this step K times; S1023: Randomly select a text chunk from the policy library {Chunk(m,n)}, record the policy name corresponding to this text chunk, and then add this text chunk to {Policy(i,k), Content(i,k)} of the original answer Response(i).
4. A fine-tuning method for a government affairs Q&A system according to claim 1, characterized in that Step S104 includes: Construct a fine-tuning dataset based on the third Q&A dataset D3, in the specific form of: Input={Question(i),{Policy(i,k),Content(i,k)}} Output={Logic(i),Answer(i)}; Use the Lora fine-tuning method to train the large language model. Through K epochs of iterative training, adjust the parameters of the model to optimize the model's understanding of the logical relationship between user questions and policy content, and finally obtain a fine-tuned large language model.
5. A fine-tuning method for a government affairs Q&A system according to claim 1, characterized in that, After step S104 includes: Answer new government affairs questions through the fine-tuned large language model. The fine-tuned large language model generates the corresponding logical reasoning Logic(i) and the final answer Answer(i) according to the input Question(i) and {Policy(i,k), Content(i,k)}.
6. A fine-tuning device for a government affairs Q&A system, applying the fine-tuning method according to any one of claims 1-5, characterized in that, Includes: An original Q&A dataset construction unit for constructing an original Q&A dataset D1 based on the questions of users and the answers of staff in the scenario; A second Q&A dataset construction unit for adding noise data to the original Q&A dataset D1 to obtain a second Q&A dataset D2; A third Q&A dataset construction unit for adding a logical reasoning process to the second Q&A dataset D2 to generate a third Q&A dataset D3; A large language model set construction unit for constructing a fine-tuned large language model according to the third Q&A dataset D3.
7. An electronic device, characterized in that, Includes: At least one processor and a memory, where: The memory is used to store computer execution instructions; At least one processor is used to execute the computer execution instructions stored in the memory, so that at least one processor executes the method described in any one of claims 1 to 5.
8. A computer storage medium, characterized in that, Computer execution instructions are stored in a computer storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Natural language processing task execution and model training method, device and equipment
CN118278527A