Language large model question answering system and method oriented to political and law field

Through a language big model question and answer system for the political and legal field, the large language model is used to process user problems and match knowledge base data, the problem of insufficient grassroots government service personnel is solved, real-time and accurate government services for residents are achieved, and the public's understanding of policies and regulations and government trust is enhanced.

CN119961390APending Publication Date: 2025-05-09AEROSPACE SCI & ENG NETWORK INFORMATION DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411836945.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

There are insufficient grassroots government service personnel at present, and it is difficult to respond to residents' government consultation needs in a timely and precise manner.

Method used

A language big-model question and answer system for the political and legal field was designed, including knowledge base, search module and knowledge generation module. The knowledge base preserves government data information on data-processed policies and regulations, conflict mediation cases and court judgment documents. The search module receives user input problems, integrates and matches information through a large language model, and generates information to be reasoned. The knowledge generation module inputs the information to be reasoned into the large language model to generate the answer information corresponding to the question.

Benefits of technology

Real-time and accurate government services for users are achieved, and can provide residents with professional interpretation of policies and regulations, legal case analysis and conflict mediation suggestions, which improves the public's understanding of policies and regulations and government trust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961390A_ABST
    Figure CN119961390A_ABST
Patent Text Reader

Abstract

The invention provides a large language model question answering system and method oriented to the political and law field. The system comprises a knowledge base, a search module and a knowledge generation module. Wherein the knowledge base is used for storing government affair data information of policy and regulation texts, contradiction mediation cases and court judgment documents after data processing; the search module is used for receiving questions input by a user and carrying out information integration; matching data information in the knowledge base with questions input by the user to generate information to be reasoned; and the knowledge generation module is used for inputting the information to be reasoned into the large language model LLM, generating answer information corresponding to the question through reasoning calculation of the large language model, and outputting the answer information to the user, thereby realizing real-time precise government affair service for the user, and solving the problems that the existing grassroots government affair service personnel are insufficient, and the resident demand is difficult to timely and precisely respond through the above system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and natural language processing, and in particular to a large language model question-answering system and method for the political and legal fields. Background Art

[0002] At present, the existing grassroots government services mainly rely on community staff to manually respond to residents' inquiries. As residents' life needs continue to enrich, the number and scope of residents' government consultations continue to increase, and problems such as insufficient grassroots government service personnel and difficulty in timely and accurate response to residents' needs continue to emerge. There is an urgent need for a self-service Q&A system that can provide residents with highly professional Q&A. The Q&A function includes the scope of application, specific provisions and implementation of policies and regulations, the process of mediation and handling of contradictions and legal disputes, reference cases and suggestions, which will help the public understand policies and regulations, enhance the public's understanding and understanding of policies and regulations, promote the construction of a rule of law society, enhance the government's credibility, and increase the public's trust and support for the government.

[0003] It can be concluded that how to solve the problem of insufficient existing grassroots government service personnel and the difficulty in responding to residents' needs in a timely and accurate manner has become one of the existing technical problems that need to be urgently solved. Summary of the invention

[0004] The present invention provides a language large model question-answering system and method for the political and legal fields, which is used to solve the problem of insufficient existing grassroots government service personnel and difficulty in responding to residents' needs in a timely and accurate manner.

[0005] In the first aspect, a language large model question answering system for the political and legal fields is provided, including: a knowledge base, a search module and a knowledge generation module; wherein:

[0006] The knowledge base is used to store government data information such as policy and regulatory texts, conflict mediation cases, and court judgments that have been processed;

[0007] The search module is used to receive questions input by users and integrate information; and to match data information in the knowledge base with questions input by users to generate information to be inferred;

[0008] The knowledge generation module is used to input the information to be inferred into the large language model LLM, and after the reasoning calculation of the large language model, generate answer information corresponding to the question and output it to the user, thereby realizing real-time and accurate government services for the user.

[0009] In one embodiment, the knowledge base is constructed by the following method:

[0010] Collect relevant texts of laws, regulations, rules and precedents in various industries, fields and regions; as well as text data of relevant political and legal cases, including cases involving legal interpretation, judicial decisions and court rulings; as well as text data of relevant government affairs processing and resolution processes;

[0011] Clean and organize the collected text data, remove outdated and duplicate information, unify the format, and save it as a local knowledge file;

[0012] Sentence processing of the local knowledge file content, segmenting long text or large documents into chunks of a set size, using a word segmentation tool to divide sentences into phrases, while ensuring that the phrases have complete and independent semantics;

[0013] The text data that has completed word segmentation and block processing is converted into a numerical vector with the help of the Embedding model;

[0014] The generated numerical vectors and knowledge points are stored in the Fasis vector database.

[0015] In one implementation, the local knowledge file content is processed by sentence segmentation, a long text or large document is segmented into chunks of a set size, and a word segmentation tool is used to segment sentences into phrases, while ensuring that the phrases have complete and independent semantics, specifically including:

[0016] Split the local knowledge file content into multiple independent knowledge points of 250 words each. Each knowledge point is used as the minimum record of the question and answer to ensure that the text size meets the length requirements of the model.

[0017] Combined with the deep learning model, the tokenizer word segmentation tool is used to perform basic text processing and divide sentences into phrases to ensure that each phrase has relatively complete and independent semantics.

[0018] In one embodiment, the search module is specifically used to receive questions input by users, and use the large language model LLM to conduct multiple rounds of conversational interactions with users according to preset examples, improve user questions until the questions reach the set expected completeness, and integrate important information in multiple rounds of conversations; and convert user questions and important information into numerical vectors through the Embedding model; and match the numerical vectors corresponding to the user questions with the numerical vectors in the Fasis vector database of the knowledge base; obtain text information of the neighborhood near the knowledge text in the matched knowledge base to prevent the complete sentence from being segmented and cut off, and merge the text data in the corresponding knowledge base and the user questions to generate information to be inferred.

[0019] In one implementation, the search module matches the numerical vector corresponding to the user question with the numerical vector in the Fasis vector database of the knowledge base, specifically including:

[0020] The similarity between the numerical vector corresponding to the user question and the numerical vector in the Fasis vector database of the knowledge base is measured by the Euclidean distance. The smaller the value of the Euclidean distance, the more similar the two vectors are, and the larger the value, the less similar the two vectors are. Find the numerical vector in the Fasis vector database of the knowledge base that is closest to the numerical vector corresponding to the user question, and obtain the n texts with the highest similarity. The Euclidean distance calculation formula is:

[0021]

[0022] Among them, x represents the numerical vector corresponding to the user's question, x i represents the corresponding vector; y represents the numerical vector in the Fasis vector database of the knowledge base, y i Represents the corresponding component vector.

[0023] In a second aspect, a language large model question answering method for the political and legal fields is provided, and the method is applied to the above-mentioned system, including:

[0024] The search module receives questions input by users and integrates information, matches data information in the knowledge base and questions input by users, and generates information to be inferred; the knowledge base stores government data information of policy and regulatory texts, conflict mediation cases, and court judgments that have been processed;

[0025] The knowledge generation module inputs the information to be inferred into the large language model (LLM). After reasoning and calculation by the large language model, it generates answer information corresponding to the question and outputs it to the user, thus realizing real-time and accurate government services for the user.

[0026] In one embodiment, the knowledge base is constructed by the following method:

[0027] Collect relevant texts of laws, regulations, rules and precedents in various industries, fields and regions; as well as text data of relevant political and legal cases, including cases involving legal interpretation, judicial decisions and court rulings; as well as text data of relevant government affairs processing and resolution processes;

[0028] Clean and organize the collected text data, remove outdated and duplicate information, unify the format, and save it as a local knowledge file;

[0029] Sentence processing of the local knowledge file content, segmenting long text or large documents into chunks of a set size, using a word segmentation tool to divide sentences into phrases, while ensuring that the phrases have complete and independent semantics;

[0030] The text data that has completed word segmentation and block processing is converted into a numerical vector with the help of the Embedding model;

[0031] The generated numerical vectors and knowledge points are stored in the Fasis vector database.

[0032] In one implementation, the local knowledge file content is processed by sentence segmentation, a long text or large document is segmented into chunks of a set size, and a word segmentation tool is used to segment sentences into phrases, while ensuring that the phrases have complete and independent semantics, specifically including:

[0033] Split the local knowledge file content into multiple independent knowledge points of 250 words each. Each knowledge point is used as the minimum record of the question and answer to ensure that the text size meets the length requirements of the model.

[0034] Combined with the deep learning model, the tokenizer word segmentation tool is used to perform basic text processing and divide sentences into phrases to ensure that each phrase has relatively complete and independent semantics.

[0035] In one implementation, the search module receives the question input by the user and integrates the information, matches the data information in the knowledge base and the question input by the user, and generates information to be inferred, specifically including:

[0036] Receive questions input by users, and use the large language model (LLM) to conduct multiple rounds of conversations with users according to preset examples, improve user questions until the questions reach the expected completeness, and integrate important information from multiple rounds of conversations;

[0037] Transform user questions and important information into numerical vectors through the Embedding model;

[0038] Match the numerical vector corresponding to the user's question with the numerical vector in the Fasis vector database of the knowledge base;

[0039] Obtain the text information of the neighborhood near the knowledge text in the matched knowledge base to prevent the complete sentence from being segmented and cut off, and merge the text data in the corresponding knowledge base and the user question to generate the information to be inferred.

[0040] In one embodiment, matching the numerical vector corresponding to the user question with the numerical vector in the Fasis vector database of the knowledge base specifically includes:

[0041] The similarity between the numerical vector corresponding to the user question and the numerical vector in the Fasis vector database of the knowledge base is measured by the Euclidean distance. The smaller the value of the Euclidean distance, the more similar the two vectors are, and the larger the value, the less similar the two vectors are. Find the numerical vector in the Fasis vector database of the knowledge base that is closest to the numerical vector corresponding to the user question, and obtain the n texts with the highest similarity. The Euclidean distance calculation formula is:

[0042]

[0043] Among them, x represents the numerical vector corresponding to the user's question, x i represents the corresponding vector; y represents the numerical vector in the Fasis vector database of the knowledge base, y i Represents the corresponding component vector.

[0044] The embodiment of the present invention provides a language large model question-answering system and method for the political and legal fields, the system comprising: a knowledge base, a search module and a knowledge generation module; wherein the knowledge base is used to store government data information of policy and regulatory texts, conflict mediation cases and court judgments that have been processed; the search module is used to receive questions input by users and integrate the information; and to match the data information in the knowledge base and the questions input by users to generate information to be inferred; the knowledge generation module is used to input the information to be inferred into a large language model LLM, generate answer information corresponding to the question through reasoning calculation of the large language model, and output it to the user, thereby realizing real-time and accurate government services for users. Through the above system, the existing problems of insufficient grassroots government service personnel and difficulty in responding to residents' needs in a timely and accurate manner are solved.

[0045] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structure particularly pointed out in the written description, claims, and drawings thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0047] Figure 1 A schematic diagram of the structure of a language large model question answering system for the political and legal fields according to an embodiment of the present invention;

[0048] Figure 2 A schematic diagram of a knowledge base construction process according to an embodiment of the present invention;

[0049] Figure 3 A schematic diagram of a process for generating question-answer information using a large language model according to an embodiment of the present invention;

[0050] Figure 4 It is a flow chart of a language large model question answering method for the political and legal fields according to an embodiment of the present invention;

[0051] Figure 5The present invention is a workflow diagram of a large language model question-answering method for the political and legal fields according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to solve the problem of insufficient grassroots government service personnel and difficulty in responding to residents' needs in a timely and accurate manner, a language large model question and answer system and method for the political and legal fields is provided.

[0053] The preferred embodiments of the present invention are described below in conjunction with the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention. In addition, the embodiments of the present invention and the features in the embodiments may be combined with each other if there is no conflict.

[0054] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0055] like Figure 1 As shown, the embodiment provides a language large model question answering system for the political and legal fields, including: a knowledge base 11, a search module 12 and a knowledge generation module 13; wherein,

[0056] The knowledge base 11 is used to store government data information of policy and regulatory texts, conflict mediation cases and court judgments that have been processed;

[0057] The search module 12 is used to receive questions input by users and integrate information; and to match data information in the knowledge base with questions input by users to generate information to be inferred;

[0058] The knowledge generation module 13 is used to input the information to be inferred into the large language model LLM, and after the inference calculation of the large language model, generate answer information corresponding to the question and output it to the user, so as to realize real-time and accurate government services for the user.

[0059] In one embodiment, Figure 2 As shown, the knowledge base 11 is constructed by the following method:

[0060] 1) Data collection: Collect relevant texts of laws, regulations, precedents, etc. in various industries, fields and regions; as well as text data of relevant political and legal cases, especially cases involving legal interpretation, judicial decisions and court rulings; as well as text data of relevant government affairs processing and resolution processes;

[0061] 2) Data cleaning and organization: Clean and organize the collected text data, remove outdated and duplicate information, unify the format, and save it as a local knowledge file to ensure data consistency and accuracy;

[0062] 3) Chunking and word segmentation: Sentence processing of local knowledge file content, dividing long text or large documents into smaller, easier-to-process fragments (chunks), splitting the original local knowledge text into several independent, shorter knowledge points, forming knowledge points of 250 words in size, each knowledge point as the minimum record of the question and answer, ensuring that the text size meets the length requirements of the model. Combined with the deep learning model, use tokenizer and other word segmentation tools to perform basic text processing, divide sentences into phrases, and ensure that each phrase has relatively complete and independent semantics;

[0063] 4) Vectorization: Convert text data into numerical vectors for computer processing and analysis. Use the Embedding model to convert text data that has completed word segmentation and block processing into numerical vectors;

[0064] 5) Knowledge storage: The generated numerical vectors and knowledge points are stored in the Fasis vector database to facilitate subsequent question-answer matching indexing.

[0065] In one embodiment, the search module 12 is used to search and process questions raised by users such as residents, and perform question-answer matching:

[0066] 1) Receive questions input by users: Use the large language model (LLM) to interact with users according to preset examples, gradually improve the questions until they reach the expected completeness, and integrate important information from multiple rounds of conversations.

[0067] For example, taking labor disputes as an example, the user asks, "The company owes me wages, what should I do?" Regarding the wage arrears, whether a labor contract has been signed is important information. The system's large language model LLM asks the user, "Have you signed a labor contract with the company?" Important information is integrated through multiple rounds of conversations.

[0068] 2) Vectorization processing: Use the Embedding model to convert questions and important information into numerical vectors for subsequent question-answer matching.

[0069] 3) Question-answer matching: Match the numerical vector corresponding to the user question with the numerical vector in the Fasis vector database of the knowledge base. The similarity between the numerical vector corresponding to the user question and the numerical vector in the Fasis vector database of the knowledge base is measured by the Euclidean distance. The smaller the value of the Euclidean distance, the more similar the two vectors are, and the larger the value, the less similar the two vectors are. Find the numerical vector in the Fasis vector database of the knowledge base that is closest to the numerical vector corresponding to the user question, and obtain the n texts with the highest similarity. The Euclidean distance calculation formula is:

[0070]

[0071] Among them, x represents the numerical vector corresponding to the user's question, x i represents the corresponding vector; y represents the numerical vector in the Fasis vector database of the knowledge base, y i Represents the corresponding component vector.

[0072] 4) Generate reasoning information: In order to ensure the integrity of the retrieved text semantic paragraphs, set the acquisition of text information near the matching knowledge text to prevent the complete sentence from being segmented and cut off. Merge the text paragraphs of the corresponding knowledge base and the user questions to generate the information to be reasoned.

[0073] In one embodiment, Figure 3 As shown, the knowledge generation module 13 is used to map the question and the matching knowledge vector into text and then input it into the large language model. After calculation by the large language model, the answer information corresponding to the question is generated and output to users such as residents:

[0074] The n most matching text paragraphs in the knowledge base are combined with the question to form the information to be inferred by the large language model. The large language model uses the Chinese language understanding and generation capabilities to infer and generate answers based on the information to be inferred. At the same time, the model returns the matching paragraphs as a reference source.

[0075] It can be seen that through the above method, a language large-model question-and-answer system for the political and legal fields can be formed to realize real-time and accurate government affairs question-and-answer for residents and other users.

[0076] The embodiment provides a language large model question and answer system for the political and legal fields, including: a knowledge base, a search module and a knowledge generation module; wherein the knowledge base is used to store government data information of policy and regulatory texts, conflict mediation cases and court judgments that have been processed; the search module is used to receive questions input by users and integrate information; and match the data information in the knowledge base and the questions input by users to generate information to be inferred; the knowledge generation module is used to input the information to be inferred into the large language model LLM, and after reasoning and calculation of the large language model, generate answer information corresponding to the question and output it to the user, so as to realize real-time and accurate government services for users. Through the above system, it is possible to provide residents with real-time and professional government information question and answer services, provide residents with functions such as policy and regulatory reading, legal case learning, handling process notification, reference case analysis and suggestions, etc., timely meet residents' government service needs, help the public enhance their understanding and awareness of policies and regulations, handle and resolve conflicts and disputes in accordance with the law, promote the construction of a rule of law society, and enhance the public's trust and support for the government.

[0077] Based on the same technical concept, the embodiment of the present application also provides a language large model question and answer method for the political and legal fields. Since the method is applied to the above-mentioned system, the implementation of the method can refer to the implementation of the system, and the repeated parts will not be repeated.

[0078] like Figure 4-5 As shown, the embodiment provides a language large model question answering method for the political and legal fields, including:

[0079] s41. The search module receives questions input by users and integrates information, matches the data information in the knowledge base and the questions input by users, and generates information to be inferred; the knowledge base stores government data information of policy and regulatory texts, conflict mediation cases and court judgments that have been processed.

[0080] In specific implementation, the search module receives questions input by users, and uses the large language model (LLM) to conduct multiple rounds of conversational interactions with users according to preset examples, perfecting the user's questions until the questions reach the expected completeness, integrating important information from multiple rounds of conversations;

[0081] Transform user questions and important information into numerical vectors through the Embedding model;

[0082] Match the numerical vector corresponding to the user question with the numerical vector in the Fasis vector database of the knowledge base: The similarity between the numerical vector corresponding to the user question and the numerical vector in the Fasis vector database of the knowledge base is measured by Euclidean distance. The smaller the value of the Euclidean distance, the more similar the two vectors are, and the larger the value, the less similar the two vectors are. Find the numerical vector in the Fasis vector database of the knowledge base that is closest to the numerical vector corresponding to the user question, and obtain the n texts with the highest similarity thereto. The Euclidean distance calculation formula is:

[0083]

[0084] Among them, x represents the numerical vector corresponding to the user's question, x i represents the corresponding vector; y represents the numerical vector in the Fasis vector database of the knowledge base, y i represents the corresponding component vector;

[0085] Obtain the text information of the neighborhood near the knowledge text in the matched knowledge base to prevent the complete sentence from being segmented and cut off, and merge the text data in the corresponding knowledge base and the user question to generate the information to be inferred.

[0086] The knowledge base is constructed by the following method:

[0087] Collect relevant texts of laws, regulations, rules and precedents in various industries, fields and regions; as well as text data of relevant political and legal cases, including cases involving legal interpretation, judicial decisions and court rulings; as well as text data of relevant government affairs processing and resolution processes;

[0088] Clean and organize the collected text data, remove outdated and duplicate information, unify the format, and save it as a local knowledge file;

[0089] Process the content of the local knowledge file by sentence division, divide long text or large documents into chunks of a set size, and use a word segmentation tool to divide sentences into phrases, while ensuring that the phrases have complete and independent semantics: divide the content of the local knowledge file into multiple independent knowledge points of 250 words in size, and each knowledge point is used as the minimum record of the question and answer to ensure that the text size meets the length requirements of the model; combine the deep learning model approach, use the tokenizer word segmentation tool to perform basic processing on the text, divide the sentences into phrases, and ensure that each phrase has relatively complete and independent semantics;

[0090] The text data that has completed word segmentation and block processing is converted into a numerical vector with the help of the Embedding model;

[0091] The generated numerical vectors and knowledge points are stored in the Fasis vector database.

[0092] S42. The knowledge generation module inputs the information to be inferred into the large language model LLM. After the large language model’s inference calculation, it generates the answer information corresponding to the question and outputs it to the user, thereby realizing real-time and accurate government services for the user.

[0093] In specific implementation, the knowledge generation module inputs the information to be inferred into the large language model LLM, and uses the n most matching text paragraphs in the above knowledge base to merge with the question to form the information to be inferred required by the large language model. The large language model uses the Chinese comprehension and generation capabilities based on the information to be inferred to generate answers, and the model returns the matching paragraphs as a reference source.

[0094] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0095] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A language large model question answering system for the political and legal fields, characterized by: include: Knowledge base, search module and knowledge generation module; among them, The knowledge base is used to store government data information such as policy and regulatory texts, conflict mediation cases, and court judgments that have been processed; The search module is used to receive questions input by users and integrate information; and to match data information in the knowledge base with questions input by users to generate information to be inferred; The knowledge generation module is used to input the information to be inferred into the large language model LLM, and after the reasoning calculation of the large language model, generate answer information corresponding to the question and output it to the user, so as to realize real-time and accurate government services for the user.

2. The system according to claim 1, characterized in that The knowledge base is constructed by the following method: Collect relevant texts of laws, regulations, rules and precedents in various industries, fields and regions; and relevant political and legal case text data, including cases involving legal interpretation, judicial decisions, and court rulings; and text data of relevant government affairs processing and resolution processes; Clean and organize the collected text data, remove outdated and duplicate information, unify the format, and save it as a local knowledge file; Sentence processing of the local knowledge file content, segmenting long text or large documents into chunks of a set size, using a word segmentation tool to divide sentences into phrases while ensuring that the phrases have complete and independent semantics; The text data that has completed word segmentation and block processing is converted into a numerical vector with the help of the Embedding model; The generated numerical vectors and knowledge points are stored in the Fasis vector database.

3. The system according to claim 2, characterized in that Sentence processing of the local knowledge file content, segmenting long text or large documents into chunks of a set size, using word segmentation tools to divide sentences into phrases, while ensuring that the phrases have complete and independent semantics, specifically including: Split the local knowledge file content into multiple independent knowledge points of 250 words each. Each knowledge point is used as the minimum record of the question and answer to ensure that the text size meets the length requirements of the model. Combined with the deep learning model, the tokenizer word segmentation tool is used to perform basic text processing and divide sentences into phrases to ensure that each phrase has relatively complete and independent semantics.

4. The system according to claim 3, characterized in that The search module is specifically used to receive questions input by users, and use the large language model LLM to conduct multiple rounds of conversation interactions with users according to preset examples, improve user questions until the questions reach the expected completeness, and integrate important information in multiple rounds of conversations; And convert user questions and important information into numerical vectors through the Embedding model; and match the numerical vector corresponding to the user question with the numerical vector in the Fasis vector database of the knowledge base; obtain the text information of the neighborhood near the knowledge text in the matched knowledge base to prevent the complete sentence from being segmented and cut off, and merge the text data in the corresponding knowledge base and the user question to generate information to be inferred.

5. The system according to claim 4, characterized in that The search module matches the numerical vector corresponding to the user question with the numerical vector in the Fasis vector database of the knowledge base, specifically including: The similarity between the numerical vector corresponding to the user question and the numerical vector in the Fasis vector database of the knowledge base is measured by the Euclidean distance. The smaller the value of the Euclidean distance, the more similar the two vectors are, and the larger the value, the less similar the two vectors are. Find the numerical vector in the Fasis vector database of the knowledge base that is closest to the numerical vector corresponding to the user question, and obtain the n texts with the highest similarity. The Euclidean distance calculation formula is: Among them, x represents the numerical vector corresponding to the user's question, x i represents the corresponding vector; y represents the numerical vector in the Fasis vector database of the knowledge base, y i Represents the corresponding component vector.

6. A language large model question answering method for the political and legal fields, the method is applied to the systems described in 1 to 5, characterized in that: include: The search module receives questions input by users and integrates information, matches data information in the knowledge base and questions input by users, and generates information to be inferred; The knowledge base stores government data information of data-processed policy and regulatory texts, conflict mediation cases, and court judgments; The knowledge generation module inputs the information to be inferred into the large language model (LLM). After reasoning and calculation by the large language model, it generates answer information corresponding to the question and outputs it to the user, thus realizing real-time and accurate government services for the user.

7. The method according to claim 6, characterized in that The knowledge base is constructed by the following method: Collect relevant texts of laws, regulations, rules and precedents in various industries, fields and regions; and relevant political and legal case text data, including cases involving legal interpretation, judicial decisions, and court rulings; and text data of relevant government affairs processing and resolution processes; Clean and organize the collected text data, remove outdated and duplicate information, unify the format, and save it as a local knowledge file; Sentence processing of the local knowledge file content, segmenting long text or large documents into chunks of a set size, using a word segmentation tool to divide sentences into phrases while ensuring that the phrases have complete and independent semantics; The text data that has completed word segmentation and block processing is converted into a numerical vector with the help of the Embedding model; The generated numerical vectors and knowledge points are stored in the Fasis vector database.

8. The method according to claim 7, characterized in that Sentence processing of the local knowledge file content, segmenting long text or large documents into chunks of a set size, using word segmentation tools to divide sentences into phrases, while ensuring that the phrases have complete and independent semantics, specifically including: Split the local knowledge file content into multiple independent knowledge points of 250 words each. Each knowledge point is used as the minimum record of the question and answer to ensure that the text size meets the length requirements of the model. Combined with the deep learning model, the tokenizer word segmentation tool is used to perform basic text processing and divide sentences into phrases to ensure that each phrase has relatively complete and independent semantics.

9. The method according to claim 8, characterized in that The search module receives questions input by users and integrates information, matches the data information in the knowledge base and the questions input by users, and generates information to be inferred, including: Receive questions input by users, and use the large language model (LLM) to conduct multiple rounds of conversations with users according to preset examples, improve user questions until the questions reach the expected completeness, and integrate important information from multiple rounds of conversations; Transform user questions and important information into numerical vectors through the Embedding model; Match the numerical vector corresponding to the user's question with the numerical vector in the Fasis vector database of the knowledge base; Obtain the text information of the neighborhood near the knowledge text in the matched knowledge base to prevent the complete sentence from being segmented and cut off, and merge the text data in the corresponding knowledge base and the user question to generate the information to be inferred.

10. The method according to claim 9, characterized in that Match the numerical vector corresponding to the user's question with the numerical vector in the Fasis vector database of the knowledge base, specifically including: The similarity between the numerical vector corresponding to the user question and the numerical vector in the Fasis vector database of the knowledge base is measured by the Euclidean distance. The smaller the value of the Euclidean distance, the more similar the two vectors are, and the larger the value, the less similar the two vectors are. Find the numerical vector in the Fasis vector database of the knowledge base that is closest to the numerical vector corresponding to the user question, and obtain the n texts with the highest similarity. The Euclidean distance calculation formula is: Among them, x represents the numerical vector corresponding to the user's question, x i represents the corresponding vector; y represents the numerical vector in the Fasis vector database of the knowledge base, y i Represents the corresponding component vector.