Large model intelligent question answering method supporting data correction
By introducing a context-aware mechanism and an LSTM model, combined with data correction methods from large models and external data sources, the problem of incomplete information in intelligent question-answering systems is solved, resulting in more accurate and personalized question-answering results.
Patent Information
- Application Number
- CN202410718301.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-11-14
AI Technical Summary
Existing intelligent question-answering systems struggle to comprehensively cover relevant information when faced with massive, complex, and noisy data from the internet, resulting in an incomplete range of candidate answers.
We introduce a context-aware mechanism and an LSTM model, combine a large model to generate preliminary answers and extract information from external data sources, and use chi-square tests and algorithm detection to correct and fuse data. We also design a user interface for feedback analysis to adjust model parameters.
It improves the accuracy and reliability of the question-answering system, enabling it to provide accurate and relevant answers to complex or fuzzy queries, reduce the spread of misinformation, and enhance the user experience.
Smart Images

Figure CN120950631A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent question answering technology, and more specifically, to an intelligent question answering method for large models that supports data correction. Background Technology
[0002] In today's information age, the rapid development of big data and artificial intelligence technologies has greatly propelled the evolution of intelligent question-answering systems. These systems aim to quickly and accurately provide answers from massive amounts of data by understanding user questions, thus meeting users' immediate information needs. However, faced with the vast, complex, and often noisy data on the internet, even the most advanced intelligent question-answering systems may encounter inaccurate, outdated, or missing information, directly impacting the reliability of the answers and the user experience. Information sources on the internet are diverse and of varying quality, containing errors, contradictions, and incomplete data. This requires intelligent question-answering systems not only to have powerful information retrieval and understanding capabilities but also to identify and correct these errors, providing more accurate answers. Existing retrieval methods struggle to cover all relevant information, resulting in incomplete candidate answers and affecting the accuracy and completeness of the final answer. Furthermore, efficiently retrieving highly relevant information from massive amounts of data remains a challenge. Therefore, it is necessary to design a large-scale intelligent question-answering method that supports data correction. Summary of the Invention
[0003] The purpose of this invention is to provide a large-scale intelligent question answering method that supports data correction, in order to solve the problem that the retrieval methods proposed in the background art are difficult to cover all relevant information, resulting in an incomplete range of candidate answers.
[0004] To achieve the above objectives, the present invention aims to provide a large-scale intelligent question-answering method that supports data correction, comprising the following steps:
[0005] S1: Users input questions into the question-answering system through an interface, while a context-aware mechanism is introduced to enable the question-answering system to understand the background of the question and the user's intent;
[0006] S2: Pass the user's question to the large model, use the large model to generate an initial answer, and extract information related to the question from relevant external data sources;
[0007] S3: Based on the initial responses and external data, perform data correction and fusion, and introduce algorithms to detect and filter misleading information;
[0008] S4: Combine the initial response with the revised external data to generate the final response;
[0009] S5: Design a user interface that allows users to evaluate the accuracy of the question-and-answer results or provide direct suggestions for correction. Automatically analyze user feedback, identify error types, and adjust model parameters or supplement training data based on the error analysis results.
[0010] As a further improvement to this technical solution, in S1, the question-answering system can save the user's dialogue history, allowing the question-answering system to understand the background of new questions and the user's intentions based on the context awareness mechanism and context content. In this way, in multi-round dialogues, the question-answering system can adjust new answers based on previous responses.
[0011] As a further improvement to this technical solution, based on the aforementioned context-aware mechanism and understanding the background and user intent of the new problem using the LSTM model, the calculation steps are as follows:
[0012] For each time step t, the basic update equation is as follows:
[0013] h t =σ(W h ·[h t-1 x t ]+b h )
[0014] Among them, h t The hidden state at the current time step t; W h The weight matrix; [h t-1 x t [] is a combined vector, including the hidden state h from the previous time step. t-1 and the input x at the current time step t b h σ is the bias term; σ is the activation function;
[0015] Forgotten Gate:
[0016] f t =σ(W f ·[h t-1 x t ]+b f )
[0017] Among them, f t For time step t; W f • is the weight matrix of the forget gate; b f For the bias term of the forget gate;
[0018] Input Gate:
[0019] i t =σ(W i ·[h t-1 x t ]+b i )
[0020] Among them, i t The output of the input gate at time step t; W i b is the weight matrix of the input gate; i This is the bias term for the input gate;
[0021] Candidate cell status:
[0022]
[0023] in, For candidate cell states at time step t; W c b is the weight matrix related to the candidate cell state; c is the bias term for candidate cell states; tanh is the hyperbolic tangent activation function;
[0024] Cell status update:
[0025]
[0026] Among them, c t f represents the final cell state at time step t; t This represents the output of the forget gate; ⊙ represents element-wise multiplication; c t-1 The cell state at time step t-1; i t The output of the input gate;
[0027] Output gate:
[0028] o t =σ(W o ·[h t-1 x t ]+b o )
[0029] Among them, o t The output of the gate at time step t; W o b is the weight matrix of the output gate; o This is the bias term for the output gate;
[0030] Hidden state:
[0031] h t =o t ⊙tanh(c t )
[0033] As a further improvement to this technical solution, the specific steps of S2 are as follows:
[0034] S21: Preprocess the user-input questions to understand the intent, keywords, and topics of the questions, and then pass the preprocessed questions to the large model;
[0035] S22: The initial answers to the large model analysis and the specific information that needs to be obtained from external data sources;
[0036] S23: Select an appropriate data source based on information needs;
[0037] S24: Construct the query statement and interface request, and execute the query against the selected data source;
[0038] S25: Collect the query results, filter and integrate them, and extract key information related to the question.
[0039] As a further improvement to this technical solution, the specific steps of S3 are as follows:
[0040] S31: Based on the chi-square test, compare the key information in the preliminary answer with the key information related to the question collected from external data to find differences and errors;
[0041] S32: When there is inaccurate or outdated information in the initial response, correct it with external data. If the data source is reliable and there are obvious errors in the response, replace it directly with external data.
[0042] S33: When the initial answer lacks detail or is incomplete, supplement it with external data, add relevant details or additional information to ensure the completeness of the answer.
[0043] As a further improvement to this technical solution, the calculation steps of S31 are as follows:
[0044] Assume there is no difference in the categorical distribution between the initial response and the external data;
[0045] Create a cross table where the horizontal rows represent the categories in the initial responses, the vertical columns represent the categories in the external data, and the values in the cells represent the frequency of each category.
[0046] The expression for the chi-square statistic is:
[0047]
[0048] Where, χ 2 This is the chi-square statistic; r is the number of rows; c is the number of columns; O ij E represents the observation frequency. ij Let i be the expected frequency corresponding to the observed frequency; i = 1, 2, ..., r; j = 1, 2, ..., c;
[0049] The expression for degrees of freedom is:
[0050] df = (r-1)(c-1)
[0051] Where df represents the degrees of freedom;
[0052] Based on the degrees of freedom and chi-square statistic, find the corresponding P-value in the chi-square distribution table, and determine the difference between the two data sources based on the P-value.
[0053] As a further improvement to this technical solution, the specific steps for correcting using external data in S32 are as follows:
[0054] S321: Check the credibility and professionalism of the data publishing organization, and check whether there are peer reviews or official certifications as quality assurance to ensure the reliability of the data source;
[0055] S322: Compare external data with data from other reliable sources to see if they are consistent, and use descriptive statistical analysis to determine the reasonableness of the data;
[0056] S323: Confirm whether the definition and classification of the external data match the context in the initial response;
[0057] S324: When it is confirmed that external data is more accurate and directly relevant, directly replace the erroneous information in the initial response and add a note at the replacement point.
[0058] As a further improvement to this technical solution, the specific steps of S33 are as follows:
[0059] S331: Identify which parts of the preliminary answer lack details or are incomplete, and determine the specific types and scope of information that need to be supplemented based on the needs of the question;
[0060] S332: Transform external data into a format that matches the initial response, ensuring consistency in units and timeframes;
[0061] S333: In the initial response, accurately insert details provided by external data and indicate the data source;
[0062] S334: Ensure that the supplementary information is logically consistent with the original content, and optimize the text of the supplementary part.
[0063] As a further improvement to this technical solution, the automatic analysis of user feedback in S5 includes analyzing the keywords, sentiment tendencies, and contextual features in the user feedback content to predict the type of error.
[0064] As a further improvement to this technical solution, the method for adjusting model parameters based on the error analysis results in S5 includes supplementing training data and active learning.
[0065] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0066] This large-model intelligent question-answering method, which supports data correction, introduces a context-aware mechanism, enabling the question-answering system to more accurately understand the deeper meaning of the question and the specific needs of the user, thus improving the personalization level and user experience. This helps to provide accurate and relevant answers even in complex or ambiguous queries. Simultaneously, by generating initial answers through a large model and combining them with external data sources for correction and fusion, not only are information sources broadened, but algorithms are also used to detect and filter misleading information, effectively reducing the spread of erroneous information and improving the accuracy and reliability of the answers. Attached Figure Description
[0067] Figure 1 This is a flowchart illustrating the overall method of the present invention. Detailed Implementation
[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Example
[0070] Please see Figure 1 As shown, a large-scale intelligent question answering method supporting data correction is provided, including the following steps:
[0071] S1: Users input questions into the question-answering system through an interface, while a context-aware mechanism is introduced to enable the question-answering system to understand the background of the question and the user's intent. In S1, the question-answering system can save the user's dialogue history, allowing the system to understand the background of new questions and the user's intent based on the context-aware mechanism and the context content. In this way, in multi-turn dialogues, the question-answering system can adjust new answers based on previous responses.
[0072] Based on context awareness and understanding of the background and user intent of new questions using the LSTM model, the computation steps are as follows:
[0073] For each time step t, the basic update equation describes how to update the hidden state at the current time step based on the state of the previous time step (i.e., the previous dialogue content) and the current input (i.e., the user's latest sentence), thereby preserving the coherence of the dialogue. The expression is:
[0074] h t =σ(W h ·[h t-1 x t ]+b h )
[0075] Among them, h t The hidden state at the current time step t; W h The weight matrix; [h t-1 x t [] is a combined vector, including the hidden state h from the previous time step. t-1 and the input x at the current time step t b h σ is the bias term; σ is the activation function;
[0076] The forget gate determines which past information is no longer important and should be "forgotten" from the cell state. This is crucial for avoiding the vanishing gradient problem and ensuring that ancient information is properly preserved or discarded. The expression is:
[0077] f t =σ(W f ·[h t-1 x t ]+b f )
[0078] Among them, f t For time step t; W f • is the weight matrix of the forget gate; b f For the bias term of the forget gate;
[0079] Input Gate:
[0080] i t =σ(W i ·[h t-1 x t ]+b i )
[0081] Among them, i t The output of the input gate at time step t; W i b is the weight matrix of the input gate; i This is the bias term for the input gate;
[0082] Candidate cell status:
[0083]
[0084] in, For candidate cell states at time step t; W c b is the weight matrix related to the candidate cell state; c is the bias term for candidate cell states; tanh is the hyperbolic tangent activation function;
[0085] Cell status update:
[0086]
[0087] Among them, c t f represents the final cell state at time step t; t This represents the output of the forget gate; ⊙ represents element-wise multiplication; c t-1 The cell state at time step t-1; i t The output of the input gate;
[0088] The input gate and candidate cell states together determine which new information should be added to the current cell state. This allows the model to selectively absorb new knowledge without worrying about indiscriminately overwriting old information;
[0089] The output gate controls which cell state information should be used to form the hidden state at the current time step, thus influencing the final output decision. This means the model can selectively reveal its internal memory, enhancing the relevance and accuracy of the output, as expressed in:
[0090] o t =σ(W o ·[h t-1 x t ]+b o )
[0091] Among them, o t The output of the gate at time step t; W o b is the weight matrix of the output gate; o This is the bias term for the output gate;
[0092] Hidden state:
[0093] h t =o t ⊙tanh(c t )
[0095] S2: Pass the user's question to the large model, use the large model to generate an initial answer, and extract information related to the question from relevant external data sources; the specific steps are as follows:
[0096] S21: Preprocess the user-input question to understand the intent, keywords, and topic of the question, and pass the preprocessed question to the large model. This can distinguish between questions of different natures such as inquiries, commands, and exclamations, identify the core words or phrases in the question, determine the domain or topic of the question based on keywords and context, provide direction for subsequent steps, and finally transform the question into a format suitable for the large model to understand and pass it to the large model.
[0097] S22: The large model analysis initially answers the questions and identifies the specific information needed from external data sources; this mainly tests the model's basic understanding of the problem.
[0098] S23: Select appropriate data sources based on information needs; based on the information needs identified by the large model, select the external data sources most likely to contain the required data, such as databases, knowledge graphs, and web search engines, and choose sources that can provide information quickly and accurately, balancing query speed and data reliability.
[0099] S24: Construct query statements and interface requests to execute queries against the selected data source; mainly based on the information requirements indicated by the large model, construct highly targeted query statements to ensure the relevance of query results, establish a connection with the selected data source through API or SDK, send the constructed query instructions, and obtain real-time or stored data.
[0100] S25: Collect query results, filter and integrate them, and extract key information relevant to the question. This mainly involves filtering out irrelevant or duplicate information from the query results, retaining only the key parts directly related to the question, combining the extracted information with the preliminary answer from the larger model to form a more complete and accurate answer, and organizing and integrating the information to make it easy for users to understand, which may include structured presentation, embedded charts, etc.
[0101] S3: Based on the initial responses and external data, perform data correction and fusion, and introduce algorithms to detect and filter misleading information; the specific steps are as follows:
[0102] S31: Based on the chi-square test, key information in the initial response is compared with key information related to the question collected from external data to identify differences and errors. Keyword matching or more sophisticated natural language processing techniques can be used to find relevant information points. Keyword matching is used to quickly locate relevant information, while natural language processing techniques are employed to deeply understand the context, ensuring the accuracy and depth of the comparison. This helps identify potential errors or omissions in the initial response. Simultaneously, the chi-square test, a statistical method, systematically compares the initial response with key information from external data sources, aiming to discover inconsistencies between the two. This method can quantify the significance of information differences.
[0103] S32: When there is inaccurate or outdated information in the initial response, correct it with external data. If the data source is reliable and there are obvious errors in the response, replace it directly with external data.
[0104] S33: When the initial response lacks detail or is incomplete, external data is used to supplement it, adding relevant details or additional information to ensure the completeness of the response. Before replacing information, the reliability of the external data source is assessed, and only data from authoritative or credible channels is adopted to avoid introducing new errors and ensure the accuracy and timeliness of the response.
[0105] The calculation steps for S31 are as follows:
[0106] Assume there is no difference in the categorical distribution between the initial response and the external data;
[0107] Create a cross table where the horizontal rows represent the categories in the initial responses, the vertical columns represent the categories in the external data, and the values in the cells represent the frequency of each category.
[0108] The expression for the chi-square statistic is:
[0109]
[0110] Where, χ 2 This is the chi-square statistic; r is the number of rows; c is the number of columns; O ij E represents the observation frequency. ij The expected frequency corresponding to the observed frequency is the theoretical frequency that each cell should have, based on the assumption that the initial answer and the external data classification distribution are indistinguishable. It is based on the proportional distribution of the row and column sums and can provide a benchmark for comparing the actual observed frequencies, thereby judging the degree of agreement between the observed values and the hypothetical model; i = 1, 2, ..., r; j = 1, 2, ..., c;
[0111] The chi-square statistic quantifies the degree of deviation between observed frequencies and expected frequencies. If the initial response does indeed show no difference in classification distribution from external data, theoretically the observed frequencies should be close to the expected frequencies. χ² 2 The value should be small. Conversely, if there is a significant difference between the two, the χ² value should be large. 2 The value will increase, by calculating χ. 2 The value can quantify the magnitude of this deviation, thereby assessing whether the difference is merely due to random fluctuations or reflects a real distributional difference;
[0112] The expression for degrees of freedom is:
[0113] df = (r-1)(c-1)
[0114] Where df represents the degrees of freedom, and the degrees of freedom reflect the degree of freedom in calculating χ. 2 Independent information content in a statistic refers to the number of frequencies that can vary independently after a certain number of observation frequencies have been determined. Degrees of freedom are a key parameter for finding the chi-square distribution table to determine the p-value; different degrees of freedom correspond to different chi-square distribution curves, which determine how we interpret χ². 2 value;
[0115] Based on the degrees of freedom and chi-square statistic, find the corresponding P-value in the chi-square distribution table, and use the P-value to determine the difference between the two data sources. The P-value is used in hypothesis testing to measure the observed χ² value. 2The probability of a value occurring under the assumed distribution is calculated. If the P-value is small, it indicates that the observed difference is unlikely to be caused by random factors. Based on the pre-set significance level, if the P-value is less than the level, the null hypothesis is rejected, that is, the preliminary answer is that there is no difference from the classification distribution of external data, and it is considered that there is a significant difference between the two; otherwise, the null hypothesis is accepted, and it is considered that the difference is not significant.
[0116] In S32, the specific steps for correction using external data are as follows:
[0117] S321: Check the credibility and professionalism of the data publishing organization, and check whether there are peer reviews or official certifications as quality assurance to ensure the reliability of the data source;
[0118] S322: Compare external data with data from other reliable sources to see if they are consistent. Use descriptive statistical analysis to determine the reasonableness of the data and ensure data consistency.
[0119] S323: Confirm whether the definition and classification of external data match the context in the initial response to ensure the applicability of the data;
[0120] S324: When it is confirmed that external data is more accurate and directly relevant, directly replace the erroneous information in the initial response and add a note at the replacement point.
[0121] Furthermore, the specific steps of S33 are as follows:
[0122] S331: Identify which parts of the initial answer lack detail or are incomplete, and determine the specific types and scope of information that need to be supplemented based on the needs of the question; be able to determine the types and scope of supplementary information to provide a clear direction for subsequent steps;
[0123] S332: Transform external data into a format that matches the initial response, ensuring consistency in units and timeframes; avoid information conflicts and misunderstandings. Consistent data formats help improve the accuracy and professionalism of the response, making it easier for readers to understand and accept.
[0124] S333: Precisely inserting details from external data and clearly citing data sources in the initial response enhances the depth and credibility of the answer. Citing data sources increases the transparency of the response, allowing readers to trace the reliability of the information, and also demonstrates respect for original sources and academic integrity.
[0125] S334: Ensure the supplementary information flows logically with the original content, and optimize the wording of the supplementary section. After adding information, ensure the overall answer remains logically coherent, avoiding abrupt or chaotic information piling up; through word optimization, seamlessly connect the new content with the original answer, improving the reading experience. Refining the language ensures the answer is not only rich in content but also clearly, fluently, and easy to understand.
[0126] S4: Combine the initial response with the revised external data to generate the final response;
[0127] S5: Design a user interface that allows users to evaluate the accuracy of question-and-answer results or provide direct correction suggestions. Automatically analyze user feedback, identify error types, and adjust model parameters or supplement training data based on the error analysis results. Automatic user feedback analysis in S5 includes analyzing keywords, sentiment, and contextual features in the feedback content to predict error types. Methods for adjusting model parameters based on error analysis results include supplementing training data and active learning. Supplementing training data involves integrating user-provided correction suggestions into the training dataset, especially for frequently occurring error types, to enhance the model's learning of these situations. Active selection involves choosing the most representative or uncertain user feedback, manually reviewing and labeling it, and using iterative training to improve model adaptability and accuracy.
[0128] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A large-scale intelligent question-answering method supporting data correction, characterized in that: Includes the following steps: S1: Users input questions into the question-answering system through an interface, while a context-aware mechanism is introduced to enable the question-answering system to understand the background of the question and the user's intent; S2: Pass the user's question to the large model, use the large model to generate an initial answer, and extract information related to the question from relevant external data sources; S3: Based on the initial responses and external data, perform data correction and fusion, and introduce algorithms to detect and filter misleading information; S4: Combine the initial response with the revised external data to generate the final response; S5: Design a user interface that allows users to evaluate the accuracy of the question-and-answer results or provide direct suggestions for correction. Automatically analyze user feedback, identify error types, and adjust model parameters or supplement training data based on the error analysis results.
2. The intelligent question-answering method for large models supporting data correction according to claim 1, characterized in that: In S1, the question-answering system can save the user's dialogue history, allowing the system to understand the background of new questions and the user's intent based on context awareness mechanisms and contextual content. In this way, in multi-turn dialogues, the question-answering system can adjust new answers based on previous responses.
3. The intelligent question-answering method for large models supporting data correction according to claim 2, characterized in that: Based on context awareness and understanding of the background and user intent of new questions using the LSTM model, the computation steps are as follows: For each time step t, the basic update equation is as follows: h t =σ(W h ·[h t-1 ,x t ]+b h ) Among them, h t The hidden state at the current time step t; W h The weight matrix; [h t-1 x t [] is a combined vector, including the hidden state h from the previous time step. t-1 and the input x at the current time step t b h σ is the bias term; σ is the activation function; Forgotten Gate: f t =σ(W f ·[h t-1 ,x t ]+b f ) Among them, f t For time step t; W f • is the weight matrix of the forget gate; b f For the bias term of the forget gate; Input Gate: i t =σ(W i ·[h t-1 ,x t ]+b i ) Among them, i t The output of the input gate at time step t; W i b is the weight matrix of the input gate; i This is the bias term for the input gate; Candidate cell status: in, For candidate cell states at time step t; W c b is the weight matrix related to the candidate cell state; c is the bias term for candidate cell states; tanh is the hyperbolic tangent activation function; Cell status update: Among them, c t f represents the final cell state at time step t; t This represents the output of the forget gate; ⊙ represents element-wise multiplication; c t-1 The cell state at time step t-1; i t The output of the input gate; Output gate: the t =σ(W o ·[h t-1 ,x t ]+b o ) Among them, o t The output of the gate at time step t; W o D is the weight matrix of the output gate; o This is the bias term for the output gate; Hidden state: h t =o t ⊙tanh(c t )。 4. The intelligent question-answering method for large models supporting data correction according to claim 3, characterized in that: The specific steps of S2 are as follows: S21: Preprocess the user-input questions to understand the intent, keywords, and topics of the questions, and then pass the preprocessed questions to the large model; S22: The initial answers to the large model analysis and the specific information that needs to be obtained from external data sources; S23: Select an appropriate data source based on information needs; S24: Construct the query statement and interface request, and execute the query against the selected data source; S25: Collect the query results, filter and integrate them, and extract key information related to the question.
5. The intelligent question-answering method for large models supporting data correction according to claim 4, characterized in that: The specific steps for S3 are as follows: S31: Based on the chi-square test, compare the key information in the preliminary answer with the key information related to the question collected from external data to find differences and errors; S32: When there is inaccurate or outdated information in the initial response, correct it with external data. If the data source is reliable and there are obvious errors in the response, replace it directly with external data. S33: When the initial answer lacks detail or is incomplete, supplement it with external data, add relevant details or additional information to ensure the completeness of the answer.
6. The intelligent question-answering method for large models supporting data correction according to claim 5, characterized in that: The calculation steps for S31 are as follows: Assume there is no difference in the categorical distribution between the initial response and the external data; Create a cross table where the horizontal rows represent the categories in the initial responses, the vertical columns represent the categories in the external data, and the values in the cells represent the frequency of each category. The expression for the chi-square statistic is: Where, χ 2 This is the chi-square statistic; r is the number of rows; c is the number of columns; O ij E represents the observation frequency. ij Let i be the expected frequency corresponding to the observed frequency; i = 1, 2, ..., r; j = 1, 2, ..., c; The expression for degrees of freedom is: df = (r-1)(c-1) Where df represents the degrees of freedom; Based on the degrees of freedom and chi-square statistic, find the corresponding P-value in the chi-square distribution table, and determine the difference between the two data sources based on the P-value.
7. The intelligent question-answering method for large models supporting data correction according to claim 6, characterized in that: In S32, the specific steps for correction using external data are as follows: S321: Check the credibility and professionalism of the data publishing organization, and check whether there are peer reviews or official certifications as quality assurance to ensure the reliability of the data source; S322: Compare external data with data from other reliable sources to see if they are consistent, and use descriptive statistical analysis to determine the reasonableness of the data; S323: Confirm whether the definition and classification of the external data match the context in the initial response; S324: When it is confirmed that external data is more accurate and directly relevant, directly replace the erroneous information in the initial response and add a note at the replacement point.
8. The intelligent question-answering method for large models supporting data correction according to claim 7, characterized in that: The specific steps of S33 are as follows: S331: Identify which parts of the preliminary answer lack details or are incomplete, and determine the specific types and scope of information that need to be supplemented based on the needs of the question; S332: Transform external data into a format that matches the initial response, ensuring consistency in units and timeframes; S333: In the initial response, accurately insert details provided by external data and indicate the data source; S334: Ensure that the supplementary information is logically consistent with the original content, and optimize the text of the supplementary part.
9. The intelligent question-answering method for large models supporting data correction according to claim 8, characterized in that: S5 automatically analyzes user feedback by analyzing keywords, sentiment, and contextual features in the feedback content to predict the type of error.
10. The intelligent question-answering method for large models supporting data correction according to claim 9, characterized in that: In S5, methods for adjusting model parameters based on error analysis results include supplementing training data and active learning.
Citation Information
Patent Citations
A web-based method for assessing the reliability of answers to community question and answer websites
CN109492076A
Dialogue model interaction method and device based on multiple tenants and storage medium
CN116881429A
Answer feedback method and device applied to large language model
CN117708292A
Dialogue response method, device and equipment of insurance intelligent question-answering system and medium
CN117709358A
Intelligent dialogue generation method, device, computer apparatus and computer storage medium
WO2021072875A1
Cited By
Question and answer method and device based on artificial intelligence
CN121480528A
Data generation method and electronic equipment
CN121525766A