GPT illusion relieving method and system based on cooperative supervision of intelligent agent and block chain

Through the collaborative supervision method of agents and blockchain, the problem of inaccurate and insufficient transparency of pre-trained big models is solved, and the reliability and transparency of question-and-answer are achieved, and the user experience and data quality are improved.

CN120278260APending Publication Date: 2025-07-08SHAANXI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510002399.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing pre-trained models such as ChatGPT have data accuracy and transparency problems when generating answers, which makes it difficult to explain its decision-making process, resulting in inaccurate and difficult to correct.

Method used

By introducing a method of collaborative supervision between agents and blockchain, a large-language model question-and-answer model collaborative system is built, and the Q&A supervision system with internal and external coordination is used to match the Q&A template in the blockchain template library, and the reliability and transparency of the answers are ensured through a multi-level regulatory mechanism.

Benefits of technology

It improves the answer reliability and transparency of the output process of the large language model, ensures the accuracy of the question and answer and the fairness of the regulatory process, reduces the answer error rate and improves user autonomy and privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278260A_ABST
    Figure CN120278260A_ABST
Patent Text Reader

Abstract

The invention provides a GPT illusion relieving method and system based on agent and block chain cooperative supervision, and the method comprises the steps: obtaining a user question, and finding out the most matched K question and answer templates in a block chain template library through an Embedding technology and a search recall sorting algorithm; if the templates are found, the integration agent integrates the templates into final question and answer templates and outputs the final question and answer templates to the user; if not, the preliminary question and answer agent generates a candidate question and answer template and evaluates the credibility of the candidate question and answer template; and the user selects whether to supervise according to the credibility. And if supervision is selected, supervising the candidate question and answer template by an external supervision mechanism according to a specified process. And after the assessment and supervision are passed, marking the candidate question and answer template as a final question and answer template by the auditing agent and the integration agent, recording the final question and answer template in a block chain template library, outputting the final question and answer template to the user, and recording a supervision result on a block chain for subsequent fine adjustment of the preliminary question and answer agent and the assessment agent. The reliability of answers to questions produced by a large model and the transparency of a decision-making process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a GPT hallucination mitigation method and system based on the collaborative supervision of agents and blockchain. Background Art

[0002] Recently, ChatGPT, an intelligent dialogue system based on pre-trained large models launched by the American artificial intelligence company OpenAI, has set off a frenzy across the network. In the process of use, some problems have emerged in ChatGPT and similar pre-trained large model technologies, and their data accuracy, model behavior transparency, and ethical compliance have become important concerns. The training data of large models may contain biases or errors, and large models will forcibly give answers even to unfamiliar questions, which results in inaccurate generated results and reduces the reliability of large models. Moreover, these large models usually work like "black boxes", making their processing processes and decision-making logics difficult to explain. If the decision-making process of large models is not transparent, it will be very difficult to determine responsibilities and take corrective measures when problems occur. Summary of the Invention

[0003] In order to overcome the above problems existing in the prior art, the purpose of the present invention is to provide a GPT hallucination mitigation method and system based on the collaborative supervision of agents and blockchain. By introducing an internal agent collaboration mode and an external blockchain during the application, fine-tuning, and output of large models, an internal and external collaborative question-and-answer supervision method for large language models can be constructed, which can improve the reliability of the answers generated by large language models, the transparency of the output process, and the accuracy of subsequent answers.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0005] A GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain includes the following steps;

[0006] Step 1: Obtain the user's question;

[0007] Step 2: Obtain the question sentence vector of the user's question and the question sentence vectors of each Q&A template in the blockchain template library through Embedding technology, and store the question sentence vectors in the Q&A template on the blockchain;

[0008] Step 3: Through the search recall ranking algorithm, if the K most matching Q&A templates can be found in the blockchain template library, the integration agent driven by the large model in the internal supervision company of the large model driven by the agent will organize these K Q&A templates into a final Q&A template and output it to the user; through the exact matching of the templates, the internal supervision company provides the first layer of protection, effectively reducing the error rate of answers caused by data loss;

[0009] Step 4: If the most matching K Q&A templates cannot be found in the blockchain template library, the large model-driven preliminary Q&A agent generates a preliminary answer based on the question, and integrates the user question and the preliminary answer into a candidate Q&A template and sends it to the large model-driven evaluation agent;

[0010] Step 5: The large model-driven evaluation agent evaluates the credibility of the candidate Q&A template from three dimensions: accuracy, consistency, and reliability. If the evaluation result is credible, or the evaluation result is generally credible or not credible and the user does not choose supervision, the large model-driven review agent marks the candidate Q&A template as the final Q&A template; the large model-driven integration agent then records the final Q&A template in the blockchain template library and outputs it to the user; the evaluation results include three cases: credible, generally credible, and not credible; the evaluation mechanism provides an early warning for hallucination problems through detailed multi-dimensional evaluation, and at the same time provides users with additional options to ensure the credibility of the answer;

[0011] Step 6: The large model-driven integration agent then records the final Q&A template in the blockchain template library and outputs it to the user;

[0012] Step 7: If the evaluation result is generally credible or not credible, and the user chooses supervision, the candidate Q&A template is sent to an external supervision agency; the external supervision agency coordinates the participation of candidate supervisors through a blockchain-driven distributed network to avoid subjectivity problems that may be caused by a single decision-making entity;

[0013] Step 8: A part of the candidate supervisors are selected as official supervisors by simple random sampling, not exceeding N members; the N is the maximum number of predetermined official supervisors;

[0014] Step 9: The official supervisors review the candidate Q&A template. If most official supervisors judge it to be incorrect, the official supervisors provide correction suggestions. The large model-driven manager sorts out the suggestions and modifies the candidate Q&A template, and resubmits it to the official supervisors. The official supervisors review the modified candidate Q&A template again and judge again whether the modified candidate Q&A template is correct; if it is incorrect, another cycle is needed until the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3; the supervision process is anonymously reviewed by the official supervisors to further weaken the bias and error propagation that may be caused by hallucinated answers and strengthen the fairness and accuracy of the answer from the outside;

[0015] Step 10: If the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3, the large model-driven review agent marks the candidate Q&A template as the final Q&A template;

[0016] Step 11: The large model-driven integration agent records the final Q&A template in the blockchain template library and outputs it to the user, and also records the supervision result on the blockchain for subsequent fine-tuning of the preliminary Q&A agent and the evaluation agent.

[0017] In step 2, the blockchain template library is a collection of multiple final Q&A templates stored on the blockchain; the Q&A template contains the user's question and the corresponding answer.

[0018] According to the method of obtaining the user question sentence vector and the question sentence vectors of each Q&A template in the blockchain template library through Embedding technology, and storing the question sentence vectors in the Q&A template on the blockchain, specifically includes:

[0019] Embedding technology captures the semantic relationships between words. In the vector space, semantically similar words are represented by close vectors; Embedding can usually represent text data as dense vectors in lower dimensions.

[0020] The questions in the Q&A template are all transformed into question sentence vectors through Embedding technology and stored on the blockchain for subsequent rapid retrieval.

[0021] In step 3, the value of K is not a fixed value, but depends on the number of the most matching Q&A templates actually searched; the internal supervision company driven by the large model includes a large model-driven preliminary Q&A agent, a large model-driven evaluation agent, a large model-driven review agent, and a large model-driven integration agent.

[0022] The internal supervision company obtains the user's question and generates candidate Q&A templates. The large model-driven evaluation agent will conduct the first supervision, decide whether to conduct the second supervision by an external supervision agency according to the result of the first supervision, finally feedback the supervision result to the internal supervision company, and finally the internal supervision company gives the user the final answer.

[0023] In step 6, the finally output Q&A result is usually presented to the user in a clear and easy-to-understand text form. These texts not only answer the user's question, but also come with some recommendations, tips or further reference information. For example, the system can directly display "The symptoms of high blood pressure are usually not obvious, but if symptoms occur, they may include headache, dizziness, palpitation, tinnitus, blurred vision, etc. Especially when the blood pressure level is high, you may feel a dull pain in the head or a blackout in front of your eyes. We recommend that you regularly monitor your blood pressure and seek medical attention in a timely manner when these symptoms appear for early intervention and treatment."

[0024] The external supervision agencies in step 7 include:

[0025] External regulatory agencies include online blockchain users, candidate supervisors, and official supervisors; all supervisors participating in the supervision and the resulting supervision processes are recorded on the blockchain;

[0026] Online blockchain users in the external regulatory agency choose whether to participate in the supervision of this candidate Q&A template. Online blockchain users who agree to participate will become candidate supervisors; the external regulatory agency includes online blockchain users, candidate supervisors, and official supervisors.

[0027] Step 8 is specifically as follows: N is the maximum number of predetermined official supervisors, and users should reach an agreement on the maximum N of official supervisors before selecting official supervisors.

[0028] Furthermore, in the supervision process, the Delphi method is used to execute group decision-making, specifically including:

[0029] The Delphi method is a group decision-making behavior, characterized by anonymity, feedback, and statistics. In essence, it is based on the professional knowledge, experience, and subjective judgment ability of numerous official supervisors;

[0030] The Delphi method adopts the form of a mailed inquiry. According to the systematic procedure, it uses the method of expressing opinions anonymously. Through multiple rounds of surveys on the views of official supervisors on issues related to the Q&A template, after repeated consultations, inductions, modifications, and technical processing, the finally summarized views that are basically consistent among official supervisors are used as the final result.

[0031] The fine-tuning process of the preliminary Q&A agent and evaluation agent in step 11 includes:

[0032] Obtain a fine-tuning dataset; the fine-tuning dataset contains multiple final Q&A templates and review results, and both the final Q&A templates and review results are stored on the blockchain to prevent tampering;

[0033] Use the fine-tuning dataset to iteratively fine-tune the preliminarily trained large model to obtain a preliminarily fine-tuned model;

[0034] Obtain a validation dataset; the validation dataset includes multiple final Q&A templates and review results;

[0035] Use the validation dataset to evaluate and validate the preliminarily fine-tuned model to adjust the model parameters and improve the accuracy and robustness of the model;

[0036] After multiple iterations of fine-tuning and validation, to obtain the preliminary Q&A agent and evaluation agent with optimized performance;

[0037] This fine-tuning process ensures the performance of the large model in a specific field and improves the accuracy and relevance of the dialogue.

[0038] GPT Hallucination Mitigation System Based on the Collaborative Supervision of Agents and Blockchains. The GPT Hallucination Mitigation System based on the collaborative supervision of agents and blockchains includes:

[0039] A data acquisition module for acquiring user questions;

[0040] A vector acquisition module, connected to the data acquisition module, for obtaining the user question sentence vector and the question sentence vectors of each Q&A template in the blockchain template library through Embedding technology, and storing the question sentence vectors in the Q&A templates on the blockchain; the blockchain template library is a collection of multiple final Q&A templates stored on the blockchain; the final Q&A templates are obtained after the candidate Q&A templates are supervised, and each Q&A template contains a user question and a corresponding answer;

[0041] A Q&A template matching module, connected to the vector acquisition module, for finding the top K most matching Q&A templates in the blockchain template library through a search recall ranking algorithm; the value of K is not a fixed value, but depends on the number of the most matching Q&A templates actually searched;

[0042] A first answer determination module, connected to the Q&A template matching module, for when the top K most matching Q&A templates can be found, organizing these K Q&A templates into a final Q&A template, which is the responsibility of an integration agent driven by a large model, and outputting the result to the user;

[0043] A candidate Q&A template determination module, connected to the Q&A template matching module, for when the top K most matching Q&A templates cannot be found in the blockchain template library, a preliminary Q&A agent driven by a large model generates a preliminary answer according to the question, and integrates the user question and the preliminary answer into a candidate Q&A template and sends it to an evaluation agent driven by a large model;

[0044] A candidate Q&A template evaluation module, connected to the candidate Q&A template determination module, for allowing an evaluation agent driven by a large model to evaluate whether the candidate Q&A template is trustworthy, generally trustworthy or untrustworthy, and the user selects whether to supervise;

[0045] A second answer determination module, connected to the candidate Q&A template evaluation module, for when the evaluation result is trustworthy, or the evaluation result is generally trustworthy or untrustworthy and the user does not choose to supervise, an audit agent driven by a large model marks the candidate Q&A template as a final Q&A template; the integration agent driven by a large model then records the final Q&A template in the blockchain template library and outputs it to the user;

[0046] The candidate supervisor determination module, connected to the candidate Q&A template evaluation module, is used to send the candidate Q&A template to an external regulatory agency when the evaluation result is generally credible or not credible and the user selects supervision; online blockchain users in the external regulatory agency select whether to participate in the supervision of this candidate Q&A template, and the online blockchain users who agree to participate will become candidate supervisors;

[0047] The official supervisor determination module, connected to the candidate supervisor determination module, is used to elect official supervisors. A part of the candidate supervisors are selected as official supervisors by simple random sampling, not exceeding N members; the N is the maximum number of predetermined official supervisors;

[0048] The candidate Q&A template supervision module, connected to the official supervisor determination module, is used for official supervisors to review the candidate Q&A template. If most official supervisors judge it to be incorrect, the official supervisors provide correction suggestions. The large model-driven review agent sorts out the suggestions and modifies the candidate Q&A template, and resubmits it to the official supervisors. The official supervisors review the modified candidate Q&A template again and judge again whether the modified candidate Q&A template is correct; if it is incorrect, another cycle is required until the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3;

[0049] The third answer determination module, connected to the candidate Q&A template supervision module, is used when the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3, the large model-driven review agent marks the candidate Q&A template as the final Q&A template; the large model-driven integration agent then records the final Q&A template in the blockchain template library and outputs it to the user, and also records the supervision result on the blockchain for subsequent fine-tuning of the preliminary Q&A agent and evaluation agent.

[0050] The beneficial effects of the present invention:

[0051] The user questions obtained in the present invention are transformed into user question vectors by Embedding, and then the most matching K Q&A templates are found in the blockchain template library through a search, recall, and ranking algorithm, improving the efficiency and accuracy of finding matching Q&A templates. A blockchain template library is constructed, and the Q&A templates are recorded on the blockchain, improving the transparency and immutability of the Q&A templates. The blockchain template library creates a trustworthy, publicly transparent long-term memory database for the base large model, and converts the authoritative and trustworthy Q&A templates after supervision into large model neuron vectors. A blockchain-driven external regulatory agency is established to perform reliable manual supervision and output quality improvement using distributed intelligence; simple random sampling is used to prevent the voting results from being biased due to the over-concentration of opinions of the same group, making the voting results more universal and improving the quality of the data; the supervision stage is based on the Nash equilibrium and the Delphi method, enhancing the fairness and integrity in the process of supervisors' evaluation and decision-making. The large model-driven preliminary Q&A agent, evaluation agent, review agent, and integration agent participate in the Q&A process, achieving the consistency of the output results. By introducing a blockchain-driven external regulatory agency to provide trustworthy feedback on the untrustworthy behaviors of the large model, the output of the existing large model is improved, realizing a regulatory intelligent dialogue based on the blockchain template library. The final Q&A templates and review results are used to train the evaluation agent and the preliminary Q&A agent, continuously learning from new data and user feedback to improve the credibility of the output and the accuracy of self-evaluation, and further improving the reliability of the problem answers produced by the large model and the transparency of the decision-making process. Description of the Drawings

[0052] Figure 1 It is a flowchart of the method for alleviating large language model hallucinations through the collaborative supervision of a large model and a blockchain provided by the present invention.

[0053] Figure 2 It is a schematic diagram of the Q&A template and the blockchain template.

[0054] Figure 3 It is a schematic diagram of the Q&A template transfer.

[0055] Figure 4 It is a schematic diagram of the system for alleviating large language model hallucinations through the collaborative supervision of a large model and a blockchain provided by the present invention. Detailed Embodiment

[0056] The present invention will be further described in detail below with reference to the accompanying drawings.

[0057] The objective of the present invention is to provide a method for alleviating large language model hallucinations through the collaborative supervision of agents and blockchains. By integrating agents and blockchains, an innovative internal and external collaborative supervision framework is constructed to enhance the reliability and security of large models. Additionally, blockchain technology is utilized to record all decision-making processes and activities of large models, providing transparency for external regulatory agencies to view and supervise. The evaluation agents and preliminary question-and-answer agents driven by large models continuously learn from new data and user feedback to improve the credibility of the output and the accuracy of the evaluation.

[0058] To make the above objectives, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] Example 1: As Figure 1 shown, this embodiment provides a method for alleviating large language model hallucinations through the collaborative supervision of agents and blockchains, including the following steps.

[0060] Step 1: The user actively poses a question in the system to obtain the user's question.

[0061] Step 2: Use Embedding technology to obtain the question sentence vectors of the user's question and each question-and-answer template in the blockchain template library, and store the question sentence vectors in the question-and-answer templates on the blockchain.

[0062] In this embodiment, as Figure 2 shown, the blockchain template library is a collection of multiple final question-and-answer templates stored on the blockchain. The final question-and-answer templates are obtained after the candidate question-and-answer templates are supervised. Each question-and-answer template contains the user's question and the corresponding answer. The candidate question-and-answer templates and the supervision process will be explained in steps 5 and 8 respectively.

[0063] Step 2 maps the user's question to the question sentence vector space in the blockchain template library, and question sentences that are semantically similar will be mapped to positions that are closer in space.

[0064] Compared with traditional methods such as the bag-of-words model and TF-IDF, Embedding can capture the semantic relationships between words. In the vector space, words that are semantically similar will be represented by closer vectors. Embedding can usually represent text data as lower-dimensional dense vectors, which helps to reduce the consumption of computing resources. In addition, the question sentence vectors in the question-and-answer templates are stored on the blockchain for subsequent rapid retrieval, which not only reduces the generation time of the sentence vectors, increases the efficiency of retrieving question-and-answer templates with a higher matching degree to the user's question, but also reduces the computing cost.

[0065] Step 2 uses the BERT model to obtain the sentence vector representation of the input question. BERT generates context-aware representations for each word and sentence through bidirectional encoding of the context, and can capture the deep connections between semantics. Further, the obtaining of the sentence vector representation of the input question includes the following steps.

[0066] (1) Tokenize the input question Q = "What are the symptoms of hypertension?" to obtain the corresponding sub-word sequence. For example:

[0067] Q = [CLS], hypertension, of, symptoms, have, which,?, [SEP];

[0068] Among them, [CLS] is the start flag of BERT, and [SEP] is the sentence separator. BERT will generate a vector representation for each sub-word.

[0069] (2) Use the pre-trained BERT model to convert each sub-word into a context-aware vector representation. Through the bidirectional encoder of BERT, for each input vocabulary, it will consider the left and right contexts of the word to obtain a context-aware word vector. Specifically, BERT converts each sub-word into a vector v1, v2,..., vn, and each vi is the representation of the corresponding sub-word.

[0070] (3) The first sub-word ([CLS]) of the BERT model is used to represent the vector representation of the whole sentence. Therefore, the sentence vector VQ can directly take the vector at the first position in the BERT output: VQ = BERT_output[0]. The finally obtained sentence vector VQ will be the high-dimensional semantic representation of the question.

[0071] Step 3: Through the search recall ranking algorithm, if the K most matching Q&A templates can be found in the blockchain template library, the large model-driven integration agent will organize these K Q&A templates into a final Q&A template and output it to the user. The value of K is not a fixed value, but depends on the number of the most matching Q&A templates actually searched.

[0072] In this embodiment, as Figure 3 shown, the blockchain-driven external regulatory agency includes online blockchain users, candidate supervisors, and official supervisors; the large model-driven intelligent agent internal regulatory company includes a large model-driven preliminary Q&A agent, a large model-driven evaluation agent, a large model-driven review agent, and a large model-driven integration agent; the internal regulatory company obtains user questions and generates candidate Q&A templates, and the large model-driven evaluation agent will conduct the first supervision, and decide whether to conduct the second supervision by the external regulatory agency according to the results of the first supervision, and finally feedback the supervision results to the internal regulatory company, and finally the internal regulatory company gives the user the final answer.

[0073] Step 3 is to compare the user question sentence vector with the question sentence vectors in each Q&A template through a search recall ranking algorithm. Once the question most relevant to the user question is found, the Q&A templates corresponding to these questions are extracted.

[0074] Step 4: If the top K most matching Q&A templates cannot be found in the blockchain template library, the large model-driven preliminary Q&A agent generates a preliminary answer based on the question and integrates the user question and the preliminary answer into a candidate Q&A template and sends it to the large model-driven evaluation agent.

[0075] Step 5: The large model-driven evaluation agent evaluates the credibility of the candidate Q&A template from three dimensions: accuracy, consistency, and reliability. If the user does not select external supervision or the evaluation agent deems the result credible, the large model-driven review agent marks the candidate Q&A template as the final Q&A template. The evaluation results include three categories: credible, generally credible, and non-credible.

[0076] Credible means meeting high standards in terms of accuracy, consistency, and reliability. Generally credible means meeting the standards in some indicators, but there may be uncertainties or fluctuations in other aspects. Non-credible means not meeting the standards in multiple key indicators or having obvious errors.

[0077] The evaluation agent determines the credibility of the Q&A template. The evaluation process and results are saved on the blockchain, and users can clearly know how the evaluation agent conducts the evaluation and the scoring basis for each answer. Although experiments show that the evaluation results of the large model-driven evaluation agent and experts are as similar as 95%, the evaluation agent still needs to rely on external regulatory agencies to ensure its credibility.

[0078] Step 6: The large model-driven integration agent then records the final Q&A template in the blockchain template library and outputs it to the user.

[0079] Step 7: If the user selects supervision and the evaluation agent deems the result generally credible or non-credible, the candidate Q&A template is sent to the blockchain-driven external regulatory agency. Online blockchain users in the external regulatory agency choose whether to participate in the supervision of this candidate Q&A template, and the online blockchain users who agree to participate will become candidate supervisors. The external regulatory agency includes online blockchain users, candidate supervisors, and official supervisors.

[0080] The user's supervision choice is closely related to their privacy protection. If users choose to participate in supervision, it means they hope to have additional control over the evaluation and review process to ensure the quality and privacy protection of each Q&A template. This choice gives users more autonomy, allowing them to participate in or supervise key steps, especially in scenarios involving sensitive information processing or high privacy requirements.

[0081] Step 8: Select a part of the candidate supervisors as official supervisors through simple random sampling, with no more than N members. The N is the maximum number of predetermined official supervisors. Before performing Step 8, the user should reach an agreement on the maximum N of the official supervisors. The higher the N, the more reliable the answers the user gets, but this will also result in higher supervision costs.

[0082] The purpose of simple random sampling is to prevent all online blockchain users who agree to participate from entering the supervision process, improve the voting efficiency, and reduce cost and resource consumption. Simple random sampling can ensure that different subgroups have the opportunity to be represented, prevent the voting results from being biased due to the over-concentration of opinions of the same group, so that the voting results are more universal, reduce self-selection bias and improve the quality of data.

[0083] Step 9: The official supervisors review the candidate Q&A templates. If most of the official supervisors judge it to be incorrect, the official supervisors provide correction suggestions. The large model-driven review agent sorts out the suggestions and modifies the candidate Q&A templates, and resubmits it to the official supervisors. The official supervisors review the modified candidate Q&A templates again and judge again whether the modified candidate Q&A templates are correct. If it is incorrect, another cycle is needed until the supervisors who judge the candidate Q&A templates to be correct account for more than 2 / 3. During the execution of Step 9, all judgment results and modification suggestions are publicly stored on the blockchain, facilitating the transparent sharing of supervision results and helping to evaluate the behavior of official supervisors later.

[0084] During the process of the official supervisors reviewing the candidate Q&A templates, the smart contract maintains the process based on the Nash equilibrium in game theory, ensuring that the supervisors execute tasks honestly and effectively, and also ensuring the fairness and transparency of the supervision process.

[0085] Further, the Nash equilibrium includes the following steps.

[0086] (1) In the case of n supervisors, it is usually represented in matrix form as a triple (SW, ∑, Π), where: SW = {a1, a2,..., a n} represents the set of strategies of n supervisors, and each strategy is designed for a specific type of candidate Q&A template.

[0087] (2) ∑ = ∑1 × ∑2 × … × ∑ n represents all possible strategy combinations, that is, each strategy ∑ k corresponds to a specific supervision task, and there are corresponding strategies r k that can be selected. The supervisors' strategy choices in each domain ∑ k can be σ k ∈∑k In addition, a strategy profile σ is defined as a vector where is the best strategy in ∑ k for (k = 1, 2…, n).

[0088] (3) Π = {π1, π2…, π n} represents a set of payoff functions, where π k : ∑ → R is the function that determines the payoff of the supervisor under a specific strategy a k for (k = 1, 2…, n), and R is the corresponding payoff.

[0089] (4) For a specific strategy σ k , if the supervisor finds that by choosing an alternative strategy σ' k to σ k while keeping all other strategies unchanged, a better result can be obtained (i.e., π k (σ' k , σ -k ) > π k ( σk , σ -k ))), then the supervisor will adjust its strategy to σ' k .

[0090] It should be noted that steps 7, 8, and 9 are all applications of the Delphi method. The Delphi method is a group decision-making behavior, characterized by anonymity, feedback, and statistics. Essentially, it is based on the professional knowledge, experience, and subjective judgment ability of many formal supervisors. The Delphi method adopts the form of a mailed questionnaire survey. According to a systematic procedure, it uses the method of anonymously expressing opinions. Through multiple rounds of surveys on the views of formal supervisors on the questions in the Q&A template, after repeated consultations, inductions, modifications, and technical processing, the views that are basically consistent among formal supervisors are finally summarized as the final result, giving full play to the role of information feedback and information control.

[0091] Step 10: If the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3, the large model-driven review agent will mark the candidate Q&A template as the final Q&A template.

[0092] Step 11: The large model-driven integration agent records the final Q&A template in the blockchain template library and outputs it to the user, and also records the supervision result on the blockchain. The template after blockchain decentralized intelligent auditing will be used for subsequent fine-tuning of the preliminary Q&A agent and evaluation agent.

[0093] Furthermore, the fine-tuning process of the preliminary Q&A agent and evaluation agent includes the following steps.

[0094] Obtain a fine-tuning data set; the fine-tuning data set contains multiple final Q&A templates and review results, and both the final Q&A templates and review results are stored on the blockchain transparently, publicly, and regulatoryly.

[0095] Use the fine-tuning data set to iteratively fine-tune the preliminarily trained large model to obtain a preliminarily fine-tuned model;

[0096] Obtain a validation data set; the validation data set includes multiple final Q&A templates and review results;

[0097] Use the validation data set to evaluate and validate the preliminarily fine-tuned model to adjust the model parameters and improve the accuracy and robustness of the model;

[0098] After multiple iterations of fine-tuning and validation, a preliminarily Q&A agent and an evaluation agent with optimized performance are obtained.

[0099] This fine-tuning process ensures the performance of the large model in a specific domain and improves the accuracy and relevance of the conversation.

[0100] The present invention realizes the transparency and reliability of the large model output by combining the method of agent collaboration and the regulatory process of blockchain technology, combining the advantages of both: using the internal collaboration mechanism of the agent to design appropriate prompt words to better reflect the knowledge involved in the large model training, and the blockchain technology constructs a blockchain template library and records the decision results of the supervisor, which can ensure the visibility inertia, transparency, and immutability of the decision-making process. The evaluation agent and the preliminary Q&A agent continuously learn from new data and user feedback to improve the credibility of the output and the accuracy of the evaluation. The decentralized intelligent regulatory process further ensures the reliability of the large model output results through the Delphi method and the Nash equilibrium.

[0101] The present invention can quickly and accurately output results in the presence of existing similar templates. During the process of searching for matching templates in the blockchain template library, Embedding technology and search recall ranking algorithms are introduced to convert both the user's question and the questions in the Q&A templates into question sentence vectors, enabling fast and accurate search in the vector space. When there are matching templates in the blockchain template library, answers can be quickly obtained through steps 1, 2, and 3, improving the answering efficiency and accuracy. The present invention can perform controllable intelligent answering based on the blockchain template library. The present invention supports operations such as searching, adding, and integrating the blockchain template library, and combines the understanding and generation capabilities of the large model to automatically achieve intelligent answering based on the blockchain template library. At the same time, the blockchain template library creates a trusted, publicly transparent long-term memory database for the basic large model, converting the authoritative and trusted Q&A templates after supervision into large model neuron vectors, realizing the transparency and immutability of the Q&A templates. The present invention continuously learns from new data and user feedback to improve the capabilities of the large model-driven evaluation agent and the preliminary Q&A agent. The large model-driven preliminary Q&A agent will continuously learn the final Q&A templates to improve the credibility of the output; the large model-driven evaluation agent will continuously learn the supervision process and results of external regulatory agencies to improve the accuracy of evaluation. These training materials of the large model are all recorded on the blockchain, ensuring the immutability and transparency of the information, thus guaranteeing the credibility of the output results and enhancing the comprehensive strength of the entire large model company.

[0102] The present invention utilizes external regulatory agencies to further ensure the reliability of the output results of the large model. Regulatory agencies are introduced to provide credible feedback on the untrustworthy behaviors of the large model, and distributed intelligence is used for reliable manual supervision and output quality improvement. Step 8 ensures the reliability of the results through simple random sampling, and step 9 is based on the Nash equilibrium, enhancing the fairness and integrity in the evaluation and decision-making process of the supervisors. The large model-driven review agent participates in the Q&A process, achieving the consistency of the output results. The present invention reduces the operating costs and improves the Q&A efficiency and quality. The present invention introduces a large model-driven intelligent agent for internal supervision of the company and a blockchain-driven external regulatory agency. The addition of the large model-driven intelligent agent improves the speed and quality of supervision, avoids redundant supervision, and helps users quickly and accurately obtain the answers to their questions, enhancing the efficiency and quality of customer service.

[0103] Embodiment 2: To execute the method corresponding to Embodiment 1 above to achieve the corresponding functions and technical effects, a GPT hallucination mitigation system based on the collaborative supervision of intelligent agents and blockchains is provided below.

[0104] As Figure 4As shown in the figure, the large model and blockchain-driven regulatory intelligent dialogue system provided in this embodiment includes: a data acquisition module 201, a vector acquisition module 202, a Q&A template matching module 203, a first answer determination module 204, a candidate Q&A template determination module 205, a candidate Q&A template evaluation module 206, a candidate supervisor determination module 207, an official supervisor determination module 207, a candidate Q&A template supervision module 209, and a second answer determination module 210.

[0105] Among them, the data acquisition module 201 is used to acquire user questions.

[0106] The vector acquisition module 202 is connected to the data acquisition module 201. The vector acquisition module 202 is used to obtain the user question sentence vector and the question sentence vector of each Q&A template in the blockchain template library through Embedding technology, and store the question sentence vector in the Q&A template on the blockchain. The blockchain template library is a collection of multiple final Q&A templates stored on the blockchain. The final Q&A template is obtained after being supervised by the candidate Q&A template. Each Q&A template contains a user question and the corresponding answer.

[0107] The Q&A template matching module 203 is connected to the vector acquisition module 202. The Q&A template matching module 203 is used to find the K most matching Q&A templates in the blockchain template library through a search recall ranking algorithm. The value of K is not a fixed value, but depends on the number of the most matching Q&A templates actually searched.

[0108] The first answer determination module 204 is connected to the Q&A template matching module 203. When the first answer determination module 204 can find the K most matching Q&A templates, it is responsible for organizing these K Q&A templates into a final Q&A template by an integration agent driven by a large model, and outputting the result to the user.

[0109] The candidate Q&A template determination module 205 is connected to the Q&A template matching module 203. When the candidate Q&A template determination module 205 cannot find the K most matching Q&A templates in the blockchain template library, a preliminary Q&A agent driven by a large model generates a preliminary answer according to the question, and integrates the user question and the preliminary answer into a candidate Q&A template and sends it to an evaluation agent driven by a large model.

[0110] The candidate Q&A template evaluation module 206 is connected to the candidate Q&A template determination module 205. The candidate Q&A template evaluation module 206 is used to let an evaluation agent driven by a large model evaluate whether the candidate Q&A template is trustworthy, generally trustworthy or untrustworthy, and the user is allowed to choose whether to supervise.

[0111] The second answer determination module 207 is connected to the candidate Q&A template evaluation module 206. When the evaluation result is credible, or the evaluation result is moderately credible or not credible and the user does not select supervision, the large model-driven review agent marks the candidate Q&A template as the final Q&A template. The large model-driven integration agent then records the final Q&A template in the blockchain template library and outputs it to the user.

[0112] The candidate supervisor determination module 208 is connected to the candidate Q&A template evaluation module 206. When the evaluation result is moderately credible or not credible and the user selects supervision, the candidate Q&A template is sent to an external supervision agency. Online blockchain users in the external supervision agency choose whether to participate in the supervision of this candidate Q&A template, and the online blockchain users who agree to participate will become candidate supervisors.

[0113] The official supervisor determination module 209 is connected to the candidate supervisor determination module 208. The official supervisor determination module 209 is used to elect official supervisors. A portion of the candidate supervisors are selected by simple random sampling to become official supervisors, with no more than N members. The N is the maximum number of predetermined official supervisors.

[0114] The candidate Q&A template supervision module 210 is connected to the official supervisor determination module 209. The candidate Q&A template supervision module 210 is used for official supervisors to review the candidate Q&A template. If the majority of official supervisors judge it to be incorrect, the official supervisors provide correction suggestions. The large model-driven review agent collates the suggestions and modifies the candidate Q&A template, and resubmits it to the official supervisors. The official supervisors review the modified candidate Q&A template again and again determine whether the modified candidate Q&A template is correct. If it is incorrect, another cycle is required until the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3.

[0115] The third answer determination module 211 is connected to the candidate Q&A template supervision module 210. When the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3, the large model-driven review agent marks the candidate Q&A template as the final Q&A template. The large model-driven integration agent then records the final Q&A template in the blockchain template library and outputs it to the user, and also records the supervision result on the blockchain for subsequent fine-tuning of the preliminary Q&A agent and evaluation agent.

[0116] Compared with the prior art, the beneficial effects of the large model and blockchain-driven supervised intelligent dialogue system provided in this embodiment are the same as those of the large model and blockchain-driven supervised intelligent dialogue method provided in Embodiment 1, and will not be elaborated here.

[0117] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0118] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. An approach for alleviating GPT hallucinations based on the collaborative supervision of agents and blockchain, characterized in that, It includes the following steps; Step 1: Obtain the user's question; Step 2: Obtain the sentence vector of the user's question and the sentence vectors of each Q&A template in the blockchain template library through Embedding technology, and store the sentence vectors of the questions in the Q&A templates on the blockchain; Step 3: Through the search recall ranking algorithm, if the most matching K Q&A templates can be found in the blockchain template library, the large model-driven integration agent within the company's large model-driven intelligent agent will organize these K Q&A templates into a final Q&A template and output it to the user; Step 4: If the most matching K Q&A templates cannot be found in the blockchain template library, the large model-driven preliminary Q&A agent will generate a preliminary answer based on the question, and integrate the user's question and the preliminary answer into a candidate Q&A template and send it to the large model-driven evaluation agent; Step 5: The large model-driven evaluation agent evaluates the credibility of the candidate Q&A template from three dimensions: accuracy, consistency, and reliability. If the evaluation result is credible, or the evaluation result is generally credible or not credible and the user does not choose supervision, the large model-driven review agent will mark the candidate Q&A template as the final Q&A template; The large model-driven integration agent will then record the final Q&A template in the blockchain template library and output it to the user; the evaluation results include three cases: credible, generally credible, and not credible; Step 6: The large model-driven integration agent will then record the final Q&A template in the blockchain template library and output it to the user; Step 7: If the evaluation result is generally credible or not credible, and the user chooses supervision, the candidate Q&A template will be sent to an external supervision agency; the external supervision agency coordinates the participation of candidate supervisors through a blockchain-driven distributed network; Step 8: Select a part of the candidate supervisors to become official supervisors through simple random sampling, not exceeding N members; N is the maximum number of predetermined official supervisors; Step 9: The official supervisors review the candidate Q&A template. If most official supervisors judge it to be incorrect, the official supervisors will provide correction suggestions. The large model-driven manager will organize the suggestions and modify the candidate Q&A template, and resubmit it to the official supervisors. The official supervisors will review the modified candidate Q&A template again and judge again whether the modified candidate Q&A template is correct; if it is incorrect, another cycle is required until the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3; the supervision process is anonymously reviewed by the official supervisors to further weaken the biases and error propagation that may be caused by hallucinated answers and strengthen the fairness and accuracy of the answers externally; Step 10: If the supervisors who judge the candidate Q&A template to be correct account for more than 2 / 3, the large model-driven review agent will mark the candidate Q&A template as the final Q&A template; Step 11: The large model-driven integration agent will then record the final Q&A template in the blockchain template library and output it to the user, and record the supervision result on the blockchain for subsequent fine-tuning of the preliminary Q&A agent and the evaluation agent.

2. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 1, wherein, In step 2, the blockchain template library is a collection of multiple final Q&A templates stored on the blockchain; The Q&A template contains a user question and the corresponding answer; Obtaining the user question sentence vector and the question sentence vectors of each Q&A template in the blockchain template library through Embedding technology and storing the question sentence vectors in the Q&A template on the blockchain specifically includes: Embedding technology captures the semantic relationships between words. In the vector space, semantically similar words are represented by close vectors; Embedding can usually represent text data as dense vectors in a lower dimension; The questions in the Q&A template are all transformed into question sentence vectors through Embedding technology and stored on the blockchain for subsequent rapid retrieval.

3. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 1, wherein, In step 3, K is not a fixed value but depends on the number of the most matching Q&A templates actually searched out; the internal supervision company driven by the large model includes a large model-driven preliminary Q&A agent, a large model-driven evaluation agent, a large model-driven review agent, and a large model-driven integration agent; The internal supervision company obtains the user question and generates candidate Q&A templates. The large model-driven evaluation agent will conduct the first supervision, decide whether to conduct the second supervision by an external supervision agency according to the results of the first supervision, finally feedback the supervision results to the internal supervision company, and finally the internal supervision company gives the user the final answer.

4. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 1, wherein In step 6, the finally output Q&A result is usually presented to the user in a clear and easy-to-understand text form. These texts not only answer the user's question but also come with some recommendations, tips, or further reference information.

5. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 1, wherein The external supervision agencies in step 7 include: The external supervision agency includes online blockchain users, candidate supervisors, and official supervisors; the supervisors participating in the supervision and the generated supervision processes are all recorded on the blockchain; The online blockchain users in the external supervision agency choose whether to participate in the supervision of this candidate Q&A template. The online blockchain users who agree to participate will become candidate supervisors; the external supervision agency includes online blockchain users, candidate supervisors, and official supervisors.

6. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 1, wherein Step 8 is specifically: N is the maximum number of predetermined official supervisors, and the user should reach an agreement on the maximum N of the official supervisors before selecting the official supervisors.

7. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 6, wherein, In the supervision process, the Delphi method is used to execute group decision-making, specifically including: The Delphi method uses the form of inquiry surveys. According to the systematic procedure, it adopts the way of expressing opinions anonymously. Through multiple rounds of surveys on the views of official supervisors on issues related to the Q&A template, after repeated consultation, induction, modification, and technical processing, finally, the views that are basically consistent among official supervisors are summarized as the final result.

8. The GPT hallucination mitigation method based on the collaborative supervision of agents and blockchain according to claim 1, wherein The fine-tuning process of the preliminary Q&A agent and the evaluation agent in step 11 includes: Obtaining a fine-tuning data set; the fine-tuning data set contains multiple final Q&A templates and review results, and both the final Q&A templates and the review results are stored on the blockchain to prevent tampering; Using the fine-tuning data set to perform iterative fine-tuning on the preliminarily trained large model to obtain a preliminary fine-tuning model; Obtain a validation dataset; the validation dataset includes multiple final Q&A templates and review results; Use the validation dataset to evaluate and validate the preliminary fine-tuning model to adjust the model parameters and improve the accuracy and robustness of the model; After multiple iterations of fine-tuning and validation, a preliminary Q&A agent and an evaluation agent with optimized performance are obtained; This fine-tuning process ensures the performance of the large model in a specific domain and improves the accuracy and relevance of the conversation.

9. A GPT hallucination mitigation system based on the collaborative supervision of agents and blockchain for implementing the method according to any one of claims 1-8, characterized in that, The GPT hallucination mitigation system for collaborative supervision by the intelligent agent and the blockchain includes: A data acquisition module is used to acquire user questions; A vector acquisition module, connected to the data acquisition module, is used to obtain the user question sentence vector and the question sentence vectors of each Q&A template in the blockchain template library through Embedding technology, and store the question sentence vectors in the Q&A template on the blockchain; the blockchain template library is a collection of multiple final Q&A templates stored on the blockchain; the final Q&A template is obtained after the candidate Q&A template is supervised, and each Q&A template contains a user question and the corresponding answer; A Q&A template matching module, connected to the vector acquisition module, is used to find the top K most matching Q&A templates in the blockchain template library through a search, recall, and ranking algorithm; the value of K is not a fixed value but depends on the number of the most matching Q&A templates actually searched; A first answer determination module, connected to the Q&A template matching module, is used to organize the top K most matching Q&A templates into a final Q&A template when the top K most matching Q&A templates can be found. This work is responsible for by an integration agent driven by a large model and outputs the result to the user; A candidate Q&A template determination module, connected to the Q&A template matching module, is used to, in the case where the top K most matching Q&A templates cannot be found in the blockchain template library, the preliminary Q&A agent driven by the large model generates a preliminary answer according to the question and integrates the user question and the preliminary answer into a candidate Q&A template and sends it to the evaluation agent driven by the large model; A candidate Q&A template evaluation module, connected to the candidate Q&A template determination module, is used to let the evaluation agent driven by the large model evaluate whether the candidate Q&A template is credible, generally credible, or not credible, and the user selects whether to supervise; A second answer determination module, connected to the candidate Q&A template evaluation module, is used to, when the evaluation result is credible, or the evaluation result is generally credible or not credible and the user does not choose to supervise, the review agent driven by the large model marks the candidate Q&A template as a final Q&A template; the integration agent driven by the large model then records the final Q&A template in the blockchain template library and outputs it to the user; A candidate supervisor determination module, connected to the candidate Q&A template evaluation module, is used to, when the evaluation result is generally credible or not credible and the user chooses to supervise, send the candidate Q&A template to an external regulatory agency; the online blockchain users in the external regulatory agency select whether to participate in the supervision of this candidate Q&A template, and the online blockchain users who agree to participate will become candidate supervisors; The official supervisor determination module, connected to the candidate supervisor determination module, is used to elect official supervisors. A portion of the candidate supervisors are selected as official supervisors through simple random sampling, with no more than N members; the N is the maximum number of predetermined official supervisors. The candidate Q&A template supervision module, connected to the official supervisor determination module, is used for official supervisors to review the candidate Q&A templates. If the majority of official supervisors judge them to be incorrect, the official supervisors provide correction suggestions. The large model-driven review agent organizes the suggestions and modifies the candidate Q&A templates, and then resubmits them to the official supervisors. The official supervisors review the modified candidate Q&A templates again and judge whether the modified candidate Q&A templates are correct again; if they are incorrect, another cycle is required until the number of supervisors who judge the candidate Q&A templates to be correct exceeds 2 / 3. The third answer determination module, connected to the candidate Q&A template supervision module, is used in the case where the number of supervisors who judge the candidate Q&A templates to be correct exceeds 2 / 3. The large model-driven review agent marks the candidate Q&A templates as the final Q&A templates; the large model-driven integration agent then records the final Q&A templates in the blockchain template library and outputs them to the user, and also records the supervision results on the blockchain for subsequent fine-tuning of the preliminary Q&A agent and evaluation agent.