Block chain-based large language model network generation optimization method
By adopting blockchain and Byzantine fault tolerance consensus mechanisms in multiple LLM networks, the problem of inability to provide efficient and trustworthy responses in the existing technology is solved, and the effect of automated collaborative work and high-quality responses is achieved.
Patent Information
- Application Number
- CN202510185259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-17
AI Technical Summary
The existing network built by multiple LLMs cannot automatically provide users with efficient and trustworthy responses, and fail to effectively prevent malicious operation of LLM devices.
Byzantine fault tolerance consensus mechanism based on blockchain is adopted, and the best response generated by all LLMs in the MLLMN is selected and packaged into blocks and linked to the blockchain for storage in a distributed manner.
The automated collaborative work of multiple LLMs has been achieved, which improves the credibility and quality of responses, prevents the impact of malicious operational behavior, and improves the efficiency of consensus processing.
Smart Images

Figure CN120163239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of artificial intelligence-generated content, and specifically to an optimization method for generating a large language model network based on blockchain. Background Art
[0002] Large language models (LLMs) have become the cornerstone of artificial intelligence (AI) and shown great potential in natural language understanding and generation tasks. These LLMs serve society in the way of AI-generated content and have been widely applied in various aspects of society, such as education, medical care, information technology, etc., due to their advanced expression and learning capabilities.
[0003] With the continuous development of this field, a large number of LLMs developed by different organizations and institutions have emerged, such as ChatGPT of OpenAI, Llama of Meta, and Wenxin Yiyan, iFlytek Spark, Doubao, Kimi, etc. developed by domestic institutions. Due to their different learning corpora, training paths, and application scenarios, the answers to the same question will naturally vary greatly. In addition, some LLMs may also have limitations and obsolescence in training data, resulting in biased content generation, low-confidence outputs, or LLM hallucinations. At the same time, a single LLM also faces challenges in coping with generalization scenarios due to its single training data. To address these issues, the collaborative work of multiple LLMs has been put on the agenda.
[0004] To jointly provide high-quality answers for UEs by multiple LLMs, relevant researchers use GPT-3, GPT-4o, Llama3-8B, Llama3-Chinese, Doubao, SparkDesk, Qwen, Kimi to jointly answer the questions of UEs, so as to avoid the answer biases and inaccuracies caused by the single training data of a single LLM. Although the above work can jointly provide answers to a question by multiple LLMs to improve the satisfaction of UEs with the answers. However, such a method does not build an automated collaborative network for multiple LLMs, and it is necessary to manually operate each LLM to obtain response responses, resulting in a long waiting time and labor costs, and it is difficult to increase the number of LLMs participating in the collaboration and expand the application scope.
[0005] To automate the collaborative work among multiple LLMs, relevant researchers have designed and developed a network that can support efficient communication among multiple LLMs (this network is called the network constructed by multiple LLMs (MLLMN)). This network has the advantages of scalability and portability, can facilitate the sharing of responses to the same question among multiple LLMs, and minimize human intervention. Although this solution can provide an efficient communication environment for multiple LLMs, this work does not consider the malicious operations of the hosting devices of LLMs, resulting in sharing unreliable responses with other LLMs in this network. At the same time, this work has not yet solved how to decide the best response in MLLMN. Summary of the Invention
[0006] In view of the above deficiencies in the prior art, the present invention provides an optimization method for generating a large language model network based on blockchain, which solves the problem that the existing network constructed by multiple LLMs cannot efficiently provide trustworthy responses for users.
[0007] To achieve the above invention purpose, the technical solution adopted by the present invention is as follows:
[0008] Provide an optimization method for generating a large language model network based on blockchain, which includes the steps:
[0009] S1. The LLM receives a question-and-answer request initiated by a user device, generates a response to the question-and-answer request as a leader, and then encrypts the question-and-answer request and the response into encrypted data.
[0010] S2. In MLLMN, according to the encrypted data, the best response among all the responses generated by all LLMs in MLLMN is selected by using Byzantine fault-tolerant consensus.
[0011] S3. The best response is fed back to the user device, and the best response is packaged into a block and linked to the blockchain, and stored on the intelligent device running the LLM in a distributed manner.
[0012] Further, step S3 further includes:
[0013] S31. The leader sends the encrypted data to the remaining LLMs in its MLLMN as followers based on the broadcast protocol through the peer-to-peer network of the blockchain.
[0014] S32. The follower decrypts the received encrypted data, generates a response according to the question-and-answer request and verifies the response quality of the leader as the voting result, and then encrypts the response and the voting result and feeds them back to the leader.
[0015] S33. The leader receives the encrypted data from the follower and decrypts it, screens the voting results indicating that the response quality of the leader is better among all the voting results, and accumulates the voting weights of the screened voting results to obtain the weight sum.
[0016] S34. Determine whether the sum of the weights is greater than or equal to a preset threshold. If so, proceed to step S35; otherwise, select a follower with the best response quality among all followers as the leader and proceed to step S31.
[0017] S35. The leader generates a proof indicating that its response has been recognized and verified, and broadcasts it to all followers.
[0018] S36. After receiving the proof, the follower verifies the validity of the proof and generates a confirmation message to send to the leader when the proof is valid.
[0019] S37. The leader determines whether the proportion of the confirmation messages it receives is greater than the preset threshold. If so, proceed to step S38; otherwise, randomly select a follower as the leader among all followers and proceed to step S31.
[0020] S38. Take the leader's response as the best response determined by the decision.
[0021] Further, the response quality of the leader is verified as follows: The follower compares the response it generates with the leader's response. When the leader's response is better, the voting result is equal to 1; when the leader's response is worse, the voting result is equal to 0.
[0022] Further, the expression for calculating the voting weight of the voting result is:
[0023]
[0024]
[0025] where, is the voting weight assigned by leader i to follower j in the r - round consensus; is the response quality weight assigned by leader i to follower j in the r - round consensus; is the reputation weight assigned by leader i to follower j in the r - round consensus; α is the proportion in the voting weight ; β is the proportion in the voting weight ; is the response quality score of leader i for follower j in the r - round consensus; is the reputation score of leader i for follower j in the r - round consensus; n is the total number of followers.
[0026] Further, the proportions α and β, and the voting weight response quality weight and reputation weight satisfy the following constraint conditions:
[0027] α + β = 1,
[0028] Further, steps S31 to S34 are the preparation stage of the Byzantine fault-tolerant consensus, and steps S35 to S38 are the submission stage of the Byzantine fault-tolerant consensus; when the Q&A request initiated by the user device enters the submission stage and there is a new Q&A request on the user device, a new round of the preparation stage of the Byzantine fault-tolerant consensus is started according to the new Q&A request.
[0029] Further, when a follower becomes a leader, its encrypted data is its response and Q&A request; when a leader becomes a follower, its encrypted data is its response and voting result.
[0030] Further, the preset threshold is equal to two-thirds.
[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0032] 1. Based on the Byzantine fault-tolerant consensus, the present invention can automatically drive the MLLMN network constructed by multiple LLMs to work, so as to prevent the user equipment UE from obtaining biased and hallucinated responses from a single LLM. At the same time, due to the Byzantine fault-tolerant ability of this consensus, it is possible to avoid the misleading and influence of potential malicious operation behaviors of the devices carrying LLMs on the MLLMN, thereby ensuring high-quality and highly credible LLM responses obtained by the UE.
[0033] 2. Based on the response ability and reputation of the LLMs, the present invention designs a weighted Byzantine fault-tolerant consensus WBFT. This design can dynamically adjust the voting rights of LLMs with poor response ability and low reputation to avoid potential malicious behaviors of such LLMs from interfering with the consensus result. In addition, this consensus is divided into two stages: preparation and submission, and the preparation stage of the next round of consensus can be processed in parallel during the submission stage of this round of consensus, thereby improving the processing efficiency of the consensus.
[0034] 3. The simulation and investigation results of the embodiments prove that the Byzantine fault-tolerant consensus WBFT of the present solution has advantages in terms of consensus success rate and latency compared with other consensuses, and the MLLMN driven by WBFT provides higher UE satisfaction with the response quality than a single LLM and the MLLMN without blockchain consensus participation. Description of the Drawings
[0035] Figure 1 It is a flowchart of an optimization method for generating a large language model network based on a blockchain.
[0036] Figure 2 It is a schematic diagram of the main steps of the Byzantine fault-tolerant consensus WBFT.
[0037] Figure 3 It is a simulation diagram of the consensus success rate in a specific example.
[0038] Figure 4 It is a simulation diagram of the consensus delay in a specific example.
[0039] Figure 5 It is a schematic diagram of the survey results of the response ability in a specific example. Specific implementation manner
[0040] The following describes the specific implementation manner of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manner. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0041] Reference Figure 1 , Figure 1 shows a flowchart of an optimization method for generating a large language model network based on a blockchain; as Figure 1 shown, this method S includes steps S1 to S3.
[0042] In step S1, the LLM receives a question-and-answer request initiated from a user device and generates a response to the question-and-answer request as a leader. Then, the question-and-answer request and the response are encrypted into encrypted data; in this solution, the leader uses its public key to encrypt and decode the data.
[0043] In step S2, in the MLLMN, according to the encrypted data, the best response among all the responses generated by all the LLMs within the MLLMN is selected using Byzantine fault-tolerant consensus; the MLLMN in this solution is a network constructed by multiple LLMs.
[0044] In step S3, the best response is fed back to the user device, and the best response is packaged into a block and linked to the blockchain, and is stored in a distributed manner on the intelligent device running the LLM.
[0045] In an embodiment of the present invention, step S3 further includes:
[0046] S31. The leader sends its encrypted data to the remaining LLMs in its MLLMN as followers based on the broadcast protocol through the peer-to-peer network of the blockchain; in steps S32 to S38, when data transmission is involved, it is all sent based on the broadcast protocol through the peer-to-peer network of the blockchain to ensure the security of data transmission.
[0047] S32. The follower decrypts the received encrypted data, generates a response according to the Q&A request, verifies the response quality of the leader as the voting result, and then encrypts and feeds back the response and the voting result to the leader; the follower in this solution encrypts and decodes the data using its private key.
[0048] Verifying the response quality of the leader is as follows: the follower compares the response it generates with the response of the leader. When the response of the leader is better, the voting result is equal to 1; when the response of the leader is worse, the voting result is equal to 0.
[0049] S33. The leader receives and decrypts the encrypted data from the follower, filters out the voting results indicating that the leader's response quality is better among all the voting results, and accumulates the voting weights of the filtered voting results to obtain the weight sum.
[0050] When implementing, the preferred expression for calculating the voting weight of the voting result in this solution is:
[0051]
[0052] where is the voting weight assigned by leader i to follower j in the r - round consensus; is the response quality weight assigned by leader i to follower j in the r - round consensus; is the reputation weight assigned by leader i to follower j in the r - round consensus; α is the proportion in the voting weight ; β is the proportion in the voting weight ; is the response quality score of leader i for follower j in the r - round consensus; is the reputation score of leader i for follower j in the r - round consensus; n is the total number of followers.
[0053] The reputation mentioned in this solution refers to whether there are malicious behaviors in the LLM and its hosting devices.
[0054] Among them, the proportions α and β and the voting weight response quality weight and reputation weight satisfy the following constraint conditions:
[0055] α + β = 1,
[0056] In each consensus of this solution, the leader will receive the answers to the questions and the voting results from other LLMs. Therefore, at the beginning of each consensus, the consensus initiator, that is, the leader, can dynamically adjust the weight allocation of the followers according to the performance of other LLMs.
[0057] S34. Determine whether the sum of weights is greater than or equal to the preset threshold. If so, proceed to step S35; otherwise, select a follower with the best response quality among all followers as the leader and proceed to step S31. When a follower becomes the leader, its encrypted data is its response and the Q&A request. When the leader becomes a follower, its encrypted data is its response and the voting result. Here, the preset threshold is equal to two-thirds.
[0058] S35. The leader generates a proof indicating that its response is recognized and verified, and broadcasts it to all followers.
[0059] S36. After receiving the proof, the followers verify the validity of the proof and generate confirmation information to send to the leader when it is valid. The proof can include the hash value of the leader's response and its timestamp, and this setting ensures the security, immutability, and traceability of the MLLMN output result.
[0060] S37. The leader determines whether the proportion of the confirmation information it receives is greater than the preset threshold. If so, proceed to step S38; otherwise, randomly select a follower as the leader among all followers and proceed to step S31.
[0061] S38. Use the leader's response as the best response determined by the decision.
[0062] As Figure 2 shown, steps S31 to S34 of this solution are the preparation stage of Byzantine fault-tolerant consensus, and steps S35 to S38 are the submission stage of Byzantine fault-tolerant consensus. When the Q&A request initiated by the user device enters the submission stage and there is a new Q&A request on the user device, a new round of the preparation stage of Byzantine fault-tolerant consensus is started according to the new Q&A request.
[0063] This solution can process other UE Q&A requests simultaneously during the consensus submission stage, that is, start a new preparation stage of consensus. By this means, the waiting delay for subsequent UE requests can be reduced.
[0064] The following details the effect of the large language model network generation optimization method of this solution in combination with specific examples:
[0065] The present invention recruited 15 volunteers to evaluate 10 common LLMs such as Llama 3.3, WizardLM 2, GPT-4o, Geimini 2Flash, Wenxin Yiyan 4.0, iFlytek Spark V4.0, Tongyi Qianwen 2.5, Doubao pro4k, Hunyuan Large, and Kimi, focusing on memory ability, daily life ability, artistic ability, logical reasoning ability, and code generation ability to obtain their response quality weights. In addition, the present invention made the reputation weights of these 10 LLMs follow a normal distribution N(0.9, 0.5).
[0066] Based on this, the present invention first compared the consensus success rates and latencies of the Byzantine Fault Tolerance consensus WBFT with the Practical Byzantine Fault Tolerance (PBFT) consensus and the Votes-as-a-Proof (VaaP) consensus through simulation. The simulation results can be referred to Figure 3 and Figure 4 .
[0067] In Figure 3 , the top three curves are the consensus success rates of the present solution under α = 0.4, β = 0.6, α = 0.5, β = 0.5, and α = 0.6, β = 0.4 respectively. The bottom curve is the consensus success rate of PBFT or VaaP. From the comparison of these four curves, it can be seen that regardless of how the response quality and reputation quality weights of each LLM in the WBFT consensus are allocated, this consensus always has a higher consensus success rate than the comparative solutions.
[0068] In Figure 4 , the top two curves are the latency simulation curves of PBFT and VaaP respectively. The bottom three curves are the latency simulation curves of the present solution under α = 0.4, β = 0.6, α = 0.5, β = 0.5, and α = 0.6, β = 0.4 respectively. From the comparison of these five curves, it can be seen that regardless of how the response quality and reputation quality weights of each LLM in the WBFT consensus are allocated, this consensus always has a lower consensus latency than the comparative solutions.
[0069] In addition, the present invention also compared the user satisfaction of a single LLM, an MLLMN without blockchain consensus participation, and a WBFT-driven MLLMN. This satisfaction was obtained from the evaluation of the above 10 LLMs by 15 volunteers in terms of five aspects: memory ability, daily life ability, artistic ability, logical reasoning ability, and code generation ability. The full score is 100, and the results are as Figure 5 shown.
[0070] In Figure 5 , the satisfaction of a single LLM is the average satisfaction of all volunteers randomly using a certain LLM; the satisfaction of the MLLMN and the WBFT-MLLMN with three different weight allocations also comes from the average satisfaction of all volunteers. This result demonstrates the superior ability of the present invention to drive the MLLMN based on the allocated weight consensus to obtain high-quality and trustworthy responses.
Claims
1. A large language model network generation optimization method based on blockchain, characterized in that: Includes steps: S1. LLM receives the question-and-answer request initiated by the user device and generates a response to the question-and-answer request as a leader, and then encrypts the question-and-answer request and the response into encrypted data; S2. In MLLMN, the best response among all responses generated by all LLMs in MLLMN is selected using Byzantine fault-tolerant consensus based on the encrypted data; S3. Feedback the best response to the user device, and package the best response into a block and link it to the blockchain, and store it in a distributed manner on the smart device running LLM.
2. The large language model network generation optimization method based on blockchain according to claim 1 is characterized in that: Step S3 further comprises: S31. The leader sends its encrypted data to the remaining LLMs in its MLLMN as followers through the peer-to-peer network of the blockchain based on the broadcast protocol; S32, the follower decrypts the received encrypted data, generates a response according to the question and answer request, and verifies the quality of the leader's response as the voting result, and then encrypts the response and voting result and feeds it back to the leader; S33, the leader receives and decrypts the encrypted data of the follower, selects the voting results indicating that the leader's response quality is better among all the voting results, and accumulates the voting weights of the selected voting results to obtain the weight sum; S34, determine whether the sum of the weights is greater than or equal to a preset threshold, if so, proceed to step S35, otherwise, select a follower with the best answer quality from all followers as the leader, and proceed to step S31; S35. The leader generates a proof that its response is recognized and verified, and broadcasts it to all followers; S36. After receiving the proof, the follower verifies the validity of the proof and generates confirmation information to send to the leader if it is valid; S37, the leader determines whether the proportion of confirmed information it has received is greater than a preset threshold, if so, proceeds to step S38, otherwise randomly selects one of all followers as the leader and proceeds to step S31; S38. Take the leader’s response as the best response made in the decision.
3. The large language model network generation optimization method based on blockchain according to claim 2 is characterized in that: The quality of the leader's response is verified as follows: the follower compares the response it generates with the leader's response and votes equal to 1 when the leader's response is better and 0 when the leader's response is worse.
4. The large language model network generation optimization method based on blockchain according to claim 2 is characterized in that: The expression for calculating the voting weight of the voting result is: in, The voting weight assigned by leader i to follower j in round r of consensus; The response quality weight assigned by leader i to follower j in round r of consensus; is the reputation weight assigned by leader i to follower j in round r of consensus; α is The voting weight The proportion of The voting weight The proportion of Score the quality of leader i's response to follower j in round r of consensus; is the reputation score of leader i for follower j in round r of consensus; n is the total number of followers.
5. The large language model network generation optimization method based on blockchain according to claim 4 is characterized in that: Proportions α and β and voting weights Answer quality weight and reputation weight The constraints that are satisfied are: α+β=1, 6. The method for generating and optimizing a large language model network based on blockchain according to any one of claims 2 to 5, characterized in that: Steps S31 to S34 are the preparation stage of the Byzantine fault-tolerant consensus, and steps S35 to S38 are the submission stage of the Byzantine fault-tolerant consensus. When the question-and-answer request initiated by the user device enters the submission stage and there is a new question-and-answer request from the user device, the preparation stage of the next round of Byzantine fault-tolerant consensus is started according to the new question-and-answer request.
7. The method for generating and optimizing a large language model network based on blockchain according to any one of claims 2 to 5, characterized in that: When a follower becomes a leader, its encrypted data is its response and question-and-answer request; when a leader becomes a follower, its encrypted data is its response and voting results.
8. The method for generating and optimizing a large language model network based on blockchain according to any one of claims 2 to 5, characterized in that: The preset threshold is equal to two thirds.