Block chain-based large model AI intelligent evaluation method, apparatus and device, and storage medium

By utilizing AI arbitration nodes and a dynamic evaluation engine on the blockchain for large-scale AI evaluation, the problems of rigid evaluation strategies and unreasonable node selection are solved, enabling the generation of efficient and in-depth evaluation reports and improving evaluation efficiency and user experience.

CN121996534APending Publication Date: 2026-05-08QINGDAO INSPUR HAIRUO ARTIFICIAL INTELLIGENCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO INSPUR HAIRUO ARTIFICIAL INTELLIGENCE CO LTD
Filing Date
2025-10-23
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing large-scale model AI evaluation technologies based on blockchain suffer from problems such as static and rigid evaluation strategies, lack of intelligent node selection mechanisms, and weak comprehensive analysis capabilities, resulting in low evaluation efficiency and poor performance.

Method used

By utilizing AI arbitration nodes and a preset node reputation database in the blockchain to identify target evaluation nodes, and combining preset static and dynamic evaluation datasets for preliminary and enhanced evaluation, the dynamic evaluation engine is used to identify weaknesses and perform multi-dimensional fusion analysis, generating an evaluation report with a comprehensive score, fine-grained capability radar chart, and improvement suggestions.

Benefits of technology

It improves the efficiency and accuracy of large-scale AI model evaluation, generates in-depth and dynamic evaluation reports, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996534A_ABST
    Figure CN121996534A_ABST
Patent Text Reader

Abstract

The invention discloses a block chain-based large model AI intelligent evaluation method, device and equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an evaluation request, and determining a target evaluation node through an AI arbitration node in a block chain, a preset node reputation library and a model type; calling each target evaluation node to perform preliminary evaluation on the target large model, recording a preliminary evaluation result in a block chain through an intelligent contract, calling an AI arbitration node to perform weakness identification on the preliminary evaluation result in the block chain to obtain a weakness identification result, and if the target large model has weakness, generating a reinforcement test question; and calling each target evaluation node and evaluating the target large model based on the enhanced test question to record an obtained to-be-processed evaluation result in the block chain, and calling a preset dynamic evaluation engine to analyze the recorded preliminary evaluation result and the to-be-processed evaluation result to obtain a fusion analysis result to generate an evaluation report. The efficiency of evaluating the large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for large-scale AI intelligent evaluation based on blockchain. Background Technology

[0002] Currently, with the rapid development of Large Language Model (LLM) technology, objective, fair, and comprehensive evaluation of its performance has become crucial. Existing technologies have attempted to leverage blockchain to address the issue of evaluation credibility. For example, a blockchain-based multi-party consensus evaluation system evaluates the model through multiple evaluators (blockchain nodes) and records the data on the chain. By utilizing the immutability and distributed consensus characteristics of blockchain, the transparency of the evaluation process and the credibility of the results are ensured.

[0003] However, after in-depth analysis, the existing solutions still have the following significant drawbacks: The evaluation strategy is static and rigid: the evaluation benchmarks and datasets used by each evaluation party are pre-set and static. This "one-size-fits-all" approach cannot perceive the unique performance of the tested model in real time during the evaluation process, and cannot conduct in-depth and dynamic targeted testing on its specific weaknesses (such as logical loopholes, factual errors, biases, etc.), resulting in insufficient evaluation depth and flexibility.

[0004] The node selection mechanism lacks intelligence: relying on traditional consensus algorithms (such as Raft, DPoS) to select nodes to participate in the evaluation, although it ensures decentralization, does not fully consider the historical evaluation quality, response speed and professional field differences of each evaluation node. This can lead to low evaluation efficiency or low-quality, mismatched nodes participating in key consensus, affecting the overall evaluation effect.

[0005] The results show weak comprehensive analysis capabilities: the final comprehensive evaluation results usually rely on simple score weighted averages or summaries, lacking an intelligent analysis layer that can deeply mine, integrate and analyze multi-party evaluation data, and identify deep behavioral patterns, consistency, robustness and bias of the model, resulting in limited evaluation insights.

[0006] As can be seen from the above, how to improve the efficiency of evaluating large models in the process of AI intelligent evaluation of large models based on blockchain is an urgent problem to be solved. Summary of the Invention

[0007] In view of this, the purpose of this invention is to provide a method, apparatus, device, and storage medium for large-scale AI intelligent evaluation based on blockchain, which can improve the efficiency of evaluating large models in the process of large-scale AI intelligent evaluation based on blockchain. The specific solution is as follows: Firstly, this application provides a blockchain-based method for AI-powered intelligent evaluation of large-scale models, including: Obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; Each target evaluation node is invoked and a preliminary evaluation of the target large model is performed based on a preset static evaluation dataset. The preliminary evaluation results are recorded in the blockchain through a smart contract. Then, the preset dynamic evaluation engine in the AI ​​arbitration node is invoked to identify weaknesses in the preliminary evaluation results in the blockchain, and weakness identification results are obtained. If the weakness identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated. Each of the target evaluation nodes is invoked and the target large model is evaluated based on the reinforcement test questions. The evaluation results to be processed are recorded in the blockchain. Then, the machine learning model in the preset dynamic evaluation engine is invoked to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, so as to obtain the fusion analysis results. Based on the fusion analysis results, an evaluation report is generated, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions. The evaluation report is then set as the target evaluation result and recorded in the blockchain.

[0008] Optionally, the step of obtaining the evaluation request corresponding to the target large model, and then determining the target evaluation node using the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model, includes: The system obtains the evaluation request corresponding to the target large model and the model type corresponding to the target large model, and determines the corresponding preset node reputation database based on the evaluation request. The system then uses the AI ​​arbitration node in the blockchain and the model type to obtain several evaluation nodes from the preset node reputation database. The preset node reputation database records the historical evaluation quality score, response speed and professional domain label of each evaluation node. Determine the node information corresponding to each of the evaluation nodes, including historical evaluation quality scores, response speed and professional domain labels, and match each of the node information with the model type to obtain the corresponding matching score. Then, set the evaluation node with the highest matching score among the various matching scores as the target evaluation node. The professional domain labels include code evaluation labels and medical text evaluation labels; the historical evaluation quality score is a quality score determined based on the accuracy and reliability of historical evaluations; and the response speed is a speed determined based on the delay time of the evaluation node in processing the evaluation task. The node performance corresponding to each of the evaluation nodes is obtained, and the preset node reputation database is updated based on the node performance to obtain a new preset node reputation database; the node performance includes the real-time load status of the node, the node response speed, and the historical evaluation quality score.

[0009] Optionally, the step of calling each target evaluation node and performing a preliminary evaluation of the target large model based on a preset static evaluation dataset, and recording the preliminary evaluation results in the blockchain via a smart contract, includes: The target evaluation node is invoked and the target large model is tested based on a preset static dataset stored locally. The model's answer and preliminary score are generated, and the model's answer and preliminary score are encapsulated into a smart contract transaction. The smart contract transaction is then stored in the blockchain. The static dataset includes preset test questions and corresponding standard answers covering several domains and ability dimensions. The consensus mechanism in the blockchain is used to detect the authenticity and integrity of the smart contract transaction data, and the detection result is obtained. If the detection result indicates that the detection is passed, the target evaluation node is called to analyze the smart contract transaction and obtain preliminary analysis results including the performance indicators of the target large model on a specific task. Based on the smart contract transactions and the preliminary analysis results, preliminary evaluation results are determined, and the preliminary evaluation results, timestamps, and node identifiers are recorded in the blockchain through smart contracts.

[0010] Optionally, the step of calling the preset dynamic evaluation engine in the AI ​​arbitration node to perform weakness identification on the preliminary evaluation results in the blockchain, and obtaining weakness identification results, if the weakness identification results indicate that the target large model has weaknesses, then generating reinforcement test questions corresponding to the weaknesses, including: A weakness identification threshold is determined, and the preset dynamic evaluation engine in the AI ​​arbitration node is invoked to identify weaknesses in the preliminary evaluation results in the blockchain, thereby obtaining the weakness identification results to be processed; the weakness identification results to be processed are the weaknesses of the target large model in preset domains, preset capabilities, or preset attributes; Determine whether each of the identified weaknesses to be processed is greater than a preset weakness identification threshold. If the identified weakness to be processed is greater than the preset weakness identification threshold, then set the identified weakness to be processed as the target weakness identification result. Using a pre-set large-scale language model and based on the target weakness identification results, reinforcement test questions, including adversarial samples and targeted challenges, are generated corresponding to the target weakness identification results. Each reinforcement test question is encapsulated as a smart contract transaction, and the smart contract transaction is distributed to each evaluation node through the blockchain network.

[0011] Optionally, the step of calling each of the target evaluation nodes and evaluating the target large model based on the reinforcement test questions, and recording the obtained evaluation results in the blockchain, includes: Call the local resources corresponding to each of the target evaluation nodes and evaluate the target large model based on the current reinforcement test question to obtain the current evaluation result to be processed, including the model answer and performance indicators, and determine whether the current evaluation result to be processed meets the preset stopping condition. If the current evaluation result to be processed meets the preset stopping condition, the operation of evaluating the target large model is stopped. If the current evaluation result to be processed does not meet the preset stopping condition, the process jumps back to the step of calling the local resources corresponding to each target evaluation node and evaluating the target large model based on the current reinforcement test questions. The preset stopping conditions include the evaluation time reaching a preset time threshold, the number of evaluations reaching a preset identification threshold, and the resource exhaustion of the evaluation node.

[0012] Optionally, the step of calling the machine learning model in the preset dynamic evaluation engine to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, to obtain fusion analysis results, includes: The preliminary evaluation results and the evaluation results to be processed are obtained from the blockchain, and the multimodal learning model in the preset dynamic evaluation engine is called to evaluate the consistency of the results of the preliminary evaluation results and the evaluation results to be processed on the same or similar questions at different times or by different nodes, so as to obtain the result consistency evaluation result. The graph neural network in the preset dynamic evaluation engine is invoked to evaluate the stability of the model output results under the inspection problem perturbation or adversarial attack on the preliminary evaluation results and the evaluation results to be processed, so as to obtain the robustness evaluation results. The causal inference model in the preset dynamic evaluation engine is invoked to evaluate the preliminary evaluation results and the evaluation results to be processed for specific groups or topics, thereby obtaining biased evaluation results; The fusion analysis results are determined based on the consistency evaluation results, the robustness evaluation results, and the bias evaluation results.

[0013] Optionally, the step of generating an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, and recording the evaluation report as the target evaluation result in the blockchain, includes: The report generation center in the AI ​​arbitration node is used to generate an evaluation report containing a comprehensive score, fine-grained capability radar chart, defect analysis and improvement suggestions based on the fusion analysis results. The evaluation report is then encapsulated as a smart contract transaction record on the blockchain. The comprehensive score is a weighted score calculated based on the weights corresponding to the fusion analysis results, used to reflect the overall performance of the target large model; the fine-grained capability radar chart is used to display the performance of the target large model in different dimensions; the defect analysis is used to describe the weaknesses of the target large model and their corresponding impacts; the improvement suggestions are optimization measures corresponding to the defect analysis; the evaluation report can be queried and verified in the blockchain through an interactive interface, and has immutability and traceability.

[0014] Secondly, this application provides a blockchain-based large-scale AI intelligent evaluation device, comprising: The evaluation node determination module is used to obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; The reinforcement test question determination module is used to call each target evaluation node and perform a preliminary evaluation of the target large model based on a preset static evaluation dataset, and record the preliminary evaluation results in the blockchain through a smart contract. Then, it calls the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain and obtain weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated. The fusion analysis result determination module is used to call each of the target evaluation nodes and evaluate the target large model based on the reinforcement test questions, so as to record the obtained evaluation results to be processed in the blockchain, and then call the machine learning model in the preset dynamic evaluation engine to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, so as to obtain the fusion analysis result. The evaluation report generation module is used to generate an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, so as to set the evaluation report as the target evaluation result and record it in the blockchain.

[0015] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned blockchain-based large-scale AI intelligent evaluation method.

[0016] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned blockchain-based large-scale AI intelligent evaluation method.

[0017] As can be seen from the above, before conducting large-scale AI intelligent evaluation based on blockchain, this application needs to obtain an evaluation request corresponding to the target large-scale model. Then, it uses the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large-scale model to determine the target evaluation node. It then calls each target evaluation node and performs a preliminary evaluation of the target large-scale model based on a preset static evaluation dataset. The preliminary evaluation results are recorded in the blockchain through a smart contract. Next, it calls the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain, obtaining the weakness identification results. If weaknesses are found... If the identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated. Each evaluation node is invoked and the target large model is evaluated based on the reinforcement test questions. The evaluation results to be processed are recorded in the blockchain. Then, the machine learning model in the preset dynamic evaluation engine is invoked to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, so as to obtain the fusion analysis results. Based on the fusion analysis results, an evaluation report including a comprehensive score, fine-grained capability radar chart, defect analysis and improvement suggestions is generated. The evaluation report is set as the target evaluation result and recorded in the blockchain.

[0018] Therefore, this application first needs to obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; secondly, call each target evaluation node and conduct a preliminary evaluation of the target large model based on the preset static evaluation dataset, and record the preliminary evaluation results in the blockchain through a smart contract, and then call the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain, and obtain the weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, then a reinforcement test question corresponding to the weakness is generated; then, call each evaluation node and evaluate the target large model based on the reinforcement test question, and record the obtained evaluation results to be processed in the blockchain, and then call the machine learning model in the preset dynamic evaluation engine to perform multi-dimensional fusion analysis on the preliminary evaluation results and the evaluation results to be processed recorded in the blockchain, and obtain the fusion analysis results; finally, based on the fusion analysis results, generate an evaluation report including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, and set the evaluation report as the target evaluation result to be recorded in the blockchain. This improves the efficiency of evaluating large models in the process of AI-powered intelligent evaluation of large models based on blockchain, thereby enhancing the user experience. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 This application discloses a flowchart of a blockchain-based large-scale AI intelligent evaluation method. Figure 2 This is a schematic diagram of a specific intelligent evaluation process disclosed in this application; Figure 3 This is a schematic diagram of a specific AI arbitration node disclosed in this application; Figure 4 This application discloses a specific flowchart of a blockchain-based method for AI intelligent evaluation of large-scale models. Figure 5 This is a schematic diagram of the interactive process of a dynamic evaluation stage disclosed in this application; Figure 6 This is a schematic diagram of the structure of a large-scale AI intelligent evaluation device based on blockchain disclosed in this application; Figure 7 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Currently, with the rapid development of large language model technology, objective, fair, and comprehensive performance evaluation has become crucial. Existing technologies have attempted to leverage blockchain to address the credibility issue in evaluation. However, after in-depth analysis, existing solutions still suffer from the following significant drawbacks: static and rigid evaluation strategies, a lack of intelligent node selection mechanisms, and weak comprehensive result analysis capabilities. Therefore, this application provides a blockchain-based intelligent evaluation method for large-scale AI models, which improves the efficiency of evaluating large models during the blockchain-based intelligent evaluation process.

[0023] See Figure 1As shown in the figure, this invention discloses a blockchain-based large-scale AI intelligent evaluation method, including: Step S11: Obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node.

[0024] In this embodiment, the flowchart of the large-scale model AI (Artificial Intelligence) intelligent evaluation based on blockchain is shown below. Figure 2 As shown: First, users initiate an evaluation request for a target large model through the evaluation request interaction module, such as a decentralized application (DApp) based on Web3 (i.e., the third generation of the Internet). Subsequently, the request is sent to the intelligent evaluation blockchain network. It's worth noting that the intelligent evaluation blockchain network consists of multiple blockchain nodes, including several ordinary evaluation nodes and at least one AI arbitration node. The AI ​​arbitration node can select a suitable ordinary evaluation node as the target evaluation node based on its maintained node reputation database. Evaluation data flows and smart contract calls circulate within the network, and the resulting structured in-depth evaluation report is stored on-chain and can be queried by users. The Blockchain Ledger records model information, all evaluation data, and the final report, ensuring data immutability and transparency.

[0025] It's worth noting that the intelligent evaluation blockchain network consists of multiple blockchain nodes, which are divided into two categories: ordinary evaluation nodes: corresponding to an evaluator, whose built-in evaluator proxy module stores static evaluation datasets for performing basic evaluation tasks. AI arbitration nodes: a special type of enhanced blockchain node that integrates a dynamic evaluation engine and an arbitration analysis engine, and maintains a node reputation database.

[0026] In this embodiment, the evaluation initiation and smart node selection are performed: First, the evaluation request interaction module is invoked to respond to the user's operation and initiate an evaluation request for the target large model to the smart evaluation blockchain network. Subsequently, the AI ​​arbitration node in the blockchain network is invoked, and based on its maintained node reputation database (which records the historical evaluation quality scores, response speeds, and professional domain tags of each ordinary evaluation node) and the type of the target large model, the most suitable multiple target evaluation nodes are intelligently selected, rather than relying solely on traditional consensus algorithms.

[0027] Specifically, the process involves obtaining evaluation requests corresponding to the target large model, and then using AI arbitration nodes in the blockchain, a corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation nodes. This can include: obtaining evaluation requests and model types corresponding to the target large model, and determining the corresponding preset node reputation database based on the evaluation requests; then using AI arbitration nodes and model types in the blockchain to obtain several evaluation nodes from the preset node reputation database; wherein, the preset node reputation database records the historical evaluation quality scores, response speed, and professional domain tags of each evaluation node; and determining the historical evaluation quality scores, response speed, and professional domain tags corresponding to each evaluation node. The system collects node information for domain labels and matches each node with the model type to obtain a matching score. The node with the highest matching score is then designated as the target evaluation node. Domain labels include code evaluation labels and medical text evaluation labels. Historical evaluation quality scores are determined based on the accuracy and reliability of historical evaluations, and response speed is determined based on the latency of the evaluation node in processing the evaluation task. The system acquires the node performance corresponding to each evaluation node and updates the preset node reputation database based on this performance, resulting in a new preset node reputation database. Node performance includes real-time node load, node response speed, and historical evaluation quality scores.

[0028] In this embodiment, the data processing flow diagram of the AI ​​arbitration node is shown below. Figure 3As shown, the AI ​​arbitration node internally integrates a dynamic evaluation engine and an arbitration analysis engine, maintains a node reputation database, and interacts with the blockchain network through a smart contract interface. The dynamic evaluation engine is responsible for real-time monitoring of initial on-chain evaluation data. This engine incorporates deep learning-based models (e.g., using a Transformer-based language model for text analysis, or a specialized knowledge graph matching model for factual verification). When a target large model is identified as having significant weaknesses or inconsistencies in a specific domain (e.g., code generation), specific capabilities (e.g., multi-step reasoning), or specific attributes (e.g., fairness, security), the dynamic evaluation engine calls the dynamic question generation module (e.g., fine-tuning an LLM to generate adversarial samples or targeted challenges) to create a new batch of reinforcement test questions targeting that weakness. The arbitration analysis engine is responsible for starting after all evaluation data (including static and dynamic evaluation results) is stored on-chain. This engine uses advanced machine learning models (e.g., combining multimodal learning, graph neural networks, or causal inference models) to perform multi-dimensional, in-depth fusion analysis of the entire chain's evaluation data. The analysis dimensions include, but are not limited to: consistency analysis (consistency of results for the same or similar problems at different times or by different nodes), robustness analysis (stability of model performance under problem perturbation or adversarial attacks), and bias analysis (model bias towards specific groups or topics). Finally, the report generation module generates a structured in-depth evaluation report, which includes a comprehensive score, a fine-grained capability radar chart, specific defect analysis, and improvement suggestions.

[0029] The node reputation database is a local database stored within the AI ​​arbitration node. It records the historical evaluation quality scores, response speeds, and specialized domain labels (e.g., code evaluation, medical text evaluation) of each ordinary evaluation node in the intelligent evaluation blockchain network. This reputation database is continuously updated based on the results of each evaluation and node performance, enabling adaptive optimization of the evaluation ecosystem. The smart contract interface interacts with the underlying blockchain network, including reading on-chain evaluation data, issuing new evaluation tasks (dynamic test questions), and writing the final in-depth evaluation report to the blockchain.

[0030] Step S12: Call each target evaluation node and perform a preliminary evaluation of the target large model based on a preset static evaluation dataset. Record the preliminary evaluation results in the blockchain through a smart contract. Then, call the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain and obtain weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, then generate reinforcement test questions corresponding to the weaknesses.

[0031] In this embodiment, the target evaluation node is used to perform a preliminary evaluation of the target large model based on the local static dataset, and the obtained preliminary evaluation data is stored on the blockchain. In addition, the dynamic evaluation engine of the AI ​​arbitration node is used to monitor and analyze the preliminary evaluation data on the blockchain in real time.

[0032] Specifically, the process involves calling various target evaluation nodes and conducting preliminary evaluations of the target large model based on a pre-set static evaluation dataset. The preliminary evaluation results are then recorded in the blockchain via smart contracts. This process can include: calling target evaluation nodes and testing the target large model based on a locally stored pre-set static dataset, generating model responses and preliminary scores, encapsulating these responses and scores into smart contract transactions, and storing these smart contract transactions in the blockchain. The static dataset includes pre-set test questions covering several domains and capability dimensions, along with corresponding standard answers. The consensus mechanism in the blockchain is used to verify the authenticity and integrity of the smart contract transactions, obtaining a verification result. If the verification result indicates that the verification is successful, the target evaluation nodes are called to analyze the smart contract transactions, obtaining preliminary analysis results including the target large model's performance metrics on a specific task. Based on the smart contract transactions and the preliminary analysis results, the preliminary evaluation results are determined, and the preliminary evaluation results, timestamps, and node identifiers are recorded in the blockchain via smart contracts.

[0033] It's worth mentioning that the dynamic evaluation engine is used to analyze the responses of the target large model based on deep learning models. If significant weaknesses or inconsistencies are identified in the model in a specific domain (such as code generation), specific capabilities (such as multi-step reasoning), or specific attributes (such as fairness), a batch of reinforcement test questions targeting those weaknesses is immediately and dynamically generated. These newly generated dynamic test questions are then encapsulated into a smart contract transaction and distributed to the target evaluation nodes. Furthermore, the target evaluation nodes receive and execute these targeted evaluations and upload the new evaluation data to the blockchain.

[0034] Specifically, the system calls the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results on the blockchain, obtaining weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, then reinforcement test questions corresponding to the weaknesses are generated. This may include: determining the weakness identification threshold and calling the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results on the blockchain, obtaining weakness identification results to be processed; the weakness identification results to be processed are the weaknesses of the target large model in preset domains, preset capabilities, or preset attributes; determining whether each weakness identification result to be processed is greater than the preset weakness identification threshold; if the weakness identification result to be processed is greater than the preset weakness identification threshold, then the weakness identification result to be processed is set as the target weakness identification result; using a preset large language model and based on the target weakness identification results, reinforcement test questions corresponding to the target weakness identification results are generated, including adversarial samples and targeted challenges, and each reinforcement test question is encapsulated as a smart contract transaction, and the smart contract transactions are distributed to each evaluation node through the blockchain network.

[0035] Step S13: Call each of the target evaluation nodes and evaluate the target large model based on the reinforcement test questions, so as to record the obtained evaluation results to be processed in the blockchain. Then, call the machine learning model in the preset dynamic evaluation engine to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, and obtain the fusion analysis results.

[0036] In this embodiment, after obtaining the reinforcement test questions, the target large model needs to be evaluated again based on the reinforcement test questions. Specifically, calling each target evaluation node and evaluating the target large model based on the reinforcement test questions to record the obtained evaluation results in the blockchain can include: calling the local resources corresponding to each target evaluation node and evaluating the target large model based on the current reinforcement test questions to obtain the current evaluation results to be processed, including the model's response and performance indicators, and determining whether the current evaluation results to be processed meet the preset stopping conditions; if the current evaluation results to be processed meet the preset stopping conditions, the operation of evaluating the target large model is stopped; if the current evaluation results to be processed do not meet the preset stopping conditions, the process jumps back to the step of calling the local resources corresponding to each target evaluation node and evaluating the target large model based on the current reinforcement test questions; wherein, the preset stopping conditions include the evaluation time reaching a preset time threshold, the number of evaluations reaching a preset identification threshold, and the resource exhaustion of the evaluation node.

[0037] In this embodiment, after all evaluation data (including static and dynamic evaluations) is stored on the blockchain, the arbitration analysis engine of the AI ​​arbitration node is activated. This engine, based on advanced machine learning models, performs multi-dimensional and in-depth fusion analysis on the entire chain of evaluation data, far exceeding simple score averaging. The analysis dimensions include consistency analysis (multiple queries of the same question), robustness analysis (performance under minor perturbations), and bias analysis.

[0038] Specifically, the machine learning model in the preset dynamic evaluation engine is invoked to perform multi-dimensional fusion analysis on the preliminary evaluation results and the evaluation results to be processed recorded in the blockchain, and the fusion analysis results are obtained. This can include: obtaining the preliminary evaluation results and the evaluation results to be processed from the blockchain, and invoking the multimodal learning model in the preset dynamic evaluation engine to evaluate the consistency of the results of the preliminary evaluation results and the evaluation results to be processed for the same or similar questions at different times or by different nodes, to obtain the result consistency evaluation result; invoking the graph neural network in the preset dynamic evaluation engine to evaluate the stability of the model output results of the preliminary evaluation results and the evaluation results to be processed under problem perturbation or adversarial attacks, to obtain the robustness evaluation result; invoking the causal inference model in the preset dynamic evaluation engine to evaluate the preliminary evaluation results and the evaluation results to be processed for specific group or topic bias, to obtain the bias evaluation result; and determining the fusion analysis result based on the result consistency evaluation result, the robustness evaluation result, and the bias evaluation result.

[0039] Step S14: Generate an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, and record the evaluation report as the target evaluation result in the blockchain.

[0040] In this embodiment, after obtaining the fusion analysis results, the present application embodiment can use the AI ​​arbitration node to generate a structured in-depth evaluation report, wherein the in-depth evaluation report includes a comprehensive score, a fine-grained capability radar chart, specific defect analysis and improvement suggestions, and then this report is recorded on the blockchain as the final comprehensive evaluation result, thereby completing this evaluation.

[0041] Specifically, an evaluation report is generated based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions. This evaluation report is then recorded on the blockchain as the target evaluation result. This process can include: utilizing the report generation center in the AI ​​arbitration node and generating an evaluation report containing a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions based on the fusion analysis results; and then encapsulating the evaluation report as a smart contract transaction record on the blockchain. The comprehensive score is a weighted score calculated based on the weights corresponding to the fusion analysis results, reflecting the overall performance of the target large model. The fine-grained capability radar chart displays the performance of the target large model across different dimensions. The defect analysis describes the weaknesses of the target large model and their corresponding impacts. The improvement suggestions are optimization measures corresponding to the defect analysis. The evaluation report can be queried and verified on the blockchain through an interactive interface and is immutable and traceable.

[0042] In this embodiment, a flowchart illustrating the process of large-scale AI intelligent evaluation based on blockchain is shown below. Figure 4 As shown: First, users can initiate an evaluation request for a target large model to the intelligent evaluation blockchain network through the evaluation request interaction module. Subsequently, the AI ​​arbitration node can intelligently filter based on its maintained node reputation database and the type of the target large model (e.g., dialogue model, code generation model, image description model, etc.) to obtain several most suitable ordinary evaluation nodes as the target evaluation nodes for this evaluation, rather than solely relying on traditional blockchain consensus algorithms such as Raft (a distributed consensus algorithm) and DPoS (a blockchain consensus algorithm). The selection criteria include whether the evaluation node's professional field matches the target model type, historical evaluation quality scores, and response speed.

[0043] Subsequently, the selected target evaluation node performs an initial evaluation of the target large model using its locally pre-set static evaluation dataset. The initial evaluation data (e.g., the model's answers to static test questions, preliminary scores, etc.) is then encapsulated into a smart contract transaction and stored on the blockchain via the smart contract interface. Next, the dynamic evaluation engine in the AI ​​arbitration node monitors new block events on the smart evaluation blockchain network in real time to acquire and analyze the initial evaluation data on-chain. Notably, the dynamic evaluation engine analyzes the target large model's answers based on a deep learning model. If it identifies significant weaknesses or inconsistencies in a specific domain, capability, or attribute (e.g., logical errors in code generation, or consistency issues in multi-step inference), it immediately and dynamically generates a batch of reinforcement test questions targeting those weaknesses. These newly generated dynamic test questions are then encapsulated into a smart contract transaction and distributed to the target evaluation node via the smart contract interface. The target evaluation node receives and executes these targeted evaluations and stores the new evaluation data (the model's answers to dynamic test questions, scores, etc.) on the blockchain again. This "dynamic analysis-generation-execution-on-chain" cycle can be executed according to preset strategies (e.g., maximum number of loops, vulnerability identification threshold, time limit, etc.) until the termination condition is met.

[0044] It's worth noting that after storing all evaluation data (including static evaluation data and dynamic evaluation data from all rounds) on the blockchain, and activating the arbitration analysis engine of the AI ​​arbitration node, advanced machine learning models are used to conduct multi-dimensional and in-depth fusion analysis of the entire chain's evaluation data. Analysis dimensions include consistency analysis (e.g., consistency of answers to the same or similar questions at different evaluation stages or from different nodes), robustness analysis (e.g., the stability of the model's output after adding small perturbations to the input question), and bias analysis (e.g., whether the model exhibits bias when handling questions involving different groups or cultural backgrounds). Finally, the AI ​​arbitration node generates a structured, in-depth evaluation report, which includes a comprehensive score, a fine-grained capability radar chart (showing the model's performance across different capability dimensions), specific defect analysis, and targeted improvement suggestions. This report will be recorded on the blockchain as the final comprehensive evaluation result, completing this evaluation.

[0045] In this embodiment, the interaction process during the dynamic evaluation phase is as follows: Figure 5As shown: AI arbitration nodes continuously monitor new block events on the intelligent evaluation blockchain network, then acquire and analyze evaluation data from ordinary evaluation nodes. Once weaknesses in the tested model are identified, the AI ​​arbitration node's dynamic evaluation engine generates dynamic test questions. These test questions are then used and published as new evaluation tasks via smart contracts. Ordinary evaluation nodes continuously monitor smart contract events and, upon receiving a new task, acquire the dynamic test questions and perform targeted evaluations on the target model. After evaluation, ordinary evaluation nodes store the new evaluation data on the blockchain via smart contracts, forming a continuously iterating dynamic evaluation closed loop.

[0046] In this embodiment, any blockchain platform supporting smart contracts can be used as the underlying framework. The dynamic evaluation engine and arbitration analysis engine in the AI ​​arbitration node can be developed and trained using mainstream deep learning frameworks such as TensorFlow (a deep learning framework) and PyTorch (an open-source machine learning library), and deployed on hardware devices with sufficient computing power, such as GPU (Graphics Processing Unit) servers. Ordinary evaluation nodes can be deployed on cloud servers or in the local environment of various evaluation institutions, and can flexibly connect to the evaluation network.

[0047] As can be seen from the above, the embodiments of this application first need to obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; secondly, call each target evaluation node and perform a preliminary evaluation of the target large model based on the preset static evaluation dataset, and record the preliminary evaluation results in the blockchain through a smart contract, and then call the preset dynamic evaluation engine in the AI ​​arbitration node to perform weakness identification on the preliminary evaluation results in the blockchain to obtain the weakness identification result. If the weakness identification result indicates that the target large model is weak, the evaluation is considered to be weak. If the model has weaknesses, reinforcement test questions corresponding to those weaknesses are generated. Then, each evaluation node is invoked to evaluate the target large model based on these reinforcement test questions. The resulting evaluation results are recorded in the blockchain. Next, a machine learning model in a pre-defined dynamic evaluation engine performs a multi-dimensional fusion analysis of the preliminary evaluation results recorded in the blockchain and the remaining evaluation results, yielding a fusion analysis result. Finally, based on the fusion analysis result, an evaluation report is generated, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions. This report is then recorded in the blockchain as the target evaluation result. This approach improves the efficiency of evaluating large models in the blockchain-based AI intelligent evaluation process, thereby enhancing the user experience.

[0048] Accordingly, see Figure 6As shown, this application also provides a blockchain-based large-scale model AI intelligent evaluation device, including: The evaluation node determination module 11 is used to obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; The reinforcement test question determination module 12 is used to call each target evaluation node and perform a preliminary evaluation of the target large model based on a preset static evaluation dataset, and record the preliminary evaluation results in the blockchain through a smart contract. Then, it calls the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain and obtain weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated. The fusion analysis result determination module 13 is used to call each of the target evaluation nodes and evaluate the target large model based on the reinforcement test questions, so as to record the obtained evaluation results to be processed in the blockchain, and then call the machine learning model in the preset dynamic evaluation engine to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, so as to obtain the fusion analysis result. The evaluation report generation module 14 is used to generate an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, so as to set the evaluation report as the target evaluation result and record it in the blockchain.

[0049] In some specific embodiments, the evaluation node determination module 11 may specifically include: The model type determination unit is used to obtain the evaluation request corresponding to the target large model and the model type corresponding to the target large model, and determine the corresponding preset node reputation database based on the evaluation request, so as to obtain a number of evaluation nodes from the preset node reputation database using the AI ​​arbitration node in the blockchain and the model type; wherein, the preset node reputation database records the historical evaluation quality score, response speed and professional domain label of each evaluation node; The matching score determination unit is used to determine the node information corresponding to each of the evaluation nodes, including historical evaluation quality scores, response speed, and professional domain labels, and to match each of the node information with the model type to obtain the corresponding matching score. Then, the evaluation node with the highest matching score is set as the target evaluation node. The professional domain labels include code evaluation labels and medical text evaluation labels. The historical evaluation quality score is a quality score determined based on the accuracy and reliability of historical evaluations. The response speed is a speed determined based on the latency of the evaluation node in processing the evaluation task. The node reputation database update unit is used to obtain the node performance corresponding to each of the evaluation nodes, and update the preset node reputation database based on the node performance to obtain a new preset node reputation database; the node performance includes the node's real-time load status, node response speed, and historical evaluation quality score.

[0050] In some specific embodiments, the reinforcement test question determination module 12 may specifically include: The model testing unit is used to call the target evaluation node and test the target large model based on a preset static dataset stored locally, generate model answers and preliminary scores, encapsulate the model answers and preliminary scores into smart contract transactions, and then store the smart contract transactions in the blockchain; wherein, the static dataset includes preset test questions and corresponding standard answers covering several domains and ability dimensions; The contract transaction analysis unit is used to use the consensus mechanism in the blockchain to detect the authenticity and integrity of the smart contract transaction data, and obtain the detection result. If the detection result indicates that the detection is passed, the target evaluation node is called to analyze the smart contract transaction and obtain preliminary analysis results including the performance indicators of the target large model on a specific task. The preliminary evaluation result determination unit is used to determine the preliminary evaluation result based on the smart contract transaction and the preliminary analysis result, and to record the preliminary evaluation result, timestamp and node identifier in the blockchain through the smart contract.

[0051] In some specific embodiments, the reinforcement test question determination module 12 may specifically include: The weakness identification unit is used to determine the weakness identification threshold and call the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain, so as to obtain the weakness identification result to be processed; the weakness identification result to be processed is the weakness of the target large model in the preset domain, preset capability or preset attribute. The target weakness identification result determination unit is used to determine whether each of the weakness identification results to be processed is greater than a preset weakness identification threshold. If the weakness identification result to be processed is greater than the preset weakness identification threshold, then the weakness identification result to be processed is set as the target weakness identification result. The reinforcement test question determination subunit is used to generate reinforcement test questions, including adversarial samples and targeted challenges, based on the target weakness identification results using a preset large language model. Each reinforcement test question is then encapsulated as a smart contract transaction and distributed to each evaluation node through the blockchain network.

[0052] In some specific embodiments, the fusion analysis result determination module 13 may specifically include: The evaluation result judgment unit is used to call the local resources corresponding to each of the target evaluation nodes and evaluate the target large model based on the current reinforcement test question to obtain the current evaluation result to be processed, including the model answer and performance indicators, and to determine whether the current evaluation result to be processed meets the preset stopping condition. The step jump unit is used to stop evaluating the target large model if the current evaluation result to be processed meets the preset stop condition, and to jump back to the step of calling the local resources corresponding to each of the target evaluation nodes and evaluating the target large model based on the current reinforcement test questions if the current evaluation result to be processed does not meet the preset stop condition. The preset stop condition includes the evaluation time reaching a preset time threshold, the number of evaluations reaching a preset recognition threshold, and the resource exhaustion of the evaluation node.

[0053] In some specific embodiments, the fusion analysis result determination module 13 may specifically include: The consistency evaluation result generation unit is used to obtain the preliminary evaluation result and the evaluation result to be processed from the blockchain, and call the multimodal learning model in the preset dynamic evaluation engine to evaluate the consistency of the preliminary evaluation result and the evaluation result to be processed on the same or similar questions at different times or by different nodes, so as to obtain the consistency evaluation result. The robustness evaluation result generation unit is used to call the graph neural network in the preset dynamic evaluation engine to evaluate the stability of the model output results under the inspection problem perturbation or adversarial attack on the preliminary evaluation results and the evaluation results to be processed, and obtain the robustness evaluation results. The biased evaluation result generation unit is used to call the causal inference model in the preset dynamic evaluation engine to evaluate the preliminary evaluation results and the evaluation results to be processed for specific groups or topics, and obtain biased evaluation results. The fusion analysis result generation unit is used to determine the fusion analysis result based on the result consistency evaluation result, the robustness evaluation result, and the bias evaluation result.

[0054] In some specific embodiments, the evaluation report generation module 14 may specifically include: The evaluation report encapsulation unit utilizes the report generation center in the AI ​​arbitration node and generates an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, a defect analysis, and improvement suggestions. The evaluation report is then encapsulated as a smart contract transaction record on the blockchain. The comprehensive score is a weighted score calculated based on the weights corresponding to the fusion analysis results, reflecting the overall performance of the target large model. The fine-grained capability radar chart displays the performance of the target large model across different dimensions. The defect analysis describes the weaknesses of the target large model and their corresponding impacts. The improvement suggestions are optimization measures corresponding to the defect analysis. The evaluation report can be queried and verified on the blockchain through an interactive interface and is immutable and traceable.

[0055] Furthermore, embodiments of this application also disclose an electronic device, Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the blockchain-based large-scale AI intelligent evaluation method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0056] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0057] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0058] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the blockchain-based large-model AI intelligent evaluation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0059] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed blockchain-based large-scale AI intelligent evaluation method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0060] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0061] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0062] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0063] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0064] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A blockchain-based AI intelligent evaluation method for large-scale models, characterized in that, include: Obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; Each target evaluation node is invoked and a preliminary evaluation of the target large model is performed based on a preset static evaluation dataset. The preliminary evaluation results are recorded in the blockchain through a smart contract. Then, the preset dynamic evaluation engine in the AI ​​arbitration node is invoked to identify weaknesses in the preliminary evaluation results in the blockchain, and weakness identification results are obtained. If the weakness identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated. Each of the target evaluation nodes is invoked and the target large model is evaluated based on the reinforcement test questions. The evaluation results to be processed are recorded in the blockchain. Then, the machine learning model in the preset dynamic evaluation engine is invoked to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, so as to obtain the fusion analysis results. Based on the fusion analysis results, an evaluation report is generated, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions. The evaluation report is then set as the target evaluation result and recorded in the blockchain.

2. The blockchain-based large-scale AI intelligent evaluation method according to claim 1, characterized in that, The process of obtaining the evaluation request corresponding to the target large model, and then determining the target evaluation node using the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model, includes: The system obtains the evaluation request corresponding to the target large model and the model type corresponding to the target large model, and determines the corresponding preset node reputation database based on the evaluation request. The system then uses the AI ​​arbitration node in the blockchain and the model type to obtain several evaluation nodes from the preset node reputation database. The preset node reputation database records the historical evaluation quality score, response speed and professional domain label of each evaluation node. Determine the node information corresponding to each of the evaluation nodes, including historical evaluation quality scores, response speed and professional domain labels, and match each of the node information with the model type to obtain the corresponding matching score. Then, set the evaluation node with the highest matching score among the various matching scores as the target evaluation node. The professional domain labels include code evaluation labels and medical text evaluation labels; the historical evaluation quality score is a quality score determined based on the accuracy and reliability of historical evaluations; and the response speed is a speed determined based on the delay time of the evaluation node in processing the evaluation task. The node performance corresponding to each of the evaluation nodes is obtained, and the preset node reputation database is updated based on the node performance to obtain a new preset node reputation database; the node performance includes the real-time load status of the node, the node response speed, and the historical evaluation quality score.

3. The blockchain-based large-scale AI intelligent evaluation method according to claim 1, characterized in that, The process of calling each target evaluation node and performing a preliminary evaluation of the target large model based on a preset static evaluation dataset, and recording the preliminary evaluation results in the blockchain via a smart contract, includes: The target evaluation node is invoked and the target large model is tested based on a preset static dataset stored locally. The model's answer and preliminary score are generated, and the model's answer and preliminary score are encapsulated into a smart contract transaction. The smart contract transaction is then stored in the blockchain. The static dataset includes preset test questions and corresponding standard answers covering several domains and ability dimensions. The consensus mechanism in the blockchain is used to detect the authenticity and integrity of the smart contract transaction data, and the detection result is obtained. If the detection result indicates that the detection is passed, the target evaluation node is called to analyze the smart contract transaction and obtain preliminary analysis results including the performance indicators of the target large model on a specific task. Based on the smart contract transactions and the preliminary analysis results, preliminary evaluation results are determined, and the preliminary evaluation results, timestamps, and node identifiers are recorded in the blockchain through smart contracts.

4. The blockchain-based large-scale AI intelligent evaluation method according to claim 1, characterized in that, The process involves calling the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results of the blockchain, obtaining weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated, including: A weakness identification threshold is determined, and the preset dynamic evaluation engine in the AI ​​arbitration node is invoked to identify weaknesses in the preliminary evaluation results in the blockchain, thereby obtaining the weakness identification results to be processed; the weakness identification results to be processed are the weaknesses of the target large model in preset domains, preset capabilities, or preset attributes; Determine whether each of the identified weaknesses to be processed is greater than a preset weakness identification threshold. If the identified weakness to be processed is greater than the preset weakness identification threshold, then set the identified weakness to be processed as the target weakness identification result. Using a pre-set large-scale language model and based on the target weakness identification results, reinforcement test questions, including adversarial samples and targeted challenges, are generated corresponding to the target weakness identification results. Each reinforcement test question is encapsulated as a smart contract transaction, and the smart contract transaction is distributed to each evaluation node through the blockchain network.

5. The blockchain-based large-scale AI intelligent evaluation method according to claim 1, characterized in that, The process of calling each of the target evaluation nodes and evaluating the target large model based on the reinforcement test questions, and recording the obtained evaluation results in the blockchain, includes: Call the local resources corresponding to each of the target evaluation nodes and evaluate the target large model based on the current reinforcement test question to obtain the current evaluation result to be processed, including the model answer and performance indicators, and determine whether the current evaluation result to be processed meets the preset stopping condition. If the current evaluation result to be processed meets the preset stopping condition, the operation of evaluating the target large model is stopped. If the current evaluation result to be processed does not meet the preset stopping condition, the process jumps back to the step of calling the local resources corresponding to each target evaluation node and evaluating the target large model based on the current reinforcement test questions. The preset stopping conditions include the evaluation time reaching a preset time threshold, the number of evaluations reaching a preset identification threshold, and the resource exhaustion of the evaluation node.

6. The blockchain-based large-scale AI intelligent evaluation method according to claim 1, characterized in that, The process involves calling the machine learning model in the preset dynamic evaluation engine to perform a multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, resulting in a fusion analysis result, including: The preliminary evaluation results and the evaluation results to be processed are obtained from the blockchain, and the multimodal learning model in the preset dynamic evaluation engine is called to evaluate the consistency of the results of the preliminary evaluation results and the evaluation results to be processed on the same or similar questions at different times or by different nodes, so as to obtain the result consistency evaluation result. The graph neural network in the preset dynamic evaluation engine is invoked to evaluate the stability of the model output results under the inspection problem perturbation or adversarial attack on the preliminary evaluation results and the evaluation results to be processed, so as to obtain the robustness evaluation results. The causal inference model in the preset dynamic evaluation engine is invoked to evaluate the preliminary evaluation results and the evaluation results to be processed for specific groups or topics, thereby obtaining biased evaluation results; The fusion analysis results are determined based on the consistency evaluation results, the robustness evaluation results, and the bias evaluation results.

7. The blockchain-based large-scale AI intelligent evaluation method according to any one of claims 1 to 6, characterized in that, The process of generating an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, and then recording the evaluation report as the target evaluation result in the blockchain, includes: The report generation center in the AI ​​arbitration node is used to generate an evaluation report containing a comprehensive score, fine-grained capability radar chart, defect analysis and improvement suggestions based on the fusion analysis results. The evaluation report is then encapsulated as a smart contract transaction record on the blockchain. The comprehensive score is a weighted score calculated based on the weights corresponding to the fusion analysis results, used to reflect the overall performance of the target large model; the fine-grained capability radar chart is used to display the performance of the target large model in different dimensions; the defect analysis is used to describe the weaknesses of the target large model and their corresponding impacts; the improvement suggestions are optimization measures corresponding to the defect analysis; the evaluation report can be queried and verified in the blockchain through an interactive interface, and has immutability and traceability.

8. A blockchain-based large-scale model AI intelligent evaluation device, characterized in that, include: The evaluation node determination module is used to obtain the evaluation request corresponding to the target large model, and then use the AI ​​arbitration node in the blockchain, the corresponding preset node reputation database, and the model type corresponding to the target large model to determine the target evaluation node; The reinforcement test question determination module is used to call each target evaluation node and perform a preliminary evaluation of the target large model based on a preset static evaluation dataset, and record the preliminary evaluation results in the blockchain through a smart contract. Then, it calls the preset dynamic evaluation engine in the AI ​​arbitration node to identify weaknesses in the preliminary evaluation results in the blockchain and obtain weakness identification results. If the weakness identification results indicate that the target large model has weaknesses, reinforcement test questions corresponding to the weaknesses are generated. The fusion analysis result determination module is used to call each of the target evaluation nodes and evaluate the target large model based on the reinforcement test questions, so as to record the obtained evaluation results to be processed in the blockchain, and then call the machine learning model in the preset dynamic evaluation engine to perform multi-dimensional fusion analysis on the preliminary evaluation results recorded in the blockchain and the evaluation results to be processed, so as to obtain the fusion analysis result. The evaluation report generation module is used to generate an evaluation report based on the fusion analysis results, including a comprehensive score, a fine-grained capability radar chart, defect analysis, and improvement suggestions, so as to set the evaluation report as the target evaluation result and record it in the blockchain.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the blockchain-based large-scale AI intelligent evaluation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs, wherein the computer programs, when executed by a processor, implement the blockchain-based large-model AI intelligent evaluation method as described in any one of claims 1 to 7.