Reward distribution method and device, computer equipment and storage medium
By registering artificial intelligence agents on the blockchain and using smart contracts for dynamic weight calculation, the problem of low flexibility in reward distribution in the metaverse environment is solved, the fairness and accuracy of reward distribution are achieved, and the continuous optimization performance of AI agents is encouraged.
Patent Information
- Application Number
- CN202510707410.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, the incentive mechanism design for artificial intelligence (AI Agent) in the metaverse environment is too static and general, lacking dynamic adaptability and personalization, resulting in low flexibility in reward distribution and an inability to accurately reflect the actual contribution of the AI Agent.
Register the artificial intelligence entity on the blockchain, obtain real-time task information and basic weights, perform summation calculations through smart contracts, dynamically adjust the dynamic weight value, calculate the target reward value based on the reward rule information, and generate dynamic weights based on historical task information and current task performance.
It achieves flexibility and accuracy in reward distribution, ensures fair and transparent reward distribution, eliminates the trust asymmetry problem, encourages AI agents to continuously optimize their performance, and tilts resources towards high-value contributors.
Smart Images

Figure CN120672031A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of blockchain, and in particular to reward distribution methods, devices, computer equipment, and storage media. Background Art
[0002] The metaverse is a complex confluence of cutting-edge technologies, including virtual reality, augmented reality, blockchain, and artificial intelligence. Within this virtual ecosystem of infinite possibilities, a multi-AI agent system, as the core architecture for complex interactions and intelligent operations, relies heavily on a sophisticated and effective incentive mechanism for efficient and coordinated operation. However, most current incentive mechanisms for AI agents used in the metaverse are relatively underdeveloped. The primary issue is that these incentive mechanisms are often overly static and generalized, lacking the necessary dynamic adaptability and personalized considerations, resulting in limited flexibility in reward allocation.
[0003] Currently, no effective solution has been proposed to the problem of low flexibility in reward allocation in related technologies. Summary of the Invention
[0004] The embodiments of the present application provide a reward distribution method, apparatus, computer device, and storage medium to at least solve the problem of low flexibility in reward distribution in related technologies.
[0005] In a first aspect, an embodiment of the present application provides a reward distribution method, the method comprising:
[0006] Registering artificial intelligence entities on the blockchain and obtaining real-time task information of each artificial intelligence entity performing a current task; the current task is published on the blockchain by a task publisher;
[0007] Obtaining reward rule information and basic weights written into a smart contract by the task publisher; the smart contract is deployed on the blockchain;
[0008] Using the smart contract, summing the basic weight and the real-time task information to obtain a dynamic weight value, and calculating a target reward value based on the reward rule information and the dynamic weight value;
[0009] Assign the target reward value to each of the artificial intelligence agents.
[0010] In some embodiments, the smart contract is used to calculate the sum of the basic weight and the real-time task information to obtain a dynamic weight value, including:
[0011] Obtaining historical task information of each of the artificial intelligence entities stored in the blockchain;
[0012] The basic weight, the historical task information, and the real-time task information are summed up using the smart contract to obtain the dynamic weight value.
[0013] In some embodiments, the summing up the basic weight, the historical task information, and the real-time task information to obtain the dynamic weight value includes:
[0014] Calculating historical contribution based on the historical task information, and calculating current contribution based on the real-time task information;
[0015] Assigning corresponding adjustment coefficients to the historical contribution and the current contribution respectively, and performing a weighted fusion calculation on the historical contribution and the current contribution based on the adjustment coefficients to obtain a dynamic performance parameter;
[0016] The basic weight and the dynamic performance parameter are summed to obtain the dynamic weight value.
[0017] In some embodiments, calculating the historical contribution based on the historical task information and calculating the current contribution based on the real-time task information include:
[0018] Obtaining the completion coefficient of each historical task performed by the artificial intelligence entity; calculating the historical contribution based on the historical task information and the completion coefficient;
[0019] Obtaining a difficulty coefficient of the current task; and calculating the current contribution based on the real-time task information and the difficulty coefficient.
[0020] In some embodiments, the calculating the target reward value based on the reward rule information and the dynamic weight value includes:
[0021] Obtaining the basic reward value of the current task written into the smart contract by the task publisher;
[0022] The target reward value is calculated according to the reward rule information using the smart contract and the basic reward value and the dynamic weight value.
[0023] In some embodiments, calculating the target reward value based on the basic reward value and the dynamic weight value includes:
[0024] Calculating the current task completion degree of the artificial intelligence agent in performing the current task, and obtaining the completion degree of similar tasks corresponding to the current task;
[0025] The task completion data is calculated based on the current task completion degree and the completion degree of similar tasks; and the target reward value is calculated based on the basic reward value, the dynamic weight value and the task completion data.
[0026] In some embodiments, obtaining real-time task information of each artificial intelligence agent performing a current task includes:
[0027] Obtaining capability attribute information of each of the artificial intelligence entities;
[0028] Based on the capability attribute information, the current task is assigned to an artificial intelligence agent that matches the current task for execution, and the real-time task information is obtained.
[0029] In a second aspect, an embodiment of the present application provides a reward distribution device, comprising:
[0030] A registration module, configured to register an AI on the blockchain and obtain real-time task information of each AI executing a current task; the current task is published on the blockchain by a task publisher;
[0031] A rule acquisition module, configured to acquire reward rule information and basic weights written into a smart contract by the task publisher; the smart contract is deployed on the blockchain;
[0032] a dynamic weight determination module, configured to utilize the smart contract to sum the basic weight and the real-time task information to obtain a dynamic weight value, and calculate a target reward value based on the reward rule information and the dynamic weight value;
[0033] An allocation module is used to allocate the target reward value to each of the artificial intelligence agents.
[0034] In a third aspect, an embodiment of the present application provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the reward distribution method as described in the first aspect above is implemented.
[0035] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the reward distribution method as described in the first aspect above.
[0036] Compared with related technologies, the reward distribution method, device, computer equipment and storage medium provided in the embodiments of the present application register artificial intelligence entities on the blockchain and obtain real-time task information of each artificial intelligence entity performing the current task; the current task is published on the blockchain by the task publisher; the reward rule information and basic weight written by the task publisher into the smart contract are obtained; the smart contract is deployed on the blockchain; the basic weight and real-time task information are summed and calculated using the smart contract to obtain a dynamic weight value, and the target reward value is calculated based on the reward rule information and the dynamic weight value; and the target reward value is distributed to each artificial intelligence entity.
[0037] Based on this, the introduction of dynamic weights injects real-time feedback into the incentive mechanism. The base weight, a quantitative indicator of the AI's initial capabilities, is combined with real-time task information to determine the dynamic weight, ensuring that reward distribution accurately reflects the AI's actual contribution. For example, AIs that excel in complex tasks receive higher dynamic weights, resulting in more rewards. This design not only incentivizes AIs to continuously optimize their performance but also directs resources toward high-value contributors. This approach effectively addresses the issue of low flexibility in reward distribution and improves its accuracy and reliability. Furthermore, the immutability of the blockchain ensures transparency and transparency of AI registration information, task execution records, and reward distribution rules. Task issuers cannot unilaterally tamper with the rules, allowing AIs to verify that their contributions are fairly assessed, eliminating the trust asymmetry that can exist in traditional centralized systems. The automated execution of smart contracts further reduces human intervention, ensuring that the reward distribution logic adheres strictly to pre-set rules, avoiding subjective bias or operational errors.
[0038] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0040] Figure 1 This is a hardware structure block diagram of a terminal according to a reward distribution method according to an embodiment of the present application;
[0041] Figure 2 is a flow chart of a reward distribution method according to an embodiment of the present application;
[0042] Figure 3 is a flow chart of another reward distribution method according to an embodiment of the present application;
[0043] Figure 4 This is a structural block diagram of a reward distribution device according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means and should not be understood as the contents disclosed in the present application being insufficient.
[0045] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.
[0046] Unless otherwise defined, technical or scientific terms used herein shall have the ordinary meaning as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "an," "the," and similar expressions used herein do not denote limitations on quantity and may refer to either the singular or the plural. The terms "comprise," "include," "have," and any variations thereof, used herein, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules (units) is not limited to the listed steps or units but may also include steps or units not listed, or may include other steps or units inherent to the process, method, product, or apparatus. The terms "connected," "connected," "coupled," and similar expressions used herein are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used herein, "plurality" means greater than or equal to two. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" may mean: A exists alone; A and B exist simultaneously; or B exists alone. The terms "first", "second", "third" and the like involved in this application are merely used to distinguish similar objects and do not represent a specific ordering of the objects.
[0047] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. Taking running on a terminal as an example, Figure 1 This is a hardware structure diagram of a terminal according to a reward distribution method of an embodiment of the present application. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0048] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the reward distribution method in the embodiments of the present application. The processor 102 executes the computer program stored in the memory 104 to perform various functional applications and data processing, that is, to implement the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0049] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the terminal's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0050] In related technologies, incentive mechanisms for AI agents often rely on fixed reward rules, failing to flexibly adjust to specific contexts, such as the AI agent's real-time performance, environmental changes, task difficulty, and urgency. For example, an AI agent might receive the same reward for completing a basic data collection task as it would for a high-risk, highly creative design task. This one-size-fits-all approach clearly fails to accurately reflect its actual contribution, leading to unfair reward distribution and failing to stimulate the AI agent's enthusiasm and creativity. Furthermore, existing incentive mechanisms often lack transparency and traceability, making it difficult to ensure the fairness and accuracy of rewards.
[0051] In order to solve the above problems, this embodiment provides a reward distribution method. Figure 2 is a flow chart of a reward distribution method according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:
[0052] Step S210: register the artificial intelligence entity on the blockchain and obtain real-time task information of each artificial intelligence entity performing the current task; the current task is published on the blockchain by the task publisher.
[0053] In the metaverse, numerous AI agents need to access blockchain networks to participate in various tasks and interactive activities. After a task publisher (such as an enterprise, developer, or individual user) publishes a task on the blockchain, each AI agent that wishes to participate completes registration through the blockchain's identity authentication system, generating a unique on-chain identity (such as a public key hash). This information, including the AI agent's identity, historical performance, capabilities, and contribution history, is recorded on the blockchain's distributed ledger. This ensures transparency and immutability, providing fundamental data support for subsequent task allocation and reward mechanisms.
[0054] Once registered, each AI can access and obtain detailed information about its current task through the blockchain's smart contract interface or decentralized application front-end. This information includes, but is not limited to, task description, objective, deadline, difficulty level, and basic requirements set by the task issuer. Due to the blockchain's real-time and transparent nature, AIs can instantly access the latest task updates, ensuring they accurately understand the requirements and make optimal decisions when executing tasks.
[0055] Step S220: Obtain the reward rule information and basic weight written into the smart contract by the task publisher; the smart contract is deployed on the blockchain.
[0056] When publishing a task, task publishers will write detailed reward rules into a smart contract deployed on the blockchain. These rules include the type of reward (such as monetary rewards, points rewards, resource rewards, etc.), the conditions for reward issuance (such as task completion, quality standards, time efficiency, etc.), the reward distribution method (such as fixed ratio distribution, distribution based on contribution, etc.), and any special reward or penalty mechanisms. The automated execution of smart contracts ensures the fairness and consistency of reward rules and reduces the possibility of human intervention.
[0057] In addition to reward rules, task issuers also set a base weight for each AI. This base weight is a quantitative measure of an AI's initial capabilities within the metaverse, representing a static assessment of its credibility, resource capacity, and historical performance. For example, if an AI's basic capabilities are assessed as "medium" (assuming a weight range of 0 to 1, with 1 being the highest), the base weight of that AI can be set to 0.5 and recorded on the blockchain. During task execution, the AI's contribution will be dynamically adjusted based on this weight and its real-time performance. This base weight serves as the fundamental parameter for reward calculations and provides a benchmark for subsequent dynamic adjustments.
[0058] In step S230, the basic weight and the real-time task information are summed up using the smart contract to obtain a dynamic weight value, and the target reward value is calculated based on the reward rule information and the dynamic weight value.
[0059] In this step, the smart contract is invoked to dynamically adjust the weights of each dimension based on the acquired real-time task information and base weights. This real-time task information specifically refers to the performance score of each AI agent in the current task, reflecting its real-time contribution. The following describes the dynamic weight calculation process: First, a mapping is maintained in the smart contract to record the base weight and dynamic weight of each AI agent. Then, during the dynamic weight calculation process, when an AI agent completes a task, an event is triggered. At this point, the smart contract calculates a new dynamic weight based on the performance score (i.e., the real-time task information) and the base weight. This dynamic adjustment mechanism fully accounts for various uncertainties during task execution, such as environmental changes, resource constraints, and unexpected situations, ensuring more reasonable and fair reward distribution.
[0060] Next, based on the adjusted dynamic weights and reward rules, the smart contract calculates the target reward for each participating AI, ensuring accuracy and fairness. The results are recorded on the blockchain for all AIs and task publishers to query and verify. Specifically, the reward rule information is first clarified, including the preset total reward pool size, reward distribution method (such as proportional distribution, step-by-step distribution, or fixed reward plus dynamic reward), upper and lower limits of individual rewards, etc.; then the dynamic weight value of each participant is obtained, which is usually dynamically calculated based on factors such as its performance and contribution in the task; and the dynamic weight values of all participants are normalized so that their sum is 1, so that the rewards can be distributed proportionally in the future; then, based on the normalized weight value and reward rule information, the target reward value of each participant is calculated by the algorithm in the smart contract, which is obtained after considering the total reward pool, reward distribution method and weight ratio; finally, the upper and lower limit constraints in the reward rule are applied to the calculated target reward value to ensure that the reward obtained by each participant is within a reasonable range. The entire process ensures the fairness and transparency of reward distribution through the automatic execution and tamper-proof characteristics of the smart contract.
[0061] Step S240: assign a target reward value to each artificial intelligence agent.
[0062] Once the target reward value is calculated, the smart contract automatically executes the reward distribution process. Based on the distribution method specified in the reward rules, the smart contract will send the corresponding rewards (such as currency, points, and resources) to the digital wallet or account of each participating AI. This process requires no human intervention, ensuring automated and immediate reward distribution. Because all reward distribution operations are recorded on the blockchain, they are highly transparent and traceable. Task issuers, other AIs, and any third party can query and verify the reward distribution results at any time, ensuring the fairness and credibility of the reward mechanism. This transparency also contributes to a healthier and more positive metaverse ecosystem, fostering healthy competition and cooperation among AIs.
[0063] In this reward distribution method, the synergy between blockchain and smart contracts creates a decentralized, transparent, and adaptive AI incentive system. Its core advantage lies in its deep integration of trust mechanisms, dynamic incentives, and resource allocation. First, the immutability of blockchain ensures the transparency of AI registration information, task execution records, and reward distribution rules. Task issuers cannot unilaterally tamper with the rules, allowing AIs to verify that their contributions are fairly assessed, thus eliminating the trust asymmetry that can exist in traditional centralized systems. The automated execution of smart contracts further reduces human intervention, ensuring that the reward distribution logic strictly adheres to pre-set rules and avoids subjective bias or operational errors. Second, the introduction of dynamic weights injects real-time feedback into the incentive mechanism. The base weight, a quantitative indicator of the AI's initial capabilities, is combined with real-time task information to determine the dynamic weight, ensuring that reward distribution accurately reflects the AI's actual contribution. For example, AIs that excel in complex tasks receive higher dynamic weights, resulting in more rewards. This design not only incentivizes AIs to continuously optimize their performance but also directs resources toward high-value contributors. Therefore, the above method effectively solves the problem of low flexibility in reward allocation and improves the accuracy and reliability of reward allocation.
[0064] In some embodiments, the above-mentioned use of smart contracts to sum the basic weight and real-time task information to obtain a dynamic weight value may also include the following steps:
[0065] Obtain the historical task information of each artificial intelligence entity stored in the blockchain; use smart contracts to sum up the basic weight, historical task information and real-time task information to obtain the dynamic weight value.
[0066] The aforementioned historical task information specifically refers to the performance scores of each AI entity performing each historical task. In this step, the calculated historical task information for each AI entity performing each historical task is recorded on the blockchain. This allows the smart contract to be invoked to retrieve each AI entity's historical task records when calculating dynamic weights. During dynamic weight adjustments, the smart contract combines historical task information with real-time task information for real-time updates and adjustments. Specifically, the smart contract can first standardize the historical and real-time task information to eliminate dimensional differences. It then calculates the dynamic weight of each AI entity by weighting the base weight, historical task information, and real-time task information according to a preset ratio.
[0067] Through the above embodiments, by combining basic weights, historical task information and real-time task information, the smart contract can dynamically calculate the weight value of the artificial intelligence entity, which helps to reflect the long-term capabilities and stability of the artificial intelligence entity, thereby improving the fairness and incentives of reward distribution and promoting the active participation and continuous optimization of the artificial intelligence entity in tasks.
[0068] In some embodiments, the above-mentioned summing up of the basic weight, historical task information, and real-time task information to obtain the dynamic weight value may further include the following steps:
[0069] Based on historical task information, the historical contribution is calculated, and based on real-time task information, the current contribution is calculated; corresponding adjustment coefficients are assigned to the historical contribution and the current contribution respectively, and based on the adjustment coefficients, the historical contribution and the current contribution are weighted and fused to obtain dynamic performance parameters; the basic weight and the dynamic performance parameter are summed to obtain the dynamic weight value.
[0070] Historical contribution can be calculated based on factors such as the number of historical tasks, completion quality, and task difficulty. For example, a formula can be defined: Historical Contribution = ∑(Historical Task Difficulty Coefficient × Historical Task Completion Quality Coefficient × Historical Time Decay Coefficient); the task difficulty coefficient is set based on the difficulty level of the task, the task completion quality coefficient is set based on the completion result of the task (such as excellent, good, or average), and the time decay coefficient is used to represent the impact of task completion time on contribution. Tasks completed earlier may have lower contribution.
[0071] Similarly, the current contribution can be calculated based on factors such as the number of real-time tasks, completion progress, and estimated completion time. For example, a formula can be defined: Current Contribution = ∑(Current Task Difficulty Coefficient × Current Task Completion Progress Coefficient × Urgency Coefficient); the Urgency Coefficient is set based on the urgency of the task (e.g., high, medium, or low).
[0072] Then, based on the historical contribution and current contribution, a corresponding adjustment coefficient is assigned. The adjustment coefficient can be flexibly set according to the actual situation. For example, if the historical contribution is high but the current contribution is low, it may indicate poor recent performance, and a lower adjustment coefficient can be assigned to the current contribution; if the current contribution continues to rise, a higher adjustment coefficient can be assigned to encourage current performance. Alternatively, the adjustment can be made based on the difficulty coefficient of the current task. The higher the difficulty coefficient, the smaller the adjustment coefficient corresponding to the historical contribution and the larger the adjustment coefficient corresponding to the current contribution.
[0073] Next, based on the adjustment coefficient, the historical contribution and current contribution are weighted and combined to obtain the dynamic performance parameter. For example, the dynamic performance parameter = historical contribution × historical adjustment coefficient + current contribution × current adjustment coefficient. Finally, the basic weight and the dynamic performance parameter are summed to obtain the dynamic weight value.
[0074] Through the above embodiment, the dynamic performance parameter is calculated based on the adjustment coefficient, which can comprehensively reflect the comprehensive performance of historical and current contributions under different weights, thereby providing a more accurate basis for subsequent dynamic weight calculation.
[0075] In some embodiments, the above-mentioned calculation of historical contribution based on historical task information and calculation of current contribution based on real-time task information may further include the following steps:
[0076] Obtain the completion coefficient of each historical task performed by the artificial intelligence body; calculate the historical contribution based on the historical task information and the historical completion coefficient; obtain the difficulty coefficient of the current task; calculate the current contribution based on the real-time task information and the difficulty coefficient.
[0077] Among them, the above-mentioned task contribution and task difficulty coefficient can be obtained by the blockchain using natural language processing or image recognition technology to perform a quality score on the task output results, calculate the task contribution, and use a recurrent neural network to perform semantic analysis on the task requirements to generate a task difficulty score; then call the smart contract to read the above data and calculate the dynamic weight.
[0078] More specifically, one way to calculate the dynamic weight is by combining multiple factors such as the basic weight of the AI, historical task performance, current task contribution, and task difficulty coefficient, as shown in the following formula:
[0079] ;
[0080] In the above formula, DW i Used to represent dynamic weight value. BW iIt is used to represent the basic weight of the AI, representing its initial ability or reputation. α and β are adjustment coefficients for historical contribution and current contribution, respectively, used to balance the impact of historical performance on the current weight and the contribution of the current task to the weight. H refers to the number of historical tasks completed by the AI. P ik It is the performance score of the artificial intelligence agent in the kth historical task, which is a value between 0 and 1. ik P is the completion coefficient of the kth historical task, reflecting the percentage or quality of task completion. ic is the performance score of the artificial intelligence agent in the current task c. c is the difficulty coefficient of the current task c. The higher the difficulty, the greater the coefficient.
[0081] Through the above embodiments, a complete dynamic weight generation mechanism is provided, which ensures the comprehensiveness of the reference factors for dynamic weight generation, thereby helping to improve the accuracy of subsequent reward distribution.
[0082] In some embodiments, the calculation of the target reward value based on the basic reward value and the dynamic weight value may further include the following steps:
[0083] Get the base reward value of the current task written into the smart contract by the task publisher; use the smart contract to calculate the target reward value based on the base reward value and dynamic weight value according to the reward rule information.
[0084] The basic reward value for the current task can be calculated using the corresponding formula based on the set basic reward benchmark and task parameters. For example, basic reward value = basic reward coefficient × time investment × resource consumption coefficient × technical difficulty coefficient × market value coefficient; each coefficient is determined based on historical data and actual scenario requirements.
[0085] The smart contract then calculates the target reward value according to the reward rule information; the reward rule information can be specifically expressed as the following formula:
[0086] R i =BR×DW i ;
[0087] In the above formula, BR is the basic reward value of the task; DW i is the dynamic weight value of the i-th artificial intelligence entity; R i is the target reward value of the i-th artificial intelligence agent.
[0088] Through the above embodiment, based on the basic reward value of the current task, a clear value benchmark is provided for subsequent calculations, ensuring the rationality and fairness of reward distribution.
[0089] In some embodiments, the calculation of the target reward value based on the basic reward value and the dynamic weight value may further include the following steps:
[0090] Calculate the current task completion degree of the artificial intelligence entity performing the current task, and obtain the completion degree of similar tasks corresponding to the current task; calculate the task completion data based on the current task completion degree and the completion degree of similar tasks; calculate the target reward value based on the basic reward value, dynamic weight value and task completion data.
[0091] In the process of calculating the target reward value, it is first necessary to quantitatively evaluate the completion of the current task performed by the artificial intelligence entity, that is, to calculate the current task completion degree. This usually involves a comprehensive consideration of the degree of achievement of various task indicators, such as the proportion of achievement of task goals, compliance with quality standards, etc.; at the same time, in order to more comprehensively measure the current task performance, it is also necessary to obtain the completion data of similar tasks corresponding to the current task. This data can be obtained by analyzing the past execution of similar tasks or the industry average level.
[0092] Next, task completion data is calculated based on the comparison between the current task's completion rate and that of similar tasks, such as the difference, ratio, or a comprehensive indicator derived from a specific algorithm. This data more accurately reflects the current task's relative performance among similar tasks. Finally, the target reward value is calculated by combining the base reward value (reflecting the fundamental value of the task itself), the dynamic weight value (dynamically adjusted based on factors such as task importance and urgency), and the task completion data. For example, target reward value = base reward value × dynamic weight value × task completion data.
[0093] Through the above embodiments, the value, importance and actual completion status of the task itself can be comprehensively considered, so that the calculation of the target reward value is more scientific and reasonable, and the artificial intelligence entity is effectively motivated to complete the task better.
[0094] In some embodiments, the above-mentioned acquisition of real-time task information of each artificial intelligence agent performing the current task may further include the following steps:
[0095] Obtain the capability attribute information of each artificial intelligence entity; based on the capability attribute information, assign the current task to the artificial intelligence entity that matches the current task to execute, and obtain real-time task information.
[0096] In this step, we obtain information about the AI's training data, model architecture, and computing power configuration. Using data analysis tools, we analyze the collected capabilities and extract key features, thus obtaining the aforementioned capability attribute information. Next, we use matching algorithms (such as cosine similarity and decision trees) to assign tasks based on the task requirements and the AI's capabilities. Based on the matching results, we assign the task to the most suitable AI.
[0097] Through the above embodiments, by obtaining the capability attribute information of each artificial intelligence entity and assigning the current task to an artificial intelligence entity that matches the task based on this information, the capabilities of the artificial intelligence entity can be more effectively utilized and the efficiency and quality of task execution can be improved.
[0098] The present application is described in detail below with reference to specific embodiments. Figure 3 is a flow chart of another reward distribution method according to an embodiment of the present application. Figure 3 As shown, the process includes the following steps:
[0099] Step S310: AI Agent Registration and Authentication. Each AI agent is registered on the blockchain and receives a unique identity. The AI agent's identity information, historical performance, capabilities, and contribution records are all stored on the blockchain, ensuring data immutability and transparency.
[0100] In step S320, the task publisher publishes the task to the blockchain. Specifically, the task publisher in the metaverse publishes the task information (including task description, required capability attributes, etc.) to the blockchain.
[0101] Step S330: Setting reward rules using smart contracts. The task publisher sets reward rules using smart contracts.
[0102] Step S340: Task matching. Task matching is performed based on the task requirements and the AI Agent's capabilities. A decentralized evaluation mechanism is introduced, whereby other AI Agents, users, or a pre-defined evaluation algorithm provide real-time evaluation of the AI Agent's performance during task execution.
[0103] Step S350: Real-time performance evaluation and weight adjustment. The calculation of dynamic weights combines multiple factors such as the AI Agent's base weight, historical task performance, current task contribution, and task difficulty coefficient.
[0104] Step S360: Reward distribution and incentive mechanism update. Specifically, rewards are automatically calculated and distributed based on the reward rules and dynamic weights set in the smart contract. Rewards can be tokens, digital assets, or other forms of incentive resources.
[0105] More specifically, suppose there is an AI agent in the metaverse, numbered A1. A1's base weight (BW A1 ) is 0.5. The adjustment coefficients (α and β) are 0.3 and 0.2 respectively. A1 has completed two tasks in the past, numbered T1 and T2. The completion coefficient of T1 (C T1 ) is 0.8, and the performance score of A1 in T1 (P A1T1) is 0.7. The completion coefficient of T2 (C T2 ) is 0.9, and the performance score of A1 in T2 (P A1T2 ) is 0.6. The current task number is T3, and the difficulty coefficient (D T3 ) is 1.2, and the performance score of A1 in T3 (P A1T3 ) is 0.8. The base reward of T3 (BR T3 ) is 100 tokens.
[0106] First, calculate the contribution of A1 in historical tasks T1 and T2, as shown in the following formula:
[0107] T1's contribution = P A1T1 ×C T1 =0.7×0.8=0.56;
[0108] T2's contribution = P A1T2 ×C T2 =0.6×0.9=0.54;
[0109] Next, calculate the dynamic weight of A1 (DW A1 ):
[0110] DW A1 =BW A1 +α×(Contribution T1+Contribution T2)+β×(P A1T3 ×D T3 )=1.022;
[0111] Finally, the reward (R A1 ):
[0112] R A1 =BR T3 ×DW A1 =100×1.022102.2 (tokens);
[0113] As can be seen from the above examples, dynamic weight adjustment enables a dynamic and personalized incentive mechanism, allowing for flexible adjustments based on the AI agent's actual performance and task requirements. The dynamic weight calculation formula comprehensively considers the AI agent's base weight, historical performance, current task contribution, and task difficulty coefficient, providing a more comprehensive reflection of the AI agent's capabilities and performance. Furthermore, all task information, AI agent performance evaluations, and reward distribution are recorded on the blockchain, ensuring data transparency and traceability. Smart contracts automatically execute reward distribution, avoiding backroom dealings and human intervention, and ensuring the fairness of the incentive mechanism.
[0114] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0115] This embodiment also provides a reward distribution device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the terms "module," "unit," "subunit," etc. may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0116] Figure 4 is a structural block diagram of a reward distribution device according to an embodiment of the present application, such as Figure 4 As shown, the device includes: a registration module 41, a rule acquisition module 42, a dynamic weight determination module 43 and an allocation module 44; wherein:
[0117] The registration module 41 is used to register the artificial intelligence entity on the blockchain and obtain the real-time task information of each artificial intelligence entity performing the current task; the current task is published on the blockchain by the task publisher; the rule acquisition module 42 is used to obtain the reward rule information and basic weight written by the task publisher into the smart contract; the smart contract is deployed on the blockchain; the dynamic weight determination module 43 is used to use the smart contract to sum the basic weight and real-time task information to obtain the dynamic weight value, and calculate the target reward value based on the reward rule information and the dynamic weight value; the allocation module 44 is used to allocate the target reward value to each artificial intelligence entity.
[0118] In some embodiments, the dynamic weight determination module 43 is also used to obtain historical task information of each artificial intelligence entity stored in the blockchain; the dynamic weight determination module 43 uses a smart contract to sum the basic weight, historical task information and real-time task information to obtain a dynamic weight value.
[0119] In some embodiments, the above-mentioned dynamic weight determination module 43 is also used to calculate the historical contribution based on historical task information, and to calculate the current contribution based on real-time task information; the dynamic weight determination module 43 assigns corresponding adjustment coefficients to the historical contribution and the current contribution respectively, and calculates the dynamic performance parameters based on the historical contribution and the current contribution based on the adjustment coefficient; the dynamic weight determination module 43 calculates the dynamic weight value based on the basic weight and the dynamic performance parameters.
[0120] In some embodiments, the dynamic weight determination module 43 is also used to obtain the completion coefficient of each historical task performed by the artificial intelligence entity; calculate the historical contribution based on the historical task information and the completion coefficient; the dynamic weight determination module 43 obtains the difficulty coefficient of the current task; and calculates the current contribution based on the real-time task information and the difficulty coefficient.
[0121] In some embodiments, the dynamic weight determination module 43 is also used to obtain the basic reward value of the current task written by the task publisher into the smart contract; using the smart contract, according to the reward rule information, the target reward value is calculated based on the basic reward value and the dynamic weight value.
[0122] In some embodiments, the dynamic weight determination module 43 is also used to calculate the current task completion degree of the artificial intelligence entity performing the current task, and obtain the completion degree of similar tasks corresponding to the current task; the dynamic weight determination module is also used to calculate the task completion data based on the current task completion degree and the completion degree of similar tasks; and calculate the target reward value based on the basic reward value, the dynamic weight value and the task completion data.
[0123] In some embodiments, the registration module 41 is further configured to obtain capability attribute information of each artificial intelligence entity; based on the capability attribute information, the current task is assigned to an artificial intelligence entity that matches the current task for execution, and real-time task information is obtained.
[0124] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination. For specific examples in this embodiment, reference can be made to the examples described in the above embodiment and optional implementations, and will not be repeated in this embodiment.
[0125] This embodiment further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0126] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0127] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0128] S1, register the artificial intelligence entity on the blockchain and obtain the real-time task information of each artificial intelligence entity performing the current task; the current task is published on the blockchain by the task publisher.
[0129] S2, obtains the reward rule information and basic weight written into the smart contract by the task publisher; the smart contract is deployed on the blockchain.
[0130] S3 uses smart contracts to sum the basic weight and real-time task information to obtain a dynamic weight value, and calculates the target reward value based on the reward rule information and the dynamic weight value.
[0131] S4, assigns target reward values to each artificial intelligence agent.
[0132] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.
[0133] In addition, in conjunction with the reward distribution method in the above embodiments, the present application embodiment may provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the reward distribution methods in the above embodiments is implemented.
[0134] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0135] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0136] Those skilled in the art should understand that the various technical features of the above-described embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0137] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A reward distribution method, characterized in that: The method comprises: Registering artificial intelligence entities on the blockchain and obtaining real-time task information of each artificial intelligence entity performing a current task; the current task is published on the blockchain by a task publisher; Obtaining reward rule information and basic weights written into a smart contract by the task publisher; the smart contract is deployed on the blockchain; Using the smart contract, summing the basic weight and the real-time task information to obtain a dynamic weight value, and calculating a target reward value based on the reward rule information and the dynamic weight value; Assign the target reward value to each of the artificial intelligence agents.
2. The reward distribution method according to claim 1, characterized in that: The method of utilizing the smart contract to sum the basic weight and the real-time task information to obtain a dynamic weight value includes: Obtaining historical task information of each of the artificial intelligence entities stored in the blockchain; The basic weight, the historical task information, and the real-time task information are summed up using the smart contract to obtain the dynamic weight value.
3. The reward distribution method according to claim 2, characterized in that: The summing up of the basic weight, the historical task information, and the real-time task information to obtain the dynamic weight value includes: Calculating historical contribution based on the historical task information, and calculating current contribution based on the real-time task information; Assigning corresponding adjustment coefficients to the historical contribution and the current contribution respectively, and performing a weighted fusion calculation on the historical contribution and the current contribution based on the adjustment coefficients to obtain a dynamic performance parameter; The basic weight and the dynamic performance parameter are summed to obtain the dynamic weight value.
4. The reward distribution method according to claim 3, characterized in that: The calculating of the historical contribution according to the historical task information and the calculating of the current contribution according to the real-time task information include: Obtaining the completion coefficient of each historical task performed by the artificial intelligence entity; calculating the historical contribution based on the historical task information and the completion coefficient; Obtaining a difficulty coefficient of the current task; and calculating the current contribution based on the real-time task information and the difficulty coefficient.
5. The reward distribution method according to claim 1, characterized in that: The calculating the target reward value based on the reward rule information and the dynamic weight value includes: Obtaining the basic reward value of the current task written into the smart contract by the task publisher; The target reward value is calculated according to the smart contract, the basic reward value and the dynamic weight value in accordance with the reward rule information.
6. The reward distribution method according to claim 5, characterized in that: The calculating the target reward value according to the basic reward value and the dynamic weight value includes: Calculating the current task completion degree of the artificial intelligence agent in performing the current task, and obtaining the completion degree of similar tasks corresponding to the current task; The task completion data is calculated based on the current task completion degree and the completion degree of similar tasks; and the target reward value is calculated based on the basic reward value, the dynamic weight value and the task completion data.
7. The reward distribution method according to any one of claims 1 to 6, characterized in that: The obtaining of real-time task information of each artificial intelligence agent performing a current task includes: Obtaining capability attribute information of each of the artificial intelligence entities; Based on the capability attribute information, the current task is assigned to an artificial intelligence agent that matches the current task for execution, and the real-time task information is obtained.
8. A reward distribution device, characterized in that: include: A registration module, used to register artificial intelligence entities on the blockchain and obtain real-time task information of each artificial intelligence entity performing the current task; The current task is published on the blockchain by the task publisher; A rule acquisition module, used to obtain the reward rule information and basic weight written into the smart contract by the task publisher; The smart contract is deployed on the blockchain; a dynamic weight determination module, configured to utilize the smart contract to sum the basic weight and the real-time task information to obtain a dynamic weight value, and calculate a target reward value based on the reward rule information and the dynamic weight value; An allocation module is used to allocate the target reward value to each of the artificial intelligence agents.
9. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the reward distribution method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the reward distribution method according to any one of claims 1 to 7 when running.
Citation Information
Cited By
Data traceability system, device and equipment based on block chain technology
CN121258549A