A method and system for NFT digital asset auction based on reinforcement learning
By using a hierarchical reinforcement learning model to predict the final price of NFTs in real time and dynamically adjust the auction strategy, the problem of existing NFT auction systems being unable to adapt to market changes is solved, auction revenue and user experience are improved, and the decentralization and transparency of the system are ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIV
- Filing Date
- 2024-11-22
- Publication Date
- 2026-05-22
AI Technical Summary
Existing NFT auction systems cannot dynamically adjust auction strategies or adapt to real-time market changes, resulting in final auction prices that are lower than market expectations. This increases bidders' costs and transaction costs, and the lack of a real-time price prediction mechanism affects auction efficiency and user satisfaction.
A hierarchical reinforcement learning model is adopted, including predictive agents and adaptive agents, to predict the final price of NFTs in real time and dynamically adjust the auction strategy. The decentralization and security of the model are ensured through blockchain consensus nodes.
It enables adaptive adjustments to the NFT auction system, increases auction revenue, reduces bidder costs, enhances user experience and auction efficiency, and ensures the system's decentralization and transparency.
Smart Images

Figure CN122072928A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of blockchain technology, specifically relating to a method and system for auctioning NFT digital assets based on reinforcement learning. Background Technology
[0002] In a narrow sense, blockchain is a chain-like data structure that combines data blocks sequentially according to time, and uses cryptography to ensure its immutability and tamper-proof nature as a distributed ledger. In a broader sense, blockchain technology utilizes a chain-like data structure to verify and store data, uses distributed node consensus algorithms to generate and update data, uses cryptography to ensure the security of data transmission and access, and uses smart contracts composed of automated script code to program and manipulate data, forming a completely new distributed infrastructure and computing paradigm. Blockchain technology provides fundamental guarantees for the security, transparency, and decentralization of digital assets.
[0003] Multimodal digital assets refer to digital assets based on different forms and types, such as cryptocurrencies, tokens, digital artworks, and virtual real estate. Multimodal digital assets are characterized by decentralization, transparency, security, and globalization, enabling trusted peer-to-peer transactions through blockchain technology. The management and trading of multimodal digital assets require flexible auction mechanisms to adapt to different asset types and trading needs.
[0004] An auction model refers to a mechanism for trading goods or services under specific rules. Common auction models include English auctions, Dutch auctions, second-price auctions (Vickrey auctions), and sealed-bid auctions. Each auction model has its applicable scenarios and advantages; by selecting the appropriate auction model, the interests of both sellers and buyers can be maximized. The introduction of reinforcement learning technology allows for the dynamic selection and optimization of auction models based on specific auction needs, thereby improving transaction efficiency and effectiveness.
[0005] Reinforcement learning is a machine learning method that learns optimal policies by interacting with the environment and receiving feedback (rewards or penalties). Reinforcement learning algorithms are widely used in complex tasks requiring decision-making, such as games, robot control, and financial trading. Its core principle is to progressively optimize the policy through a trial-and-error process and by maximizing long-term rewards. Reinforcement learning models can adaptively adjust auction strategies to ensure optimal auction results in different situations, improving the intelligence and efficiency of the auction process.
[0006] Smart contracts are automated scripts deployed on a blockchain that automatically execute contract terms when predetermined conditions are met. Key characteristics of smart contracts include automated execution, immutability, and transparency. Through smart contracts, the auction process for digital assets can be automated and executed without trust, ensuring fairness and efficiency in transactions. The application of smart contracts in multimodal digital asset auctions can automate the entire process from user input to auction algorithm matching.
[0007] An oracle is a mechanism that introduces external data into a blockchain, solving the problem of connecting the blockchain with the outside world. Oracles can provide smart contracts with reliable external data, such as market prices, weather data, and sporting event results. In multimodal digital asset auctions, oracles can provide real-time market data and other relevant information, ensuring the accuracy and timeliness of data during the auction process.
[0008] Patent CN118134613A proposes a model trading method based on hierarchical reinforcement learning in personalized federated learning, applicable to multi-user auction model scenarios. This method establishes a bilateral auction trading platform to incentivize model trading among participants, prioritizing the protection of privacy and security in local datasets. A hierarchical multi-agent reinforcement learning algorithm optimizes users' bidding strategies, balancing competition and cooperation, ultimately achieving stable participant returns and optimal overall social welfare.
[0009] Patent CN117350410A proposes a microgrid group collaborative operation optimization method based on multi-agent federated reinforcement learning, applicable to the field of distribution network operation and control technology. This method constructs a microgrid group interactive optimization operation model and a Markov decision model, employing an attention-based multi-agent federated reinforcement learning method for optimization management. While protecting the privacy of all parties in the microgrid group, it solves the problems of gradient instability and low learning efficiency in traditional federated reinforcement learning methods, achieving collaborative optimization of multi-party strategies.
[0010] Patent CN111932293A proposes a blockchain-based method for secure resource transactions. It utilizes an iterative dual-auction mechanism to protect the privacy of both buyers and sellers, maximizing benefits for both parties. By storing transaction information using blockchain technology, it solves the problem of information leakage risks associated with traditional centralized systems, achieving decentralization and ensuring the immutability of transaction records.
[0011] The paper "A reinforcement learning model for the reliability of blockchainoracles" proposes a Bayesian reinforcement learning oracle reliability (BLOR) mechanism to identify trustless and cost-effective oracles. BLOR learns oracle behavior by developing a Bayesian cost-dependent reputation model and utilizes reinforcement learning (knowledge gradient algorithm) to guide the learning process. While the BLOR mechanism performs well in selecting oracles, it still has some shortcomings and problems. First, BLOR relies on the Bayesian cost-dependent reputation model and knowledge gradient algorithm, which requires significant computational resources and time, potentially leading to inefficiencies in large-scale applications.
[0012] The paper "Utilizing Deep Reinforcement Learning and Q-Learning algorithms for Improved Ethereum Cybersecurity" explores and develops deep reinforcement learning (Deep RL) and Q-Learning algorithms to improve Ethereum's network security. The research adopts a design science research paradigm from information systems research, constructing deep reinforcement learning and Q-Learning algorithms to enhance Ethereum's network security. This framework uses a policy based on model-free deep deterministic policy gradient (DDPG) to find the optimal incentive in a continuous action space to persuade users to reduce energy consumption. While this research has made some progress in improving Ethereum's network security, these algorithms may encounter performance bottlenecks when handling high-dimensional and dynamic network security problems.
[0013] The main problems with existing technologies in multimodal digital asset auctions are: traditional auction systems typically rely on a single auction model, such as English or Dutch auctions, and cannot dynamically select the optimal auction model based on different auction needs, making it difficult to achieve optimal results in a volatile market environment. Existing auction systems lack intelligent optimization mechanisms and cannot adjust auction strategies based on real-time data and market changes, thus affecting auction efficiency and user satisfaction. Furthermore, user privacy and data security are not adequately protected during the auction process, posing risks of information leakage and unfair competition. Existing technologies also fall short in automating the auction process, relying heavily on manual intervention, which impacts transaction efficiency and transparency. Summary of the Invention
[0014] The main technical problem to be solved by this invention is:
[0015] 1) Existing NFT auction systems are based on fixed rules and lack dynamic adjustment capabilities, making them unable to adapt to real-time market changes. This results in final auction prices falling below market expectations, reducing liquidity and increasing costs for bidders. The lack of a real-time price prediction mechanism during dynamic auctions leads to a poor auction experience for sellers.
[0016] 2) Current NFT auction systems rely on fixed smart contract rules for dynamic parameter adjustments, making it difficult to adjust key parameters (such as bid increments and auction durations) in real time according to the bidding situation. Therefore, they cannot achieve optimal auction results when demand fluctuates drastically.
[0017] 3) Traditional forecasting methods are usually based on historical data and are independent of the auction system. They cannot provide price forecasts in real time during the auction process, which affects the accuracy of prices and leads to a decrease in auction efficiency and user satisfaction.
[0018] The technical solution adopted in this invention is as follows:
[0019] A reinforcement learning-based method for auctioning NFT digital assets includes the following steps:
[0020] Deploy a hierarchical reinforcement learning model in a blockchain auction smart contract, the hierarchical reinforcement learning model including a predictive agent and an adjusting agent;
[0021] During the auction process, the final price of the NFT is predicted in real time by a predictive agent, and the seller's auction strategy is dynamically adjusted in real time by adjusting the agent, thereby increasing auction revenue and reducing bidder costs.
[0022] Furthermore, the predictive agent predicts key auction parameters, including the starting price and the reserve price, and dynamically predicts the final auction price; the adjusting agent adjusts the auction parameters according to the real-time auction status, raising the price during periods of high bidding activity and lowering the increment during periods of bidding stagnation, thus enabling the auction strategy to be adaptively adjusted.
[0023] Furthermore, the update of the hierarchical reinforcement learning model requires consensus from all consensus nodes in the blockchain. If a malicious update proposal is made, other nodes will reject the proposal to prevent the model from being updated improperly. When a consensus node proposes a model update, it must be accompanied by the aggregate signature of all nodes to ensure decentralization and security through the blockchain's consensus protocol.
[0024] Furthermore, the creation phase of the auction process includes:
[0025] Set and adjust the initial auction parameters, including the starting price P. s Reservation price P r Temporary buyout price P band bidding increment P m ;
[0026] The seller submits auction information, which is then converted into the initial state s0 and used as input for the predictive agent;
[0027] The predictive agent first generates a temporary action a. temp1 To adjust P s and P r The temporary action then updates the state, forming a temporary state s. temp1 And pass it on to the adjusting agent;
[0028] Adjust the agent to generate another temporary action a temp2 To adjust P b and P m The newly generated temporary state s temp2 And return it to the predictive agent;
[0029] The predictive agent outputs the predicted price P and combines it with other temporary parameters to generate the final action a0.
[0030] Furthermore, the bidding phase of the auction process includes:
[0031] The current state s t Input is fed into the adjusting agent to generate a temporary action a. temp To adjust P b P m And auction extension time T e Then update the status s temp And pass it on to the predictive agent;
[0032] The predictive agent adjusts the price P based on the updated data and outputs the final action a. t .
[0033] Furthermore, the settlement phase of the auction process includes:
[0034] Based on the collected data from the bidding and creation phases, the hierarchical reinforcement learning model is updated using a proximal policy optimization algorithm.
[0035] The NFT auction problem can be represented as a quadruple (S, A, P, R):
[0036] State space S: Represents all possible states of the auction environment, including the individual state spaces of the predicting agent and the adjusting agent;
[0037] Operational Space A: This includes the operational spaces of the predictive agent and the adjusting agent. The predictive agent achieves accurate parameter prediction through a three-layer fully connected network, while the adjusting agent is responsible for quickly responding to market changes and dynamically adjusting auction parameters.
[0038] State transition function P: At time step t, the state transition function P(s) t+1 |s t ,a t Describe the action a being performed. t After from state s t Transition to s t+1 The probability of;
[0039] The reward function R is used to evaluate the predictive agent and adjust the agent's behavior in a specific state. For the predictive agent, the reward function is defined based on the similarity between the actual transaction price and the predicted final price. If the auction is successful, R represents the similarity between the actual price and the final price. If the auction fails, a penalty coefficient F is applied, such that R = -F.
[0040] Furthermore, the hierarchical reinforcement learning model is trained using a proximal policy optimization method, and the auction policy is updated based on the observed rewards and transitions. Over time, the auction policy converges to the optimal policy to maximize the cumulative expected return and achieve Nash equilibrium.
[0041] An NFT digital asset auction system based on reinforcement learning includes a blockchain and a hierarchical reinforcement learning model deployed in an auction smart contract on the blockchain. The hierarchical reinforcement learning model includes a predictive agent and an adjusting agent. During the auction process, the predictive agent predicts the final price of the NFT in real time, and the adjusting agent dynamically adjusts the seller's auction strategy in real time, thereby increasing auction revenue and reducing bidder costs.
[0042] The beneficial effects of this invention are as follows:
[0043] 1. This invention proposes a novel auction framework for NFT auctions in the Web 3.0 digital economy. By employing a hierarchical reinforcement learning (HRL) structure, it can predict the final price of NFTs in real time during the auction process and dynamically optimize auction parameters, thereby increasing auction revenue and reducing bidder costs. The upper-layer prediction agent of the novel auction framework predicts the final price in real time, while the lower-layer adjustment agent adjusts parameters based on the real-time auction status, increasing the price during periods of high bidding activity and decreasing the increment during periods of stagnation. This adaptive adjustment of the auction strategy enhances user experience and profitability.
[0044] 2. The novel auction framework introduces a hierarchical reinforcement learning algorithm into smart contracts, automating the entire process from price prediction to dynamic adjustment of auction parameters. While ensuring decentralization and transparency, dynamic parameter optimization improves the final auction price and reduces transaction costs for bidders.
[0045] 3. By utilizing game theory analysis, the new auction framework can ensure the strategic balance between sellers and bidders, thereby achieving Nash equilibrium, effectively reducing non-optimal results caused by strategy changes, and increasing the fairness and competitiveness of the auction process.
[0046] 4. The new auction framework is trained on a large NFT transaction dataset and can dynamically adjust the auction strategy under different market conditions, which significantly improves the system's adaptability in a changing market environment, making it more robust and flexible, and ensuring optimal results in diverse auction scenarios. Attached Figure Description
[0047] Figure 1 This is a diagram illustrating the three stages of the Non-Fungible Token Auction System (NFTAS).
[0048] Figure 2 This is a schematic diagram of the creation stage of the novel auction algorithm of this invention.
[0049] Figure 3 This is a schematic diagram of the bidding stage of the novel auction algorithm of this invention. Detailed Implementation
[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0051] I. Design of the NFT Auction System NFTAS
[0052] like Figure 1 As shown, the NFT (Non-Fungible Token) auction process in the Non-Fungible Token Auction System (NFTAS) can be divided into three stages: the creation stage, the bidding stage, and the settlement stage. The specific components of NFTAS will be explained in detail in Section 1.1, and the process of each stage will also be described later.
[0053] 1.1 NFTAS Components
[0054] The NFTAS system consists of a set of interconnected components that work together to ensure the efficient operation of the auction process. Each smart contract or supporting infrastructure plays a specific role, including:
[0055] User Interface (UI): The UI is the primary platform for interaction between sellers and bidders, acting as a bridge between users and other components, supporting the creation of auctions and the submission of bids. Through integration with the novel auction algorithm of this invention, the UI can also provide real-time updates on the auction status.
[0056] Auction Smart Contract (ASC): As a core component of NFTAS, ASC manages all three phases of NFTAS and ensures smooth auction execution through the coordination of multiple smart contracts. ASC manages auction creation, maintains the current highest bid, verifies bid eligibility, and initiates the settlement phase at the end of the auction. ASC allows sellers to grant permissions to a novel auction algorithm to dynamically adjust auction parameters and predict the final auction price of NFTs in real time, thereby improving auction prices and optimizing the user experience.
[0057] Payment Smart Contract (PSC): The PSC manages all payment transactions related to the auction. It is responsible for locking up funds submitted by bidders, verifying their ability to pay, and releasing the funds to the seller after the auction. Simultaneously, the PSC also handles the return of deposits to unsuccessful bidders, ensuring the security and transparency of the entire transaction process.
[0058] NFT Smart Contract (NFTSC): The NFTSC is responsible for the lifecycle management of the NFT. It locks the NFT during the creation phase to ensure that the NFT cannot be transferred or withdrawn during the auction process; during the settlement phase, the NFTSC securely transfers ownership of the NFT to the highest bidder.
[0059] InterPlanetary File System (IPFS): IPFS is used to store metadata related to NFTs and transaction data during the auction process, providing data support for updating new auction algorithms. The use of IPFS not only ensures the security, transparency, and immutability of NFT data, but also reduces the cost of on-chain storage.
[0060] 1.2 Creation Phase
[0061] During the creation phase, the seller initiates an auction request through the UI. The seller needs to provide the following NFT auction information:
[0062] ID: A unique identifier for an NFT.
[0063] Starting price (P) s The initial price of an auction is usually set at a low price to attract bids.
[0064] Reserve price (P) r ): The seller's minimum acceptable price. If the bid does not reach the reserve price, the auction may fail.
[0065] Bid Increment (P) m ): The minimum amount that must be increased in each round of bidding to ensure that each bid increases the amount based on the existing amount.
[0066] Temporary buyout price (P) b): Allows bidders to instantly purchase NFTs at a price that is set as a time-limited price during the auction period.
[0067] Auction Duration (D): The total time during which the auction is open for bidding.
[0068] Metadata (M): Contains detailed information about the NFT, including title, description, creation date, owner history, creator information, and the NFT's token standard type.
[0069] After receiving this information, the UI will invoke ASC to create an auction transaction on the blockchain. Upon receiving the request, ASC will perform the following two main operations:
[0070] 1) Locking NFT: ASC calls NFTSC to lock the NFT and transfer its ownership from the seller to NFTSC, preventing the NFT from being transferred or withdrawn during the auction.
[0071] 2) Storing NFT metadata: ASC stores auction request information on IPFS and generates a unique Content Identifier (CID) to ensure data integrity and accessibility.
[0072] 1.3 Bidding Stage
[0073] During the bidding phase, potential bidders participate in the auction by submitting bids through the UI. After a bid is submitted, the PSC verifies the bidder's ability to pay and locks the bid amount. Upon successful verification, the PSC notifies the ASC to register the bid. The ASC then confirms the validity of the bid, ensuring it is higher than the current highest bid, and records the new bid amount and the bidder's address in the auction status. Afterwards, the ASC instructs the PSC to refund the previous highest bid.
[0074] The novel auction algorithm of this invention dynamically adjusts the auction strategy and predicts the possible final auction price. This predicted final price is only visible to the seller, so that the seller can adjust their expectations and strategies according to market changes.
[0075] 1.4 Settlement Stage
[0076] When the auction ends or the bidder selects the "provisional buyout price" (P b The settlement phase begins when an NFT is purchased.
[0077] At this stage, the auction smart contract (ASC) performs several key operations to ultimately determine the auction result:
[0078] 1) Payment to Seller: ASC instructs the payment smart contract (PSC) to transfer the locked funds from the highest bidder's account to the seller to complete the payment process.
[0079] 2) Transfer to buyer: ASC calls the NFT smart contract (NFTSC) to securely transfer ownership of the NFT from NFTSC to the highest bidder, ensuring a safe change of ownership.
[0080] 3) Saving and updating the auction algorithm model: ASC stores auction results on IPFS to maintain data transparency and facilitate auditing. Furthermore, auction data is periodically used to update the auction algorithm model. The updated model is also stored in IPFS and periodically uploaded to the NFT Auction System (NFTAS) via IPFS, thereby improving the system's efficiency and adaptability.
[0081] II. Design of a Novel Auction Algorithm
[0082] 2.1 Overview of Motivation and Methods
[0083] The novel auction algorithm is designed to improve the final price of NFTs while reducing bidding costs through a new hierarchical reinforcement learning (HRL) algorithm. This algorithm continuously monitors the auction smart contract (ASC) associated with the NFT auction, predicts the final selling price of the NFT, and dynamically adjusts the seller's auction strategy in real time. This adjustment includes modifying auction parameters to optimize the final transaction outcome for both the seller and the bidders.
[0084] The HRL method is adopted because NFT auction systems (NFTAS) possess unique complexities that traditional deep reinforcement learning (DRL) methods struggle to handle effectively. In the DRL framework, the combination of high-dimensional state and operation spaces with the uncertainties of traditional auction dynamics makes it difficult for a single agent to learn the optimal strategy. Furthermore, the NFT auction process faces numerous uncertainties, such as unpredictable bidder behavior, fluctuating NFT values, and the need to adapt to rapidly changing market conditions. These factors make it difficult for a single DRL agent to effectively learn ideal strategies in different NFT auction scenarios.
[0085] Furthermore, the NFT auction process is inherently multi-layered, involving decision-making at multiple stages. The entire process begins with setting initial auction parameters (such as the starting price and reserve price) and then dynamically adjusts these parameters in response to real-time bidding and changes in market sentiment. Traditional DRL methods typically employ a flat learning approach, which struggles to manage such multi-scale decision-making needs. This approach lacks flexibility and adaptability in managing high-level strategic planning and nuanced tactical adjustments.
[0086] The novel auction algorithm employs the HRL method to overcome the aforementioned limitations by decomposing the problem into two specialized subtasks. Each task is managed by a dedicated agent at a different level. In the novel auction algorithm, the upper-level prediction agent (A... p) is responsible for predicting key auction parameters (such as the starting price P) s and reservation price P r ), and dynamically predict the final auction price; the lower-level adjustment agent (A a The responsible entity is to continuously adjust the bidding increment (P). m ), temporary buyout price (P) b And the auction extension period. Among them, the forecast agent A... p Also known as a predictive agent, adjusting agent A a Also known as a modulated agent. This hierarchical design is particularly suitable for NFTAS:A p Focusing on long-term strategic planning, and A a It can quickly adapt to dynamic changes during the auction process. By simplifying the learning process and improving the model's adaptability and responsiveness, this algorithm significantly improves the overall performance of NFTAS.
[0087] Furthermore, the novel auction algorithm is designed to be completely decentralized and highly secure. Model updates require consensus from all consensus nodes in the blockchain; any malicious update proposals will be rejected by other nodes, preventing improper model updates. The algorithm's update parameters are finalized as a transaction, completing the decentralized model update process and ensuring system decentralization and security throughout. When a consensus node proposes a model update, it must include the aggregated signatures of all nodes to ensure decentralization and security through the blockchain's consensus protocol.
[0088] 2.2 Structure of Hierarchical Reinforcement Learning Model
[0089] The hierarchical design of the novel auction algorithm includes two main agents: the upper-level prediction agent (A) p ) and the lower-level adjustment agent (A a These two agents dynamically predict the final auction price and optimize the auction parameters through specific interaction methods.
[0090] 2.2.1 Creation Phase
[0091] During the creation phase, the seller submits auction information I, which is then converted into the initial state s0, as A. p The input. This stage involves setting and adjusting several initial auction parameters, including the starting price (P). s ), reservation price (P) r ), temporary buyout price (P) b ) and bidding increment (P m A p First, generate a temporary action a. temp1 To adjust P s and P rSubsequently, this temporary action updated the state, creating a temporary state s. temp1 and pass it to A a A a Generate another temporary action a temp2 To adjust P b and P m The newly generated temporary state s temp2 Return to A p Finally, A p Output the predicted final price P, and combine it with other temporary parameters to generate the final action a0. These other temporary parameters refer to s0, s... temp1 wait.
[0092] The creation phase process is as follows: Figure 2 As shown, where P' s 、P' r 、P' b 、P' m These represent the adjusted starting price, reserve price, temporary buyout price, and bid increment for the auctioned product, respectively. The critic network refers to a neural network used to update the agent.
[0093] 2.2.2 Bidding Stage
[0094] During the bidding phase, the two agents continue to interact to adapt to the dynamic changes in the auction process. Current state s t Input to A a Generate temporary action a temp Adjust P b P m Auction extension time (T) e Then update the status s. temp And pass it to A p A p Based on the updated data, further adjust P and output the final action a. t During this phase, the collected bidding data, along with data from the creation phase, will be used to update the model for the new auction algorithm.
[0095] The bidding process is as follows: Figure 3 As shown, where P' s 、P' r T' e P' and P' represent the adjusted starting price, reserve price, auction extension time, and predicted final price of the auctioned product, respectively.
[0096] 2.2.3 Model Update
[0097] During the settlement phase, the agent bases its decisions on the collected data set {(s)} i ,a i ,r i,s i+1 )|i∈N + For each policy denoted as , 1 ≤ i, the Proximal Policy Optimization (PPO) algorithm is used for periodic updates to find the optimal policy, further optimizing the system's adaptability and dynamic adjustment capabilities. Where s i Let s represent the state s at the i-th time step, and a i Let a and r represent the actions at the i-th time step. i Let r and s represent the reward at the i-th time step. i+1 Let s represent the state at time step i+1, and N be the state of the time step i+1. + Represents a positive integer.
[0098] The hierarchical structure of the novel auction algorithm clearly divides the task, A p A is responsible for long-term strategic planning. a It is responsible for short-term tactical adjustments, thereby simplifying the complex decision-making process, enhancing the system's adaptability and responsiveness, and ultimately improving auction results, achieving higher bidding returns and participation.
[0099] Formalizing the problem in the new auction algorithm:
[0100] In the NFT auction process, the system dynamically adjusts its strategy based on the current state and results after each round of bidding until the auction ends. Therefore, the NFT auction problem conforms to the characteristics of the Markov Decision Process (MDP) framework. This problem can be represented as a quadruple (S, A, P, R):
[0101] State space (S): Represents all possible states of the auction environment, including A p and A a The individual state space. The state space is defined as S = {s p ,s a |s p ∈R 10 ,s a ∈R 10}. Among them, s p A represents p The individual state space, s a A represents p Let R be the individual state space, and R represent real numbers.
[0102] Operating space (A): A p and A a The operation space is defined as a p =(P s ,P r ,P) and a a =(P m ,P b ,T e ), where P sP represents the starting price of the auctioned product. r P represents the reservation price, and P represents the predicted final price. m P represents the incremental bid. b Indicates the temporary buyout price, T e This indicates an extension of the auction period. A p Accurate parameter prediction is achieved through a three-layer fully connected network (FCN), while A a This is responsible for responding quickly to market changes and dynamically adjusting auction parameters.
[0103] State transition function (P): At time step t, the state transition function P(s) t+1 |s t ,a t ) describes the execution of action a t After from state s t Transition to s t+1 The probability of this is determined by external random factors such as bidder behavior, NFT valuation, and market sentiment.
[0104] Reward function (R): The reward function evaluates the agent's behavior in a specific state. For A p The reward function is defined based on the similarity between the actual transaction price and the predicted final price. If the auction is successful, R(s) p ,a p R(s) represents the similarity between the actual price and the final price; if the auction fails, a penalty coefficient F is applied, such that R(s) = 0. p ,a p ) = -F.
[0105] Considering that bidding in NFTAS will incur additional fees, increasing bidding costs, A a The reward function is defined as follows:
[0106] R(s a ,a a )=(w1*b+w2*c)*e x
[0107] Where w1 and w2 are weighting factors, b is the bidding increment, c is the bidding cost, and x = -λ*t is a time-related exponential decay factor (λ>0 controls the decay rate). This is done to balance the final revenue while minimizing costs and improving the user experience.
[0108] III. Game Theory Analysis
[0109] In the context of NFT auctions, this invention explores the theoretical proof that a pioneer strategy converges to a Nash equilibrium strategy. By building a model using game theory and reinforcement learning, it demonstrates that in NFT auctions, the pioneer strategy can achieve stability, ensuring that no party (including sellers and bidders) has an incentive to unilaterally deviate from its strategy—a core condition for Nash equilibrium.
[0110] 3.1 Problem Setting
[0111] In the context of NFT auctions, the interaction between sellers and bidders can be viewed as a dynamic game. Sellers optimize the auction process by adjusting auction parameters, while bidders submit their bids based on these parameters. Pioneers, representing the sellers, learn algorithms to develop optimal strategies for adjusting auction parameters and predict bidders' reactions. This setup stems from the idea behind novel auction models.
[0112] 3.2 Nash Equilibrium in a Single-Agent Scenario
[0113] A novel auction model uses an HRL (Hierarchical Reinforcement Learning)-based algorithm to control the dynamic parameters and bidding behavior of the auction. We focus on how this algorithm learns to achieve the optimal strategy for Nash equilibrium. Nash equilibrium is achieved when all participants, given the strategies of others, cannot improve their own payoffs through unilateral strategy adjustments.
[0114] In Markov decision-making, the optimal policy π represents the Nash equilibrium policy, where:
[0115] 1) For the seller, strategy π maximizes expected auction revenue while taking into account the dynamic adjustment of auction parameters.
[0116] 2) For bidders, strategy π maximizes the expected net return (i.e., NFT valuation minus the winning bid price) and adjusts the bidding strategy based on the auction status and the behavior of other bidders.
[0117] 3.3 Reinforcement learning converges to Nash equilibrium
[0118] The auction algorithm employs the HRL framework, allowing for dynamic adjustment of auction parameters and prediction of the optimal bidding strategy. This method consists of two layers:
[0119] 1) Advanced Prediction Agent (Ap): Set key fixed auction parameters and predict the final price at the end of the auction.
[0120] 2) Low-level adjustment agent (Aa): Dynamically adjusts auction parameters based on real-time bidding conditions.
[0121] A reinforcement learning agent is trained using the PPO (Proximal Policy Optimization) method, and the policy is updated based on observed rewards and transitions. Over time, the policy π... θConverging to the optimal policy π * To maximize the cumulative expected return:
[0122]
[0123] Where τ represents the state and behavior trajectory from the start to the end of the auction, γ is the discount factor, E represents the expectation, T represents the total time step of the auction, and R(s) t ,a t ) represents the reward function.
[0124] According to reinforcement learning theory, under conditions of sufficient exploration and appropriate learning rate adjustment, the pioneer's policy will converge to the optimal policy π. * This achieves Nash equilibrium because:
[0125] a) The seller's strategy of adjusting auction parameters maximized the participants' income.
[0126] b) The bidder's strategy (predicted and controlled by the pioneer) maximizes its own utility under the auction parameters set by the seller.
[0127] Therefore, neither the seller nor the bidder can increase their own profits by unilaterally changing their strategies, thus failing to meet the Nash equilibrium condition.
[0128] The innovative aspects of this invention include:
[0129] 1) This invention proposes a design scheme for an NFT auction system (NFTAS), which includes three stages: creation, bidding, and settlement. It achieves an efficient and secure NFT auction process through the collaborative work of a series of smart contract components. NFTAS includes a user interface (UI), an auction smart contract (ASC), a payment smart contract (PSC), an NFT smart contract (NFTSC), and an InterPlanetary File System (IPFS), ensuring the transparency and immutability of auction data while reducing on-chain storage costs.
[0130] 2) This invention proposes a novel auction algorithm based on hierarchical reinforcement learning (HRL). This algorithm predicts the agent (A)... p ) and Adjustment Agent (A a Through collaboration with other technologies, the algorithm dynamically predicts the final price of NFTs and optimizes auction parameters to achieve higher bidding returns and user engagement. The HRL structure introduced by the algorithm enables hierarchical management and solves the multi-scale decision-making problem of NFT auction systems.
[0131] 3) The HRL algorithm of this invention dynamically adjusts parameters such as bid increment and reserve price in NFT auctions to adapt to real-time market fluctuations, thereby optimizing seller revenue and improving the bidder participation experience. A p The agent is responsible for long-term strategic forecasting, Aa The agent then makes short-term tactical adjustments to simplify the learning process and enhance the system's responsiveness.
[0132] 4) The HRL algorithm of this invention is fully decentralized, and model updates are verified through consensus nodes to prevent malicious updates. Simultaneously, the algorithm update process ensures system security through the blockchain consensus protocol, and all update parameters are achieved through transaction records, improving the adaptability and overall performance of the NFT auction system.
[0133] 5) This invention proposes a specific implementation process for NFTAS, including the storage and locking of NFT metadata during the creation phase, dynamic bid adjustment during the bidding phase, and fund payment and NFT ownership transfer during the settlement phase. By storing auction data in IPFS and regularly updating the algorithm model, NFTAS achieves transparent auction data storage and sustainable system optimization.
[0134] 6) Overall, this invention addresses the auction needs of the NFT market by designing a decentralized auction system based on hierarchical reinforcement learning. It balances system security, responsiveness, and auction revenue optimization, and achieves efficient operation of NFT auction scenarios while ensuring data transparency and immutability.
[0135] Another embodiment of the present invention provides an NFT digital asset auction system based on reinforcement learning, comprising a blockchain and a hierarchical reinforcement learning model deployed in an auction smart contract on the blockchain. The hierarchical reinforcement learning model includes a predictive agent and an adjusting agent. During the auction process, the predictive agent predicts the final price of the NFT in real time, and the adjusting agent dynamically adjusts the seller's auction strategy in real time, thereby increasing auction revenue and reducing bidder costs. The specific working processes of the predictive and adjusting agents in the hierarchical reinforcement learning model can be referred to the corresponding processes in the foregoing method embodiments.
[0136] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the steps of the method of the present invention.
[0137] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) storing a computer program that, when executed by a computer, implements the various steps of the method of the present invention.
[0138] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and to implement it accordingly. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification; the scope of protection of the present invention is defined by the claims.
Claims
1. A reinforcement learning-based NFT digital asset auction method, characterized in that, Includes the following steps: Deploy a hierarchical reinforcement learning model in a blockchain auction smart contract, the hierarchical reinforcement learning model including a predictive agent and an adjusting agent; During the auction process, the final price of the NFT is predicted in real time by a predictive agent, and the seller's auction strategy is dynamically adjusted in real time by adjusting the agent, thereby increasing auction revenue and reducing bidder costs.
2. The method according to claim 1, characterized in that, The predictive agent predicts key auction parameters, including the starting price and the reserve price, and dynamically predicts the final auction price; the adjusting agent adjusts the auction parameters according to the real-time auction status, raising the price during periods of high bidding activity and lowering the increment during periods of stagnant bidding, thus enabling the auction strategy to be adaptively adjusted.
3. The method according to claim 1, characterized in that, The update of the hierarchical reinforcement learning model requires consensus from all consensus nodes in the blockchain. If a malicious update proposal is made, other nodes will reject the proposal to prevent the model from being updated improperly. When a consensus node proposes a model update, it must be accompanied by the aggregate signature of all nodes to ensure decentralization and security through the blockchain's consensus protocol.
4. The method according to claim 1, characterized in that, The creation phase of the auction process includes: Set and adjust the initial auction parameters, including the starting price P. s Reservation price P r Temporary buyout price P b and bidding increment P m ; The seller submits auction information, which is then converted into the initial state s0 and used as input for the predictive agent; The predictive agent first generates a temporary action a. temp1 To adjust P s and P r The temporary action then updates the state, forming a temporary state s. temp1 And pass it on to the adjusting agent; Adjust the agent to generate another temporary action a temp2 To adjust P b and P m The newly generated temporary state s temp2 And return it to the predictive agent; The predictive agent outputs the predicted price P and combines it with other temporary parameters to generate the final action a0.
5. The method according to claim 1, characterized in that, The bidding phase of the auction process includes: The current state s t Input is fed into the adjusting agent to generate a temporary action a. temp To adjust P b P m And auction extension time T e Then update the status s temp And pass it on to the predictive agent; The predictive agent adjusts the price P based on the updated data and outputs the final action a. t .
6. The method according to claim 1, characterized in that, The settlement phase of the auction process includes: Based on the collected data from the bidding and creation phases, the hierarchical reinforcement learning model is updated using a proximal policy optimization algorithm. The NFT auction problem can be represented as a quadruple (S, A, P, R): State space S: Represents all possible states of the auction environment, including the individual state spaces of the predicting agent and the adjusting agent; Operational Space A: This includes the operational spaces of the predictive agent and the adjusting agent. The predictive agent achieves accurate parameter prediction through a three-layer fully connected network, while the adjusting agent is responsible for quickly responding to market changes and dynamically adjusting auction parameters. State transition function P: At time step t, the state transition function P(s) t+1 |s t ,a t Describe the action a being performed. t After from state s t Transition to s t+1 The probability of; The reward function R is used to evaluate the predictive agent and adjust the agent's behavior in a specific state. For the predictive agent, the reward function is defined based on the similarity between the actual transaction price and the predicted final price. If the auction is successful, R represents the similarity between the actual price and the final price. If the auction fails, a penalty coefficient F is applied, such that R = -F.
7. The method according to claim 1, characterized in that, The hierarchical reinforcement learning model is trained using a proximal policy optimization method. The auction policy is updated based on the observed rewards and transitions. Over time, the auction policy converges to the optimal policy to maximize the cumulative expected return and achieve Nash equilibrium.
8. A reinforcement learning-based NFT digital asset auction system, characterized in that, This includes blockchain and a hierarchical reinforcement learning model deployed in auction smart contracts on the blockchain. The hierarchical reinforcement learning model includes a predictive agent and an adjusting agent. During the auction process, the predictive agent predicts the final price of the NFT in real time, and the adjusting agent dynamically adjusts the seller's auction strategy in real time, thereby increasing auction revenue and reducing bidder costs.
9. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 7.