A large-scale cross-chain dynamic access management system and method
By introducing multi-signal authentication of gateway clusters and strengthening learning punishment mechanisms, the problems of low cross-link entry efficiency and insufficient security are solved, and the stability and security management of large-scale cross-chain dynamic access is realized, and the compliance and reliability of cross-chain transactions are improved.
Patent Information
- Application Number
- CN202211547318.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-05
AI Technical Summary
In cross-chain application scenarios with high security requirements, the existing cross-chain access methods have problems such as different authentication standards, low access efficiency and insufficient security, especially in large-scale cross-chain dynamic access environments.
A multi-signal access chain authentication mechanism based on gateway clusters is designed to verify the compliance of cross-chain transactions through cross-chain gateways and relay chains, listen to cross-chain events in real time, and introduce a punishment mechanism based on reinforcement learning to reward and punish non-compliance operations, and add a request status list and access list to maintain the stability and security of cross-chain transactions.
It improves the efficiency and security of cross-link entry, reduces management complexity, enhances the network's fault tolerance, prevents network attacks, mobilizes the enthusiasm of gateway nodes through reward and punishment mechanisms, and ensures the compliance and reliability of cross-chain transactions.
Smart Images

Figure CN116015736B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of blockchain, and specifically relates to a large-scale cross-chain dynamic access management system and method. Background Art
[0002] Blockchain technology is an integrated application of computer technologies such as distributed data storage, peer-to-peer transmission, distributed consensus algorithms, and encryption algorithms. Blockchain technology does not rely on an additional third-party management agency or hardware facilities, and there is no central control. Except for the self-contained blockchain itself, through distributed accounting and storage, each node realizes self-verification, transmission, and management of information. In the business application scenarios where the business forms are becoming increasingly complex, there is a lack of a unified interconnection mechanism between chains, which greatly restricts the liquidity of the value of digital assets on the blockchain. Therefore, the cross-chain demand has emerged. Cross-chain refers to the realization of trusted interoperability between different ledgers by connecting relatively independent blockchain systems. However, information exchange involves data synchronization between chains and corresponding cross-chain calls, and the implementation is more complex. The interoperability barriers between various blockchain applications are extremely high, and cross-chain information sharing cannot be effectively carried out.
[0003] Traditional cross-chain access mostly adopts a chain-to-chain architecture, where different chains are directly connected through their respective cross-chain gateways, and the cross-chain gateways assume the roles of information interaction and verification. However, this method is only applicable to scenarios where there is a trust foundation among cross-chain participants or the security requirements are not so high. For application scenarios with high security requirements, the technology still needs to be improved. For example, in the patent application CN114499898A, a blockchain cross-chain secure access method and device are disclosed, which is a basic patent proposed for the problem of unsmooth blockchain cross-chain transaction processes. The method is used to run between multiple application chains and at least one relay chain. Each application chain connects to the relay chain through a multi-node gateway and uses a one-time session key for communication. However, it lacks corresponding processing for non-compliant transactions and does not consider the need to adapt the gateway interface according to the characteristics of specific application chains, which is not conducive to cross-chain transaction security and access management. Summary of the Invention
[0004] To solve the problems of application chain access authentication, inconsistent standards, low access efficiency, etc. under large-scale cross-chain dynamic access, the present invention provides a large-scale cross-chain dynamic access management system and method, designs an access chain authentication mechanism based on multi-signature of a gateway cluster, real-time monitors cross-chain events during the full access cycle of the access application chain, and verifies the compliance of cross-chain transactions through cross-chain gateways and relay chains. At the same time, a penalty mechanism linked to the compliance of cross-chain process operations is designed based on reinforcement learning to reward and punish non-compliant operation objects. In addition, a request status list and an access list are added to provide stable cross-chain services for large-scale blockchain dynamic access while maintaining normal cross-chain management.
[0005] A large-scale cross-chain dynamic access management system according to the present invention includes a gateway interface module, a gateway monitoring module, a distribution module, a signature module, a relay chain verification module, a reward and punishment module, and an execution module;
[0006] The gateway interface module is used to adapt to different blockchains;
[0007] The gateway monitoring module is used to monitor cross-chain events on the corresponding transfer chain by the gateway node and confirm whether the cross-chain events exist;
[0008] The signature module is used to sign the cross-chain events confirmed by the cross-chain gateway;
[0009] The distribution module is used to distribute cross-chain events from the gateway to the relay chain;
[0010] The relay chain verification module is used to verify whether the cross-chain object is compliant. If the verification fails, the relay chain returns a failure receipt and error information to the source chain, and at the same time enables the reward and punishment mechanism;
[0011] The reward and punishment module is used to perform gradient updates on the transaction rules of the cross-chain object, strengthen the dynamic interaction between the reinforcement learning and the relay chain verification rules according to the cross-chain event information, dynamically adjust the compliance of the cross-chain object of the application chain based on the maximization of the feedback reward expectation, and promote the successful access of the cross-chain event;
[0012] The execution module is used to execute transactions on the destination chain; after the cross-chain object passes the verification, it is submitted to the target chain for execution, which represents the successful access of the cross-chain object. At the same time, the destination chain adds the cross-chain object record to the management contract access list and sends a receipt to the source chain; in addition, according to the execution result, the object status is fed back to the associated relay chain, and the relay chain transaction management contract is updated accordingly.
[0013] Further, the reward and punishment module includes an environment information acquisition module, a model building module, a policy update module, and a reward and punishment feedback module;
[0014] The environment information acquisition module is used to acquire application chain transaction information; the information includes: cross-chain event information, relay chain verification rules, and relay chain verification results;
[0015] The model building module is used to build a model based on the reinforcement learning policy gradient algorithm, and learn the strategy to make the cross-chain object pass the relay chain verification successfully through the reward guidance obtained by interacting with the surrounding environment;
[0016] The policy update module is used to train the model based on the reinforcement learning policy gradient algorithm, design the reinforcement learning parameters according to the actual situation, and update the parameters according to the existing gradient increment formula, and cycle on this basis to obtain an optimal expected policy;
[0017] The reward and punishment feedback module is oriented towards the goal of reinforcement learning, enabling the model to continuously adjust the object information towards the corresponding relay chain verification rules.
[0018] Further, the gradient increment formula is as follows:
[0019]
[0020] where θ is the network parameter, τ is the data trajectory, R(τ) is the sum of rewards at each stage, is the gradient value of the objective function π θ (a|s) in the current model state, and P θ (τ) is the probability of a data trajectory occurring.
[0021] Further, after the reward and punishment feedback module adjusts the policy action, based on the original receipt volume of the source chain, when the receipt volume increases, it is considered that after the adjustment action of the reinforcement learning model, the number of compliant objects passing through the relay chain verification increases, thereby increasing the probability of this action occurring; conversely, when the receipt volume decreases, it is considered that after the adjustment action of the reinforcement learning model, the number of compliant objects passing through the relay chain verification decreases, thereby reducing the probability of this action occurring. Based on the above large-scale cross-chain dynamic access management system, the present invention designs a large-scale cross-chain dynamic access management method, which includes the following steps:
[0022] Step 1: Application chain 1 initiates an object access request. When the chain management contract throws a cross-chain event, it carries the address of application chain 2, and at the same time lists the cross-chain object in the request access status table in the management contract;
[0023] Step 2: After the corresponding cross-chain gateway 1 listens and confirms the existence of the event, it receives the cross-chain event and signs the event, queries it through a distributed hash table in the cross-chain gateway cluster, and sends the event and the signature to relay chain A that manages application chain 1;
[0024] Step 3: Relay chain A verifies the cross-chain event and the attached signature. If the verification fails, relay chain A returns a failure receipt and related error information to application chain 1, and at the same time enables the reward and punishment mechanism, and application chain 1 adjusts and resends until relay chain A passes the verification;
[0025] Step 4: After relay chain A passes the verification, it sends the cross-chain event and the event proof to the corresponding cross-chain gateway 2;
[0026] Step 5: After cross-chain gateway 2 receives the cross-chain event and the event proof, it signs them, and then queries through a distributed hash table in the cross-chain gateway cluster according to the destination chain address of the cross-chain event, and sends the event and the signature to relay chain B that manages application chain 2;
[0027] Step 6: The relay chain B verifies the cross-chain event and the signature. If the verification fails, the relay chain B returns a failure receipt and relevant error information to the application chain 1, and at the same time enables the reward and punishment mechanism, and the application chain 1 adjusts and resends;
[0028] Step 7: After being verified by the relay chain B, it is submitted to the application chain 2 for execution. At the same time, the cross-chain object record is added to the access list of the chain management contract of this chain, indicating successful access; in addition, a receipt is sent to the application chain 1 to clear the record of this object in the request status list of the application chain 1 management contract.
[0029] Furthermore, the reward and punishment mechanism specifically includes the following steps:
[0030] Step a: For specific cross-chain events of the application chain, use the environmental information acquisition module to dynamically and real-time collect cross-chain event data, such as corresponding relay chain verification rules, cross-chain gateway signature information, relay chain verification results, etc.;
[0031] Step b: Use the model building module to build a model based on the reinforcement learning policy gradient algorithm according to the obtained cross-chain event data, and apply the policy theoretically to the event;
[0032] Step c: Use the reward and punishment feedback module to obtain the model adjustment action effect of the previous round of cross-chain events based on the original model state;
[0033] Step d: According to the previous round of cross-chain event model state and model adjustment action reward, use the policy update module to update the corresponding policy parameters, so as to cyclically control the agent's action policy and apply it.
[0034] The beneficial effects of the present invention are as follows: The present invention provides a large-scale cross-chain dynamic access management system and method. Through the data connection between the gateway interface module, gateway monitoring module, signature module, distribution module, relay chain verification module, reward and punishment module, and execution module, the access efficiency can be effectively improved; based on the gateway cluster, the fault tolerance ability is enhanced while preventing network attacks, and a gateway adaptation interface is set to enhance the interaction ability; the present invention monitors cross-chain events in real time, verifies the compliance of cross-chain transactions through the cross-chain gateway and relay chain, and enhances its security; the present invention designs a punishment mechanism to reward and punish non-compliant operation objects, mobilize the enthusiasm of the gateway nodes while reducing the security risk; by adding a request status list and an access list, the cross-chain transaction status is updated in real time, reducing its management complexity. Description of the Drawings
[0035] Figure 1 is the structural diagram of the cross-chain dynamic access management system described in the present invention;
[0036] Figure 2It is a schematic diagram of the cross-chain transaction processing flow of the present invention;
[0037] Figure 3 This is a flowchart of the reward and punishment mechanism based on reinforcement learning. DETAILED DESCRIPTION
[0038] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments in conjunction with the accompanying drawings.
[0039] like Figure 1 As shown, a large-scale cross-chain dynamic access management system described in the present invention includes a gateway interface module, a gateway monitoring module, a distribution module, a signature module, a relay chain verification module, a reward and punishment module, and an execution module; the main contents of each module are as follows:
[0040] (1) Gateway interface module:
[0041] Used to adapt to different blockchains. The cross-chain gateway connects to specific types of blockchains and forwards cross-chain messages. In order to improve the interaction capabilities with the application chain, a gateway plug-in mechanism is set up to implement specific interfaces, and the interaction interface is adapted to different blockchains;
[0042] (2) Gateway monitoring module:
[0043] Used for gateway nodes to monitor cross-chain events on the corresponding blockchain. The source chain initiates an access request, and the chain management contract throws a cross-chain event. Each gateway node needs to monitor and confirm whether the cross-chain event exists. Only after confirming the event exists can it be signed and distributed;
[0044] (3) Signature module:
[0045] Used by the cross-chain gateway to sign the received cross-chain events so that they can be sent to the relay chain as one of the audit contents;
[0046] (4) Distribution module:
[0047] Used by the gateway to distribute cross-chain events to the relay chain. The distribution module is responsible for the specific transfer objects of cross-chain transactions. Each cross-chain gateway maintains a distributed hash table for the entire network to record the management service relationship between the application chain and the relay chain. After signing the cross-chain event, the event and signature are sent to the relay chain that manages the source chain by querying the distributed hash table. When the gateway joins the routing network, it broadcasts all application chain information managed by the relay chain, and other gateways update the distributed hash table based on the broadcast information;
[0048] (5) Relay chain verification module:
[0049] It is used to verify whether the cross-chain object is compliant on the relay chain. In a multi-relay-chain cross-chain architecture, the relay chain verifies the cross-chain objects sent by the application chain, and each cross-chain object joins the relay chain on an equal footing. The relay chain verification content includes the attached signature of the gateway and the transaction rules of cross-chain events. If the verification fails, the relay chain returns a failure receipt and error information to the source chain, and at the same time enables a reward and punishment mechanism;
[0050] (6) Reward and Punishment Module:
[0051] It is used to perform gradient updates on the transaction rules of cross-chain objects. For non-compliant cross-chain objects caused by various factors, a punishment mechanism is introduced on the relay chain, and the reliability of the process is ensured through economic incentives based on the reinforcement learning algorithm.
[0052] Reinforcement learning dynamically adjusts the compliance of the application chain cross-chain object based on the dynamic interaction between cross-chain event information and relay chain verification rules, and maximizes the expected feedback reward to promote cross-chain access;
[0053] (7) Execution Module:
[0054] It is used to execute transactions on the destination chain. After the cross-chain object passes the verification, it is submitted to the target chain for execution, indicating that the cross-chain object has successfully accessed. At the same time, the destination chain adds the cross-chain object record to the management contract access list and sends a receipt to the source chain.
[0055] In addition, according to the execution result, it feeds back the object status to the associated relay chain, and the relay chain transaction management contract makes corresponding updates.
[0056] The reward and punishment module includes an environment information acquisition module, a model building module, a policy update module, and a reward and punishment feedback module. The specific content of each module is as follows:
[0057] (1) Environment Information Acquisition Module:
[0058] It is used to acquire application chain transaction information, and the information includes: cross-chain event information, relay chain verification rules, relay chain verification results, etc.;
[0059] (2) Model Building Module:
[0060] It is used to build a model based on the reinforcement learning policy gradient algorithm, and learn a strategy that enables cross-chain objects to pass relay chain verification through the reward guidance obtained by interacting with the surrounding environment;
[0061] (3) Policy Update Module:
[0062] It is used to train the model based on the reinforcement learning policy gradient algorithm. According to the gradient increment formula of this algorithm as follows:
[0063]
[0064] where θ is the network parameter, τ is the data trajectory, and R(τ) is the sum of rewards at each stage. is the gradient value of the objective function π θ (a|s) under the current model state, and P θ (τ) is the probability of a data trajectory occurring. Design the reinforcement learning parameters according to the actual situation, update the parameters based on the existing gradient increment formula, and loop on this basis to obtain an optimal expected policy;
[0065] (4) Reward and punishment feedback module:
[0066] This reward and punishment mechanism is oriented towards the goal of reinforcement learning, enabling the model to continuously adjust the object information towards the corresponding relay chain verification rules. After adjusting the policy action, based on the original receipt quantity of the source chain, when the receipt quantity increases, it is considered that after the adjustment action of the reinforcement learning model, the number of compliant objects passing through the relay chain verification increases, thus increasing the probability of this action occurring; conversely, when the receipt quantity decreases, it is considered that after the adjustment action of the reinforcement learning model, the number of compliant objects passing through the relay chain verification decreases, thus reducing the probability of this action occurring.
[0067] As Figure 2 described above, based on the above management system, the present invention provides a large-scale cross-chain dynamic access management method, including the following steps:
[0068] Step 1: Application chain 1 initiates an object access request. When the chain management contract throws a cross-chain event, it carries the address of application chain 2, and at the same time lists the cross-chain object in the request access status table in the management contract;
[0069] Step 2: After the corresponding cross-chain gateway 1 listens and confirms the existence of the event, it receives the cross-chain event and signs the event, queries it through a distributed hash table in the cross-chain gateway cluster, and sends the event and signature to relay chain A of management application chain 1;
[0070] Step 3: Relay chain A verifies the cross-chain event and the attached signature. If the verification fails, relay chain A returns a failed receipt and relevant error information to application chain 1, and at the same time enables the punishment mechanism. Application chain 1 adjusts and resends until relay chain A verifies successfully;
[0071] Step 4: After relay chain A verifies successfully, it sends the cross-chain event and the event proof to the corresponding cross-chain gateway 2;
[0072] Step 5: After cross-chain gateway 2 receives the cross-chain event and the event proof, it signs them, then queries through a distributed hash table in the cross-chain gateway cluster according to the destination chain address of the cross-chain event, and sends the event and signature to relay chain B of management application chain 2;
[0073] Step 6: The relay chain B verifies the cross-chain event and the signature. If the verification fails, the relay chain B returns a failure receipt and relevant error information to the application chain 1, and at the same time enables a penalty mechanism, and the application chain 1 adjusts and resends.
[0074] Step 7: After being verified by the relay chain B, it is submitted to the application chain 2 for execution. At the same time, the cross-chain object record is added to the access list of the chain management contract of this chain, indicating successful access. In addition, a receipt is sent to the application chain 1 to clear the record of this object in the request status list of the application chain 1 management contract.
[0075] For non-compliant access objects caused by various factors, a penalty mechanism is introduced on the relay chain to ensure the stability of the process through economic incentives. The present invention designs a penalty mechanism based on reinforcement learning. The reinforcement learning dynamically interacts according to the cross-chain event information and the relay chain verification rules, and dynamically adjusts the compliance of the application chain object based on maximizing the expected feedback reward, so as to promote the achievement of the cross-chain access purpose.
[0076] Among them, as Figure 3 shown, the reward and punishment mechanism specifically includes the following steps:
[0077] Step a: For a specific application chain cross-chain event, use the environmental information acquisition module to dynamically and real-time collect cross-chain event data, such as corresponding relay chain verification rules, cross-chain gateway signature information, relay chain verification results, etc.
[0078] Step b: Use the model building module to build a model based on the reinforcement learning policy gradient algorithm according to the obtained cross-chain event data, and apply the policy theoretically to the event.
[0079] Step c: Use the reward and punishment feedback module to obtain the model adjustment action effect of the previous round of cross-chain event based on the original model state.
[0080] Step d: According to the previous round of cross-chain event model state and the model adjustment action reward, use the policy update module to update the corresponding policy parameters, so as to cycle and control the agent action policy and apply it.
[0081] Taking the cross-chain transfer between different central bank digital currencies as an example, commercial banks 1 and 2 act as cross-chain gateways to sign the transaction, and central banks C and D act as relay chain validators to run relay chain nodes. The purpose of the transaction of central bank A is to transfer funds to central bank B. Central bank A sends a transfer request to commercial bank 1 that adapts to the business, and records the transaction status of this transaction in the system of central bank A; after commercial bank 1 confirms the existence of this transaction, it signs the transaction, then queries the bank subordination table to determine the sending object of the next link, and sends the transaction information and the signature to central bank C that manages central bank A; central bank C verifies the identification of the transfer bank and the transfer information. If the transfer process and the transaction itself comply with the current inter-bank transaction regulations, it means passing this inspection, and sends the transfer business to the adapted commercial bank 2.
[0082] After commercial bank 2 confirms the existence of this transaction, it signs the transaction again, then queries the bank subordination table and the receiving bank to determine the sending object of the next link, and sends the transaction information and the signature to central bank D that manages central bank B; central bank D verifies the second signature of the transfer business and other information. If the inspection objects all comply with the inter-bank transaction regulations at this stage, it means passing the inspection. Central bank D sends the transfer amount to the destination central bank B; central bank B receives this transfer, and at the same time records in the system that this transaction has been completed, and sends a transaction completion receipt to central bank A; after receiving the receipt, central bank A marks the status of this transfer business in the system as completed.
[0083] It is possible that the two central banks acting as validators fail to pass the verification of this transaction. No matter where the transaction stagnates in the intermediate central bank, the current transaction needs to be immediately intercepted back to central bank A, including the feedback of the intermediate central bank and the information of the handling commercial bank. Central bank A rectifies the transfer business to be carried out according to the transaction rules for specific problems and feedback, such as reasonably modifying the transfer amount, checking whether the process of the handling personnel is compliant, etc., and then resubmits the transfer request on the transaction chain. The subsequent transactions are adjusted according to the above requirements; after several rounds, review the transaction receipts received by central bank A. If the receipts increase, it means that the rectification is feasible. If no receipt is received or the business is returned midway, it means that the regularization has no effect, and it needs to be adjusted according to the process again and resubmitted until all the transferred businesses sent are successfully completed.
[0084] The above description is only the preferred solution of the present invention, and is not used as a further limitation of the present invention. All equivalent changes made by using the content of the specification and drawings of the present invention are within the protection scope of the present invention.
Claims
1. A large-scale cross-chain dynamic access management system, characterized in that The system includes a gateway interface module, a gateway monitoring module, a distribution module, a signature module, a relay chain verification module, a reward and punishment module, and an execution module; The gateway interface module is used to adapt to different blockchains; The gateway monitoring module is used for the gateway node to monitor the cross-chain events on the corresponding transmission chain to confirm whether the cross-chain event exists; The signature module is used by the cross-chain gateway to confirm the existence of the cross-chain event signature; The distribution module is used by the gateway to distribute cross-chain events to the relay chain; The relay chain verification module is used to verify whether the cross-chain object is compliant. If the verification fails, the relay chain returns a failure receipt and error information to the source chain, and enables a reward and punishment mechanism; The reward and punishment module is used to perform gradient updates on the transaction rules of cross-chain objects. Reinforcement learning dynamically adjusts the compliance of cross-chain objects on the application chain based on the dynamic interaction between cross-chain event information and relay chain verification rules based on the expectation of maximizing feedback rewards, thus facilitating the successful access of cross-chain events. The execution module is used to execute transactions on the destination chain; After the cross-chain object is verified, it is submitted to the target chain for execution, which means that the cross-chain object is successfully connected. At the same time, the destination chain adds the cross-chain object record to the management contract access list and sends a receipt to the source chain. In addition, the object status is fed back to the associated relay chain based on the execution result, and the relay chain transaction management contract is updated accordingly. The reward and punishment module includes an environment information acquisition module, a model building module, a strategy update module and a reward and punishment feedback module; The environment information acquisition module is used to obtain application chain transaction information; the information includes: cross-chain event information, relay chain verification rules, and relay chain verification results; The model building module is used to build a model based on the reinforcement learning policy gradient algorithm, and learns the strategy of making cross-chain objects compliant and pass the relay chain verification through the reward guidance obtained by interacting with the surrounding environment; The policy update module is used to train a model based on the reinforcement learning policy gradient algorithm, design reinforcement learning parameters according to actual conditions, and update parameters according to the existing gradient increasing formula, and then repeat the cycle to obtain an optimal expected strategy. Wherein, the gradient increasing formula is as follows: where θ is the network parameter, τ is the data trajectory, and R(τ) is the sum of rewards at each stage. is the objective function π under the current model state θ (a|s) gradient value, P θ (τ) is the probability of a data trajectory occurring; The reward and punishment feedback module is guided by the goal of reinforcement learning, so that the model adjustment object information can continuously tend towards the corresponding relay chain verification rules.
2. The large-scale cross-chain dynamic access management system according to claim 1, characterized in that, After the reward and punishment feedback module adjusts the strategy action, the original source chain receipt volume is used as a benchmark. When the receipt volume increases, it is considered that after the reinforcement learning model adjustment action occurs, the number of compliant objects verified by the relay chain increases, thereby increasing the probability of the action occurring; conversely, when the receipt volume decreases, it is considered that after the reinforcement learning model adjustment action occurs, the number of compliant objects verified by the relay chain decreases, thereby reducing the probability of the action occurring.
3. A large-scale cross-chain dynamic access management method, characterized in that Based on the large-scale cross-chain dynamic access management system according to any one of claims 1-2, the method comprises the following steps: Step 1: Application chain 1 initiates an object access request. The chain management contract throws a cross-chain event with the address of application chain 2, and lists the cross-chain object in the request access status table in the management contract; Step 2: After the corresponding adapted cross-chain gateway 1 monitors and confirms the existence of the event, it receives the cross-chain event, signs the event, queries it through a distributed hash table in the cross-chain gateway cluster, and sends the event and the signature to the relay chain A that manages Application Chain 1; Step 3: The relay chain A verifies the cross-chain event and the attached signature. If the verification fails, the relay chain A returns a failure receipt and relevant error information to Application Chain 1, and at the same time enables a reward and punishment mechanism. Application Chain 1 adjusts and resends until the relay chain A passes the verification; Step 4: After the relay chain A passes the verification, it sends the cross-chain event and the event proof to the corresponding adapted cross-chain gateway 2; Step 5: After receiving the cross-chain event and the event proof, the cross-chain gateway 2 signs them, and then queries them through a distributed hash table in the cross-chain gateway cluster according to the destination chain address of the cross-chain event, and sends the event and the signature to the relay chain B that manages Application Chain 2; Step 6: The relay chain B verifies the cross-chain event and the signature. If the verification fails, the relay chain B returns a failure receipt and relevant error information to Application Chain 1, and at the same time enables a reward and punishment mechanism. Application Chain 1 adjusts and resends; Step 7: After the relay chain B passes the verification, it submits it to Application Chain 2 for execution, and at the same time adds the cross-chain object record to the access list of the chain management contract, indicating successful access; in addition, it sends a receipt to Application Chain 1 to clear the record of this object in the management contract request status list of Application Chain 1.
4. A large-scale cross-chain dynamic access management method according to claim 3, characterized in that The reward and punishment mechanism specifically includes the following steps: Step a: For a specific application chain cross-chain event, use the environment information acquisition module to dynamically and real-time collect cross-chain event data, such as the corresponding relay chain verification rules, cross-chain gateway signature information, relay chain verification results, etc.; Step b: Use the model building module to build a model based on the reinforcement learning policy gradient algorithm according to the obtained cross-chain event data, and apply it to the event according to the theoretical strategy; Step c: Use the reward and punishment feedback module to obtain the model adjustment action effect of the previous round of cross-chain event based on the original model state; Step d: According to the previous round of cross-chain event model state and the model adjustment action reward, use the policy update module to update the corresponding policy parameters, so as to cycle and control the agent's action strategy and apply it.