A blockchain-based illegal data on-chain identification method and device
By designing smart contracts and a reputation-based random selection algorithm on the blockchain, combined with incentive mechanisms and game theory analysis, the problems of node bias and low credibility in the blockchain illegal data identification model are solved, achieving efficient identification and management of illegal data and improving the credibility and fairness of the identification results.
Patent Information
- Application Number
- CN202211431918.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-11-16
AI Technical Summary
Existing blockchain illegal data identification models suffer from node bias and low credibility, making it difficult to identify data fairly and impartially, and difficult to filter out malicious nodes, leading to frequent illegal data upload incidents.
This paper proposes a blockchain-based method for identifying illegitimate data. By generating smart contracts and identification node pool smart contracts, a reputation-based random selection algorithm is used to select identification nodes. Combined with incentive mechanisms and game theory analysis, the method ensures that identification nodes identify data fairly and impartially.
It improves the credibility and accuracy of blockchain-based illegal data identification, reduces illegal data uploads, and achieves effective management and incentive mechanisms for identification nodes, ensuring the fairness and reliability of identification results.
Smart Images

Figure CN115801402B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of blockchain technology, and more specifically, to a method and apparatus for identifying illegitimate data uploaded to the blockchain. Background Technology
[0002] Blockchain is essentially a decentralized distributed ledger, composed of interconnected blocks that record data. It features decentralization, tamper-proofing, and anonymity. While its tamper-proof nature provides a reliable method for information storage, these very characteristics have attracted malicious users who exploit them to upload harmful information and engage in other illegal activities. These security incidents have significantly hampered the development of blockchain technology, necessitating its regulation.
[0003] One direction is to utilize blockchain technology itself to control the uploading of illicit data to the blockchain. The first step is to identify illicit data about to be uploaded using blockchain technology. How to use blockchain technology to identify illicit data uploading is a pressing issue. Strengthening the identification of illegal activities using blockchain and establishing an effective blockchain risk identification system are of great significance to the long-term development of blockchain technology.
[0004] In recent years, many scholars have identified illegitimate data through on-chain data analysis. However, the applicant's analysis reveals that existing identification models are mostly single-chain models, where nodes are often biased, making it difficult to fairly and impartially identify one's own data, thus lacking credibility. Furthermore, these models struggle to filter out malicious nodes, further reducing their reliability. Therefore, there is an urgent need to propose a more reliable identification method to reduce the occurrence of illegal data uploads to the blockchain.
[0005] To address the aforementioned issues, this application proposes a blockchain-based method and apparatus for identifying illegitimate data uploaded to the blockchain. Summary of the Invention
[0006] This invention provides a blockchain-based method and apparatus for identifying illegitimate data uploaded to the blockchain. It solves the problem in related technologies that two nodes with discontinuous jumps in the chain can correctly decrypt the gradient information of the intermediate node through collusion, and the problem in the chain structure that subsequent user nodes have to wait for the transmission of the preceding nodes, which slows down the convergence speed of the model and causes the uplink communication time to be too long.
[0007] Firstly, this application provides a blockchain-based method for identifying illegitimate data on-chain. The method is applied to an identification chain within the blockchain and includes: generating a smart contract; negotiating with the chain to be identified and filling identification parameters into the smart contract based on the negotiation result; sending a token to the chain to be identified, causing the chain to return data information stored in the InterPlanetary File System and a hash value generated by the token to the identification chain; selecting identification nodes from ordinary nodes; and broadcasting the received hash value to the identification nodes, enabling each identification node to download the data information using the hash value and identify the data information.
[0008] Furthermore, the smart contract includes at least an identification node pool smart contract and an identification smart contract.
[0009] Furthermore, the responsibilities of the identification node pool smart contract include at least: registering ordinary nodes, transitioning the state of ordinary nodes, and selecting identification nodes.
[0010] Furthermore, the negotiation between the chain being identified and the smart contract, and the filling of identification parameters into the smart contract based on the negotiation result, includes: negotiating an identification treaty with the chain being identified, and filling the identification parameters into the identification smart contract based on the negotiation result, wherein the identification parameters include at least the number of identification nodes.
[0011] Furthermore, the identification parameters also include: identification time and identification cost.
[0012] Furthermore, the step of selecting the identification node from the ordinary nodes includes: using a reputation-based node random selection algorithm and selecting the identification node from the ordinary nodes according to a seed; wherein the seed includes at least a hash value composed of the total reputation value of the nodes in the identification node pool, the timestamp, and the block height.
[0013] Furthermore, the status of the ordinary node includes online and offline, and ordinary nodes that are online are preferentially selected as identification nodes.
[0014] Furthermore, the identification node is reviewed based on its behavioral history, and then processed according to the review results.
[0015] Furthermore, the step of reviewing the identification node based on its behavioral history and processing the identification node according to the review result includes: the review result includes the degree of maliciousness; processing the identification node according to the degree of maliciousness, wherein the processing includes at least one of the following methods: deducting the reputation value of the identification node; kicking the identification node out of the node pool.
[0016] Secondly, this application also provides a blockchain-based illegal data on-chain identification device, the device including a memory and a processor, the memory for storing a computer program, and when the computer program is executed by the processor, implementing the steps in the above-mentioned blockchain-based illegal data on-chain identification method.
[0017] This invention provides a blockchain-based method for identifying illegitimate data uploaded to the blockchain. It employs a reputation-based random selection algorithm and implements this method using an Ethereum smart contract. The smart contract proposed in this application not only identifies illegitimate data uploaded to the blockchain but also manages the identification nodes. By incentivizing the identification nodes, they are required to fairly and impartially identify data information, thereby reducing the occurrence of illegal data uploads to the blockchain. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Appendix Figure 1 This is a schematic diagram of the overall architecture of the recognition model provided in an embodiment of the present invention;
[0020] Appendix Figure 2 A schematic diagram of the blockchain-based illegal data on-chain identification process provided in an embodiment of the present invention;
[0021] Appendix Figure 3 A schematic diagram of a reputation-based random selection algorithm provided in an embodiment of the present invention;
[0022] Appendix Figure 4 This is a schematic diagram of node state transitions provided in an embodiment of the present invention;
[0023] Appendix Figure 5 This is a schematic diagram of the identification state transition provided in an embodiment of the present invention.
[0024] Appendix Figure 6 This is a schematic diagram of gas consumption for each interface provided in an embodiment of the present invention.
[0025] Appendix Figure 7 This is a schematic diagram comparing algorithms under the same reputation value, provided in an embodiment of the present invention.
[0026] Appendix Figure 8 This is a schematic diagram comparing algorithms under different reputation values provided in an embodiment of the present invention.
[0027] Appendix Figure 9This is a schematic diagram illustrating reputation growth provided in an embodiment of the present invention.
[0028] Appendix Figure 10 This is a schematic diagram illustrating reputation decline provided in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] In the embodiments of this specification, routing nodes and business nodes can be collectively referred to as blockchain nodes in the blockchain network. Nodes are designed as anonymous participants, and can be enterprises or individuals seeking to generate revenue by identifying illegitimate data. They can only be selected as identification nodes through a selection algorithm. In our identification model, the incentive function rewards or punishes identification nodes based on their behavior; to obtain higher returns, identification nodes must act honestly. The blockchain node can be a server connected to the blockchain network or a user terminal connected to the blockchain network; the specific form of the blockchain node is not limited here. In a blockchain system, a smart contract can be understood as code that can be understood and executed by all nodes in the blockchain.
[0031] Blockchain is essentially a decentralized distributed ledger composed of interconnected blocks that record data. Each block's header contains a cryptographic hash of the previous block, establishing a connection between the two. Blockchain features decentralization, tamper-proofing, and anonymity. However, this tamper-proofing and anonymity have attracted malicious users who exploit these features to upload harmful information and engage in other illegal activities. These security incidents have significantly hampered the development of blockchain technology, necessitating its regulation.
[0032] In recent years, many scholars have identified illicit data information by analyzing on-chain data. Matzutt et al. identified illegal fund transfers and uploaded content on the blockchain, finding that most of the 1600+ files contained text or images with clearly illegal content. Goldsmith et al. identified and analyzed blockchain hacker subnetworks, finding that hackers exchanged BTC for specific funds and categorized them into different hacker groups based on network characteristics. Spagnuolo et al. proposed a modular framework, Bitlodine, for blockchain identification, investigating CryptoLocker ransomware and accurately quantifying ransom payments and victim information. Matzutt et al. identified and analyzed blockchain-stored data, categorizing its content and finding some illegal and irregular information. These papers experimentally confirmed the existence of illicit data in some blockchains.
[0033] However, existing identification models are mostly single-chain models, where nodes are often biased, making it difficult to fairly and impartially identify one's own data, thus lacking credibility. Furthermore, these models struggle to filter out malicious nodes, further reducing their reliability. Therefore, there is an urgent need to propose a more reliable identification model to reduce the occurrence of illegal data uploads to the blockchain.
[0034] This application provides an on-chain identification method. The main difference lies in that the identification model is primarily based on the blockchain itself, utilizing blockchain smart contracts for identification. Blockchain itself is decentralized, providing a foundation of trust for both parties. Smart contracts are a key technology of blockchain, enabling automatic execution of contract content under certain conditions. This allows for trusted transactions without the involvement of third-party institutions. Li et al. proposed a blockchain-based solution to address the problem of easily forged and tampered house construction operation records, utilizing the tamper-proof characteristics of blockchain to achieve trusted recording of operations. Yong et al. proposed a blockchain-based solution to address issues such as falsified vaccine records, utilizing the tamper-proof characteristics of blockchain and smart contract technology to achieve trusted recording of vaccine data. Omar et al. designed a general framework using blockchain smart contracts to solve the problem of accurately tracking epidemic prevention materials during the pandemic. Zhu et al. proposed a method for trusted recording of infectious disease information using blockchain, greatly facilitating the trusted tracing of infectious disease pathways. Zhou et al. designed a witness model using blockchain smart contract technology to achieve trusted recording of cloud service violations. Based on this, using blockchain technology for trusted identification of illegal blockchain data, i.e., on-chain identification, has certain feasibility. In on-chain identification models, consensus is typically achieved through protocol voting and is managed and executed automatically by algorithms. Dursun et al. proposed an on-chain identification governance model. This model utilizes policy-based management and decentralized identity technology to achieve identification governance of the blockchain system. While simple and user-friendly, this model cannot filter honest nodes, resulting in low model credibility. Bao et al. proposed a multi-supervised permissioned blockchain model that supports identification and auditing, and experimentally demonstrated its security and efficiency. Although this model is secure and efficient, it struggles to guarantee node credibility. While the aforementioned identification governance models achieve identification governance of the blockchain, on the one hand, some models lack methods for selecting nodes, leading to low node credibility; on the other hand, these models lack incentive mechanisms, and nodes may make contrary identifications due to self-interest, resulting in low credibility of the identification results. These two factors contribute to the overall low credibility of the models.
[0035] To address the aforementioned issues, this application proposes a blockchain-based method for identifying illegitimate data uploaded to the blockchain. This method is applied to... Figure 1 The recognition model architecture shown.
[0036] In one embodiment, the illegitimate data on-chain identification model can have two blockchains: an identification blockchain (IBC) for identifying illegitimate data about to be uploaded and an identified blockchain (IDBC) for preparing to upload the data. Illegitimate data about to be uploaded is identified through a decentralized voting process by adding an identification blockchain. Nodes are designed as anonymous participants, such as businesses or individuals seeking revenue by identifying illegitimate data. They are selected as identification nodes only through a selection algorithm. In this invention, the incentive function rewards or punishes identification nodes based on their behavior; to obtain higher returns, identification nodes must act honestly. The identification model uses the Nash equilibrium principle of game theory to analyze and prove the credibility of the identification nodes.
[0037] The following will describe the specific implementation methods. Figure 2 The process of identifying illegitimate data on the blockchain, as shown below, is explained in detail.
[0038] Step 10: Generate a smart contract.
[0039] In one embodiment, the identification model can consist of two smart contracts: a Smart Contract of Identification Node Pool (INPSC) and a Smart Contract of Identification (ISC). The IBC can have at least two types of nodes: identification nodes (IN) and normal nodes (NN).
[0040] In one embodiment, the IBC first generates the INPSC and ISC. The INPSC has three main responsibilities: registering the NN, transitioning the NN's state, and selecting the IN. The blockchain, as a trusted party, provides a platform for businesses or individuals wishing to govern the blockchain or generate revenue through identification services; they need to register as NNs.
[0041] In one embodiment, the identification model sets up a basic smart contract to manage ordinary nodes, providing a set of interfaces for any IBC node to register to the identification node pool. Furthermore, the status of registered NNs can change to "online" or "offline," with online NNs being prioritized for selection. NNs in the identification node pool are managed according to their registration order.
[0042] Step 20: Negotiate with the identified chain and fill the identification parameters into the smart contract according to the negotiation result.
[0043] In one embodiment, before identification begins, the IBC needs to negotiate an identification treaty with the IDBC, including the identification time T. Identification And identification cost F Identification And so on. Among these, the most important thing between the IBC and IDBC is to negotiate the number of INs. The more INs, the higher the reliability of the identification. However, increasing the number will also lead to a higher identification cost. Based on the negotiation results, the IBC fills these parameters into the ISC.
[0044] Step 30: Send a token to the identified chain, so that the identified chain returns the data information stored in the InterPlanetary File System and the hash value generated by the token to the identification chain.
[0045] In one embodiment, when an IDBC is about to upload a piece of data to the blockchain, it needs to identify it before uploading. To prevent data retransmission issues caused by interruptions during this data transmission, the IDBC uploads the data to the InterPlanetary File System (IPFS) and sends the returned hash to the IBC. The IBC broadcasts the hash to each identification node (IN), which downloads the data and identifies it using the hash. During the identification process, T... Identification IBC will return a corresponding token based on the IN voting results.
[0046] In one embodiment, if the data is not illegal, a portion of the identification fee F will be waived. Reward Ultimately, IDBC will pay F to IBC. Identification or F Identification -F reward .
[0047] Step 40: Select the identification node from the ordinary nodes.
[0048] In one embodiment, the IBC's incentive mechanism is to incentivize INs based on their reputation; INs earn revenue by providing identification services. The more INs there are, the more reliable and trustworthy the model becomes. To further enhance trustworthiness, N INs form an identification committee to participate in the identification process. Together, they identify illegitimate data and receive identification fees from the IBC as a reward.
[0049] In one embodiment, INs are selected using a reputation-based node random selection algorithm, and these INs form an identification committee. This algorithm is implemented by INPSC on the IBC, and the ISC is responsible for calling the algorithm. INPSC designs two interfaces called by the ISC for selecting Number INs from the identification node pool. The "Request" interface is mainly used to record parameters required by the selection algorithm interface. The "Select" interface is used to select Number NNs as INs.
[0050] In one embodiment, such as Figure 3 As shown, the reputation-based node random selection algorithm selects NNs with high reputation values based on the seed, but it does not rule out selecting individual NNs with lower reputations. The higher the reputation, the greater the probability of being selected. When the reputation value of an NN is below 0, it will no longer be selected as an IN. Furthermore, a new seed is generated for each selected potential identification node. The new seed is generated based on the previous seed and a timestamp. This process is repeated until the required Number of potential identification nodes are selected. At the beginning of the algorithm, it is first checked whether there are enough available NNs, and the number of NNs should be much greater than IN to ensure the fairness, impartiality, and randomness of the algorithm's output.
[0051] In one embodiment, the hash value composed of the total reputation value of nodes in the identification node pool, timestamp, and block height is used as the seed, as shown in Formula (1).
[0052] seed=hash(trust,block.timestamp,block.height) (1)
[0053] Considering the difficulty of using the hash value composed of the total reputation value, timestamp, and block height of all nodes in the identification pool as the seed, and the difficulty for outsiders to know the total reputation value of the nodes in the pool, it can be proven that the reputation-based random selection algorithm is random. That is, IDBC cannot manipulate the selection result to gain an advantage in the identification committee. Using a reputation-based random selection algorithm ensures that the majority of INs selected for the identification committee are honest and independent nodes. Because these INs are anonymous, IDBC has difficulty knowing which nodes have entered the committee, thus ensuring that the INs do not side with IDBC's interests.
[0054] In one embodiment, the algorithm is designed to favor nodes with high reputation scores and incorporates randomness. This enhances the credibility of IBC identification and establishes mutual trust between IBC and IDBC. IDBC sends the data stored in IPFS and the hash value generated by the token to IBC. It's worth noting that IBC sends a token to IDBC. IBC then broadcasts this hash to all INs on the identification committee. Each IN downloads the data using this hash value.
[0055] Step 50: Broadcast the received hash value to the identification node, so that each identification node can download the data information through the hash value and identify the data information.
[0056] In one embodiment, ISC will start timing T from the moment the first IN report data is found to be illegal or non-compliant. Report During this time period, the ISC will accept reports from other INs. When the timer ends, if the ISC receives no fewer than 'Illegal' reports from the Identification Committee, the IBC will automatically confirm that the data is illegal or non-compliant. The 'Illegal' value is also the result of negotiation between the IBC and IDBC; the IBC will include it as an important parameter in the ISC. It's important to note that the 'Illegal' value must be greater than half of the 'Number' value. The larger the 'Illegal' value, the more reliable the IBC's identification result will be. For example, if there are 5 identification nodes in the Identification Committee, at least 3 INs with 'Illegal' values need to report illegal or non-compliant data for the IBC to confirm that the data is illegal. It's crucial to note that the ISC will only accept identification reports from committee members; reports from nodes outside the committee will be considered illegal and rejected. Furthermore, an IN cannot report multiple times within the same reporting period. In a sense, these 'Number' identification nodes constitute a Number-Man game. In this game, each IN aims to maximize its gain.
[0057] In one embodiment, the blockchain illegitimate data identification process will end in two scenarios. One scenario is when the identification time T... Identification The transaction has concluded, and the data is not invalid. Alternatively, the data may be invalid. Depending on these different scenarios, IBC, IN, and IDBC receive the corresponding fees from the contract.
[0058] One of the most common forms of game theory is the strategy form of multi-player games. This definition consists of a set of participants (identifying nodes), a set of strategy profiles, and a payoff function. The interaction between INs constitutes the basic type of dynamic game with complete information in game theory; complete information means that the incentive functions of each IN are known to the others. In one embodiment, we define the IN game as follows:
[0059] Definition 1, Identifying the Game: This is a game with Number INs, denoted by (INC, ST, EX).
[0060] Where INC = {IN1,IN2,...,IN} Number} is a set of Number INs. Each IN is selected using a reputation-based random selection algorithm, and these INs form an identification committee.
[0061] ST = ST1 × ST2 × ... × ST Number It is a set of IN action policy configurations, where ST h It is the identification node IN h A set of action strategies for recognition. h You can choose any action AC h ∈ST h Each IN may take a different action in each recognition process, i.e., AC = {AC1, AC2, ..., AC...} Number}。 (h=1,2,...,Number).
[0062] EX = {EX1, EX2, ..., EX} Number} is a set of activation functions, EX h It is the identification node IN h Implement specific action strategies AC h The excitation function is given by (h = 1, 2, ..., Number), EX. h ={R h ,P h}, where R and P represent reputation incentives, respectively.
[0063] The identification process generally involves two actions: reporting illegitimate data to the IBC and remaining silent while not reporting illegitimate data to the IBC. In the game of Number INs, illegitimate data INs are placed into the report set INR, and identification nodes that do not report illegitimate data are placed into the silent set INS. If... Then IN h ∈INR; Then IN h ∈INS(h=1,2,...,Number). These operations determine whether the data is invalid. Data is considered invalid only if the number of elements in INR is greater than or equal to Illegal; see Definition 2 for details. Otherwise, IBC recognizes it as normal data. The identification process is now complete.
[0064] Definition 2, Illegal Data Confirmation: When ||INR||>=Illegal, where That is, if there are no less than Illegal INs reporting the data information, then IBC considers this data information as illegal data information. Otherwise, IBC considers this data information not as illegal data information.
[0065] Here, Number should be as large as possible. The larger Number is, the higher the credibility of the recognition result. In addition, Illegal < number and Illegal > 2 to more fairly and credibly identify data information. When IN conducts identification, it needs to pay a corresponding deposit each time it reports to IBC. If this data information is not identified as illegal data information, then IN will not be able to recover the deposit and its own credibility value will also be deducted. According to the above analysis, the detailed incentive function design is as shown in Definition 3.
[0066] Definition 3. Incentive function: Design the value of the incentive function according to the final recognition result.
[0067] When the recognition result is that this data information is illegal and违规 information:
[0068]
[0069] When the recognition result is that this data information is not illegal and违规 information:
[0070]
[0071]
[0072] Among them, (2) and (5) are the credibility reward formulas when the recognition node successfully recognizes. (3) and (4) are the credibility penalty functions when the recognition node makes a wrong recognition. Among them, μ and θ are credibility factors, responsible for controlling the credibility growth and reduction of the recognition node in different situations. It should be noted that the advantage of doing this is that it can make the recognition node with a high credibility value reduce the evil-doing rate, because every time it does something evil, it will reduce a large amount of credibility value, thus reducing its income. At the same time, it can also encourage the recognition node with a low credibility value to correctly identify the data, because every time it makes a correct recognition, it will increase a relatively large amount of credibility value, thus increasing its income. According to game theory, if an IN knows the actions that other INs will take in the future, then it will choose a behavior strategy that maximizes the benefit according to the actions taken by other INs, which is called the best response. Therefore, the best response of IN is defined as follows.
[0073] Definition 4. The best response of the recognition node INh:
[0074] Let AC -h ={AC1,AC2,...,AC h-1 ,AC h+1 ,...,AC number} is the set of other IN actions that do not have INh. Then IN h The optimal response strategy for other nodes is Make EX h (AC h AC -h )≥EX h (AC h AC -h ), thus making IN h To obtain the maximum benefit.
[0075] The Nash equilibrium can be viewed as a stable state among the Number of identification nodes. In this state, no IN will choose any other action strategy. At this point, to obtain the maximum benefit, each IN will adopt its own optimal action strategy.
[0076] Definition 5. Nash Equilibrium Point: This is a special IN action point AC = (ACh, AC-h) such that every IN action AC... h All of these are actions taken by other identification nodes (AC). -h The optimal response strategy. and You can get EX h (AC h AC -h )≥EX h (AC h AC -h ).
[0077] In one embodiment, based on the above definition, the following theorem can be derived.
[0078] Theorem 1: In the node identification game, there are exactly two Nash equilibria.
[0079] Proof: According to definitions 1 and 2, in a game with Number INs, Number ≥ 3, Both Number and Illegal are integers.
[0080] For a set of policy configuration files This means that ||INR|| = Number > Illegal, i.e., the data is invalid. According to the activation function designed in Definition 3, for The benefits he can obtain are If an IN option chooses another action, i.e., chooses silence instead of reporting, the final identification status of the data will not be changed. This is because ||INR||=Number-1≥Illegal, and the benefit of this IN option is... According to Definition 5, this policy profile is a Nash equilibrium point.
[0081] Similarly, for another set of policy configuration files This means that ||INS||=Number>Number-Illegal≥2, therefore the identification result of this data information is normal data information. From the activation function defined in 3, we know that for... The benefits it can obtain are If an individual IN chooses a different action strategy, namely reporting illegal data, the identification result of that data will not change. This is because ||INS||=Number-1>Number-Illegal-1≥2, and the payoff of that IN is... According to Definition 5, this action strategy configuration is also a Nash equilibrium point.
[0082] Other action strategies are combinations of actions, including both reporting illegal data and remaining silent. This means... and At this point, there are two scenarios: the data is invalid or it is valid. When the data is invalid, ||INR|| ≥ Illegal. It can change the action to report the data. However, the identification result of the data will not change because ||INR||+1>Illegal, thus making IN... h Increased income, from Increase to On the other hand, when this data information is normal data information, ||INS||>Number-Illegal, It can change its action to not report that the data is illegal, i.e., remain silent. However, the identification status will not change, ||INS||>Number-Illegal+1. This makes IN... h Increase his income, from Increase to These examples demonstrate that these IN action strategies are not at a Nash equilibrium.
[0083] Therefore, in the identification game between IN, there are exactly two Nash equilibria, i.e. and Taking five identification nodes as an example, i.e., Number = 5, according to Definition 2, Table 1 shows the previously defined activation functions. The value elements in Table 1 are vectors of the corresponding activation functions, which can be represented as (AC1, AC2, AC3, AC4, AC5). According to Theorem 1, the Nash equilibrium points of IN are [(+,+), (+,+), (+,+), (+,+), (+,+)] and [(+,+), (+,+), (+,+), (+,+), (+,+)].
[0084] From the above analysis, it can be seen that for a rational identification node that wants to maximize its gains while minimizing risk, it must behave as follows during the identification process: If it considers the data to be illegal, IN knows that most other nodes are more likely to report it to gain more revenue. Therefore, this node will also report the data as illegal. Conversely, if the data is legitimate, this node knows that other nodes will not report it to minimize losses, as each IN has to pay a deposit for each report. Thus, when the data is legitimate, all INs tend to remain silent, achieving a Nash equilibrium. Similarly, if the data is invalid, IN's profit-seeking nature will cause it to choose to report the data, thus reaching another Nash equilibrium point, i.e.
[0085] The above analysis shows that in order to maximize its benefits, IN must choose to be an honest node, that is, it must fairly and impartially identify the data information.
[0086] Table 2. Five-player game and its incentive functions
[0087]
[0088] Reputation-based random selection algorithms ensure that the majority of INs selected are trustworthy and independent. The incentive function design allows for fairer and more impartial identification of INs. However, a penalty mechanism is still needed to reduce malicious or irrational INs. All interactions between INs and smart contracts are stored on IBC, allowing for the review and punishment of INs based on their behavioral history.
[0089] In one embodiment, the identification nodes are audited based on their behavioral history, and lazy identification nodes are penalized. Lazy identification nodes refer to those INs who are unwilling to actively perform their duties. This is because some identification nodes, when faced with incentives, choose to be lazy, meaning they choose not to report data to the IBC regardless of whether the data is illegal. In this way, while avoiding penalties for identification errors, some identification revenue can still be obtained.
[0090] In one embodiment, node reputation values can be used to better manage these INs (Information Not Reported). INs that fail to report illegitimate data will have their reputation deducted. INs that report illegitimate data but ultimately find it to be legitimate will also have their reputation deducted. When a node's reputation value is 0, that node will no longer be selected as an IN. Furthermore, the final incentive function is also linked to the reputation value. In this way, lazy INs are reduced, and each IN is incentivized to fairly and impartially identify data.
[0091] Randomization algorithms can ensure that witnesses are largely independent, and the incentive function can also motivate witnesses to make the right judgment.
[0092] In one embodiment, since all interactions with the smart contract are uploaded to the blockchain and can be viewed, it is possible to examine a node by reviewing its historical behavior to determine if it is malicious. Therefore, nodes can be audited to detect malicious and irrational nodes and remove them from the node pool.
[0093] Some nodes choose to report at specific times, such as reporting illegal data in the last few minutes of a given time period. While these nodes may not initially gain any benefit, and could even incur significant costs, their behavior demonstrates cheating to other nodes. It implies that other nodes should also report during this time period, thus colluding to make identification inaccurate, and allowing them to reap substantial benefits.
[0094] In one embodiment, to avoid the above situation, a smart contract can be used to review the node and determine whether it is a malicious node based on its specific behavior. The specific punishment measures are determined according to the severity of its malicious behavior. For minor malicious behavior, a portion of its reputation value is deducted; for severe malicious behavior, it is removed from the battery, preventing it from gaining benefits through identification. The applicant of this invention has further analyzed and verified the technical solution of this invention through experiments. Specifically, as follows:
[0095] Based on the illegitimate data identification model and incentive function design, the applicant conducted related experiments using Ethereum smart contracts. Our illegitimate data identification model includes two types of smart contracts: INPSC and ISC. In this section, the applicant elaborates on these two types of contracts, describes the detailed functionality of the interfaces within the smart contracts, and demonstrates the various states of IN. Subsequently, the applicant experimented with the transaction costs of these interfaces on the Ethereum testnet.
[0096] The interfaces designed in both smart contracts were named Figure 4 and Figure 5The text above the arrow. The text format is role:interface, meaning only this role or smart contract can call this interface. Smart contracts can implement specific functions for specific roles through an inspection mechanism, a feature of the programming language provided by Ethereum. Here, IBC represents the identifying chain, IDBC represents the identified chain, IN represents the identifying node, ISC represents the generated identifying smart contract, and INPSC represents the identifying node pool smart contract.
[0097] Figure 4 This demonstrates the four states of a node defined in a smart contract: "Online," "Offline," "Quasi-identified Node," and "Identified Node." The state transition mechanism of IN is as follows: Figure 4 As shown.
[0098] After registering in INPSC, a blockchain node becomes a regular node in the identification node pool. At this point, it can choose between "offline" and "online" states. NNs in the "offline" state do not participate in the selection algorithm. Only NNs in the online state have a chance to be selected as "prospective identification nodes" through a reputation-based random selection algorithm. During this time, NNs need to constantly check their status on IBC. Checking their status is free, satisfying the requirement of regular nodes to check their status frequently anytime, anywhere. After the status changes to "prospective identification node," a timer window appears. The NN must confirm before the window closes to become an IN (Information Node), at which point its status becomes "identification node." If it chooses to refuse, the random selection algorithm will be re-executed. Calling the ISC interface "INconfirm" represents confirming the selection to become an IN. Before the data information identification process ends, the IN has the right to choose "INrelease" to voluntarily exit the ISC. Additionally, IBC can also disband the identification committee using "ResetIN."
[0099] To reduce malicious nodes, the applicant set a reputation value attribute for the NameNode (NN) to measure its behavior. Each node is assigned an initial reputation value, Rinit, upon registration. This can be a predetermined constant or a value dependent on certain interests; here, we set the initial value Rinit = 50. Furthermore, because some NNs do not frequently check their status, they may fail to confirm their inclusion as a candidate node within the specified time when the ISC selects them. In this case, their status becomes "quasi-identified node." To return a node to an "online" state, the "Reverse" function is used.
[0100] Figure 5 This displays the status transitions of the data information recognition results by ISC. There are five recognition statuses: "Idle," "Initial," "Recognized," "Illegal," and "Completed," as shown below. Figure 5The circles in the diagram indicate state transition paths. The arrows represent the transition paths. The two squares in the diagram represent the corresponding roles in the smart contract. At the end of the identification process, they can each withdraw their income.
[0101] Smart contracts on IBC cannot run on their own. State transitions must be triggered by certain interfaces and incur corresponding costs. We designed an interface for IN to modify its state, which can be used when necessary. For example, when the identification service ends normally, IBC can end the identification process and receive revenue through the "IBCEnd" interface, while a portion of the identification fee will be distributed to different INs through an incentive function. Similarly, when data is identified as invalid, it can end the identification process through "IBCEndIllegalandWithdraw". In this case, the identification fee will be distributed to both IBC and IN, and the identification state will change from "Identification" to "Completed".
[0102] To test the illegitimate data on-chain identification model, the applicant conducted related simulation experiments on smart contracts on Rinkeby. The applicant set up several accounts on Rinkeby to simulate different roles: IDBC, IBC, and IN. The applicant used the interface by allocating "Ether" to each account and paying different fees according to the incentive function. For the experiment, a basic INPSC was first deployed, and several accounts were registered to the identification node pool. Then, IDBC sent the hash value returned by IPFS to initiate IBC data identification. Afterwards, all possible scenarios were tested to test and verify all interfaces. The results show that the applicant's identification model basically meets the identification requirements.
[0103] The credibility of the identification model is proven by game theory and guaranteed by a reputation-based random selection algorithm, with its credibility supported by the technology of the blockchain itself. Therefore, the applicant primarily analyzes some performance information from the experimental study. Performance refers to the complexity of each interface in the contract. This determines the corresponding fees that IN needs to pay in IBC. Because IN needs to execute certain procedures defined in the methods, the more complex the interface, the higher the fees it needs to pay, i.e., the higher the cost. This is measured by Ethereum's definition of "gas". The identification fee is the product of gas consumption and the price per unit. Therefore, the cost is similar whether on the mainnet or in the simulation experiment. Thus, the applicant recorded all gas consumption for each operation of the identification model.
[0104] Figure 6The simulation results of the identification model are presented. It can be seen that compared to IDBC and IN, IBC requires more cost throughout the identification cycle, while IDBC and IN consume less through interface usage. This aligns with our initial design intent. In most cases, the primary purpose of IBC is to identify illegitimate data that IDBC is about to upload to the blockchain. IN consumes less gas but receives more gas in return. IDBC pays less gas, a small percentage compared to the identification service fee. These cost considerations are sufficient to convince IN and IDBC to participate in the identification model. Furthermore, these gas costs are based on smart contracts, and it is still possible to further optimize the complexity to reduce gas consumption.
[0105] In a comparative experiment of node random selection algorithms, different algorithms were applied to smart contracts, which were then deployed to the Ethereum test network Rinkeby. Five nodes were placed on the test network, each with either the same or different reputation values, and 20, 40, 60, 80, and 100 experiments were conducted. The experiments tested the effectiveness of different algorithms in selecting three nodes from five nodes. The experimental results are as follows: Figure 7 and Figure 8 As shown.
[0106] Figure 7 This shows the number of times nodes with the same reputation value were selected. Figure 8 This shows the number of times each of the five nodes was selected when their reputation scores were 100, 80, 60, 40, and 20. Figure 7 As can be seen, when node reputation values are the same, ordinary nodes are selected as identification nodes approximately the same number of times. However, when selecting nodes with reputation values of 100 and 80, the RBRSA is 23% and 13% higher than the algorithm, respectively; when selecting nodes with reputation values of 40 and 20, the RBRSA is 11% and 27% lower than the algorithm, respectively. The two algorithms select nodes with a reputation value of 60 approximately the same number of times. Figure 8 As shown in the diagram. Although the greedy algorithm is better at selecting honest nodes compared to the other two algorithms, it will always choose the nodes with the highest reputation values when faced with different reputation scores. On the one hand, nodes with low reputation values will no longer be selected, lacking fairness; on the other hand, if nodes with high reputation values group together, it will reduce the credibility of the recognition model. Therefore, this paper designs a node selection algorithm that selects more reliable nodes compared to the other two algorithms.
[0107] In the incentive mechanism experiment, we assume that nodes continuously identify data correctly or incorrectly, and demonstrate the effect of the incentive mechanism under different μ parameters by controlling the parameter μ. Figure 7 and Figure 8 As shown.
[0108] Figure 9 This graph shows the reputation growth effect when a node correctly identifies data information 50 times consecutively. As can be seen from the graph, the reputation value increases with the number of correct identifications. However, the higher the reputation value, the slower the reputation value growth. Furthermore, the convergence speed of the reputation value increases with the increase of the μ value. The convergence speed at μ = 0.3 is 16% faster than at μ = 0.2. Figure 10 This graph shows the reputation decline when a node misidentifies data 20 times consecutively. As the graph shows, the reputation value decreases with increasing misidentifications. However, the lower the reputation value, the slower the decline. Furthermore, the faster the reputation value converges, the higher the μ value becomes. Therefore, the incentive mechanism can reduce malicious nodes among high-reputation nodes because it imposes more penalties on them; conversely, it can increase honest nodes among low-reputation nodes because it rewards them more.
[0109] This application proposes a blockchain-based model for identifying illegitimate data uploads and designs an incentive function for each identification node. To maximize their profits, identification nodes must provide honest identification services, a principle we demonstrate using game theory. Finally, the model is implemented using an Ethereum smart contract. The smart contract not only identifies illegitimate data uploaded to the blockchain but also manages the identification nodes. Experimental studies test the performance of the model's interface and verify its feasibility. Incentivizing identification nodes ensures they identify data fairly and impartially.
[0110] Based on a general inventive concept, this invention also provides a blockchain-based illegal data on-chain identification device. The device includes a memory and a processor. The memory is used to store a computer program. When the computer program is executed by the processor, it implements the steps in the above-described blockchain-based illegal data on-chain identification method.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying illegitimate data on the blockchain, characterized in that, The method is applied to an identification chain in a blockchain, and the method includes: Generate smart contracts; Negotiate with the identified chain and fill the identification parameters into the smart contract based on the negotiation results; Send a token to the identified chain, so that the identified chain returns the data information stored in the InterPlanetary File System and the hash value generated by the token to the identification chain; Select the identification node from the ordinary nodes; The received hash value is broadcast to the identification node, so that each identification node can download the data information through the hash value and identify the data information.
2. The blockchain-based method for identifying illegitimate data on-chain as described in claim 1, characterized in that, The smart contract includes at least an identification node pool smart contract and an identification smart contract.
3. The blockchain-based method for identifying illegitimate data on-chain as described in claim 2, characterized in that, The responsibilities of the smart contract for the identification node pool include at least: registering ordinary nodes, transitioning the state of ordinary nodes, and selecting identification nodes.
4. The blockchain-based method for identifying illegitimate data on-chain as described in claim 2, characterized in that, The process of negotiating with the identified chain and filling the identification parameters into the smart contract based on the negotiation result includes: The identification treaty is negotiated between the chain being identified and the chain being identified. Based on the negotiation results, the identification parameters are filled into the identification smart contract, wherein the identification parameters include at least the number of identification nodes.
5. The blockchain-based method for identifying illegitimate data as described in claim 4, characterized in that, The identification parameters also include: identification time and identification cost.
6. The blockchain-based method for identifying illegitimate data on-chain as described in claim 3, characterized in that, The process of selecting an identification node from ordinary nodes includes: The identification node is selected from ordinary nodes using a reputation-based node random selection algorithm and based on a seed; The seed includes at least a hash value composed of the total reputation value of nodes in the identification node pool, timestamps, and block heights.
7. The blockchain-based method for identifying illegitimate data on-chain as described in claim 6, characterized in that, The status of the ordinary nodes includes online and offline, and ordinary nodes that are online are preferentially selected as identification nodes.
8. The blockchain-based method for identifying illegitimate data on-chain as described in claim 1, characterized in that, The identification node is reviewed by examining its behavioral history, and then processed based on the review results.
9. The blockchain-based method for identifying illegitimate data on-chain as described in claim 8, characterized in that, The step of reviewing the identification node based on its behavioral history and processing the identification node according to the review results includes: The audit results include the degree of wrongdoing; Based on the degree of malice, the identified node is processed, and the processing includes at least one of the following methods: deducting the reputation value of the identified node; or kicking the identified node out of the node pool.
10. A blockchain-based device for identifying illegitimate data uploaded to the blockchain, characterized in that, The device includes a memory and a processor, the memory being used to store a computer program, which, when executed by the processor, implements the blockchain-based method for identifying illicit data on-chain as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Credible threat intelligence identification method and device based on blockchain consensus mechanism
CN112039840A
Illegal node identification method, computer equipment and storage medium
CN113886124A