Probabilistic Relays for Efficient Propagation in Blockchain Networks
By building an interface correlation-based data transmission model in the Bitcoin network, selectively relay data packets solve the problem of inefficiency in existing networks under high concurrent transactions, and achieve faster and safer data transmission.
Patent Information
- Application Number
- JP2024015335
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-08-21
- Filing Date
- 2024-02-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2038-06-25
AI Technical Summary
When existing Bitcoin networks process large amounts of transactions, the inefficiency of transmitting and distributing data packets is caused by queue backlog and network bottlenecks, and they are unable to effectively handle highly concurrent transaction requests.
By establishing interrelated interfaces between nodes, a data transmission model based on interfaces is built, selectively relay data packets based on correlation coefficients between interfaces, to reduce unnecessary data transmission and avoid the impact of malicious nodes.
It improves the speed and efficiency of data packets in the network, reduces the traffic between nodes, avoids performance bottlenecks caused by queue backlog, and enhances the security and reliability of the network.
Smart Images

Figure 0007675876000015 
Figure 0007675876000016 
Figure 0007675876000017
Abstract
Description
[Technical field]
[0001] This specification generally relates to computer-implemented methods and systems suitable for implementation in nodes of a blockchain network. Improved blockchain node structures, network architectures, and protocols for handling large numbers of transactions and large blocks of transactions are described. The present invention is particularly suited for, but not limited to, use with the Bitcoin blockchain. [Background technology]
[0002] In this document, we use the term "blockchain" to include all forms of electronic, computer-based distributed ledgers. These include consensus-based blockchain and transaction chain technologies, permissioned and un-permissioned ledgers, shared ledgers, and variations thereof. The most widely known application of blockchain technology is the Bitcoin ledger, although other blockchain implementations have been proposed and developed. Although Bitcoin may be referred to herein for convenience and explanation, it should be noted that the present invention is not limited to use with the Bitcoin blockchain, and alternative blockchain implementations and protocols are within the scope of the present invention. The term "user" may refer herein to a person or a processor-based resource.
[0003] A blockchain is a peer-to-peer electronic ledger implemented as a computer-based decentralized distributed system, composed of blocks of transactions. Each transaction is a data structure that encodes the transfer of control of digital assets between participants in the blockchain system and contains at least one input and at least one output. Each block contains a hash of the previous block so that blocks are chained together to form a permanent and immutable record of all transactions that have been written to the blockchain since its inception. Transactions contain small programs, known as scripts, embedded in their inputs and outputs. Scripts specify how and by whom the output of a transaction can be accessed. In the Bitcoin platform, those scripts are written using a scripting language based on Stack.
[0004] In order for a transaction to be written to the blockchain, it must be "validated". Network nodes (miners) perform the task of making sure that each transaction is valid, and invalid transactions are rejected from the network. A software client installed on the node performs this validation task by running its locking and unlocking scripts on the unspent transaction (UTXO). If the execution of the locking and unlocking scripts evaluates to TRUE, the transaction is valid and the transaction can be written to the blockchain. Thus, for a transaction to be written to the blockchain, it must i) be validated by the first node that receives the transaction, and if the transaction is validated, the node relays it to other nodes in the network, ii) be added to a new block constructed by miners, and iii) be mined, i.e., appended to the public ledger of past transactions.
[0005] While blockchain technology is most widely known for its use in implementing cryptocurrencies, digital entrepreneurs are beginning to explore the use of both the cryptographic security system on which Bitcoin is based, and the data that can be held on the blockchain, to implement new systems. It would be highly advantageous if blockchain were used for automated tasks and processes that are not limited to the field of crypto assets. Such solutions could leverage the benefits of blockchains (e.g., permanent tamper-resistant recording of events, distributed processing, etc.) while broadening their applications.
[0006] According to Blockchain.info [Bitcoin transaction levels are available through Blockchain Luxembourg SARL and can be retrieved from http: / / blockchain.info], in April 2017, the average number of Bitcoin transactions per block was around 2000 units. In the near future, the constraint on the maximum block size may be relaxed and the number of transactions per block may increase significantly.
[0007] Competitive cryptocurrencies need to propagate large volumes of unconfirmed transactions as fast as possible. By comparison, Visa's electronic account settlement has a peak capacity of 56,000 transactions per second. [VISA transaction levels are summarized in a document published by Visa in June 2015, "Visa Inc. at a Glance," available at http: / / usa.visa.com / dam / VCOM / download / corporate / media / visa-fact-sheet-Jun2015.pdf.]
[0008] The current three-step messaging protocol for the exchange of new transactions in the Bitcoin network by inventory is insufficient to handle the rapid distribution of transaction volumes several orders of magnitude larger than the current standard (~5 transactions per second [Decker, Christian and Wattenhofer, Roger (2013). Information propagation in the bitcoin network. IEEE Thirteenth International Conference on Peer-to-Peer Computing (P2P), 2013]).
[0009] Today's Bitcoin network is centered around mining in terms of computational work. With a significant increase in transaction volume, this is not always feasible. The solution described herein allows the Bitcoin network to handle the propagation of large volumes of transactions.
[0010] Known methods of sending data packets or transactions through the Bitcoin network, such as distributing new transactions by three-step messaging, result in slow propagation and distribution of data packets across the network. During the preparation phase, queues arise in and out of nodes. Summary of the Invention
[0011] Overall, the present invention resides in a novel approach for processing and propagating increased transaction volumes that are several orders of magnitude greater than current blockchain capacities. This can be achieved by supporting faster propagation and distribution of data packets across the network by reducing communication between nodes. Long queues in and out of nodes caused by bottlenecks are suppressed by selectively relaying data packets according to correlations between interfaces.
[0012] This method allows nodes to operate adaptively (i) independent of the number of interfaces connected to peer nodes, and (ii) independent of changes related to the node's interfaces, thereby taking into account new connections, lost connections, and malicious nodes, and maintaining the integrity of the network.
[0013] In this way, there is no limit to the number of interfaces on a node that can be managed, and the method adapts to the network and node environment, so network performance and size are not limited. Inefficient transmissions are minimized and malicious nodes are avoided.
[0014] In this manner, the performance of the blockchain network is improved and the blockchain network or the overlay network that interfaces with the blockchain network is improved.
[0015] Thus, according to the present invention there is defined a method as defined in the appended claims.
[0016] Thus, there is provided a computer-implemented method for a node of a blockchain network, the node having a number of interfaces connecting to peer nodes, the method comprising: determining a correlation matrix having correlation coefficients representative of correlations between data processed at each interface of said node; receiving data at a receive interface of said node; Selecting at least one other interface from a plurality of other interfaces of the node and relaying the received data from the at least one other interface, the other interface being selected according to a set of correlation coefficients of the receiving interfaces; It is desirable to provide such a method, comprising:
[0017] The data may correspond to an object such as a transaction or a block.
[0018] An indicator is derived from the correlation matrix, and data is relayed if the correlation between the receiving interface and the at least one other interface is lower than the indicator. Alternatively, relaying can occur if the correlation is higher than the indicator.
[0019] The indicator can be a threshold. The indicator can represent a level of correlation between the interface receiving the data and a number of interfaces, such as an average value. The average value can be an average of the correlation coefficients between the receiving interface and the other interfaces.
[0020] The indicator is used to determine a metric that sets the criteria for selecting which of the plurality of other interfaces to select for relaying data.
[0021] The metrics may be viewed as instructions used to select from which interface data should be relayed. The instructions may be set according to or dependent on thresholds or indicators.
[0022] The metric may be used to rank the correlation of (respective) interfaces between a node and one or more other nodes. This may provide the advantage of being able to better understand and respond to the quality of relaying provided by nodes peering in the network, as described in more detail below. This not only allows for a more granular understanding and appropriate response to network behavior, but also provides advantages in terms of detecting and responding to malicious behavior. The present invention may thus provide a more efficient and more secure network.
[0023] Each node in the network may be associated with a correlation matrix based solely on the transaction flow received. To avoid the propagation of malicious information, correlation information may not be exchanged between nodes or peers in the network.
[0024] As an example, the instructions may require that an interface have a correlation index that is below the indicator. The instructions can determine which interface to relay data to depending on whether the correlation index is above or below the indicator.
[0025] An example of data relay when the index falls below the indicator may be a normal steady-state network condition. In 'steady-state' conditions, the number of interfaces and connected peers does not change and the correlation matrix does not change.
[0026] An example of data relay when the index exceeds the indicator may be a change of state, such as the addition of a new peer node connection to the node. Different instructions may be applied to different interfaces.
[0027] The data to be distributed may include data received from the blockchain network, including data from miners, peer nodes, and full nodes. The correlation between data processed at each interface of the node may take into account overlaps of data packets or objects received at the node. The indicator may be the average or median correlation index of each interface of the node. The indicator may define a point between the lowest and highest correlated interfaces. Data or objects may be relayed from an interface when the correlation index falls below the indicator. Interfaces with a low degree of correlation may be prioritized since it can be safely assumed that highly correlated interfaces will get the data.
[0028] Nodes process data through interfaces, which includes sending and receiving data or objects, such as transactions or blocks.
[0029] The data may be in a network packet representing a serialized transaction and an identification representing a connection to an adjacent or peer node. The interface may be a logical interface ID representing a TCP / IP connection to a sending / receiving peer.
[0030] The node can construct the correlation matrix by monitoring (i) the data identifier of each packet of data processed through each interface and (ii) the same transaction processed through a pair of interfaces, and then determine the correlation coefficient between any two interfaces.
[0031] The correlation matrix may have m(m-1) elements, as follows:
number
[0032] The correlation matrix may have m(m-1) elements, as follows:
number
[0033] The indicator may be determined by determining, for each interface connected to a peer node, a set of correlation coefficients derived from the correlation matrix, the set having the correlation coefficients between each interface, and deriving an average or median from the set.
[0034] The indicator is used by the instruction set to determine which interface to relay the data to.
[0035] The indicator for determining the number of interfaces to relay data may further be based on at least one of a reset time, which is the time from node startup or start-up, and a change time, which is the time between change events including at least one of a new peer node connecting to an interface, a termination of a connection to an interface, and an interface connecting to a malicious node. The reset time may be used for new nodes. The change time may be used for existing nodes.
[0036] At node startup, the node may connect with peer nodes, such as nodes adjacent or directly connected to the node, and relay data via all interfaces during the reset time period during which the correlation matrix is built, and after the period has elapsed, the node relays all objects from the node via interfaces having a correlation index below the indicator.
[0037] Upon detection of a change event, the correlation matrix may be reset and re-determined, which may occur when a periodic reset is required.
[0038] Upon detecting a change event, the node can relay all objects from the node through an interface that has a correlation index above the indicator.
[0039] Upon detecting the disconnection of a peer node from the interface, the correlation matrix may be reset and re-determined, which may occur when a periodic reset is required.
[0040] Upon detecting a connection to a new peer node, the peer node is connected to the interface, which may relay data through all other interfaces for the duration of the reset time and / or the change time.
[0041] The node can generate raw data and the number of interfaces selected to process and transmit the raw data. This can be augmented by increasing the indicator to a number between the current or nominal number of interfaces selected for relaying and the total number of interfaces.
[0042] It is also desirable to provide a computer-readable storage medium having computer-executable instructions that, when executed, configure a processor to perform any of the claimed methods.
[0043] It is also desirable to provide an electronic device having an interface device, one or more processors coupled to the interface device, and a memory coupled to the one or more processors, the memory storing computer-executable instructions that, when executed, configure the one or more processors to perform any of the claimed methods.
[0044] It is also desirable to provide a node of a blockchain network, said node being configured to perform any of the claimed methods.
[0045] It is also desirable to provide a blockchain network with claimed nodes.
[0046] The present invention is suitable for use with, but not limited to, the Bitcoin (BTC) blockchain.
[0047] It is also desirable to provide a supernode of a blockchain network, the supernode having a plurality of claimed nodes and a shared storage entity for storing the blockchain, the shared storage entity being either a common storage node, a distributed storage, or a combination thereof, and blocks assembled by the plurality of nodes are sent to and stored in the shared storage entity, whereby the shared storage entity holds the blockchain. The shared storage entity has a storage capacity of at least 100 gigabytes.
[0048] It is also desirable to provide a blockchain network having a plurality of supernodes as claimed, said plurality of supernodes being connected on said blockchain network, a shared storage entity of each supernode being configured to store a copy of the blockchain, said blockchain network having at least 10 supernodes.
[0049] Such an improved solution has now been devised.
[0050] These and other aspects of the invention will be apparent from and will be elucidated with reference to the embodiments described hereinafter, which will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief description of the drawings]
[0051] [Figure 1] The overall structure of the block is shown. [Diagram 2] We present an improved architecture of the Bitcoin network in terms of an operational diagram showing the steps from when a user submits a transaction to the blockchain. [Diagram 3] 1 shows a graph illustrating an example of the aggregate size of transactions waiting for confirmation in MEMPOOL. [Figure 4]Shows multiple nodes linked to an internal centralized storage facility. [Diagram 5] It represents a configuration in which each node is part of both a distributed MEMPOOL and a distributed storage facility. [Figure 6] Represents network packets being sent and received serially at the application level through nodes on the Bitcoin network. [Figure 7] It represents the rules that a single transaction between nodes across the Bitcoin network follows. [Figure 8] 1 is an example of a correlation matrix showing correlation coefficients of interfaces connected to peer nodes. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0052] Described herein is a solution to the problem of processing and storing large gigabyte-sized blocks.
[0053] <Type of Blockchain Network and Validation Node> A blockchain network may be described as a peer-to-peer, open-membership network in which anyone can join without an invitation or consent from other members. Distributed electronic devices that run instances of the blockchain protocol on which the blockchain network operates may participate in a blockchain network. Such distributed electronic devices may be called nodes. The blockchain protocol may be, for example, the Bitcoin protocol or other cryptocurrency.
[0054] The electronic devices that run the blockchain protocol and form the nodes of the blockchain network may be of various types, including, for example, computers such as desktop computers, laptop computers, tablet computers, servers, computer farms, mobile devices such as smartphones, wearable computers such as smart watches, or other electronic devices.
[0055] The nodes of a blockchain network are coupled to one another using suitable communication technologies, which may include wired and wireless communication technologies. Often, a blockchain network is implemented at least in part over the Internet, and some of the nodes may be located in geographically dispersed locations.
[0056] Currently, nodes keep a global ledger of all transactions on the blockchain, which are grouped into blocks, each of which contains a hash of the previous block in the chain. The global ledger is a distributed ledger, and each node may store a full or partial copy of the global ledger. Transactions by nodes that affect the global ledger are verified by other nodes to ensure that the global ledger remains valid. The details of implementing and operating a blockchain network, such as one that uses the Bitcoin protocol, are well understood by those skilled in the art.
[0057] Each transaction typically has one or more inputs and one or more outputs. Script embedded in the inputs and outputs specifies how and by whom the transaction's outputs can be accessed. A transaction's output may be an address to which a value is moved as a result of the transaction. That value is then associated with that output address as an Unspent Transaction Output (UTXO). Subsequent transactions may then reference that address as an input in order to use or distribute that value.
[0058] Nodes can be of various types or categories depending on their functionality. It has been suggested that there are four basic functions associated with a node: wallet, mining, full blockchain maintenance, and network routing. Variations of those functions may exist. A node may have more than one of those functions. For example, a "full node" provides all four functions. For example, a lightweight node, such as may be implemented with a digital wallet, may feature only wallet and network routing functions. Rather than keeping a full blockchain, a digital wallet may keep track of block headers, which serve as an index when querying blocks. Nodes communicate with each other using a connection-oriented protocol such as TCP / IP (Transmission Control Protocol).
[0059] A further type or category of node may be provided, namely, merchant nodes (sometimes referred to herein as "M nodes"). M nodes are designed to focus on fast propagation of transactions. M nodes may or may not hold the complete blockchain and do not perform mining functions. In that sense, M nodes are similar to lightweight nodes or wallets. However, M nodes include additional functionality to enable fast propagation of transactions. The operational focus of M nodes is the fast validation and propagation of unconfirmed transactions to other M nodes, from which unconfirmed transactions are quickly pushed to other nodes in the blockchain network. To facilitate this function, M nodes are allowed more incoming connections, especially outgoing connections, which may otherwise be allowed to nodes under the governance protocol.
[0060] M nodes may be collectively referred to as the Merchant Network (or "M-net"). The word "merchant" may be interpreted to mean "specialized." M nodes may be incorporated into a blockchain network. Each M node is a specialized node on the blockchain network that meets certain hardware and performance features that ensure it can perform the functions of the M node. That is, the M-net may be considered a sub-network distributed within and throughout the blockchain network. M nodes may be arranged and configured to perform one or more specialized functions or services.
[0061] For the M-net to run reliably and be able to provide services at a particular security level, M nodes need to maintain a proper view of the entire M-net, and therefore an efficient routing protocol needs to be in place. Whenever an M node receives an initiation transaction, it needs to broadcast it to several other M nodes in addition to other nodes. In the context of M-net, this means finding a solution to the multiple salesman problem (MTSP). There is a plethora of solutions to address this problem, and any one of them may be used in M-net. Each M node performs some state-of-the-art form of routing optimization.
[0062] In some implementations, M-net is implemented as a decentralized IP multicast type network, i.e., multicast may be used to enable rapid dissemination of incoming transactions to the blockchain network, ensuring that transactions are quickly broadcast throughout the M-net, allowing all M nodes to then forward the transaction to other nodes in the blockchain.
[0063] Multicast network architecture allows simultaneous delivery of data to a group of destination nodes without data duplication for each node interested in receiving the information. When a node wants to receive a multicast transmission, it joins a multicast group (registration phase) and is then able to receive all data sent to the multicast group. IP multicast can accommodate a larger population of receivers by not requiring a priori knowledge of how many receivers exist, and the network infrastructure is used efficiently by requiring the source to send a packet only once. Due to the nature of multicast networks, the use of connection-oriented protocols (such as TCP) is not practical due to the simultaneous communication with a larger number of other nodes. Therefore, connectionless protocols are used.
[0064] Some blockchain networks, such as Bitcoin, use TCP for inter-node communication. Data packets sent with TCP have an associated sequence number that is used for ordering.
[0065] In addition to this, the TCP protocol involves a three-way handshake procedure both when establishing and terminating a connection. Packets sent via TCP have associated overhead, they have associated sequence numbers, and there is a three-way handshake protocol. In establishing a connection, 128-136 bytes are sent, while closing a connection takes 160 bytes. Thus, the handshake for packet transmission can take up to 296 bytes.
[0066] Furthermore, when a node receives a new transaction, it notifies other nodes of an inventory (INV) message containing the hash of the transaction. The node receiving the INV message checks whether the hash of the transaction has been previously confirmed. If not, the node requests the transaction by sending a GETDATA message. The time required to send a transaction from node A to node B is T1 = validity confirmation + TCP(INV + GETDATA + TX). Here, TCP() indicates the overhead introduced by the TCP handshake procedure with respect to time.
[0067] <TCP Protocol and Three-Way Handshake> The current peer-to-peer Bitcoin protocol defines 10 data messages and 13 control messages. In this regard, the transfer of data or data packets related to object transfer may refer to individual transactions or blocks.
[0068] A complete list of the messages used by the Bitcoin P2P protocol is known. For the complete list, refer to the Bitcoin Developer Reference available from http: / / bitcoin.org / en / developer-reference.
[0069] A subgroup of messages is related to the request or distribution of objects. These messages are BLOCK, GETDATA, INV, TX, MEMPOOL, and NOTFOUND.
[0070] The BLOCK message transmits a single serialized block and may be sent for two different reasons: either a node always sends it in response to a GETDATA message requesting a block with inventory type MSG_BLOCK (provided the node has that block available for relaying), or alternatively, nodes or miners may send unsolicited block messages broadcasting their newly mined blocks to their peers.
[0071] The GETDATA message requests one or more data objects from another node. Usually the objects are requested by an inventory previously received using an INV message. The response to a GETDATA message can be a TX message, a BLOCK message, or a NOTFOUND message.
[0072] The GETDATA message cannot be used to request arbitrary data, such as past transactions that are no longer in the memory pool or relay set. The GETDATA message should be used to request objects from nodes that previously advertised them.
[0073] The INV message (inventory message) transmits one or more inventories of objects known to the sending peer. It can be sent unsolicited to announce a new transaction or block, or in response to a GETBLOCKS or MEMPOOL message. The receiving peer can compare the inventory from the INV message with a pre-stored inventory to request objects that have not yet been seen.
[0074] The TX message transmits a single transaction. It is sent in response to a GETDATA message requesting a transaction using the inventory along with the identification (ID) of the requested transaction.
[0075] A MEMPOOL message requests the IDs of transactions that the receiving node has validated but that have not been published in a block, i.e., transactions that exist in the receiving node's memory pool.
[0076] The response to this message is one or more IV messages containing the transaction ID. A node sends as many INV messages as needed to browse its complete memory pool. A full node can use the MEMPOOL message to quickly gather most or all of the unconfirmed transactions available on the network. A node can set a filter before sending a MEMPOOL to only accept transactions that match the filter.
[0077] A NOTFOUND message is a response to a GETDATA message requesting an object that is not available for relaying at the receiving node. For example, a node may remove a spent transaction from an older block, thus making transmission of such a block impossible.
[0078] The messages subgroup is concerned with the reliability and efficiency of objects. These messages are FEEFILTER, PING, PONG, and REJECT.
[0079] The FEEFILTER message is a request to the receiving peer not to relay the transaction in the INV message if the fee rate is below a specified value. The MEMPOOL restriction provides protection against attacks and spam transactions that have a low fee rate and are unlikely to be included in a mined block. The receiving peer may choose to ignore the message and not filter the transaction in the INV message.
[0080] The PING message serves to verify that the receiving peer is still connected. If a TCP / IP error occurs when sending the PING message (e.g., a connection timeout), the sending node can assume that the receiving node has disconnected. The response to a PING message is a PONG message. The message contains a nonce.
[0081] A PONG message responds to a PING message and informs the PINGing node that it is still operational. By default, Bitcoin Core disconnects any client that does not respond to a PING message within 20 minutes. To allow nodes to track latency, a PONG message sends back the same nonce received in the corresponding PING message.
[0082] The REJECT message informs the receiving node that one of its previous messages has been rejected. Examples of reasons for rejecting a message include: - the message was not decrypted; The block is not valid, i.e. an invalid proof of work or an invalid signature was provided; The transaction is not valid, i.e. an output value greater than the input or an invalid signature is provided, The block is using a version that is no longer supported, The connecting node uses a protocol version that the rejecting node does not support, The transaction uses the same inputs as a previously rejected transaction (double spend), The transaction did not have sufficient fee or priority to be relayed or mined There is.
[0083] As an example, nodes 'i' and 'j' on the Bitcoin network communicate using the following steps: 1. Node i sends an INV message containing a list of transactions. 2. Node j responds with a GETDATA message requesting a subset of the transactions it had previously seen. 3. Node i sends the requested transaction.
[0084] Here, the method seeks to optimize the protocol of the blockchain network to at least improve the distribution of data.
[0085] FIG. 6 represents a practical scenario where data, in the form of network packets, is sent and received serially at the application level according to primitives provided by the operating system.
[0086] If transaction x fits into a single Ethernet / IP packet, its transmission to m peers requires the buffering of m distinct outgoing packets. Both the incoming and outgoing network packets contain, among other information, Serialized transactions, · Logical interface ID representing the TCP / IP connection to the sending / receiving peer Includes.
[0087] The expected time for an incoming transaction to be processed is the input queue L i The expected time for a processed transaction to be successfully transmitted depends on the average length (in packets) of the output queue L o It depends on the average length of
[0088] Therefore, the efficient relay of transactions is L i and L o However, the probabilistic model for selective relaying of transactions to peers depends on reducing both the value of L odirectly affects and induces L i also affects.
[0089] In the current Bitcoin implementation, INV and GETDATA message packets are queued in the I / O buffer just like transactions, severely impacting send and receive latency.
[0090] <Efficient Transaction Propagation - Probabilistic Relay> If node i were allowed to send new transactions directly, without the use of inventory exchange, the transactions would spread through the network at a faster rate, but without some regulation the network would become flooded.
[0091] Thus, the present invention uses a mechanism for selective relay of data or objects from a node to a peer node to avoid transmission of huge amounts of unnecessary transactions. Thus, the present invention provides improved network efficiency and reduces the amount of resources required by the network. The mechanism can be a probabilistic model.
[0092] The mechanism or probabilistic model for relay transmission is based on the assumption, see Fig. 7, that three nodes i, j and k are part of the Bitcoin network and are connected to each other. Node i is directly connected to node j and node k. Node j and node k are indirectly connected through node i or through the Bitcoin network.
[0093] Node i is shown with two interfaces a and b that process data in the form of transaction r. A transaction is initiated at node i and goes through various stages as it propagates across the Bitcoin network between nodes. The stages include: A first phase r of processing the transaction for transmission from interface a to a peer node so that it is received by node j. 1 , A second step r where node j processes the transaction for transmission via Bitcoin so that it is received by node k. 2 , and A third phase r in which node k processes the transaction for transmission to peer node i, which receives the transaction on interface b. 3 Includes:
[0094] For the avoidance of doubt, 1 , r 2 , and r 3 are the same transaction, and the subscripts indicate relays.
[0095] It is assumed that: If a node receives from interface (b) the same transaction that it processed for transmission from another interface (a), then the two interfaces share a degree of correlation. · If a transaction originating at node i reaches node j through a given input interface j, then a second transaction originating at node i will reach j through the same interface with high probability.
[0096] Referring again to FIG. 7, as an example of relay correlation, nodes j and k are peers of node i. If i generates a new transaction and relays it to node j, and later the same transaction is received at node i from node k, then j and k share a logical path through the network. Thus, the originating node does not need to relay the same information to both of its peer nodes. For clarity, node i does not need to send transaction r to node k because node k has a high probability of receiving it. Similarly, conversely, node k does not need to send transaction r to node i because node k has a high probability of receiving it. The logic is valid in both directions.
[0097] <Relationship between node interfaces> Figure 8 is an example of a local correlation matrix C that can be determined for node i having five interfaces a, b, c, d, and e connected to the pianode. The matrix represents the degree of correlation for incoming traffic. The formation and application of this matrix are described below. The values shown in Figure 8 are for illustrative purposes only.
[0098] As an example, each node constructs a correlation matrix C by determining a coefficient c that represents the correlation between transactions received from interfaces a and b. Such coefficients are determined for all pairs of interfaces. ab Using the list of transaction IDs received from each interface, let t
[0099] and t a and t b be the number of transactions received from interfaces a and b, respectively, and let t ab be the number of double transactions received from both a and b. The correlation coefficient between interfaces a and b is given by Equation 1:
Equation
[0100] By convention, for the coefficient c ab the lexicographical order of its indices a < b is assumed. That is, a is assigned '0' (a = 0), b is assigned '1' (b = 1), c is assigned '2' (c = 2), and so on, continuing in the same way between interface IDs and numerical values. Considering the correlation coefficients for the set of interfaces {a, b, d, e}, interfaces a and b are more correlated than interfaces d and e when c ab > c de .
[0101] Since the elements on the main diagonal are not significant, the matrix size can be reduced to m(m - 1) elements, as shown in Figure 8.
[0102] <Assigning values to node interfaces> The overall correlation index for interface a is:
number
[0103] Using the correlation matrix in Figure 8 as an example, the correlation index c of interface a a can be expressed as the sum of the correlation coefficients between each interface. c a =c ab +c ac +c ad +c ae =0.2+0.8+0.2+0.2=1.4
[0104] This metric can be used to rank the correlation of individual interfaces and understand the quality of relaying provided by node peers. As an example, if the correlation index of a is significantly higher than the average, then the transactions received from a are highly redundant.
[0105] In contrast, if the correlation index of a is significantly lower than the average, then either (i) the transactions received from a are somewhat unique, or (ii) the peer nodes connected to a are behaving maliciously. Malicious behavior is discussed in more detail below.
[0106] It should be noted that each node builds its own correlation matrix based only on the received transaction flow: to avoid the propagation of malicious information, no correlation information is exchanged between nodes or peers.
[0107] <Indicator> Considering an incoming transaction from interface a, a node min ,m max], and then relay to the number of peers m* in the
[0108] The current value m* is the correlation index of interface a, which contains the m-1 correlation coefficients {c a} depends on the current distribution of {c a}=[c 0a ,c 1a ,···,c am-1 ] That is, {c a} is a list or set of correlation coefficients, the set having the coefficients for interface a. The correlation index is calculated as the sum of the coefficients, i.e., c a =c ab +c ac +c ad +c ae It is.
[0109] The number of interfaces selected for relaying from interface a is m*(a) based on the set {c a Metric θ of each element of i (a) Depends on the calculation:
number
[0110] Metric θ i (a) acts as a switch or selector, indicating whether the interface will relay data or not. For example, an interface will relay data if its metric is '1' and will not relay data if its metric is '0'.
[0111] Bar C a Group a metric θ by defining it as the average of the coefficients in i (a) is the corresponding correlation coefficient c ai Bar C a contributes to m*(a) if it is lower than:
number
[0112] θ i (a) Alternatively, instead of the average, a}. The metric θ i (a) We also used statistical analysis to find the a}, for example, the metric θ i (a) is the corresponding correlation coefficient c ai The group a} contributes to m*(a) if it is at least one standard deviation below the mean of
[0113] As an example, an incoming transaction from interface a in Figure 8 is a}, the packets are relayed to the m*(a) smallest correlation interfaces. a A subset of {c* a} is defined as
[0114] Returning to the example in Figure 8, when {ca}={0.2,0.8,0.2,0.2},
number
[0115] Therefore, the number of interfaces selected for relaying from interface a, m*(a), is '3'. Metric θ i (a) The coefficients in the subset below the mean or indicator that determines the value of c* a}, interfaces b, d, and e are selected for relaying.
[0116] The coefficient 'cutoff' point or level used to determine the metric in the above example was the mean of the coefficients '0.35'. This threshold or indicator is the metric θ i (a) Then m*(a) is determined.
[0117] Overall, therefore, the method can be said to relay data from the interfaces according to indicators, which may be used to determine, for each interface, a metric that determines whether it will relay data.
[0118] Although indicators were used to determine the metrics in the above examples, other factors may affect the metrics for each interface. The metrics for each interface may be calculated according to changes to the correlation matrix.
[0119] <Response to changes> As noted above, the determination of which interface should be used to relay information assumes steady state conditions in which the matrix is not altered.
[0120] Further details of the protocol are now given below with regard to a description of how the protocol can respond to changes, for example when a peer node joins the network and forms a new connection with a node, or when a peer node leaves the network and is no longer connected to a node's interface.
[0121] When responding to changes, data can be relayed based on the metric, which takes into account the indicator and therefore the correlation index.
[0122] <Start> When node i starts up, it initializes m peer connections with other nodes in a blockchain, such as the Bitcoin network. See, for example, the Bitcoin Developer Reference for full details. Initially, node i does not have any information on the correlation between data passing through its interface, and has a limited set of data that represents the data passing through or flowing through it. Thus, relaying of complete transactions will take some time to execute.
[0123] During this period T bootup Between them, the metric θ i (a) contributes to m*(a) by setting each interface to '1' and relaying data from all interfaces. Thus, the metric allows data to be relayed regardless of the indicator over a given period of time.
[0124] Period T bootup The length of is a function f(m) of m, i.e. the more connections there are, the more time is needed to build an accurate correlation matrix. The following function is proposed, purely by way of example: f 1 (m):=m f 2 (m):=m 2
[0125] T bootup After,node i performs selective relaying according to the existing nodes,shown below.
[0126] <Existing node> The connections of a general node j change over time to account for changes either due to (i) other nodes joining the network, (ii) other nodes leaving the network, and / or (iii) other nodes behaving maliciously. Malicious nodes are selectively blacklisted and the corresponding connections are closed.
[0127] Therefore, any change in the entire network graph is detected and a quantity T change It can be parameterized by T change Once every, the local correlation matrix of node j needs to be updated.
[0128] The updating of the correlation matrix of the node may include at least one of the following: · Cycle reset when correlation matrix is reset Each new node relays a complete transaction in time. Tbootup Then, node j performs selective relaying on its new m* values for each interface that connects to a peer node. - Updates that update the correlation matrix For a given interface α, the selected set {c* a}, the β highest correlation interfaces (0<β <m min ) is {c* a} are exchanged with the β minimal correlation interfaces not in
[0129] That is, the data is usually divided into a subset {c* a} to the m*(a) smallest correlation interfaces in the matrix, but when the matrix is updated, the subset {c* a}, interfaces with low correlation indexes are added to the subset instead of interfaces with high correlation indexes in the subset. This can ensure the integrity of the flow of data across the network during changes.
[0130] Returning to the example in Figure 8, typically, c d =c da +c db +c dc +c de=0.2+0.4+0.4+0.6=1.6 {c d}={0.2,0.4,0.4,0.6}
[0131] Bar C d Group d}, the metric θ i (d) is the corresponding correlation coefficient c di Bar C d contributes to m*(d) if lower than:
number
[0132] In that case,
number
[0133] Typically, the number of interfaces selected for relaying from interface d, m*(d), is '3'. i (d) The coefficients in the subset below the mean or indicator that determines the value of c* d}, and interfaces a, b and c are selected for relaying. For the avoidance of doubt, the coefficient corresponding to interface e, i.e., {0.6}, is d Not within}.
[0134] When an update occurs relative to an interface d, the pair {c* d It happens that there are 2(β) highest correlation interfaces in}. They are interfaces b and c, both with a value of 0.4.
[0135] There is only one interface that qualifies as the β-minimum correlated interface that is not in {c*d}, and that interface is e. Therefore, only one 'swap' can be performed, and the choice of which interface to swap is either b or c, since they both have the same coefficient of 0.4. As an example, the choice of which interface is selected for swapping can be made according to a lexicographical priority.
[0136] This interface then becomes a}, which means that the lowest coefficients are replaced by the remaining coefficients in the set {c d}={0.4,0.4,0.6}. The lowest coefficients correspond to interfaces b and c, and therefore data is relayed to peer nodes connected to those interfaces rather than interface a.
[0137] If a peer disconnects, its coefficient in the correlation matrix becomes invalid. change At the end of the period, a periodic reset as above is required. Then, node j performs selective relaying on its new m* values for each connecting interface.
[0138] When a new peer b joins, the interface it is connected to is {c* a}. As an example, suppose an interface with b is selected for relaying for some time T join For every γ incoming transactions, a relay may be randomly selected for relaying.
[0139] T join is T bootup and / or T change This relay helps node b to build its own correlation matrix. T joinAt the end of the process, b receives the updated values {c* a} is selected for relay according to
[0140] Node i may need to check if its peer j is still alive because no incoming traffic was received from its interface. temp A temporary relay request of can be sent to j.
[0141] A node issuing a new transaction, i.e., the originating node, needs to carefully select a set of selected peers for relaying to ensure distribution in the network. For example, a number m** of nodes on any interface a can be selected as the first relays, where m*(a) <m**<mである。
[0142] The complete list of parameters is detailed below in Table 1. [Table 1]
[0143] <Malicious node> Malicious nodes aim to make the probabilistic model for transaction distribution ineffective. Malicious nodes can function or behave in any of the following ways: A malicious node does not propagate transactions that should be propagated. Honest nodes connected to this malicious node may be able to obtain those transactions from other honest (or malicious) nodes. However, full distribution of transactions in the network is not necessary, as the set of miners can still receive them and include them in newly mined blocks. Malicious nodes propagate the same legitimate transaction multiple times. Using a lookup table, the receiving node can keep track of previously received transactions. Thus, fraud is easily detectable and malicious peers are removed. Malicious nodes propagate invalid transactions. If the receiving node performs transaction validation, the fraud is easily detectable and the malicious peer is removed. Malicious nodes generate and propagate a huge number of dummy transactions. Receiving nodes can respond differently to a huge number of valid incoming transactions from a peer. In response, (i) the receiver asks the sender to reduce the transaction relay rate. If the problem persists, the sending peer is removed; (ii) the receiver can demand (and check) a minimum transaction fee for relayed transactions. This makes attacks expensive for malicious nodes. If the minimum transaction fee is not respected, the sending peer is removed; and / or (iii) the sending peer is simply removed.
[0144] Based on the above attacks and probabilistic models, the following properties can be inferred. Transaction relay rates depend on both the processing power and bandwidth availability of individual nodes. For this reason, no default maximum rate is enforced. Malicious nodes cannot simply stay silent. They must forward valid transactions to maintain their connection.
[0145] <Message> New message types may be introduced to support implementation of the methods herein, including the probabilistic relay of data or objects from node interfaces on the Bitcoin network. Additionally, some of the current message types detailed above with respect to the current peer-to-peer Bitcoin protocol are no longer used, while message types not described below are unchanged for the Bitcoin P2P protocol.
[0146] <Data message> The message GETDATA, which requests one or more data objects from another node, is no longer used or has been deprecated. Similarly, the message INV, which sends one or more inventories of objects known to the sending peer, is no longer used or has been deprecated. The message NOTFOUND as a response to GETDATA is no longer used or has been deprecated.
[0147] The MEMPOOL message, already introduced above, requests IDs of transactions that the receiving node has verified as valid but that have not been published in a block, i.e. transactions that are present in the receiving node's memory pool. The response to this message is one or more IV messages containing the transaction IDs. The node sends as many INV messages as are needed to reference its complete memory pool. A node may disable this feature and thus not be required to send its local MEMPOOL.
[0148] <control message> PING and PONG messages to verify that peers are still connected are removed. Silent nodes, such as malicious nodes, are managed as above. FEEFILTER messages remain unchanged. However, peers that do not respect the fee threshold are removed.
[0149] New messages TEMP and FLOW are introduced. TEMP messages are used when a node is asked to check if a peer is still alive because no incoming traffic is received from its interface (for an arbitrary short time T temp A temporary relay request TEMP is sent to the peers connected to the interface. The receiving peer starts relaying, but within a given time window T emp If it does not comply with the rules, it will be removed from the queue. The FLOW message is used when a node requests a peer to change its transaction relay rate (number of transactions per second) according to local flow control. The rate may be increased, decreased, or temporarily suspended (relay rate = 0). If the receiving peer does not respect the new relay rate, it is deleted. A new FLOW request is needed to change the current relay rate. If the receiving peer does not receive a second FLOW request to resume relaying, it deletes the sender. The MINFEE message is used when a node is asked to accept transactions filtered by a minimum transaction fee. Minimum values can be set for individual transactions and / or for the sum of fees if multiple transactions fit into a single IP packet. If a receiving peer does not respect the transaction fee restrictions, it will be dropped.
[0150] It should be noted that the above embodiments illustrate rather than limit the invention, and that those skilled in the art can design many alternative embodiments without departing from the scope of the invention as defined by the appended claims. In the claims, any reference signs in parentheses shall not be construed as limiting the scope of the claims. The words "comprising" and "comprises" and the like do not exclude the presence of elements or steps other than those listed in any claim or the specification as a whole. In this specification, "comprising" means "includes or consists of" and "comprising" means "including or consisting of". A single reference of an element does not exclude a plurality of references of such elements and vice versa. The invention may be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In a device claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The mere fact that certain measures are recited in mutually different claims does not indicate that a combination of these measures cannot be used to advantage.
[0151] <Summary> Today's Bitcoin network is centered around mining in terms of computational work. With a significant increase in transaction volume, this is not always feasible. The solution described herein allows the Bitcoin network to handle the propagation of large volumes of transactions.
[0152] The present invention can provide a method of distributing data packets in addition to, or preferably as an alternative to, known methods of sending data packets or transactions across a blockchain network such as the Bitcoin network, for example distributing new transactions using three-step messaging.
[0153] The present invention can support faster propagation and distribution of data packets across a network by reducing communication between nodes.
[0154] Furthermore, long queues caused by bottlenecks at nodes are suppressed by selectively relaying data packets according to correlations between interfaces.
[0155] The overall number of packet transmissions is reduced while maintaining a safe level of information redundancy.
[0156] Finally, the invention can provide an adaptation means to accommodate changes related to the nodes' interfaces so that new connections, interruptions in connections, and malicious nodes are taken into account, thereby maintaining the integrity of the network. This means that there is no limit to the number of interfaces on a node that can be managed, and the performance and size of the network is not limited, as the method adapts to the network and node environment. Inefficient transmissions are minimized and malicious nodes are avoided.
Claims
1. 1. A computer-implemented method for a node of a blockchain network, the node having a number of interfaces connecting to peer nodes, the method comprising: determining a correlation matrix having correlation coefficients representative of correlations between data processed at each interface of said node; receiving data corresponding to a transaction at a receive interface of said node and determining a correlation index to determine whether said transaction is unique; selecting at least one other interface from a plurality of other interfaces of the node and relaying the received data from the at least one other interface, the other interface being selected according to the set of correlation coefficients of the receiving interfaces; The method according to claim 1,
2. an indicator is derived from the correlation matrix; if a correlation between the receiving interface and the at least one other interface is lower than the indicator, the data is relayed. The method of claim 1.
3. the indicator is used to determine a metric that sets criteria for selecting which of the plurality of other interfaces is selected for relaying data. The method of claim 2.
4. The data is present in network packets representing serialized transactions and identification representing connections to adjacent or peer nodes. The method of claim 1.
5. The node: (i) a data identifier for each packet of data processed through each interface; and (ii) the same transaction processed through a pair of interfaces; constructing said correlation matrix by monitoring the correlation coefficients between any two interfaces; 5. The method according to any one of claims 1 to 4.
6. The correlation matrix having m(m-1) elements is: [0010] is used to determine the correlation index of interface a as follows: m is the number of interfaces connected to the peer node; 6. The method according to any one of claims 1 to 5.
7. The correlation matrix having m(m-1) elements is: [0025] are used to determine a set of correlation coefficients for interface a as follows:
7. The method according to any one of claims 1 to 6.
8. The indicator is determining, for each interface connected to a peer node, a set of correlation coefficients derived from said correlation matrix, the set comprising correlation coefficients between each interface; deriving a mean or a median from said set; Determined by, The method according to claim 2 or 3.
9. the number of interfaces selected for relaying from the interface depends on a metric derived from the set of correlation coefficients for the interfaces; [0030] is determined from a is the interface, m is the number of interfaces connected to peer nodes, m*(a) is the number of nodes selected for the node's relay interface, and θ is the number of nodes in the set {c a The average correlation coefficient of the set of interfaces in a ), where: [0045] That is, 9. The method according to any one of claims 1 to 8.
10. Relaying data is (i) a reset time, which is the time since node startup or power-up; (ii) a change time, which is the time between change events including at least one of a new peer node connecting to an interface, a termination of a connection to an interface, and an interface connecting to a node being classified or determined to be malicious; Further based on at least one of 10. The method according to any one of claims 1 to 9.
11. Upon node start-up, the node connects with peer nodes and relays data over all interfaces during the reset time during which the correlation matrix is constructed, and after the time period has elapsed, the node relays all data if the correlation between the receiving interface and the other interfaces is lower than an indicator derived from the correlation matrix. The method of claim 10.
12. Upon detecting a change event, the correlation matrix is reset and re-determined.
12. The method according to claim 10 or 11.
13. Upon detecting a change event, the node relays all objects from the node through interfaces if the correlation between the receiving interface and the other interfaces exceeds an indicator derived from the correlation matrix.
12. The method according to claim 10 or 11.
14. Upon detecting a disconnection of a peer node from an interface, the correlation matrix is reset and re-determined.
12. The method according to claim 10 or 11.
15. Upon detecting a connection between a new peer node and an interface, the interface: (i) the duration of the reset time; and / or (ii) the time of the change; relaying data through all interfaces between 12. The method according to claim 10 or 11.
16. A computer readable storage medium having computer executable instructions which, when executed, configure a processor to perform a method according to any one of claims 1 to 15.
17. an interface device, one or more processors coupled to the interface device, and a memory coupled to the one or more processors; The memory stores computer executable instructions that, when executed, configure the one or more processors to perform a method according to any one of claims 1 to 15. Electronic devices.
18. A node of a blockchain network, The node is configured to perform a method according to any one of claims 1 to 15.
19. A blockchain network comprising the nodes according to claim 18.
Citation Information
Patent Citations
Computerized system including participation locality-aware overlay module, and method implemented by computer
JP2005094773A
Measurement-based construction of locality-aware overlay networks
US20050060406A1