Method for securing generation of a blockchain
Patent Information
- Application Number
- PCT/EP2026/058625
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058625_01102026_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR SECURING A BLOCKCHAIN STRUCTURE
[0002] TECHNICAL FIELD
[0003] The present invention relates to securing, through machine learning, the development of a blockchain by a network of node devices, called "peers".
[0004] STATE OF PRIOR ART
[0005] Blockchain is a distributed ledger technology that allows data to be stored securely, transparently, and immutably, using cryptographically linked blocks of data. Each block of data contains a timestamp and a link to the previous block, thus forming a continuous chain that is easily detectable for any alteration.
[0006] For the development of a blockchain, a set of node devices, called "peers", executes a consensus process, which is a data sharing mechanism that ensures security, integrity and decentralization of data transfers and allows peers to agree on the content of the blockchain through message exchanges via a communication network.
[0007] The consensus process is a distributed communication process that allows, for example, secure data sharing, ensuring data integrity in a decentralized manner, and thus enabling the content of the blockchain to be immutable and identical between peers.
[0008] As an example, a well-known consensus process is Proof of Authority (PoA). In this case, block validation is performed by trusted peers. This type of consensus process is widely used in private networks due to its low energy requirements and the speed of transactions it enables, compared to other consensus processes (Proof of Work (PoW), Proof of Stake, etc.). The blockchain aims to store data shared between peers and verified by peers in blocks organized in an ordered manner. All peers must have their data stored in the same order. The purpose of the consensus process is therefore for all peers to reach a common agreement on the data in the blockchain (the data is organized into blocks) and on its order within the blockchain.The number of blocks available to an even-numbered device at a given time t is called the "block chain height".
[0009] All peers therefore have data stored in an identical and ordered manner. The goal is for the set of peers using consensus to ensure that each peer maintains a ledger containing reliable data whose integrity can be guaranteed and whose ordering is the same. The number of these ordered and validated blocks held by each peer at time t is its "block chain height" at time t.
[0010] To build the blockchain, peers establish peer-to-peer (P2P) connections. For example, Fig. 1 schematically illustrates a communication system 100 in which four peers, P111, P2112, P3113, and P4114, interact in a network to build the blockchain. In Fig. 1, each peer has established a P2P connection with every other peer in the communication system 100.
[0011] Each peer Pl 111, P2 112, P3 113, P4 114 has an ordered list of connections with the other peers in the communication system 100. For example:
[0012] - the ordered list of peer Pl 111 lists the other peers of the communication system 100 in the following order: P2 112, P3 113, P4 114;
[0013] - the ordered list of peer P2 112 lists the other peers of the communication system 100 in the following order: P3 113, P4 114, Pl 111;
[0014] - the ordered list of peer P3 113 lists the other peers of the communication system 100 in the following order: P4 114, P111, P2 112; and
[0015] - the ordered list of peer P4 114 lists the other peers of the communication system 100 in the following order: Pl 111, P2 112, P3 113.
[0016] For each peer, the ordered list to be used is established by default. As in the illustrative example above, the ordered list is generally offset from one peer to the next, thus distributing requests among the peers at the start of blockchain creation. Subsequently, asynchronies between actions performed by peers reduce, or even eliminate, this peer-to-peer distribution effect.
[0017] In more detail, the ordered list defines the order in which a given peer requests data from other peers in the communication system during the consensus process; that is, a sequence of requests to other peers during the execution of the consensus process. The ordered list is used circularly. Thus, considering a peer that is missing data blocks, this peer will attempt to retrieve one or more missing data blocks from the first peer in its ordered list and then proceed to the second peer in its ordered list if the first peer does not have all the missing data blocks, and so on.
[0018] When peer A requests a node from peer B as part of the consensus process, peer A sends peer B a notification message (e.g., a "hello" message) containing, among other things, a P2P connection identifier, the signature of the genesis block of the blockchain, the hash of the data, the amount of blockchain space available to peer A, and a list of peers known to peer A. The signature is used, for example, to guarantee both integrity and authenticity. The hash is used, for example, to guarantee data integrity. The amount of blockchain space available to peer A identifies the blocks available to peer A. Upon receiving this message, peer B performs standard checks, particularly authentication checks of peer A.Note that the signature is not a hash but a sign of a hash (here it is the signature of the hash of the genesis block), the hash being a function in which data is entered and which returns as output data which always has the same length.
[0019] Peer B sends an acknowledgment containing, among other things, the P2P connection identifier, a hash of the blockchain's genesis block, and the chain height available to Peer B. Assuming that Peer B's available chain height is greater than Peer A's, Peer A can then send Peer B a request to receive hashes of blocks that are missing from Peer A's chain (e.g., via a "getblockhashes" message). Upon receiving this message, Peer B performs standard checks, particularly cryptographic checks, and if the request is valid, Peer B returns a message providing the requested hashes for the missing block(s) (e.g., via a "blockhashes" message). Upon receiving this message, Peer A also performs standard checks, particularly cryptographic checks, on the received hashes.Peer A, provided the received message is valid, can then request the retrieval of missing data, block by block (e.g., via a "getblock" message). Upon receiving a message requesting one or more missing blocks, peer B performs standard checks, particularly cryptographic checks, and, provided the request is valid, provides protocol information enabling the retrieval of the block in question, such as the encoding type applied to the data in the block and a link to a server (e.g., a web server) where the block can be retrieved. For example, connections to servers can be made using the SSH ("Secure Shell") protocol.In this case, peer A, who wants to send their data, creates a temporary server in which they store the encrypted blocks and sends all the data to connect to this server to peer B and to decrypt the data thus stored.
[0020] As schematically illustrated in Fig. 1, after execution of the CONS 120 proof-of-authority consensus process, peers Pl 111, P2 112, P3 113, P4 114 converged to the same block chain height and validated blocks B0 130 0 (genesis block), B1 130 1, B2 130-2,.., Bn 130_n.
[0021] Consensus processes can be subject to cyberattacks by malicious node devices that aim to destabilize the development of the blockchain, or even to harm the integrity of its data.
[0022] It is therefore desirable to provide a solution that increases the security of blockchain development.
[0023] DESCRIPTION OF THE INVENTION
[0024] To this end, a method is proposed for securing the creation of a blockchain using node devices, called peers, in a communication network. The method comprises the following steps, from a preparatory phase prior to the creation of the blockchain:
[0025] - to set up the peer network to carry out simulations, with and without cyberattacks, in the context of developing test blockchains;
[0026] - collect, independently for each peer, information relating to events occurring within the framework of exchanges in the communication network taking place during the simulations, the information collected being associated with a label indicating whether the information collected is obtained during a cyberattack against the peer in question, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle;
[0027] - to train, independently for each pair, at least one model, using the batches as input data in said training and the associated labels to determine a classification output, said at least one model for each pair being in the form of at least one model trained to be capable of detecting an ongoing cyberattack and / or at least one model trained to be capable of detecting a cyberattack in preparation
[0028] where, over N simulation execution cycles, N-Ki+1 batches B(z), with z = 0,..,N-Ki, are created for training against an ongoing cyberattack, the information collected for the Ki execution cycles being used as input to said model, and the label indicating whether a cyberattack is ongoing within the information collected in at least one of the Ki execution cycles is used as output and
[0029] where over N execution cycles by simulation N-K2-K2 '+1 batches B( / ), with j = 0,..,NK.2-K.2 are created for training relative to a cyberattack in preparation, the information collected for the K2 execution cycles being used as input to said model and the label indicating whether a cyberattack is in progress within the information collected in at least one of the following K2 ' cycles is used as output;
[0030] the process further comprising the following steps of a use phase subsequent to the preparatory phase:
[0031] - set up the peer network for the development of the blockchain, by activating at least one trained model on each peer;
[0032] - collect, independently for each peer, information relating to events occurring within the framework of exchanges in the communication network during the usage phase, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle;
[0033] - inject the input batches of said model associated with each pair and obtain at output a classification relating to a cyberattack detection. According to a particularity of the invention, each pair has at least two models in the form of at least one model trained to be able to detect a cyberattack in progress and at least one model trained to be able to detect a cyberattack in preparation.
[0034] According to another feature of the invention, each model has been trained to detect one type of attack from a predefined set of attack types, each attack being associated with at least one model trained to be able to detect an ongoing cyberattack and / or at least one model trained to be able to detect a cyberattack in preparation.
[0035] According to another feature of the invention, at least one of said models trained to be able to detect an ongoing cyberattack is a deep learning model of the GCN type.
[0036] According to another feature of the invention, at least one of said models trained to be able to detect a cyberattack in preparation is a deep learning model of the T-GCN type.
[0037] According to another feature of the invention, each model has a loss function of the binary cross-entropy type.
[0038] Another object of the invention relates to a communication system comprising a network of node devices, called peers, the peers being configured to develop a blockchain, in a communication network, the system comprising electronic circuitry configured to perform the following steps, from a preparatory phase in which the peer network is set up to carry out simulations, with and without cyberattacks, in the context of developing test blockchains:
[0039] - collect, independently for each peer, information relating to events occurring within the framework of exchanges in the communication network taking place during the simulations, the information collected being associated with a label indicating whether the information collected is obtained during a cyberattack against the peer in question, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle;- to train, independently for each pair, at least one model, using the batches as input data in said training and the associated labels to determine a classification output data, said at least one model for each pair being in the form of at least one model trained to be able to detect an ongoing cyberattack and / or at least one model trained to be able to detect a cyberattack in preparation;
[0040] where, over N simulation execution cycles, N-Ki+1 batches B(z), with z = 0,..,N-Ki, are created for training against an ongoing cyberattack, the information collected for the Ki execution cycles being used as input to said model, and the label indicating whether a cyberattack is ongoing within the information collected in at least one of the Ki execution cycles is used as output and
[0041] where over N simulation execution cycles N-K2-K2'+l batches B( / ), with j = 0,..,NK.2-K2 are created for training relative to a cyberattack in preparation, the information collected for the K2 execution cycles being used as input to said model and the label indicating whether a cyberattack is in progress within the information collected in at least one of the following K2' cycles is used as output;
[0042] the electronic circuitry being further configured to perform the following steps of a usage phase subsequent to the preparatory phase where the peer network is set up for the development of the blockchain, by activating on each said peer at least one trained model:
[0043] - collect, independently for each peer, information relating to events occurring within the framework of exchanges in the communication network during the usage phase, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle;
[0044] - inject the batches into the input of said model associated with each pair and obtain as output a classification relating to a cyberattack detection.
[0045] According to a particular feature of the invention, each peer has at least two models, consisting of at least one model trained to detect an ongoing cyberattack and at least one model trained to detect a cyberattack in preparation. Advantageously, a cyberattack can be detected efficiently, taking into account the history and evolution of observed information, for an ongoing attack against the peer in question or one in preparation, a countermeasure action can then be applied and thus increase the security of the blockchain development.
[0046] BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The features of the invention mentioned above, as well as others, will become clearer upon reading the following description of at least one exemplary embodiment, said description being made in relation to the accompanying drawings, among which:
[0048] [Fig. 1] schematically illustrates a communication system adapted and configured for the development of a blockchain, according to the state of the art;
[0049] [Fig. 2] schematically illustrates a communication system adapted and configured for the development of a blockchain, according to the present invention;
[0050] [Fig. 3 A] schematically illustrates the communication system, in a preparatory phase;
[0051] [Fig. 3B] schematically illustrates an algorithm for training at least one model by machine learning, within the framework of the preparatory phase;
[0052] [Fig. 4] schematically illustrates an algorithm for using said at least one model after training; and
[0053] [Fig. 5] schematically illustrates an example of a hardware platform suitable for use in the communication system.
[0054] DETAILED DESCRIPTION OF IMPLEMENTATION METHODS
[0055] Figure 2 schematically illustrates the communication system 100 adapted and configured for the creation of a blockchain, according to the present invention. Compared to Figure 1, each pair P111, P2112, P3113, P4114 is equipped here with at least one MOD model, developed, for example, using machine learning. Each MOD model aims to secure the creation of the blockchain by being able to detect a type of cyberattack for which the MOD model in question has been trained.
[0056] Each pair Pl 111, P2 112, P3 113, P4 114, for example, is equipped with several models (n+1 models, n > 0) MOD 200_0,..,200_n. Each MOD 200_0,..,200_n model is dedicated to a specific type of cyberattack from a predefined set of cyberattack types simulated during a preparatory phase. The set of MOD 200_0,..,200_n models with which each pair Pl 111, P2 112, P3 113, P4 114 is equipped allows, in particular, the detection of a range of cyberattack types, thus significantly improving the security of blockchain development.
[0057] For example, for at least one type of cyberattack among said predefined set of cyberattack types, and preferably for each type of cyberattack among said predefined set of cyberattack types, each pair Pl 111, P2 112, P3 113, P4 114 is equipped with two MOD models, a first MOD model having been trained to detect a cyberattack in progress and a second MOD model having been trained to detect a cyberattack in preparation (prediction).
[0058] A cyberattack in preparation is a cyberattack that is being set up against one or more other peers (for whom the said cyberattack is therefore in progress) and which presents a risk of propagation to other peers of the communication system 100, as for example in the case of an Eclipse type cyberattack.
[0059] To build the blockchain, various types of consensus processes can be used, such as: Proof of Authority (PoA), Proof of Delegated Authority (PoDA), Proof of Elapsed Time (PoET), and Proof of Stake (PoS). Other hybrid or custom consensus processes based on one or more of the above-listed processes can also be used. The Practical Byzantine Fault Tolerance (PBFT) algorithm and its variants can also be used in the context of this disclosure. To be able to detect a cyberattack, each MOD 200_0,..,200_n model is pre-trained during the preparation phase, as detailed below in relation to Figures 3A and 3B.
[0060] Each pair Pl 111, P2 112, P3 113, P4 114 is equipped with a MON 300 data monitoring device or function configured to collect information to enable the training of each MOD 200_0,..,200_n model and its subsequent use once trained. As schematically illustrated in Fig. 3A, when training has not yet been performed, the MOD 200_0,..,200n models are not present on any of the pairs Pl 111, P2 112, P3 113, P4 114. As detailed below, training requires prior collection COLL 340 of a set of information obtained by each MON 300 data monitoring device or function.
[0061] These collected data are then processed to allow training of each MOD 200_0,..,200_n model independently for each pair Pl 111, P2 112, P3 113, P4 114. As schematically illustrated in Fig. 3A, each pair Pl 111, P2 112, P3 113, P4 114 transmits its collected data to a respective processing platform PP1 351, PP2352, PP3 353, PP4354. Each processing platform PP1 351, PP2 352, PP3 353, PP4 354 is then configured to train each MOD model 200_0,..,200_n specifically for a given pair Pl 111, P2 112, P3 113, P4 114 with data collected from that pair. Alternatively, each pair Pl 111, P2 112, P3 113, P4 114 has adequate processing and memory resources to process the collected data and perform its own training of each MOD model 200_0,..,200_n, without relying on an external processing platform.In another embodiment, all peers Pl 111, P2 112, P3 113, P4 114 transmit their collected data to the same PP processing platform. This PP processing platform then processes the collected data from peers Pl 111, P2 112, P3 113, P4 114 separately and trains each MOD model 200_0,..,200_n independently for each peer Pl 111, P2 112, P3 113, P4 114. This embodiment simply allows for the sharing of processing and memory resources.
[0062] Then, in step 301, the communication system 100 is put in place for the purpose of further training each MOD model 200_0,..,200_n.
[0063] In step 302, simulations of the development of test blockchains under real-world conditions are carried out. In other words, the blocks used contain test data, potentially without any particular meaning (but of a size consistent with the data (blocks) that are supposed to be shared subsequently and trying to be very varied), but the CONS 120 consensus process is carried out by peers Pl 111, P2 112, P3 113, P4 114 on these test data blocks under real-world conditions of use of the communication system 100. The simulations include simulations without cyberattack, as well as simulations with cyberattacks, for each of the targeted cyberattack types (ze, for each of the cyberattack types for which at least one so-called MOD model 200_0,..,200_n must be created).Simulations can be triggered with test blockchains already started, with different blockchain heights for the Pl 111, P2 112, P3 113, P4 114 peers. One or more blocks can be added during each simulation.
[0064] The simulations preferentially include phases of operation of the communication system 100 where consensus is reached, phases of operation of the communication system 100 where consensus leads to failure, phases of hardware failures and operational recovery following these hardware failures, situations of adding one or more peers and situations of removing one or more peers.
[0065] Simulations involving cyberattacks include, for example, the implementation of cyberattacks according to one or more of the following types of cyberattacks:
[0066] Botnet-type cyberattacks:
[0067] One or more entities (ze, node devices, potentially on the same machines as the peers) are deployed in the communication network without being part of the set of peers intended to participate in the blockchain creation. These entities replicate a botnet-type cyberattack. In other words, these entities send a multitude of messages to peers P111, P2112, P3113, and P4114 in a time-coordinated manner.
[0068] Eclipse-type cyberattacks:
[0069] One or more entities (ze, node devices, potentially on the same machines as the peers) are deployed in the communication network without being part of the set of peers intended to participate in building the blockchain. These entities replicate an Eclipse-type cyberattack. In other words, these entities block a subset of peers P111, P2112, P3113, and P4114 (not all peers) on the communication network. For example, a multitude of SYN packets are sent to block the connections of one or more specific peers, partially isolating them from the rest of the communication network. By "spamming" one or more peers to the point of partially blocking their exchanges on the communication network, the affected peer(s) exhibit a difference in blockchain height.Denial of Service (DoS) and / or Distributed DoS (DDOS) type cyberattacks:
[0070] One or more entities (ze, node devices, potentially on the same machines as the peers) are deployed in the communication network without being part of the set of peers intended to participate in the blockchain creation. These entities replicate a DoS or DDoS cyberattack. For example, the entities open a multitude of sockets in the communication network and perform large, random connection flows to peers P111, P2, P3, P4, and P14. Timejacking cyberattacks:
[0071] A subset of peers Pl 111, P2 112, P3 113, and P4 114 (not all peers) replicates a time-hijacking cyberattack. The peer(s) in question force a change in the clock reference used for timestamping operations during the consensus process, creating a fork in the blockchain.
[0072] Man-in-the-middle cyberattacks:
[0073] One or more entities (ze, node devices, potentially on the same machines as the peers) are deployed in the communication network without being part of the set of peers intended to participate in the blockchain creation. These entities replicate a man-in-the-middle cyberattack. In other words, these entities will perform a proxy function that alters certain packets transmitted between peers P111, P2112, P3113, P4114, for example by altering the hashes.
[0074] Cyberattacks by flooding messages ("Spam" in English):
[0075] One or more entities (ze, node devices, potentially on the same machines as the peers) are deployed in the communication network without being part of the set of peers intended to participate in the blockchain creation. These entities replicate a message flooding cyberattack. In other words, these entities flood the communication network with messages, for example, using UDP (User Datagram Protocol) packets by successively flooding different UDP ports.
[0076] In step 303, information is collected during simulations by the MON 300 data monitoring device or function of each peer Pl 111, P2 112, P3 113, P4 114. Each peer Pl 111, P2 112, P3 113, P4 114 thus collects information during simulations, including during simulations with a cyberattack, even when the peer Pl 111, P2 112, P3 113, P4 114 is not directly attacked. A label is attached to the collected information to indicate whether the information in question is being collected while a simulated cyberattack is in progress against the peer Pl 111, P2 112, P3 113, P4 114 in question.Thus, if, for example, during a cyberattack simulation at time T, peer Pl 111 experiences a cyberattack but peer P2 112 does not, then the label applied to the data collected by peer Pl 111 during time T indicates an ongoing cyberattack, and the label applied to the data collected by peer P2 112 during time T does not indicate an ongoing cyberattack. And if, at a later time T', the cyberattack has spread to peer P2 112, then the label applied to the data collected by peer P2 112 during time T' indicates an ongoing cyberattack. In a cyberattack simulation, the label preferentially indicates the type of cyberattack. Since the collection of this information is distributed, peers Pl 111, P2 112, P3 113, and P4 114 are therefore informed in a coordinated manner about which label to apply.
[0077] Thus, during these simulations, information relating to events occurring within the framework of exchanges in the communication network taking place during the simulations, including during the development of the test blockchains, is collected.
[0078] During cyberattack simulations, for example of the Botnet type, each time a predefined event occurs (e.g., receiving a message, triggering a cryptographic verification...), the peer in question collects the following information: - address (typically, IP address) of the device that caused the event (e.g., source of a received message);
[0079] - timestamp of the event;
[0080] - connection frequency of the device that caused the event to the peer in question; - size of the packets exchanged (received and / or transmitted) with the device in question;
[0081] - signatures; and
[0082] - blockchain height. During cyberattack simulations, for example of the Eclipse type, each time a predefined event occurs (e.g., receiving a message, triggering a cryptographic verification...), the peer in question collects the following information: - address (typically, IP address) of the device that initiated the event (e.g., source of a received message);
[0083] - timestamp of the event;
[0084] - connection frequency of the device that initiated the event to the peer in question; - packet headers; and
[0085] - chain height of blocks.
[0086] During cyberattack simulations, for example of the DOS and / or DDoS type, each time a predefined event occurs (e.g., receiving a message, triggering a cryptographic verification, etc.), the peer in question collects the following information:
[0087] - address (typically, IP address) of the device that caused the event (e.g., source of a received message);
[0088] - timestamp of the event;
[0089] - exchange latency;
[0090] - connection frequency of the device that initiated the event to the peer in question; - connection time;
[0091] - size of packets exchanged (received and / or transmitted) with the device in question; and - height of blockchain.
[0092] During cyberattack simulations, for example by time hijacking, each time a predefined event occurs (e.g., receiving a message, triggering a cryptographic verification, etc.), the peer in question collects the following information:
[0093] - address (typically, IP address) of the device that caused the event (e.g., source of a received message);
[0094] - timestamp of the event;
[0095] - difference in timestamp compared to the content of messages received from the device that caused the event, if applicable;
[0096] - exchange latency; and
[0097] - blockchain height. During cyberattack simulations, for example man-in-the-middle attacks, each time a predefined event occurs (e.g., receiving a message, triggering a cryptographic verification...), the peer in question collects the following information: - address (typically, IP address) of the device that caused the event (e.g., source of a received message);
[0098] - timestamp of the event;
[0099] - exchange latency;
[0100] - packet headers; and
[0101] - chain height of blocks.
[0102] During cyberattack simulations, for example by message flooding, each time a predefined event occurs (e.g., receiving a message, triggering a cryptographic verification, etc.), the peer in question collects the following information:
[0103] - address (typically, IP address) of the device that caused the event (e.g., source of a received message);
[0104] - timestamp of the event;
[0105] - connection frequency of the device that initiated the event to the peer in question; - quantity of packets exchanged (received and / or transmitted) with the device in question; and - block chain height.
[0106] Additional information derived from the collected information, such as statistics (mean, variance), can be added during or after the execution of all or part of the set of simulations, by each MON 300 data monitoring device or function, and integrated into the collected data.
[0107] Data collection is organized in cycles. Subsequently, and similarly, detections by each MOD 200_0,..,200_n model are also organized in cycles. These are called execution cycles (generally referred to as "timesteps" in Anglo-Saxon terminology). Execution cycles define temporal checkpoints to pace the simulations and data collection, and subsequently, during the usage phase, the operation of each MOD 200_0,..,200_n model.
[0108] Specifically, events occurring during the same execution cycle are buffered together, allowing them to be manipulated (e.g., cleaned up, preprocessed) and reorganized by cycle. This facilitates the training of each MOD 200_0,..,200_n model, and subsequently, the detection of cyberattacks.
[0109] At the end of an execution cycle, the number of sessions per collection cycle can, for example, be added to the collected information. A session is a communication protocol concept that allows computer systems to communicate coherently and interactively for the duration of a given interaction. A session is a temporary interaction between two or more endpoints in a communication network. It begins with the establishment of a connection and ends with its closure. Sessions allow for the interactive exchange of data between the devices or services involved. Sessions are generally stateful.
[0110] Each simulation ends, for example, when an operator determines they have collected enough information and sends an instruction to peers Pl 111, P2 112, P3 113, P4 114 to terminate it. Alternatively, each simulation ends when a predetermined number of execution cycles is reached. In another variant, each simulation ends when a predetermined number of execution cycles without data collection (no relevant event occurring) is reached.
[0111] In an optional step 304, the collected data are transferred to a processing platform to have more resources available for processing and training each MOD 200_0,..,200_n model (independently for each of the peers Pl 111, P2 112, P3 113, P4 114), as already explained above in relation to Fig. 3A.
[0112] In an optional step 305, the collected data is cleaned and / or the data is preprocessed to improve the quality of the data presented to each MOD 200_0,..,200_n model for training.
[0113] For example, cleaning can involve removing non-significant data (e.g., beyond an acceptable margin relative to a standard deviation) or incorrectly recorded data (e.g., typing error, for example) from the collected data.
[0114] For example, preprocessing might involve normalizing or standardizing values, or performing numerical encoding of categorical information (e.g., one-hot encoding). Another example is including padding data to ensure dimensionality consistency in the collected and formatted descriptive data. For instance, at each point in time when information is collected on any peer, a dummy descriptor is added for each of the other peers. This ensures that the event history for any peer has the same number of elements as for any other peer, making event histories easier to manipulate and inject into the MOD 200 model.
[0115] Within each peer, the collected data is for example reorganized by execution cycle, and within each execution cycle, the collected data is grouped, chronologically, by peer with respect to which at least one event occurred during said execution cycle.
[0116] An example would be the following, for data collected by a peer identified by ID3 that has interacted with other peers identified by ID1, ID2, ID4 during execution cycles "Timestep 1", "Timestep 2". Note that each event corresponding to a data point or set of data (identified here by dataX where X is a number) is associated with a timestamp (identified here by h dataX) representing the instant at which the event in question occurred.
[0117] -Timestep 1:
[0118] ID1:
[0119] -datai, h datal
[0120] -data2, h_data2
[0121] -data3, h_data3
[0122] ID2:
[0123] -datai, h datal
[0124] -data2, h_data2
[0125] -data3, h_data3
[0126] -data4, h_data4
[0127] -Timestep 2:
[0128] ID1:
[0129] -datai, h datal
[0130] -data2, h_data2
[0131] ID4: -datai, h datal
[0132] Note, as already indicated above, that there may be execution cycles where no relevant event has occurred, and therefore during which no data has been collected.
[0133] Thus, each peer obtains a chronological description over N execution cycles (“N timesteps”) of data sets representing the actions of all other peers in the communication system 100.
[0134] In step 306, the collected data is distributed among different models. This step 306 is performed when several MOD models (200_0, ..., 200_n) need to be trained for each pair, meaning a model dedicated to each type of cyberattack that was simulated in step 302.
[0135] The information collected during simulations without cyberattacks is shared by all MOD 200_0,..,200_n models of the same peer, while the information collected during simulations with cyberattacks is reserved only for the MOD 200_0,..,200_n model dedicated to the specific type of cyberattack for that peer. In step 307, the input and output data are prepared for each MOD 200_0,..,200_n model.
[0136] The purpose of training each MOD 200 model is to detect a cyberattack against the peer Pl 111, P2 112, P3 113, or P4 114 for which that MOD 200 model is intended. The MOD 200 model can, for example, be specifically designed to detect an ongoing cyberattack or to detect a cyberattack in preparation (i.e., a cyberattack underway on another peer that will likely spread to the peer in question).
[0137] The training is performed in batches of several successive execution cycles over the N execution cycles covered by the data collection. Each batch differs from any other batch in that there is a lag of at least one cycle between the two batches considered.
[0138] In an embodiment specifically adapted for detecting ongoing cyberattacks, N-Ki+1 batches B(z) (with z = 0, ..., N-Ki) are created from the information collected over N execution cycles, where each batch B(z) consists of the information collected in the Ki execution cycles between the zth and αth cycles (therefore, there is a one-cycle offset between two successive batches). Then, for each batch B(z), the information collected for the Ki execution cycles considered is used as input to at least one MOD model, and the label indicating whether a cyberattack is in progress within the information collected in at least one of the Ki execution cycles in question is used as output. For example, for a sampling period of 1 ms, Ki might be chosen between 100 and 500. For example, for a sampling period of 1 second, Ki might be chosen between 3 and 50.
[0139] In an embodiment adapted specifically for detecting cyberattacks in preparation, N-K2-K2'+1 batches B( / ) (with j = 0, ..., NK.2-K.2') are created from the information collected over N execution cycles, where each batch B( / ) consists of the information collected in the K2+K2' execution cycles between the i-th cycle and the K2+K2'+j-th cycle (therefore, there is a one-cycle offset between two successive batches). Then, for each batch B( / ), the information collected for the K2 execution cycles considered is used as input to at least one MOD model, and the label indicating whether a cyberattack is in progress within the information collected in at least one of the following K2' execution cycles is used as output. For example, for a sampling period of 1ms, K2 will be chosen between 100 and 500 and K2' will be chosen between 50 and 250.For example, for a sampling period of 1 second, K2 will be chosen between 3 and 50 and K2' will be chosen between 2 and 30. K2' will be chosen lower than K2. K2 and K2' are of the same order, the time dependence being limited.
[0140] In machine learning, the loss function quantifies the margin of error between a prediction of the considered machine learning model and the actual expected target output. In one particular embodiment, each MOD 200_0,..,200_n machine learning model has a binary cross-entropy loss function, since the output is a choice between a cyberattack (ongoing or being prepared, depending on the machine learning model) or no cyberattack. In step 308, the training of each MOD 200_0,..,200_n model is performed independently for each pair. For each MOD 200_0,..,200_n model, the training is performed, for example, using the data sets, as described above in relation to step 307. Each MOD 200_0,..,200_n is for example trained to detect that a cyberattack is in progress or being prepared.
[0141] One or more models from among the MOD 200_0,..,200_n models are, for example, pre-trained with data obtained from simulations carried out in another communication system or collected during real cyberattacks suffered by another communication system, this data then being formatted in the same way as the information collected during the simulations of step 302 or even the cleaning and / or preprocessing of step 305. Such pre-training makes it possible to increase the accuracy of the MOD 200_0,..,200_n model by enriching the panel of data available for its training.
[0142] A deep learning model can be used, for example of the GCN (Graph Convolutional Network) type for the detection of ongoing cyberattacks and of the T-GCN (Temporal-GCN) type for the detection of cyberattacks in preparation.
[0143] A reinforcement learning model is used, for example, where each peer is an agent, the communication system 100 is the environment and the rewarding function is representative of a successful detection rate of cyberattacks, such as the ratio between the number of cyberattacks detected and the number of cyberattacks carried out, multiplied by the difference between the time the cyberattack started and the time it was detected.
[0144] Reinforcement learning relies on interactions between the agent and its environment, where the agent executes a policy of actions based on a model. This model receives reward signals that evaluate the quality of its decisions. The agent then adjusts its policy based on this feedback, seeking to maximize the cumulative reward. In other words, the agent learns to choose actions that lead to optimal results, based on predefined criteria and the experience gained from its interactions with its environment. Initially, in simulations within a controlled environment, the reward function relies on explicit labels, offering a positive reward for an appropriate action (such as cutting a connection or issuing an alert) and a penalty for false positives or delayed responses.The peer set Pl 111, P2 112, P3 113, P4 114 collects metrics (such as IP address, latency, connection frequency, packet size, and blockchain height) while planned cyberattacks (such as DDoS, botnet, and Eclipse attacks) occur at various intervals to diversify the learning process. Once the agent is trained in this context, it is deployed passively in an unlabeled environment to observe its detections using a monitoring system.
[0145] Subsequently, for the MOD 200_0,..,200_n model usage phase, the reward function is adjusted to rely solely on observable indicators (such as positive latency and throughput changes, and anomaly detection), incorporating a penalty for unnecessary actions. Based on the results of this reward function, the agent updates its action policy (for example, to transmit a cyberattack detection alert or to perform a connection termination), thus ensuring a gradual transition to an autonomous operational environment.
[0146] Fig. 4 schematically illustrates an algorithm for using at least one so-called MOD model 200_0,..,200_n (use phase) after training for the development of a blockchain.
[0147] In a 401 step, the communication system 100 is set up for the purpose of building the blockchain in question. Each MOD model 200 0, ... ,200_n that has been trained as previously described is activated on all peers Pl 111, P2 112, P3 113, P4 114 (each peer has its own model(s) specifically trained for the peer in question).
[0148] In a 402 step, the blockchain creation process begins. Peers Pl 111, P2 112, P3 113, P4 114 execute a message exchange protocol, as during the simulations, when they seek consensus in the blockchain creation process.
[0149] In a 403 step, each peer Pl 111, P2 112, P3 113, P4 114 collects information relating to events occurring within the communication network during the creation of the blockchain. Each peer collects this information independently and for its own purposes, that is, to feed input data into its own MOD model(s) 200_0,..,200n.
[0150] In a 404 step, each peer Pl 111, P2 112, P3 113, P4 114 injects the collected information into its own MOD 200_0,..,200_n model(s). Each MOD 200_0,..,200_n model outputs, for example, a classification which indicates, according to said trained model MOD 200_0,..,200_n, whether a cyberattack, in progress or in preparation, is detected or not.
[0151] In step 405, each pair P111, P2 112, P3 113, P4 114 determines whether at least one of its own MOD 200_0,..,200_n models detects a cyberattack. If so, step 406 is performed; otherwise, the process loops back to step 403.
[0152] In step 406, a countermeasure action is taken to counter or minimize the detected cyberattack.
[0153] A countermeasure action is, for example, to generate a message or alarm signal, so that a third-party operator or device can take action against the detected cyberattack.
[0154] Another countermeasure action is, for example, to analyze the collected data to identify one or more devices responsible for the detected cyberattack, and to add an identifier (e.g., IP address) of the device(s) to a blacklist. This blacklist identifies any device with which any connection is prohibited. The contents of this blacklist can be shared among peers (P111, P2, P3, P4, P14), if necessary via a backup channel.
[0155] Another countermeasure action, for example, is for peers Pl 111, P2 112, P3 113, and P4 114 to migrate to a backup communication network. Switching to such a backup communication network is particularly beneficial when detecting botnet, Eclipse, DoS, or DDoS attacks, or spam attacks.
[0156] For example, it is also possible to apply the countermeasure after a predefined number of cyberattack detection reports have been reached among peers Pl 111, P2 112, P3 113, P4 114.
[0157] The algorithm in Fig. 4 is executed until the blockchain is fully generated. Fig. 5 schematically illustrates an example of a DISP 500 device hardware platform, which is suitable for use in the communication system 100. This example of a DISP 500 device hardware platform is suitable for implementing each pair Pl 111, P2 112, P3 113, P4 114. This example of a DISP 500 device hardware platform is suitable for implementing the PP 350 processing platform.
[0158] The hardware platform of the DISP 500 device then comprises, connected by a communication bus 510: a processor or CPU (for "Central Processing Unit") 501, or a cluster of such processors, such as GPUs ("Graphics Processing Units"); a random access memory (RAM) 502; a read-only memory (ROM) 503, or a rewritable memory of the type EEPROM ("Electrically Erasable Programmable ROM"), for example of the Flash type; a data storage device, such as a hard disk drive (HDD) 504, or a storage media reader, such as an SD card reader (for "Secure Digital");a set of input and / or output interfaces, such as 505 communication interfaces, enabling communication and, in particular, for peers to exchange messages during the execution of the CONS 120 consensus process in the development of blockchains.
[0159] The 501 processor is capable of executing instructions loaded into RAM 502 from ROM 503, external memory (not shown), storage media such as an SD card or HDD, or a communication network. When the DISP 500 hardware platform is powered on, the 501 processor can read instructions from RAM 502 and execute them. These instructions form a computer program that causes the 501 processor to implement the steps, behaviors, and algorithms described herein in relation to the device (peers, processing platforms PP1 351, PP2352, PP3 353, PP4354) in question.
[0160] All or part of the steps, behaviors, and algorithms described herein can be implemented in software by executing a set of instructions by a programmable machine, such as a DSP (Digital Signal Processor) or a processor, or implemented in hardware by a dedicated machine or component (chip) or a dedicated set of components (chipset), such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). Generally, the DISP 500 device hardware platform comprises electronic circuitry arranged and configured to implement the steps, behaviors, and algorithms described herein in relation to the device (pair, PP 350 processing platform) in question.
Claims
25 DEMANDS 1. A method for securing the creation of a blockchain by node devices, called peers (111, 112, 113, 114), in a communication network, the method comprising the following steps, from a preparatory phase, prior to the creation of the blockchain: - to set up (301) the peer network (111, 112, 113, 114) to carry out (302) simulations, with and without cyberattacks, in the context of developing test blockchains; - collect (303), independently for each peer (111, 112, 113, 114), information relating to events occurring within the framework of exchanges in the communication network occurring during the simulations, the information collected being associated with a label indicating whether the information collected is obtained during a cyberattack against the peer (111, 112, 113, 114) in question, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle; - perform training (308), independently for each pair, on at least one model (200_0, 200_n), using the batches as input data in said training and the associated labels to determine a classification output, said at least one model for each pair being in the form of at least one model trained to be capable of detecting an ongoing cyberattack and / or at least one model trained to be capable of detecting a cyberattack in preparation where, over N simulation execution cycles, N-Ki+1 batches B(z), with z = 0,..,N-Ki, are created for training against an ongoing cyberattack, the information collected for the Ki execution cycles being used as input to said model, and the label indicating whether a cyberattack is ongoing within the information collected in at least one of the Ki execution cycles is used as output and where over N simulation execution cycles N-K2-K2'+1 batches B( / ), with j = 0, ..., NK.2-K.2 are created for training in relation to a cyberattack in preparation, the information collected for the K2 execution cycles being used as input to said model and the label indicating whether a cyberattack is in progress within the information collected in at least one of the following K2' cycles is used as output; the method further comprising the following steps of a usage phase subsequent to the preparatory phase: - set up (401) the peer network (111, 112, 113, 114) for the development of the blockchain, by activating on each peer (111, 112, 113, 114) said at least one trained model (200_0,..,200_n); - collect (403), independently for each peer, information relating to events occurring within the framework of exchanges in the communication network during the usage phase, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle; - inject (404) the input batches of said model associated with each pair and obtain as output a classification relating to a cyberattack detection.
2. The method according to claim 1, wherein each pair has at least two models in the form of at least one model trained to be able to detect an ongoing cyberattack and at least one model trained to be able to detect a cyberattack in preparation.
3. The method according to claim 1 or 2, wherein each model (200_0,..,200_n) has been trained to detect one type of attack from a predefined set of attack types, each attack being associated with at least one model trained to be able to detect an ongoing cyberattack and / or at least one model trained to be able to detect a cyberattack in preparation.
4. A method according to any one of claims 1 to 3, wherein at least one said model (200_0,..,200_n) trained to be able to detect an ongoing cyberattack is a deep learning model of the GCN type.
5. A method according to any one of claims 1 to 4, wherein at least one of said models trained to be able to detect a cyberattack in preparation is a deep learning model of the T-GCN type.
6. A method according to any one of claims 1 to 4, wherein each model (200_0,..,200_n) has a loss function of the binary cross-entropy type.
7. Communication system (100) comprising a network of node devices, called peers (111, 112, 113, 114), the peers (111, 112, 113, 114) being configured to develop a blockchain, in a communication network, the system comprising electronic circuitry configured to perform the following steps, from a preparatory phase in which the peer network (111, 112, 113, 114) is set up (301) to carry out (302) simulations, with and without cyberattacks, in the context of developing test blockchains: - collect (303), independently for each peer (111, 112, 113, 114), information relating to events occurring within the framework of exchanges in the communication network occurring during the simulations, the information collected being associated with a label indicating whether the information collected is obtained during a cyberattack against the peer (111, 112, 113, 114) in question, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle; - perform training (308), independently for each pair, on at least one model (200_0, 200_n), using the batches as input data in said training and the associated labels to determine a classification output, said at least one model for each pair being in the form of at least one model trained to be capable of detecting an ongoing cyberattack and / or at least one model trained to be capable of detecting a cyberattack in preparation where, over N simulation execution cycles, N-Ki+1 batches B(z), with z = 0,..,N-Ki, are created for training against an ongoing cyberattack, the information collected for the Ki execution cycles being used as input to said model, and the label indicating whether a cyberattack is ongoing within the information collected in at least one of the Ki execution cycles is used as output and where over N simulation execution cycles N-K2-K2'+l batches B( / ), with j = 0,..,NK.2-K2 are created for training relative to a cyberattack in preparation, the information collected for the K2 execution cycles being used as input to said model and the label indicating whether a cyberattack is in progress within the information collected in at least one of the following K2' cycles is used as output; the electronic circuitry being further configured to perform the following steps of a usage phase subsequent to the preparatory phase where the peer network (111, 112, 113, 114) is set up (401) for the development of the blockchain, by activating on each peer (111, 112, 113, 114) said at least one trained model (200_0,..,200_n): - collect (403), independently for each peer, information relating to events occurring within the framework of exchanges in the communication network occurring during the usage phase, the information being collected by cycle and grouped into successive batches, each batch comprising several successive cycles, two successive batches being offset by one cycle; - inject (404) the input batches of said model associated with each pair and obtain as output a classification relating to cyberattack detection 8. Communication system according to claim 7, wherein each pair has at least two models in the form of at least one model trained to be able to detect an ongoing cyberattack and at least one model trained to be able to detect a cyberattack in preparation.