Secure data transmission system and method in distributed network
Through machine learning, the generation of counterfeit data and introduction of differential privacy noise, combined with hash chain and multi-path transmission, the problems of data eavesdropping and single point of failure in distributed networks are solved, and the privacy, integrity and reliability of data are achieved.
Patent Information
- Application Number
- CN202510408563.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
In distributed networks, data is easily eavesdropped by intermediate nodes during transmission. Existing encryption technology cannot effectively prevent semi-honest nodes from stealing private information, and a single point of failure during transmission has a great impact, so data integrity and privacy are difficult to guarantee.
Machine learning is used to generate counterfeit data and introduce differential privacy noise. The data blocks are encrypted and dynamically arranged through the hash chain, combining multi-path transmission and Reed-Solomon erasure coding to ensure the privacy and integrity of the data, and identify potential threats through path identification and hash values.
Effectively confusing real data and counterfeiting data, prevent eavesdropping and tampering, improve the reliability and stability of data transmission, reduce the impact of single point of failure, and ensure the integrity and privacy of data.
Smart Images

Figure CN120263471A_ABST
Abstract
Description
Technical Field
[0001] This solution belongs to the technical field of data transmission, and specifically relates to a secure data transmission system and method in a distributed network. Background Technique
[0002] With the rapid development of distributed network technology, the security of data transmission has become a crucial issue. In a distributed network, when a source node transmits target data to a target node, it usually needs to be transmitted through multiple relay nodes to achieve efficient data distribution and sharing. However, the security of these relay nodes cannot be fully guaranteed, especially in an open or semi-trusted network environment. For example, in edge computing, the Internet of Things, or distributed storage systems, relay nodes may be managed by different entities, and their behaviors are difficult to fully predict and control. Some relay nodes may exhibit "semi-honest" behaviors, that is, they will perform data transmission tasks according to the protocol regulations, but at the same time will try to eavesdrop on or analyze the target data during transmission to obtain privacy information. This behavior poses a serious threat to the confidentiality, integrity, and privacy of the target data.
[0003] Currently, although many encryption technologies are applied to data transmission in distributed networks, most methods assume that all nodes in the network are fully trusted. However, in practical applications, this assumption does not always hold. For example, traditional symmetric encryption and asymmetric encryption technologies can protect the confidentiality of data during transmission to a certain extent, but they cannot prevent semi-honest nodes from stealing privacy information after data decryption. In addition, although security protocols such as SSL / TLS are widely used in network communication, in a distributed network, these protocols may not be able to fully prevent the eavesdropping behavior of intermediate nodes. Summary of the Invention
[0004] The purpose of this solution is to provide a secure data transmission system and method in a distributed network to solve the problem that data is easily eavesdropped by intermediate nodes during transmission.
[0005] To achieve the above purpose, this solution provides a secure data transmission method in a distributed network, including the following steps:
[0006] S10: The source node analyzes the target data using a model constructed by a machine learning method, marks the sensitive fields in the target data, and generates forged data with a similarity to the target data exceeding a preset similarity threshold based on the distribution and quantity of the sensitive fields;
[0007] S20: The source node and the target node pre-share the initial key K0. The source node generates the permutation key K according to the number of times of transmitting the target data to the target node through the hash chain K n =SHA3(K n-1 )n , the source node, according to the data type of the target data and K n cuts the target data and the forged data into variable-length first data blocks, and according to K n dynamically arranges the first data blocks;
[0008] S30: The source node encrypts the arranged first data blocks and transmits the first data blocks to the target node through different paths;
[0009] S40: After receiving the encrypted first data blocks, the target node decrypts the first data blocks, and then obtains the number of times K that the source node transmits the target data according to the first data blocks n , according to K n and the data type of the first data, restores the arrangement order of the first data blocks, and extracts the target data according to the arrangement order.
[0010] This solution also provides a secure data transmission system in a distributed network that uses the secure data transmission method in the distributed network.
[0011] The principle and technical effect of this solution are as follows: First, by generating forged data similar to the target data and introducing differential privacy noise, this solution effectively confuses the real data and the forged data, so that even if an eavesdropper intercepts the data, it is impossible to distinguish which are the real data and which are the forged data, thus protecting the privacy of the target data; at the same time, this solution encrypts the data to ensure that even if an eavesdropper intercepts the data, it is impossible to decrypt and obtain the real content; and, this solution generates a dynamic arrangement key through a hash chain to dynamically arrange the data blocks, further increasing the randomness of the data, making it difficult for an eavesdropper to infer the original data from the order of the data blocks, thus solving the problem that data is easily eavesdropped by intermediate nodes during the transmission process.
[0012] Second, by means of multi-path transmission, this solution divides the data into multiple data blocks and transmits them to the target node through different paths. Even if some paths fail or are attacked, other paths can still transmit data normally, thus significantly reducing the impact of single-point failures on data transmission and improving the availability and stability of this solution.
[0013] Furthermore, this solution generates an arrangement key through a hash chain, and each transmitted data block is associated with a specific arrangement key. This method can record the transmission path and order of the data. In case of data leakage or other security incidents, it is possible to quickly trace the transmission path of the data, locate the source of the problem, and facilitate the adoption of corresponding security measures.
[0014] Finally, the forged data in this solution plays a confusing role during the transmission process. Even if the data is leaked, it is difficult for attackers to distinguish between real data and forged data. At the same time, the data is cut and encrypted before transmission, further reducing the risk of data storage. Even if the storage device is attacked or the data is stolen, due to the data being confused and encrypted, it is difficult for attackers to obtain valuable information, thus reducing the risk of data storage in this solution.
[0015] In summary, this solution solves the problem that data is easily eavesdropped by intermediate nodes during the transmission process, and can effectively protect data privacy and security, while ensuring data integrity and availability.
[0016] Furthermore, the machine learning model is a confusion model or a variational autoencoder. When the source node generates forged data, differential privacy noise is introduced to make the statistical feature deviation between the forged data and the target data less than 5%. The model construction includes the following steps:
[0017] S11: Construct a confusion neural network structure including an input layer, multiple hidden layers, and an output layer;
[0018] S12: Use the target data to train the confusion neural network, and optimize the network parameters by minimizing the reconstruction error. The calculation formula of the reconstruction error is shown in the following formula (1):
[0019]
[0020] where, x i is the input target data, is the forged data output by the confusion neural network, and N is the number of target data for training the confusion neural network.
[0021] By introducing differential privacy noise when generating forged data, it can effectively prevent attackers from reverse - inferring sensitive information of the target data from the forged data. The differential privacy technology protects the privacy of the target data and reduces the risk of data tampering by adding noise, making it impossible for attackers to distinguish whether a single data element is used to generate forged data. In addition, the introduction of differential privacy noise not only protects privacy but also enhances data integrity. Even if an intermediate node attempts to tamper with the data, the target node can detect the tampering behavior by verifying the consistency of statistical features when restoring the data. The lightweight GAN can generate forged data highly similar to the target data, and the generated samples are indistinguishable from the real data in terms of statistical features. Moreover, the lightweight GAN and VAE can generate high - quality forged data while having low computational costs and the number of model parameters. This enables the source node to efficiently generate forged data in resource - constrained environments without significantly increasing the computational burden.
[0022] Furthermore, when the source node transmits the first data block to the target node through different paths, it selects more than 3 non-intersecting paths and encodes the data block into a K + n redundant form by combining Reed-Solomon erasure code to ensure that the complete information can be restored from any K paths of data.
[0023] Through Reed-Solomon erasure code, the data is encoded into an n + k redundant form. Even if up to k data blocks are lost or damaged during transmission, the complete original data can still be restored from the remaining data blocks. This significantly improves the reliability and fault tolerance of data transmission. By introducing redundant data during the encoding process, the reliability and security of the data can be improved without significantly increasing the transmission bandwidth. The generation and transmission of redundant data are completed based on the encoding algorithm and do not introduce additional transmission delays. Moreover, the decoding process of Reed-Solomon erasure code is based on matrix operations, and the lost data blocks can be quickly restored through efficient algorithms (such as Gaussian elimination). This enables the proposed solution to quickly restore data and reduce downtime in the face of data loss or path failures.
[0024] Furthermore, when the source node cuts the target data and the forged data into first data blocks of variable length, it also dynamically adjusts the block size according to the real-time network bandwidth. When the bandwidth is higher than the threshold, 1 - 2KB blocks are used; when the bandwidth is lower than the threshold, 256 - 512B blocks are used; for text data, the block boundary is aligned with the sentence or paragraph end mark; for image data, the block boundary is aligned with the JPEG / PNG encoding block.
[0025] By dynamically adjusting the block size, the source node can optimize data transmission according to network conditions, reduce unnecessary delays, and improve the overall transmission efficiency. Moreover, in the case of low bandwidth or unstable network, smaller blocks can reduce the size of a single data packet, thereby reducing the packet loss rate and improving the reliability of data transmission. Combined with Reed-Solomon erasure code, even if some data blocks are lost, the proposed solution can still restore the complete information through redundant data, further improving the fault tolerance of the solution. Furthermore, through specific block boundary alignment, the generated forged data is more similar to the target data in structure, increasing the eavesdropping difficulty of intermediate nodes.
[0026] Furthermore, when the source node encrypts the data, at the transport layer, the source node uses the QUIC protocol to encrypt the communication metadata; at the application layer, the source node encrypts each first data block using the ChaCha20-Poly1305 algorithm, and the encryption key is generated through elliptic curve Diffie-Hellman negotiation; the source node and the target node generate a permutation key K n and perform dynamic permutation on the first data block.
[0027] This multi-layer encryption and dynamic arrangement mechanism significantly enhances the confidentiality of data, ensuring the integrity and authenticity of data during transmission. By dynamically adjusting the block size and selecting the optimal path, this solution optimizes the data transmission efficiency and reduces latency. At the same time, combining Reed-Solomon erasure code and multi-path transmission improves the fault tolerance ability to ensure reliable data transmission. In addition, the introduction of differential privacy noise further enhances privacy protection, making it difficult for intermediate nodes to distinguish between real data and forged data, thus effectively preventing eavesdropping and tampering behaviors. This comprehensive security mechanism provides an efficient, reliable, and secure solution for data transmission in distributed networks, significantly improving the overall performance and security of this solution.
[0028] Further, when the source node dynamically arranges the first data block, it is based on the permutation key K n to generate a pseudo-random sequence, and dynamically adjusts the generation interval of the pseudo-random sequence according to the network latency, and then uses the Fisher-Yates
[0029] shuffle algorithm to disorder the first data block. When arranging, starting from the last first data block, randomly select a first data block within the unprocessed first data blocks in turn to exchange with the first data block at the current position until all first data blocks are processed.
[0030] This solution significantly enhances the security of data transmission by generating a pseudo-random sequence based on the permutation key and using the Fisher-Yates shuffle algorithm to dynamically disorder the data block, effectively preventing data from being predicted and tampered with. At the same time, dynamically adjusting the generation interval of the pseudo-random sequence according to the network latency optimizes the use of computing resources, reducing the permutation frequency to reduce overhead in high-latency network environments, and maintaining a high permutation frequency in low-latency environments to ensure data randomness, thus achieving a balance between performance and security. In addition, the disordered first data block is more fault-tolerant during transmission, facilitating the local retransmission mechanism and further improving the reliability and efficiency of data transmission.
[0031] Further, the source node randomly selects n first data blocks from the first data blocks of the forged data as trap data blocks, and inserts trap identifiers containing unique hash values into the trap data blocks. The unique hash value is generated by formula (2), and formula (2) is as follows:
[0032] H i = SHA3(K n ||BlockID i ) (2),
[0033] where BlockID i is the number or identifier uniquely identifying the first data block;
[0034] The source node assigns a unique path identifier to each path and adds the path identifier to the header of each first data block;
[0035] After receiving the first data block, the destination node records the path identifier in the header of each first data block. The destination node performs a hash operation on the trap data block to obtain a verification hash H i ′, and compares H i ′ with H in the trap data block i . If they are different, the path identifiers of the trap data block headers where H i and H i are different are used as dangerous path identifiers. The relay nodes on the path nodes are obtained through the dangerous path identifiers as dangerous nodes. A path where any relay node is different from the dangerous node is obtained as a safe path. The first data blocks with the same header path identifier as the dangerous path identifier are used as dangerous data blocks. Retransmission information is generated based on the dangerous data blocks and sent to the source node through the safe path.
[0036] This solution uses unique hash values and path identifiers to accurately identify potential security threats, ensuring the integrity and reliability of the first data. At the same time, this solution optimizes path selection, dynamically responds to security threats, and promotes the reliable transmission of the first data through the retransmission mechanism. In addition, this solution also enhances the anti-attack ability by concealing the transmission path identifier and using steganography, improves the resilience of this solution to attacks, and can adapt to different network environments and security requirements, thus providing a flexible and efficient secure data transmission solution for various distributed networks.
[0037] Furthermore, the source node and the destination node store the relay node sequence corresponding to each path identifier and record the network addresses and topological structures of each relay node in the relay node sequence. After the destination node uses the path identifier as a dangerous path identifier, it obtains the relay node sequence as the first sequence through the path identifier, calculates the dangerous association degree of each relay node in the first sequence, uses the nodes with a dangerous association degree greater than the preset maximum dangerous value as suspected nodes, generates verification information based on the suspected nodes, and sends the verification information to the source node through the safe path. After receiving the verification information, the source node generates forged data as verification data based on the retransmission information and sends the verification data to the destination node through the suspected nodes. After receiving the verification information, the destination node compares H i and H i ′. If the comparison result is inconsistent, the suspected nodes are removed from the relay node sequence, and the relay node sequence is synchronized to the source node through the safe path; if the comparison result is consistent, the preset maximum dangerous value is adjusted, and the suspected nodes are obtained again.
[0038] This solution significantly improves the security and reliability of data transmission in a distributed network by meticulously recording and analyzing the relay node sequences corresponding to each path identifier and their network topologies. By calculating the risk correlation degree of relay nodes to identify suspected nodes, and using this information to generate verification data, and then detecting and confirming potential security threats by comparing hash values, it not only optimizes the use of network resources, enhances the resilience and adaptability of this solution to attacks, but also strengthens the maintainability of this solution by dynamically adjusting security policies and timely updating relay node sequences. In addition, this solution ensures data integrity and promotes reliable data transmission by accurately detecting and removing suspicious nodes, thus providing an efficient and dynamic secure data transmission solution for the distributed network. Brief Description of the Drawings
[0039] Figure 1 It is a flowchart of the secure data transmission method in the distributed network in the embodiment of the present invention.
[0040] Figure 2 It is a flowchart of the source node and the target node sending verification information to exclude suspected nodes in the embodiment of the present invention. Detailed Embodiment
[0041] The following will clearly and completely describe the concept and technical effects of the present invention in combination with embodiments to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention:
[0042] As Figure 1 shown, a secure data transmission method in a distributed network includes the following steps:
[0043] S10: The source node analyzes the target data using a model constructed by a machine learning method, marks the sensitive fields in the target data, and generates forged data with a similarity to the target data exceeding a preset similarity threshold based on the distribution and quantity of the sensitive fields;
[0044] S20: The source node and the target node pre-share the initial key K0. The source node generates the permutation key K n =SHA3(K n-1 ) according to the number of times of transmitting the target data to the target node. The source node cuts the target data and the forged data into variable-length first data blocks according to the data type of the target data and K n , and dynamically permutes the first data blocks according to K n ; n
[0045] S30: The source node encrypts the first data block after arrangement and transmits the first data block to the target node through different paths;
[0046] S40: After receiving the encrypted first data block, the target node decrypts the first data block, and then obtains the number of times K that the source node transmits the target data according to the first data block n , according to K n and the data type of the first data, restore the arrangement order of the first data block, and extract the target data according to the arrangement order.
[0047] Among them, the machine learning model is a confusion model or a variational autoencoder. When the source node generates forged data, differential privacy noise is introduced to make the statistical characteristics deviation between the forged data and the target data less than 5%; The steps of constructing the model are as follows:
[0048] S11: Construct a confusion neural network structure including an input layer, multiple hidden layers and an output layer;
[0049] S12: Use the target data to train the confusion neural network, and optimize the network parameters by minimizing the reconstruction error. The calculation formula of the reconstruction error is shown in the following formula (1):
[0050]
[0051] Among them, x i is the input target data, is the forged data output by the confusion neural network, and N is the number of target data for training the confusion neural network.
[0052] Among them, when the source node transmits the first data block to the target node through different paths, more than 3 non-intersecting paths are selected, and the data block is encoded into a K + n redundant form in combination with Reed-Solomon erasure code to ensure that the complete information can be restored from any K paths of data.
[0053] Among them, when the source node cuts the target data and the forged data into variable-length first data blocks, the block size is also dynamically adjusted according to the real-time network bandwidth. When the bandwidth is higher than the threshold, 1-2KB blocks are used, and when the bandwidth is lower than the threshold, 256-512B blocks are used; For text data, the block boundary is aligned with the sentence or paragraph end mark; For image data, the block boundary is aligned with the JPEG / PNG encoding block.
[0054] Among them, when the source node encrypts data, at the transport layer, the source node uses the QUIC protocol to encrypt communication metadata; at the application layer, the source node encrypts each first data block using the ChaCha20-Poly1305 algorithm, and the encryption key is generated through elliptic curve Diffie-Hellman negotiation; the source node and the destination node generate a permutation key K n , and perform dynamic permutation on the first data block.
[0055] Among them, when the source node performs dynamic permutation on the first data block, based on the permutation key K n generate a pseudo-random sequence, dynamically adjust the generation interval of the pseudo-random sequence according to network latency, and then use the Fisher-Yates shuffle algorithm to scramble the first data block; during permutation, starting from the last first data block, randomly select a first data block in the unprocessed first data blocks in turn to exchange with the first data block at the current position until all first data blocks are processed.
[0056] Among them, the source node randomly selects n first data blocks as trap data blocks from the first data blocks of the forged data according to the number of paths, inserts trap identifiers containing unique hash values into the trap data blocks, and the unique hash value is generated by formula (2), and formula (2) is as follows:
[0057] H i =SHA3(K n ||BlockID i ) (2),
[0058] Among them, BlockID i is the number or identifier that uniquely identifies the first data block;
[0059] The source node assigns a unique path identifier to each path and adds the path identifier to the head of each first data block;
[0060] After receiving the first data block, the destination node records the path identifier at the head of each first data block. The destination node performs a hash operation on the trap data block to obtain a verification hash H i ′, compare H i ′ with H i in the trap data block, use the path identifiers at the heads of the trap data blocks where H i and H i ′ are different as dangerous path identifiers, obtain the relay nodes on the path nodes as dangerous nodes through the dangerous path identifiers, obtain the paths where any relay node is different from the dangerous node as safe paths, regard the first data blocks with the same head path identifier as the dangerous data blocks, generate retransmission information according to the dangerous data blocks, and send the retransmission information to the source node through the safe path.
[0061] Among them, the source node and the target node store the relay node sequence corresponding to each path identifier, and record the network address and topological structure of each relay node in the relay node sequence; after the target node uses the path identifier as the dangerous path identifier, as Figure 2 shown, obtain the relay node sequence as the first sequence through the path identifier, calculate the dangerous association degree of each relay node in the first sequence, use the nodes with the dangerous association degree greater than the preset maximum dangerous value (generally 60%, which can be specifically set by the administrator) as suspected nodes, generate verification information based on the suspected nodes, and send the verification information to the source node through the secure path; after the source node receives the verification information, generate forged data as verification data according to the retransmission information, and send the verification data to the target node through each suspected node respectively; after the target node receives the verification information, compare the H i and H i ′, if the comparison result is inconsistent, remove the suspected nodes from the relay node sequence according to the path mark, and synchronize the relay node sequence to the source node through the secure path; if the comparison result is consistent, adjust the preset maximum dangerous value, and obtain the suspected nodes again; until the H i and H i ′ of the verification information are inconsistent.
[0062] Among them, when calculating the dangerous association degree of the relay node, it is measured according to factors such as the position of the relay node in the network and the traffic carried. For example, the core node has a higher importance and can be given a higher weight; the edge node has a lower importance and the weight is relatively small. The importance of the node can be determined by the administrator's scoring or analysis based on the network topology structure. At the same time, analyze the association degree of the relay node with known security events. For example, if the node is associated with multiple high-risk security events, its security event association degree is higher. The security event association information can be obtained through the security event database or security analysis tools. The specific implementation formula is as shown in formula (3) below:
[0063] Dangerous association degree = (node importance × γ + security event association degree × δ) / (γ + δ) (3)
[0065] Among them, γ and δ are weight coefficients, which are used to balance the influence of node importance and security event association degree on the dangerous association degree. They can be set according to the network security policy and the administrator's experience.
[0066] This embodiment also includes a secure data transmission system in a distributed network that uses the secure data transmission method in the distributed network.
[0067] In specific implementation, in the data center (source node) of an e-commerce enterprise, a lightweight generative adversarial network model is used to analyze the user order data (target data) to be transmitted on the same day. This model quickly identifies sensitive fields such as user ID numbers and bank card numbers. Based on the distribution and quantity of these sensitive fields, counterfeit data is generated, and differential privacy noise is introduced during the generation process to ensure that the statistical feature deviation between the counterfeit data and the target data is less than 5%. For example, for the address information in user orders, the counterfeit data will generate seemingly real but actually fictional addresses, and these addresses are extremely similar to the real addresses in terms of statistical features such as regional distribution and format.
[0068] The data center (source node) and the server cluster (target node) pre-share an initial key. The source node generates a permutation key through a hash chain according to the number of times of transmitting user order data to the target node on the same day. Suppose it is the 100th time to transmit data on the same day, and the source node generates a specific permutation key accordingly.
[0069] According to the data type (text type) of the user order data, the source node dynamically adjusts the block size in combination with the real-time network bandwidth. Since the network bandwidth in the morning is higher than the threshold, the order data and the counterfeit data are cut into first data blocks of 1-2 KB in size, and the block boundaries are aligned with the end of sentences or paragraphs to facilitate subsequent processing. For example, an order text containing detailed user purchase information is split into appropriate-sized data blocks in units of sentences.
[0070] The source node generates a pseudorandom sequence based on the permutation key and dynamically adjusts the generation interval of the pseudorandom sequence according to the real-time network latency. Since the network latency in the morning is low, the generation interval is short. Using the Fisher-Yates shuffle algorithm, starting from the last first data block, a data block is randomly selected from the unprocessed data blocks in turn and exchanged with the data block at the current position until all data blocks are processed, completing the dynamic permutation of the first data blocks.
[0071] At the transport layer, the source node uses the QUIC protocol to encrypt the communication metadata to ensure the security of communication connection information. At the application layer, the ChaCha20-Poly1305 algorithm is used to encrypt each first data block, and the encryption key is generated through elliptic curve Diffie-Hellman negotiation.
[0072] The source node selects 4 non-intersecting paths and encodes the data blocks into a redundant form in combination with the Reed-Solomon erasure code to ensure that the complete information can be restored from any 3 paths. For example, a group of order data blocks are encoded and transmitted through four different network lines respectively. Even if one line fails or the data is stolen, the data can be restored through the other three lines.
[0073] The server cluster (target node) decrypts the encrypted first data block after receiving it. According to the first data block, the number of times the source node transmits the target data is obtained, and the arrangement order of the first data block is restored in combination with the data type, and the original user order data is successfully extracted.
[0074] The source node randomly selects some data blocks from the first data block of the forged data as trap data blocks (assuming there are 4 trap data blocks, namely trap data blocks 1-4), and inserts trap identifiers (B1, B2, B3, B4 respectively) containing unique hash values. After the target node receives the data block, it performs a hash operation on the trap data block to obtain verification hashes (H1′, H2′, H3′, H4′ respectively), and compares them with the hash values (H1, H2, H3, H4 respectively) in the trap data block. It is found that H1′ of trap data block 1 is inconsistent with H1, and the head path identifier of trap data block 1 is used as the dangerous path identifier.
[0075] The target node obtains the relay node on the path node as the dangerous node through the dangerous path identifier, and obtains a path different from the dangerous node of any relay node from the stored path information as the safe path (the other 3 paths).
[0076] The target node takes the first data block with the same head path identifier and dangerous path identifier as the dangerous data block, generates retransmission information according to the dangerous data block, and sends it to the source node through the safe path. At the same time, the target node calculates the danger correlation degree of each relay node in the relay node sequence corresponding to the dangerous path, and takes the node with the danger correlation degree greater than the preset maximum danger value (set to 60%) as the suspected node (assuming that the eligible ones are relay node A and relay node B), and generates verification information and sends it to the source node through the safe path. After the source node receives the verification information, it generates forged data as verification data and sends it to the target node through relay node A and relay node B respectively. The target node compares the hash values of the verification data and finds that the hash value of the verification data sent through relay node A is inconsistent (the hash value of the verification data sent through relay node B is consistent), removes relay node A from the relay node sequence, and synchronizes the relay node sequence to the source node through the safe path.
[0077] The above are only embodiments of the present invention, and common knowledge such as specific structures and characteristics well known in the art are not described in detail here. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners and the like described in the specification can be used to interpret the content of the claims.
Claims
1. A secure data transmission method in a distributed network, characterized in that, It includes the following steps: S10: The source node analyzes the target data using a model constructed by a machine learning method, marks the sensitive fields in the target data, and generates forged data with a similarity to the target data exceeding a preset similarity threshold based on the distribution and quantity of the sensitive fields; S20: The source node and the destination node pre-share an initial key K0. The source node generates a permutation key K n = SHA3(K n-1 ) according to the number of times of transmitting the target data to the destination node. The source node cuts the target data and the forged data into variable-length first data blocks according to the data type of the target data and K n , and dynamically permutes the first data blocks according to K n . n =SHA3(K n-1 ) n , n , n S30: The source node encrypts the arranged first data block and transmits the first data block to the target node through different paths; S40: After receiving the encrypted first data block, the target node decrypts the first data block, and then obtains the number of times K that the source node transmits the target data according to the first data block n , according to K n and the data type of the first data, restore the arrangement order of the first data block, and extract the target data according to the arrangement order.
2. The secure data transmission method in a distributed network according to claim 1, characterized in that: The machine learning model is a confusion model or a variational autoencoder. When generating forged data, the source node introduces differential privacy noise to make the statistical feature deviation between the forged data and the target data less than 5%. The steps for constructing the model include: S11: Construct a confusion neural network structure including an input layer, multiple hidden layers, and an output layer; S12: Use the target data to train the confusion neural network, and optimize the network parameters by minimizing the reconstruction error. The calculation formula for the reconstruction error is shown in the following formula (1): Among them, x i is the target data of the input, is the forged data output by the obfuscation neural network, and N is the number of target data for training the obfuscation neural network.
3. A secure data transmission method in a distributed network according to claim 2, characterized in that: When the source node transmits the first data block to the target node through different paths, it selects more than 3 non-intersecting paths and encodes the data block into a K + n redundant form in combination with Reed-Solomon erasure code to ensure that the complete information can be restored from any K paths of data.
4. A secure data transmission method in a distributed network according to claim 3, characterized in that: When the source node cuts the target data and the forged data into first data blocks with variable lengths, it also dynamically adjusts the block size according to the real-time network bandwidth. When the bandwidth is higher than the threshold, it uses 1 - 2KB blocks, and when the bandwidth is lower than the threshold, it uses 256 - 512B blocks; for text data, the block boundaries are aligned with the sentence or paragraph end marks; for image data, the block boundaries are aligned with the JPEG / PNG encoding blocks.
5. A secure data transmission method in a distributed network according to claim 4, characterized in that: When the source node encrypts data, at the transport layer, the source node uses the QUIC protocol to encrypt communication metadata; at the application layer, the source node encrypts each first data block using the ChaCha20-Poly1305 algorithm, and the encryption key is generated through elliptic curve Diffie-Hellman negotiation; the source node and the destination node generate a permutation key K n , and perform dynamic permutation on the first data block.
6. A secure data transmission method in a distributed network according to claim 5, wherein: When the source node dynamically arranges the first data block, it is based on the permutation key K n Generate a pseudo-random sequence, dynamically adjust the generation interval of the pseudo-random sequence according to the network latency, and then use the Fisher-Yates shuffle algorithm to scramble the first data block; during the arrangement, start from the last first data block, and randomly select a first data block in the unprocessed first data blocks in turn to exchange with the first data block at the current position until all the first data blocks are processed.
7. A secure data transmission method in a distributed network according to claim 6, characterized in that: The source node randomly selects n first data blocks from the first data blocks of the forged data as trap data blocks according to the number of paths, and inserts trap identifiers containing unique hash values into the trap data blocks. The unique hash value is generated by the following formula (2), and formula (2) is shown as follows: H i = SHA3(K n || BlockID i ) (2), Among them, BlockID i is the number or identifier that uniquely identifies the first data block; The source node assigns a unique path identifier to each path and adds the path identifier to the head of each first data block; After receiving the first data block, the target node records the path identifier of each first data block header, and the target node performs a hash operation on the trapped data block to obtain the verification hash H i ', and compares H i ' with the H in the trapped data block i . If they are different, the path identifiers of the trapped data block headers where H i and H i are different are used as dangerous path identifiers. Relay nodes on the path nodes are obtained through the dangerous path identifiers as dangerous nodes, a path where any relay node is different from the dangerous node is obtained as a safe path, the first data blocks with the same header path identifier as the dangerous path identifier are used as dangerous data blocks, retransmission information is generated based on the dangerous data blocks, and the retransmission information is sent to the source node through the safe path.
8. A secure data transmission method in a distributed network according to claim 7, characterized in that: The source node and the destination node store the relay node sequences corresponding to each path identifier, and record the network addresses and topologies of the respective relay nodes in the relay node sequences; after the destination node uses the path identifier as the dangerous path identifier, it obtains the relay node sequence as the first sequence through the path identifier, calculates the dangerous correlation degrees of the respective relay nodes in the first sequence, uses the nodes with dangerous correlation degrees greater than the preset maximum dangerous value as suspected nodes, generates verification information based on the suspected nodes, and sends the verification information to the source node through a secure path; after receiving the verification information, the source node generates forged data as verification data based on the retransmission information, and sends the verification data to the destination node through the suspected nodes; after receiving the verification information, the destination node compares the H i and H i ', if the comparison results are inconsistent, the suspected nodes are removed from the relay node sequence, and the relay node sequence is synchronized to the source node through the secure path; if the comparison results are consistent, the preset maximum dangerous value is adjusted, and the suspected nodes are obtained again.
9. A secure data transmission system in a distributed network, characterized in that, A secure data transmission method in a distributed network described in any one of claims 1 - 8 is used.