Big data secure transmission method combining block chain and machine learning
By combining blockchain and machine learning methods, dynamically matching encryption rules and using Merkel tree and blockchain transaction information verification, the security and efficiency problems in big data transmission are solved, adaptive encryption and active defense are realized, and the security and efficiency of data transmission are improved.
Patent Information
- Application Number
- CN202510912709.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-03
AI Technical Summary
The existing technology is difficult to take into account both security and efficiency in big data transmission. The encryption algorithm at rest lacks adaptability, the centralized verification mechanism is fragile, and blockchain technology has storage and efficiency bottlenecks and insufficient dynamic defense capabilities.
Combining blockchain and machine learning, through data transmission risk identification model and abnormal transmission behavior identification model, dynamically match encryption rules, use Merkel tree and blockchain transaction information to verify data integrity, and build a dynamic security architecture.
It realizes the security and efficiency improvement of big data transmission, dynamically perceives transmission risks, provides adaptive encryption strength and active defense, ensures data integrity and reliability, and is suitable for efficient transmission of various data types.
Smart Images

Figure CN120434045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the interdisciplinary field of information security and distributed computing, and in particular to a method for securely transmitting big data that combines blockchain and machine learning. Background Art
[0002] Big data is generally defined as a collection of data characterized by volume, variety, and velocity. This massive scale, combined with its high complexity and real-time nature, poses significant challenges to traditional data processing technologies.
[0003] As application scenarios become increasingly broad and in-depth, the scale, complexity, and transmission requirements of big data have led to prominent security and efficiency issues, becoming a key bottleneck restricting its further development. Specifically, existing big data transmission technologies face the following major challenges: 1. Inherent limitations of static encryption algorithms: Although static encryption algorithms (such as AES and RSA) are widely used for data protection, they have inherent flaws in dynamic, large-scale data transmission environments. The core problem is that they rely on predefined fixed keys and encryption modes, lack the adaptive ability to adapt to real-time threats during transmission (such as increasingly intelligent man-in-the-middle attack variants), and cannot effectively detect and respond to data theft or tampering during transmission.
[0004] 2. Vulnerability of centralized verification mechanisms: The current mechanism for ensuring data integrity is highly dependent on audits by trusted third-party organizations (such as Certificate Authorities (CAs). This centralized architecture presents a single point of failure risk: once the central server is attacked by a distributed denial of service (DDoS) attack or fails, the entire verification system will be paralyzed, causing data loss or service interruption.
[0005] 3. Bottlenecks in direct application of blockchain technology: Blockchain is being explored as an alternative due to its decentralized and tamper-proof features, but it faces significant bottlenecks when applied to large-scale data transmission: a. Storage and efficiency bottlenecks: When processing GB-level or even TB-level big data (such as medical images and satellite remote sensing data), the consensus mechanism (such as Proof-of-Work PoW) of full-node storage of complete blocks will introduce severe processing delays, and transmission efficiency may drop to less than one-tenth of traditional methods; b. Lack of dynamic defense capabilities: Blockchain mainly focuses on the verification and recording of transactions. It essentially lacks the ability to actively perceive risks in transmission channels (such as the network layer), and cannot analyze transmission traffic patterns in real time to intelligently identify abnormal behaviors (such as DDoS attacks and signs of data leaks). It is limited to recording transactions after the fact, and it is difficult to provide active real-time defense responses during the transmission process.
[0006] In summary, existing technology systems struggle to effectively balance the two core requirements of security and efficiency when addressing the challenges of big data transmission. Traditional static encryption and centralized authentication architectures present serious security risks and low reliability. While blockchain technology improves security, its high storage overhead and inherent latency make it difficult to meet the efficiency requirements of high-throughput data transmission, and it also lacks the dynamic intelligent protection capabilities of transmission channels. Therefore, developing a secure big data transmission method that combines blockchain and machine learning to improve both the security and efficiency of big data transmission has become a pressing technical challenge. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a big data security transmission method that combines blockchain and machine learning to improve the security and efficiency of big data transmission.
[0008] The present invention is implemented as follows: a method for secure transmission of big data combining blockchain and machine learning, comprising the following steps: Step S1: Create a data transmission risk identification model, train the data transmission risk identification model, and then deploy it to the transmission end; create an abnormal transmission behavior identification model, train the abnormal transmission behavior identification model, and then deploy it to the receiving end; Step S2: setting an encryption rule set including standard encryption rules, intermediate encryption rules, and advanced encryption rules, and pre-setting the encryption rule set into the transmitting end and the receiving end; Step S3: The transmission end obtains big data to be transmitted, the data type of which is structured data, semi-structured data or unstructured data, and performs data sharding and preprocessing on the big data to obtain a plurality of data blocks; Step S4: The transmission end obtains real-time security monitoring data, inputs the real-time security monitoring data into a deployed data transmission risk identification model to obtain a transmission risk identification result, matches a corresponding encryption rule from a preset encryption rule set based on the transmission risk identification result, encrypts each data block based on the matched encryption rule to obtain a data encryption block, and compresses each data encryption block to obtain a data compression block; the encryption rule is a standard encryption rule, an intermediate encryption rule, or an advanced encryption rule; Step S5: The transmitting end calculates the data fingerprint of each of the data compression blocks, constructs a Merkle tree based on each of the data fingerprints, uploads the root hash value of the Merkle tree to the blockchain, obtains blockchain transaction information, and sends the blockchain transaction information and each of the data compression blocks to the receiving end in sequence; Step S6: The receiving end receives the transmitted blockchain transaction information and each data compression block in real time, and performs security defense through the deployed abnormal transmission behavior recognition model during the receiving process; Step S7: After the receiving end performs integrity verification on each data compression block using the blockchain transaction information, it decompresses each data compression block to obtain a data encryption block, decrypts the data encryption block using the preset encryption rule set to obtain a data block, and concatenates each data block to obtain big data, thereby completing the transmission of the big data. Step S8: The receiving end records the big data reception log in real time, encrypts the big data reception log into an encrypted log using the encryption rule set, and uploads it to the blockchain, and iteratively optimizes the deployed abnormal transmission behavior recognition model using the big data reception log.
[0009] Furthermore, in step S1, creating a data transmission risk identification model and training the data transmission risk identification model and deploying it to the transmission end are specifically as follows: Create a data transmission risk identification model based on the data input preprocessing layer, risk feature extraction layer, risk feature fusion layer and risk prediction output layer; The data input preprocessing layer is constructed based on a network traffic preprocessing module, a system log preprocessing module, a network device status preprocessing module, a user behavior preprocessing module, a vulnerability scanning preprocessing module, and a threat intelligence preprocessing module; the network traffic preprocessing module, the system log preprocessing module, the network device status preprocessing module, the user behavior preprocessing module, the vulnerability scanning preprocessing module, and the threat intelligence preprocessing module are all constructed based on standardized units and embedded units; The network traffic preprocessing module is used to standardize and preliminarily embed network traffic data to obtain a network traffic embedding vector; the system log preprocessing module is used to standardize and preliminarily embed system log data to obtain a system log embedding vector; the network device status preprocessing module is used to standardize and preliminarily embed network device status data to obtain a network device status embedding vector; the user behavior preprocessing module is used to standardize and preliminarily embed user behavior data to obtain a user behavior embedding vector; the vulnerability scanning preprocessing module is used to standardize and preliminarily embed vulnerability scanning data to obtain a vulnerability scanning embedding vector; the threat intelligence preprocessing module is used to standardize and preliminarily embed threat intelligence data to obtain a threat intelligence embedding vector; The risk feature extraction layer is constructed based on a network traffic feature extraction module, a system log feature extraction module, a network device status feature extraction module, a user behavior feature extraction module, a vulnerability scanning feature extraction module, and a threat intelligence feature extraction module; the network traffic feature extraction module, the system log feature extraction module, the network device status feature extraction module, the user behavior feature extraction module, the vulnerability scanning feature extraction module, and the threat intelligence feature extraction module are all constructed based on a convolutional neural network unit, a gated recurrent unit, and a splicing unit; The network traffic feature extraction module is used to extract network traffic risk features from the network traffic embedding vector; the system log feature extraction module is used to extract system log risk features from the system log embedding vector; the network device status feature extraction module is used to extract network device status risk features from the network device status embedding vector; the user behavior feature extraction module is used to extract user behavior risk features from the user behavior embedding vector; the vulnerability scanning feature extraction module is used to extract vulnerability scanning risk features from the vulnerability scanning embedding vector; the threat intelligence feature extraction module is used to extract threat intelligence risk features from the threat intelligence embedding vector; The risk feature fusion layer is used to fuse network traffic risk features, system log risk features, network device status risk features, user behavior risk features, vulnerability scanning risk features, and threat intelligence risk features through a multi-head self-attention mechanism unit to obtain a risk fusion feature; The risk prediction output layer is constructed based on the risk item identification module, the risk level classification module and the result output module; The risk item identification module is used to infer the risk fusion features to obtain the risk item probability distribution; the risk level classification module is used to infer the risk fusion features to obtain the risk level probability distribution; the result output module is used to output the transmission risk identification result carrying the risk item and risk level based on the risk item probability distribution and the risk level probability distribution; the risk level is low risk, medium risk or high risk; Setting the risk optimization function of the data transmission risk identification model to adopt Adam optimizer; The risk loss function of the data transmission risk identification model is set as: ; in, represents the loss value of the risk loss function; Represents the risk term cross entropy loss, using classification cross entropy loss; represents the risk level cross entropy loss, using classification cross entropy loss; Both represent weight coefficients; Obtain a large amount of historical security monitoring data, including at least network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data; The network traffic data is cleaned at least by removing noise data and duplicate data, including data standardization of unified time format and standardized numerical fields, including feature extraction of session features and traffic features; the system log data is cleaned at least by formatting logs and removing irrelevant logs, including data standardization of unified log levels and unified time formats, including feature extraction of event features and text features; the network device status data is cleaned at least by removing abnormal status data and duplicate data, including data standardization of unified device identification and standardized numerical fields, including feature extraction of performance features and status features; the user behavior data is cleaned at least by formatting logs and removing irrelevant logs, including data standardization of unified log levels and unified time formats, including feature extraction of event features and text features. The data cleaning includes at least removing invalid behavior data and removing duplicate data, including data standardization of unified user identification and unified time format, and extracting behavioral features and time features; the vulnerability scanning data includes at least removing invalid vulnerability data and removing duplicate data, including data standardization of unified vulnerability level and unified time format, and extracting vulnerability features and risk features; the threat intelligence data includes at least removing invalid intelligence data and removing duplicate data, including data standardization of unified threat type and unified time format, and extracting threat features and credibility features, so as to complete the preprocessing of each historical security monitoring data; Constructing a risk data set after labeling each of the preprocessed historical safety monitoring data, including at least risk items and risk levels; dividing the risk data set into a first training set, a first validation set, and a first test set in a ratio of 8:1:1 by a stratified sampling method; The data transmission risk identification model is trained using the first training set, risk optimization function, and risk loss function until the preset first early stopping condition is met. The trained data transmission risk identification model is verified and tested using the first validation set and the first test set, respectively. The data transmission risk identification model that passes the test is deployed to the transmission end.
[0010] Furthermore, in step S1, creating an abnormal transmission behavior recognition model and training the abnormal transmission behavior recognition model and deploying it to the receiving end are specifically as follows: An abnormal transmission behavior recognition model is created based on the abnormal feature extraction layer, the abnormal feature fusion layer and the abnormal prediction output layer; The abnormal feature extraction layer is constructed based on a multi-scale temporal channel, a spatial dependency channel, a behavioral pattern channel, and a feature aggregation module; the multi-scale temporal channel is used to extract multi-scale temporal features from the input network transmission data stream through a bidirectional gated recurrent unit and a dilated convolutional layer; the spatial dependency channel is used to extract spatial dependency features from the input network transmission data stream through a graph convolutional network unit; the behavioral pattern channel is used to extract behavioral pattern features from the input network transmission data stream through a self-attention mechanism and a statistical feature extractor; the feature aggregation module is used to aggregate multi-scale temporal features, spatial dependency features, and behavioral pattern features through a gated feature fusion unit to obtain contextual features; The abnormal feature fusion layer is used to perform cross-modal fusion of the context features through a hierarchical attention mechanism and a spatiotemporal compression module to obtain a comprehensive behavior representation; The anomaly prediction output layer is used to map the comprehensive behavior representation to the abnormal behavior probability distribution and classify and output the transmission behavior recognition results; Setting the behavior optimization function of the abnormal transmission behavior identification model to adopt Lookahead optimizer and RAdam optimizer; Setting the behavior loss function of the abnormal transmission behavior recognition model to adopt classification cross entropy loss; Obtain a large amount of historical network transmission data flow including at least source IP address, destination IP address, port, communication protocol, packet size, transmission interval and identification bit distribution; Performing preprocessing on each of the historical network transmission data streams, including at least data cleaning, feature extraction and conversion, data dimensionality reduction, data bucketing, and data formatting; After preprocessing, each of the historical network transmission data streams is annotated with at least abnormal behaviors to construct a behavior dataset; the behavior dataset is divided into a second training set, a second validation set, and a second test set in a ratio of 8:1:1 by a stratified sampling method; The abnormal transmission behavior recognition model is trained using the second training set, behavior optimization function, and behavior loss function until the preset second early stopping condition is met. The trained abnormal transmission behavior recognition model is verified and tested using the second verification set and the second test set, respectively. The abnormal transmission behavior recognition model that passes the test is deployed to the receiving end.
[0011] Furthermore, in step S2, the standard encryption rule is specifically: Obtain the current date string, perform MD5 hash calculation on the date string to obtain a hash value H1, extract the first 8 bytes from the hash value H1 as the dynamic key K1; convert the large data to be encrypted into a UTF-8 encoded byte array B1, generate an 8-byte dynamic salt value Y1, append the dynamic salt value Y1 to the front of the byte array B1 to obtain augmented data S1, call the dynamic key K1 using the DES algorithm to encrypt the augmented data S1 to obtain ciphertext data C1, and perform Base64 encoding on the ciphertext data C1 to obtain a data encryption block; The intermediate encryption rules are specifically as follows: Obtain the current date string, perform a SHA-256 hash calculation on the date string to obtain a 32-byte hash value H2, use the first 16 bytes of the hash value H2 as the dynamic key K2, and the last 16 bytes as the dynamic key K3; convert the large data to be encrypted into a UTF-8 encoded byte array B2, divide the byte array B2 into sub-arrays B21 and B22, call the dynamic key K2 using the AES-128 algorithm to encrypt the sub-array B21 to obtain an encrypted block EA1, call the dynamic key K3 using the Blowfish algorithm to encrypt the sub-array B22 to obtain an encrypted block EB1, concatenate the encrypted blocks EA1 and EB1 to obtain concatenated data P1, cyclically shift the concatenated data P1 by 5 bits in units of bytes to obtain transformed data T1, perform a position swap operation on every two consecutive bytes in the transformed data T1 to obtain ciphertext data C2, and perform Base32 encoding on the ciphertext data C2 to obtain a data encryption block; The advanced encryption rules are specifically as follows: Get the current date string, perform SHA-512 hash calculation on the date string to obtain a 64-byte hash value H3, use the first 32 bytes of the hash value H3 as the dynamic key K4, the middle 16 bytes as the dynamic key K5, and the last 16 bytes as the dynamic salt value Y2; convert the large data to be encrypted into a UTF-8 encoded byte array B3, divide the byte array B3 into sub-arrays B31, B32, and B33, use the dynamic key K4 to encrypt the sub-array B31 through the AES-256 algorithm to obtain the encrypted block EA2, and use the 3DES algorithm to call the encrypted block EA2. The subarray B32 is encrypted using the dynamic key K5 to obtain an encrypted block EB2. A byte-by-byte XOR operation is performed on the subarray B32 and the dynamic salt value Y2 to obtain an encrypted block EC2. The encrypted blocks EA2, EB2, and EC2 are concatenated to obtain concatenated data P2. The upper 4 bits and the lower 4 bits of each byte in the concatenated data P2 are swapped to obtain transformed data Q. The transformed data Q is divided into 16-byte blocks. Each of the 16-byte blocks is reordered in a preset reverse order to obtain ciphertext data C3. The ciphertext data C3 is Base16 encoded to obtain a data encryption block.
[0012] Furthermore, the step S3 is specifically as follows: The transmission end obtains big data to be transmitted, where the data type of the big data is structured data, semi-structured data or unstructured data; the big data is segmented based on a preset file size to obtain a number of data blocks, and each of the data blocks is preprocessed, including at least data cleaning and data standardization.
[0013] Furthermore, the step S4 is specifically as follows: The transmission end obtains real-time security monitoring data including at least network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data, performs preprocessing on each of the real-time security monitoring data including at least data cleaning and data standardization through a streaming computing engine, and inputs the preprocessed real-time security monitoring data into a deployed data transmission risk identification model. The data transmission risk identification model performs parallel inference on multiple GPUs to output a transmission risk identification result containing risk items and risk levels; parsing the transmission risk identification result, and matching the standard encryption rule from the encryption rule set when the risk level is low risk; matching the intermediate encryption rule from the encryption rule set when the risk level is medium risk; and matching the advanced encryption rule from the encryption rule set when the risk level is high risk; Based on the matching encryption rules, each data block is encrypted in parallel to obtain a corresponding data encryption block, each data encryption block is named based on the current date string and the data block number, and each data encryption block is compressed using the DEFLATE algorithm to obtain a data compression block; the encryption rule is a standard encryption rule, an intermediate encryption rule or an advanced encryption rule.
[0014] Furthermore, the step S5 is specifically as follows: The transmitting end calculates the data fingerprint of each of the data compression blocks through the SM3 algorithm, constructs a Merkle tree based on the data fingerprint of each of the data compression blocks, uploads the root hash value of the Merkle tree to the blockchain, waits for the root hash value to be passed by the blockchain consensus, obtains the blockchain transaction information feedback including at least the transaction hash, block height and block hash, and sends the blockchain transaction information and each data compression block to the receiving end in sequence through a secure communication protocol.
[0015] Furthermore, the step S6 is specifically as follows: The receiving end receives each of the transmitted data compression blocks in real time, and during the receiving process, collects a real-time network transmission data stream including at least the source IP address, destination IP address, port, communication protocol, packet size, transmission interval, and identification bit distribution, performs preprocessing on each of the real-time network transmission data streams through a streaming computing engine, including at least data cleaning and data standardization, and inputs each of the preprocessed real-time network transmission data streams into a deployed abnormal transmission behavior recognition model. The abnormal transmission behavior recognition model performs parallel inference on multiple GPUs to output transmission behavior recognition results; The receiving end analyzes the transmission behavior identification result. When the transmission behavior identification result carries abnormal transmission behavior, the receiving end stops receiving the data compression block and pushes an early warning notification to a pre-associated management terminal for security defense.
[0016] Furthermore, the step S7 is specifically as follows: The receiving end obtains a root hash value from the blockchain through the blockchain transaction information, performs integrity verification on each data compression block based on the root hash value, decompresses each data compression block through the DEFLATE algorithm to obtain a data encryption block, matches encryption rules from a preset encryption rule set based on the encoding format of the data encryption block, decrypts and verifies each data encryption block based on the matched encryption rule and the date character string carried in the file name of the data encryption block to obtain a data block, and splices each data block based on the data block number carried in the file name of the data encryption block to obtain big data, thereby completing the transmission of big data.
[0017] Furthermore, the step S8 is specifically as follows: The receiving end records in real time a big data reception log including at least real-time security monitoring data, reception time, and blockchain transaction information, encrypts the big data reception log into an encrypted log using the advanced encryption rules in the encryption rule set, and uploads the encrypted log to the blockchain; constructs an incremental data set using the big data reception log, annotates the incremental data set for abnormal behavior, and iteratively optimizes the deployed abnormal transmission behavior identification model.
[0018] The advantages of the present invention are: 1. Create and train a data transmission risk identification model and deploy it to the transmission end; create and train an abnormal transmission behavior identification model and deploy it to the receiving end; set an encryption rule set including standard encryption rules, intermediate encryption rules and advanced encryption rules, and pre-set it into the transmission end and the receiving end; the transmission end obtains big data of the data type to be transmitted, which is structured data, semi-structured data or unstructured data, and performs data segmentation and preprocessing on the big data to obtain several data blocks; the transmission end obtains real-time security monitoring data, inputs the real-time security monitoring data into the data transmission risk identification model to obtain a transmission risk identification result, matches the corresponding encryption rule from the encryption rule set based on the transmission risk identification result, encrypts each data block based on the matched encryption rule to obtain a data encryption block, compresses each data encryption block to obtain a data compression block, and calculates the data fingerprint of each data compression block, constructs a Merkle tree based on each data fingerprint, uploads the root hash value of the Merkle tree to the blockchain, obtains blockchain transaction information, and sends the blockchain transaction information and each data compression block to the receiving end in sequence; the receiving end receives During the process, security defense is carried out through the abnormal transmission behavior identification model. After the integrity of each data compression block is verified through blockchain transaction information, each data compression block is decompressed to obtain a data encryption block, which is decrypted using the encryption rule set to obtain a data block. The data blocks are spliced to obtain big data to complete the transmission of big data, and the big data reception log is recorded in real time. The big data reception log is encrypted into an encrypted log using the encryption rule set and uploaded to the blockchain. The deployed abnormal transmission behavior identification model is iteratively optimized through the big data reception log; that is, through machine learning models (data transmission risk identification model, abnormal transmission behavior identification model), transmission risks (transmission end) and abnormal behaviors (receiving end) are dynamically perceived in real time to achieve adaptive encryption strength adjustment and active defense; at the same time, the blockchain only stores the root hash value of the Merkle tree constructed by data fingerprints. Under the premise of ensuring data integrity and source credibility, the efficiency bottleneck of directly uploading big data to the chain is avoided, thereby greatly improving the security and efficiency of big data transmission under the dual protection of dynamic threat protection and lightweight trusted verification.
[0019] 2. By combining machine learning risk assessment models (data transmission risk identification models and abnormal transmission behavior identification models) with blockchain, a dynamic security architecture is constructed, which provides multi-layer defense: dynamic matching of encryption rules (such as standard, intermediate or advanced) based on real-time security monitoring data to adapt to different risk levels; at the same time, the receiving end uses the deployed abnormal behavior model to defend against attacks in real time. This integration significantly improves security and adaptability, reduces the risk of data leakage or tampering, and is particularly suitable for processing highly sensitive big data.
[0020] 3. By calculating data fingerprints, constructing Merkle trees, and uploading the root hash value to the blockchain, the integrity of big data during transmission is ensured; the distributed ledger nature of the blockchain ensures that the root hash value cannot be tampered with, and integrity verification (using blockchain transaction information) provides end-to-end verifiability, which allows any data tampering to be quickly detected, preventing man-in-the-middle attacks or data corruption, and greatly improving the reliability of big data transmission.
[0021] 4. Through data sharding, preprocessing and data compression, the processing efficiency of big data is optimized, and bandwidth requirements and transmission delays are reduced: data sharding allows parallel processing, and compression reduces data volume. Combined with risk-based encryption rule selection, unnecessary high-intensity encryption (such as using advanced encryption only when the risk is high) is avoided, effectively improving transmission speed and system throughput, which is particularly suitable for real-time transmission scenarios of massive big data.
[0022] 5. By recording big data reception logs, uploading them to the blockchain in encrypted form, and using them to optimize the abnormal transmission behavior identification model, a closed-loop feedback mechanism is introduced. The machine learning model can be iteratively optimized based on actual log data to improve the accuracy of abnormal behavior detection and defense capabilities. At the same time, the encrypted uploading of big data reception logs to the blockchain ensures the reliability and auditability of big data reception logs, enabling the system to adapt to new threats and improve security performance in the long term.
[0023] 6. By supporting multiple data types (structured, semi-structured and unstructured data) and providing a universal security framework through a preset encryption rule set; that is, from data acquisition to decryption and splicing, the entire process is applicable to various big data scenarios (such as enterprise data transmission and IoT device communication) and is highly versatile.
[0024] 7. By deeply integrating blockchain and machine learning technologies, an intelligent and dynamic big data security transmission system has been built. The core advantage lies in the adaptive encryption strength adjustment driven by real-time risk assessment (such as dynamically matching multi-level encryption rules according to transmission risks), and the use of blockchain to ensure that the entire link data is tamper-proof and traceable (through Merkle tree root hashing and log encryption and storage). At the same time, with the help of the closed-loop optimization mechanism of the machine learning model, the defense accuracy is continuously improved (based on the abnormal behavior identification model of the receiving-end log iteration). It not only significantly enhances the transmission security and integrity of structured / semi-structured / unstructured mixed data, but also greatly optimizes transmission efficiency (data sharding compression and lightweight verification) and system resource utilization (layered encryption reduces computing power consumption). Ultimately, while effectively resisting complex network attacks, it provides a scalable, low-cost, and highly compliant integrated solution for big data applications in multiple fields.
[0025] 8. By integrating six key security data sources, including network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data, it covers multiple dimensions such as network transmission activities, system operation status, device health, user operations, system vulnerabilities, and external threat intelligence. It greatly expands the scope of risk identification and significantly improves the model's ability to capture complex, hidden, or cross-data source coordinated attacks, avoiding risk assessment blind spots caused by data silos.
[0026] 9. By designing independent preprocessing modules (including standardization and embedding units) and feature extraction modules (based on CNN+GRU+splicing units) for each type of data source, the most appropriate preprocessing (such as specific data cleaning and standardization methods) and feature extraction (such as CNN capturing spatial patterns and GRU capturing temporal dependencies) can be performed for the characteristics of different data types (such as time series, log text, and status indicators). This avoids information loss or noise introduction caused by mixed processing of different types of data, and can generate high-quality, risk-oriented embedding vectors and risk features, significantly improving the data quality and feature validity of the input model, laying a solid foundation for subsequent fusion and prediction, and is a key prerequisite for the high accuracy of the model.
[0027] 10. The risk feature fusion layer uses a multi-head self-attention mechanism, which can dynamically learn the importance weights and correlations between risk features from different data sources. It is particularly good at capturing long-distance dependencies and nonlinear interactions. The "multi-head" design allows the model to simultaneously focus on related information from different representation subspaces. This enables the data transmission risk identification model to intelligently fuse risk features from six different sources and understand the complex interactions between them (for example, a user's suspicious behavior combined with a device's known vulnerabilities and related attack patterns in external threat intelligence), thereby generating more comprehensive and accurate risk fusion features, significantly improving the ability to identify complex threats.
[0028] 11. The risk prediction output layer simultaneously performs two tasks: risk item identification (what is the risk) and risk level classification (low, medium, and high risk). It shares the underlying risk fusion features, utilizes the correlation between tasks, and shares feature representations. This is more efficient than training two independent models and can promote each other. The output result (risk item + level) is more operationally instructive than a single risk item or a single level, and the risk loss function allows the importance of the two tasks to be flexibly adjusted through the weight coefficients α and β. That is, the data transmission risk identification model can provide richer risk information (not only knowing what risks there are, but also their severity) in a single inference, which facilitates the subsequent adoption of differentiated security measures (such as strengthening encryption levels, alarms, blocking, and audit enhancements). Multi-task learning helps improve the generalization ability of the model.
[0029] 12. By specifying the specific steps of data cleaning, standardization and feature extraction for each type of historical security monitoring data (such as removing noise / duplicate data, unifying formats / identifiers, extracting session features / event features, etc.), data quality and consistency are ensured; stratified sampling of labeled data (8:1:1 division into training / validation / test sets) helps to maintain a balanced data distribution and make transmission risk identification results more reliable; the Adam optimizer (adaptive learning rate, suitable for complex non-convex optimization) and early stopping strategy (based on validation set performance to prevent overfitting) are used; after training, the validation set is tuned and the test set is evaluated to ensure that the model performance meets the standards before deployment. The entire process design is scientific and standardized, which maximizes the validity of the training data, the stability of the model training, and the generalization performance and reliability of the final deployed model.
[0030] 13. By building a modular deep learning model (data transmission risk identification model) based on multi-source heterogeneous data (network traffic, system logs, device status, user behavior, vulnerability scanning and threat intelligence), using professional data preprocessing (standardization and embedding) and feature extraction (combination of CNN and GRU), combined with a multi-head self-attention mechanism to achieve risk feature fusion, and using a multi-task learning framework to simultaneously output risk items and levels (low / medium / high risk), and finally deployed at the transmission end, it significantly improves the comprehensiveness, accuracy, real-time and operability of data transmission risk identification, providing an efficient and reliable intelligent solution for active defense.
[0031] 14. The abnormal feature extraction layer adopts a multi-channel design (multi-scale temporal channel, spatial dependency channel, and behavioral pattern channel), and combines it with a feature aggregation module to integrate contextual features. The multi-scale temporal channel uses a bidirectional gated recurrent unit (BiGRU) and dilated convolution to capture transmission behavior characteristics at different time granularities. The spatial dependency channel uses a graph convolutional network (GCN) to process network topology relationships (such as dependencies between IP addresses). The behavioral pattern channel uses a self-attention mechanism and a statistical feature extractor to identify abnormal patterns in behavioral patterns. The abnormal feature fusion layer further implements cross-modal fusion through a hierarchical attention mechanism and a spatiotemporal compression module. This multi-dimensional feature extraction and fusion mechanism overcomes the limitations of traditional single feature models (such as temporal models only), and can fully cover the temporal, spatial, and behavioral pattern characteristics of network transmission data. The feature aggregation module reduces feature redundancy and enhances the robustness and generalization ability of the model.
[0032] 15. The behavior optimization function uses a combination of the Lookahead optimizer and the RAdam optimizer, and the behavior loss function uses categorical cross-entropy loss. The Lookahead optimizer stabilizes the training process through a forward-looking mechanism, and the RAdam optimizer (Rectified Adam) adaptively adjusts the learning rate to prevent gradient vanishing / explosion. The loss function optimizes the model output for classification tasks. An early stopping condition (second early stopping condition) is introduced during training, and the data set is partitioned using stratified sampling (8:1:1 ratio). This optimization strategy significantly improves training efficiency. The combination of Lookahead and RAdam reduces the number of training iterations, while the early stopping mechanism prevents overfitting and conserves computing resources. Stratified sampling ensures balanced data distribution, reduces sample bias (such as insufficient abnormal behavior samples), and improves the performance of the abnormal transmission behavior recognition model on validation and test sets.
[0033] 16. Historical network transmission data streams undergo multi-step preprocessing, including data cleaning, feature extraction and conversion, data dimensionality reduction, data bucketing, and data formatting. A behavioral dataset is then constructed in conjunction with abnormal behavior annotation. This optimization is performed based on network transmission characteristics (such as source IP address, destination IP address, and packet size), ensuring the high quality and consistency of input data. For example, data dimensionality reduction (such as PCA or t-SNE) reduces noise and the curse of dimensionality, data bucketing (such as grouping based on transmission intervals) enhances feature interpretability, and data formatting adapts to model input requirements. This directly improves the training effect of the abnormal transmission behavior recognition model and reduces the false alarm rate caused by the data. At the same time, the stratified sampling method (8:1:1) ensures the representativeness of the dataset division, making the abnormal transmission behavior recognition model more reliable after deployment.
[0034] 17. By building a multi-channel collaborative abnormal transmission behavior recognition model (integrating multi-scale time series channels, spatial dependency channels, and behavioral pattern channels), combined with feature aggregation and cross-modal fusion mechanisms, the comprehensiveness and accuracy of abnormal transmission behavior detection are significantly improved. The use of Lookahead and RAdam dual optimizers and classification cross-entropy loss function, combined with stratified sampling data partitioning and early stopping mechanisms, greatly optimizes training efficiency and model generalization capabilities. At the same time, a systematic data preprocessing process (including cleaning, dimensionality reduction, bucketing, etc.) ensures input quality. After rigorous verification, the model is directly deployed to the receiving end, achieving low-latency real-time anomaly recognition, providing efficient, reliable, and scalable intelligent protection capabilities for secure transmission in the blockchain environment.
[0035] 18. By setting standard encryption rules to dynamically generate an MD5 hash key (K1) and a random salt value (Y1) based on the current date, the key's timeliness and cracking resistance are ensured while eliminating the predictability of encrypted data, effectively defending against rainbow table attacks and pattern analysis. At the same time, the DES algorithm combined with a dynamic salt value appending mechanism ensures efficient processing of large data encryption, while the final Base64-encoded output provides cross-platform compatibility, enabling the solution to balance security strength, operational efficiency, and system versatility.
[0036] 19. By setting intermediate encryption rules to dynamically generate dual keys (K2 and K3) based on the current date, the segmented data is encrypted in parallel using the AES-128 and Blowfish encryption algorithms, and multi-layer obfuscation is achieved through cyclic shifting and byte swapping, significantly improving system security. Dynamic keys can prevent the risk of long-term key leakage, the hybrid algorithm design reduces the probability of single-point attack, and post-processing operations enhance the ciphertext's anti-analysis capabilities. At the same time, data block strategies and Base32 encoding optimize processing efficiency and compatibility, providing efficient and reliable end-to-end protection for big data applications while ensuring encryption strength.
[0037] 20. By setting advanced encryption rules, date-driven dynamic key generation (SHA-512 hash-derived multi-level keys and salt values) and hybrid encryption mechanisms (AES-256, 3DES and XOR salt value operations to parallel process data blocks), combined with multiple data transformations (byte bit swapping, block reordering) and Base16 standardized output, efficient and crack-resistant big data encryption is achieved, significantly improving security (dynamic keys reduce long-term exposure risks, hybrid algorithms enhance anti-attack redundancy, and bit transformations destroy data patterns) while taking practicality into account (block parallel processing optimizes computing efficiency, and Base16 encoding ensures cross-platform compatibility), forming an encryption system that is dynamic, obfuscated, and robust, suitable for big data scenarios with high security requirements.
[0038] 21. By pre-sharding any type of big data to be transmitted (including structured, semi-structured or unstructured data) according to preset file sizes and performing pre-processing operations such as data cleaning and standardization, the coordinated optimization of transmission efficiency, data processing quality and system robustness is achieved; the sharding mechanism and pre-processing are integrated into the transmission link to effectively reduce network bandwidth pressure and transmission delay, while eliminating data errors and format differences in advance through cleaning and standardization, which not only significantly improves the accuracy of subsequent analysis results, but also greatly reduces the resource overhead of the receiving end; the independence of shard processing further ensures that single point failures do not affect the overall task, so that the solution can flexibly adapt to diverse application scenarios while ensuring high reliability, ultimately providing efficient, reliable and scalable technical support for big data transmission and application.
[0039] 22. The streaming computing engine integrates multi-source security monitoring data in real time, and the risk identification model based on multi-GPU parallel reasoning accurately assesses the transmission risk level. It also adaptively matches differentiated encryption rules based on the dynamic risk level (low / medium / high) to achieve intelligent resource scheduling. At the same time, it adopts parallel encryption, structured naming and DEFLATE compression technology to significantly improve data processing efficiency while ensuring high security (such as enabling strong encryption in high-risk scenarios), forming an innovative secure transmission system with real-time response capabilities, computing performance optimization, flexible resource allocation and full-link closed-loop protection.
[0040] 23. By integrating data compression, the SM3 national secret algorithm, the Merkle tree, and blockchain technology, an efficient, secure, and reliable data transmission and verification mechanism has been constructed: First, the SM3 algorithm is used to generate a high-strength tamper-proof fingerprint of the data compression block, and the Merkle tree structure is used to improve the efficiency of massive data verification to logarithmic levels; then, the root hash value is anchored to the blockchain, and decentralized evidence storage and timestamp authentication are achieved with the help of a consensus mechanism to ensure the credibility of the data source and its historical traceability; finally, combined with secure communication protocols (such as HTTPS / TLS), compressed blocks are incrementally transmitted, and blockchain evidence credentials (transaction hash, block height, etc.) are transmitted synchronously, significantly optimizing bandwidth utilization (reducing transmission costs), strengthening end-to-end data integrity and non-repudiation, and providing the receiving end with lightweight verification capabilities without the need for third-party intervention. This is particularly suitable for cross-domain trusted exchange scenarios of highly sensitive and large-volume data.
[0041] 24. The streaming computing engine performs millisecond-level pre-processing on multi-dimensional network transmission data streams (including source / destination IP, protocol features, etc.), and combines it with an anomaly recognition model based on multi-GPU parallel reasoning to accurately detect threats. Once abnormal transmission behavior is detected, data reception is immediately interrupted and an alert is automatically pushed, significantly improving the system's response speed and defensive initiative. While reducing the false alarm and missed alarm rates, it effectively blocks potential attack chains, combining the comprehensive advantages of high-precision detection, optimized resource utilization, and automated security protection.
[0042] 25. By comprehensively utilizing the trusted root hash value of the blockchain to ensure the integrity of the data source, combined with DEFLATE efficient compression to save transmission bandwidth, and through block processing to achieve parallel verification, decompression and decryption; at the same time, the date string carried in the file name is used to dynamically generate the decryption key to improve security, and accurate splicing is achieved based on the data block number. Finally, through multiple verification mechanisms (root hash verification, post-decryption verification) and flexible encryption rule matching, while ensuring tamper-proof and strong encryption throughout the entire data transmission process, the reliability, efficiency and scalability of big data transmission are significantly improved.
[0043] 26. By building a highly reliable closed-loop security system, and by real-time recording and encrypted uploading of big data receiving logs containing security monitoring data and blockchain transaction information, an unalterable data evidence is formed; at the same time, based on the incremental data set, abnormal behaviors are dynamically labeled and the abnormal transmission behavior recognition model is optimized, which significantly enhances the data tamper resistance, real-time response capability, intelligent analysis accuracy (adaptive incremental learning model) and resource efficiency, can prevent the risk of man-in-the-middle attacks and data leakage, and continuously improve the accuracy of security protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0045] Figure 1 This is a flow chart of a method for secure transmission of big data that combines blockchain and machine learning. DETAILED DESCRIPTION
[0046] The technical solution in the embodiments of the present application has the following overall idea: through a machine learning model, dynamic perception of transmission risks and real-time identification of abnormal behaviors are achieved to realize adaptive encryption strength adjustment and active defense; at the same time, the blockchain is used to store only the root hash value of the Merkle tree constructed by data fingerprints, while ensuring data integrity and source credibility, avoiding the efficiency bottleneck of directly uploading big data to the chain, thereby improving the security and efficiency of big data transmission under the dual protection of dynamic threat protection and lightweight trusted verification.
[0047] Please refer to Figure 1 As shown, a preferred embodiment of the present invention is a method for securely transmitting big data that combines blockchain and machine learning, comprising the following steps: Step S1: Create a data transmission risk identification model, train the data transmission risk identification model, and then deploy it to the transmission end; create an abnormal transmission behavior identification model, train the abnormal transmission behavior identification model, and then deploy it to the receiving end; Step S2: setting an encryption rule set including standard encryption rules, intermediate encryption rules, and advanced encryption rules, and pre-setting the encryption rule set into the transmitting end and the receiving end; Step S3: The transmission end obtains big data to be transmitted, the data type of which is structured data, semi-structured data or unstructured data, and performs data sharding and preprocessing on the big data to obtain a plurality of data blocks; Step S4: The transmission end obtains real-time security monitoring data, inputs the real-time security monitoring data into a deployed data transmission risk identification model to obtain a transmission risk identification result, matches a corresponding encryption rule from a preset encryption rule set based on the transmission risk identification result, encrypts each data block based on the matched encryption rule to obtain a data encryption block, and compresses each data encryption block to obtain a data compression block; the encryption rule is a standard encryption rule, an intermediate encryption rule, or an advanced encryption rule; Step S5: The transmitting end calculates the data fingerprint of each of the data compression blocks, constructs a Merkle tree based on each of the data fingerprints, uploads the root hash value of the Merkle tree to the blockchain, obtains blockchain transaction information, and sends the blockchain transaction information and each of the data compression blocks to the receiving end in sequence; Step S6: The receiving end receives the transmitted blockchain transaction information and each data compression block in real time, and performs security defense through the deployed abnormal transmission behavior recognition model during the receiving process; Step S7: After the receiving end performs integrity verification on each data compression block using the blockchain transaction information, it decompresses each data compression block to obtain a data encryption block, decrypts the data encryption block using the preset encryption rule set to obtain a data block, and concatenates each data block to obtain big data, thereby completing the transmission of the big data. Step S8: The receiving end records the big data reception log in real time, encrypts the big data reception log into an encrypted log using the encryption rule set, and uploads it to the blockchain, and iteratively optimizes the deployed abnormal transmission behavior recognition model using the big data reception log.
[0048] In step S1, creating a data transmission risk identification model and training the data transmission risk identification model and deploying it to the transmission end are specifically as follows: Create a data transmission risk identification model based on the data input preprocessing layer, risk feature extraction layer, risk feature fusion layer and risk prediction output layer; The data input preprocessing layer is constructed based on a network traffic preprocessing module, a system log preprocessing module, a network device status preprocessing module, a user behavior preprocessing module, a vulnerability scanning preprocessing module, and a threat intelligence preprocessing module; the network traffic preprocessing module, the system log preprocessing module, the network device status preprocessing module, the user behavior preprocessing module, the vulnerability scanning preprocessing module, and the threat intelligence preprocessing module are all constructed based on a normalization unit and an embedding unit; the normalization unit uses Z-score normalization to process numerical data; the embedding unit uses word embedding or positional embedding to process sequence or categorical data, and its specific structure is a linear fully connected layer (dimension 128) plus a ReLU activation function, outputting a unified embedding vector; The network traffic preprocessing module is used to standardize and preliminarily embed network traffic data to obtain a network traffic embedding vector; the system log preprocessing module is used to standardize and preliminarily embed system log data to obtain a system log embedding vector; the network device status preprocessing module is used to standardize and preliminarily embed network device status data to obtain a network device status embedding vector; the user behavior preprocessing module is used to standardize and preliminarily embed user behavior data to obtain a user behavior embedding vector; the vulnerability scanning preprocessing module is used to standardize and preliminarily embed vulnerability scanning data to obtain a vulnerability scanning embedding vector; the threat intelligence preprocessing module is used to standardize and preliminarily embed threat intelligence data to obtain a threat intelligence embedding vector; The risk feature extraction layer is constructed based on a network traffic feature extraction module, a system log feature extraction module, a network device status feature extraction module, a user behavior feature extraction module, a vulnerability scanning feature extraction module, and a threat intelligence feature extraction module. Each of these modules is constructed based on a convolutional neural network unit, a gated recurrent unit, and a concatenation unit. That is, each module comprises a convolutional neural network unit (CNN) and a gated recurrent unit (GRU) to form a dual-channel feature extraction path. The convolutional neural network unit is used to extract local spatial features (e.g., two CNN layers, a convolution kernel size of 3x3, and ReLU activation). The gated recurrent unit is used to extract time series features (e.g., one GRU layer with 128 hidden units). The two-channel outputs are concatenated into a comprehensive feature vector (dimension 256) through a concatenation unit, i.e., the corresponding risk feature. The network traffic feature extraction module is used to extract network traffic risk features from the network traffic embedding vector; the system log feature extraction module is used to extract system log risk features from the system log embedding vector; the network device status feature extraction module is used to extract network device status risk features from the network device status embedding vector; the user behavior feature extraction module is used to extract user behavior risk features from the user behavior embedding vector; the vulnerability scanning feature extraction module is used to extract vulnerability scanning risk features from the vulnerability scanning embedding vector; the threat intelligence feature extraction module is used to extract threat intelligence risk features from the threat intelligence embedding vector; The risk feature fusion layer is used to fuse network traffic risk features, system log risk features, network device status risk features, user behavior risk features, vulnerability scanning risk features, and threat intelligence risk features through a multi-head self-attention mechanism unit to obtain a risk fusion feature. The multi-head self-attention mechanism unit includes a multi-head attention calculation (8 heads, key-query-value mechanism) and a residual connection, and outputs a fused feature vector (dimension 256), namely the risk fusion feature, which is used to capture the global dependency between the outputs of each module in the risk feature extraction layer (for example, vulnerability scanning data may strengthen the risk signal of threat intelligence data), reduce data redundancy and noise, and improve the discriminability of risk features. The residual connection ensures gradient stability. The risk prediction output layer is constructed based on the risk item identification module, the risk level classification module and the result output module; The risk item identification module is used to infer the risk fusion features to obtain the risk item probability distribution; the risk level classification module is used to infer the risk fusion features to obtain the risk level probability distribution; the result output module is used to output the transmission risk identification result carrying the risk item and risk level based on the risk item probability distribution and the risk level probability distribution; the risk level is low risk, medium risk or high risk; the risk item identification module outputs the risk item probability distribution by connecting the output layer (full connection layer dimension 10, corresponding to the number of risk item categories) through a fully connected layer (dimension 128) plus Dropout (dropout rate 0.3) and a softmax activation function; the risk level classification module outputs the risk level probability distribution by connecting the output layer (full connection layer dimension 3, corresponding to the risk level category) through a fully connected layer (dimension 64) plus a softmax activation function. The risk optimization function of the data transmission risk identification model is set to use the Adam optimizer (learning rate 0.001, beta1=0.9, beta2=0.999), which is suitable for multi-task and non-convex optimization problems and accelerates training convergence; The risk loss function of the data transmission risk identification model is set as: ; in, represents the loss value of the risk loss function; Represents the risk term cross entropy loss, using categorical cross entropy loss (Categorical Cross-Entropy Loss); represents the risk level cross entropy loss, using classification cross entropy loss; Both represent weight coefficients; ; ; in, Indicates the real risk item label; represents the predicted probability of the risk item; Indicates the actual risk level label; represents the predicted probability of risky wind turbines; Obtain a large amount of historical security monitoring data, including at least network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data; Network traffic data includes: a. Traffic characteristics: including data packet size, traffic rate, and transmission protocol type (such as TCP, UDP, HTTP, HTTPS, etc.); for example, HTTP traffic for normal web browsing is usually small data packets with high-frequency interactions, while file downloads are characterized by large data packets with longer duration; b. Recording of the source IP address and destination IP address, as well as the port number, of the data packet; by analyzing the flow of traffic, abnormal communication paths can be discovered; for example, if an internal server is found to frequently interact with an unknown external IP address for a large amount of data, it may indicate a data leakage risk; c. Session information: the duration and connection status of each network session (such as established, ongoing, interrupted, etc.); abnormal sessions that have not been disconnected for a long time may indicate the existence of potential malicious connections.
[0049] System log data includes: a. Host logs: operating system logs of servers and terminal devices, which record events such as system startup, shutdown, user login / logout, permission changes, and software installation / uninstallation; for example, frequent failed login attempts may indicate a brute force attack; b. Application logs: logs generated by various network applications (such as databases, web servers, mail servers, etc.); for example, web server logs will record the pages visited by users, request methods (GET, POST, etc.), and return status codes (200 for normal, 404 for page not found, and 500 for internal server error); if 500 errors occur frequently, it may indicate an application vulnerability; c. Security device logs: log records of security devices such as firewalls, intrusion detection systems (IDS), and intrusion prevention systems (IPS); firewall logs will record information on allowed or denied traffic, and IDS / IPS logs will record detected attack behaviors and attack types (such as SQL injection, cross-site scripting attacks, etc.).
[0050] Network device status data includes: a. Device performance indicators: CPU usage, memory usage, interface bandwidth utilization, etc. of network devices (such as routers and switches); if the CPU or memory usage of the device remains high for a long time, it may affect network performance and even cause device failure; b. Device configuration information: The device's network interface configuration (such as IP address, subnet mask, VLAN configuration, etc.), routing table information, etc.; by monitoring changes in configuration information, unauthorized configuration tampering can be discovered in a timely manner; c. Device operation status: The device's online / offline status, port status (such as whether the port is open, whether a fault has occurred, etc.); for example, if a critical network port is found to be suddenly closed, it may affect network connectivity.
[0051] User behavior data includes: a. User identity information: login account, user role, permission level, etc.; by analyzing the match between user identity and behavior, abnormal behavior can be discovered; for example, low-privilege users accessing high-privilege resources may indicate the risk of permission abuse; b. User operation records: various user operations on the network, such as file access, application usage, network resource access, etc.; for example, users frequently visit external suspicious websites or download large numbers of files from unknown sources, which may indicate abnormal user behavior; c. User access pattern: user access time, access frequency, access path, etc.; if users are found to frequently access certain sensitive resources outside of working hours, further investigation may be required.
[0052] Vulnerability scanning data includes: a. Vulnerability detection results: By regularly scanning the network system for vulnerabilities, the types of vulnerabilities found (such as operating system vulnerabilities, application vulnerabilities, configuration vulnerabilities, etc.), the vulnerability level (high risk, medium risk, low risk), and the device or system where the vulnerability is located are recorded; b. Vulnerability repair status: The time and method of vulnerability repair and the verification results after repair are recorded; by continuously monitoring the vulnerability repair status, it can be ensured that the network system eliminates security risks in a timely manner.
[0053] Threat intelligence data includes: a. External threat intelligence: external threat intelligence obtained from security vendors, industry alliances and other channels, including malicious IP address lists, malicious domain name lists, malware signature libraries, etc.; by comparing this intelligence with internal network data, potential external threats can be discovered in a timely manner; b. Internal threat intelligence: threat intelligence generated within the enterprise through security incident analysis, auditing and other means, such as abnormal behavior patterns of internal users, malware propagation paths in the internal network, etc.; this intelligence helps enterprises better deal with internal threats.
[0054] The network traffic data is cleaned at least by removing noise data and duplicate data, including data standardization of unified time format and standardized numerical fields, including feature extraction of session features and traffic features; the system log data is cleaned at least by formatting logs and removing irrelevant logs, including data standardization of unified log levels and unified time formats, including feature extraction of event features and text features; the network device status data is cleaned at least by removing abnormal status data and duplicate data, including data standardization of unified device identification and standardized numerical fields, including feature extraction of performance features and status features; the user behavior data is cleaned at least by formatting logs and removing irrelevant logs, including data standardization of unified log levels and unified time formats, including feature extraction of event features and text features. The data cleaning includes at least removing invalid behavior data and removing duplicate data, including data standardization of unified user identification and unified time format, and extracting behavioral features and time features; the vulnerability scanning data includes at least removing invalid vulnerability data and removing duplicate data, including data standardization of unified vulnerability level and unified time format, and extracting vulnerability features and risk features; the threat intelligence data includes at least removing invalid intelligence data and removing duplicate data, including data standardization of unified threat type and unified time format, and extracting threat features and credibility features, so as to complete the preprocessing of each historical security monitoring data; Constructing a risk data set after labeling each of the preprocessed historical safety monitoring data, including at least risk items and risk levels; dividing the risk data set into a first training set, a first validation set, and a first test set in a ratio of 8:1:1 by a stratified sampling method; The data transmission risk identification model is trained using the first training set, risk optimization function, and risk loss function until the preset first early stopping condition is met. The trained data transmission risk identification model is verified and tested using the first validation set and the first test set, respectively. The data transmission risk identification model that passes the test is deployed to the transmission end.
[0055] In step S1, creating an abnormal transmission behavior recognition model and training the abnormal transmission behavior recognition model and deploying it to the receiving end are specifically as follows: An abnormal transmission behavior recognition model is created based on the abnormal feature extraction layer, the abnormal feature fusion layer and the abnormal prediction output layer; The abnormal feature extraction layer is constructed based on a multi-scale temporal channel, a spatial dependency channel, a behavioral pattern channel, and a feature aggregation module; the multi-scale temporal channel is used to extract multi-scale temporal features from the input network transmission data stream through a bidirectional gated recurrent unit and a dilated convolutional layer; the bidirectional gated recurrent unit (Bi-GRU) extracts bidirectional long-term and short-term dependencies (such as burst traffic duration) in the network transmission data stream, and the dilated convolutional layer (dilation=1,2,4) captures periodic patterns of different granularities (such as DDoS attack pulses) from the network transmission data stream, and outputs multi-scale temporal features including bidirectional long-term and short-term dependencies and periodic patterns; the spatial dependency channel is used to extract spatial dependencies from the input network transmission data stream through a graph convolutional network unit. Features; that is, constructing an IP-port topology graph through a graph convolutional network unit (GCN), learning inter-node communication patterns through an adjacency matrix, and identifying unconventional connection patterns (such as abnormal port access in scanning behavior); the behavior pattern channel is used to extract behavior pattern features from the input network transmission data stream through a self-attention mechanism and a statistical feature extractor; the self-attention mechanism is used to focus on key behavior fragments (such as abnormal TCP flag sequences), and the statistical feature extractor is used to calculate non-stationary indicators such as entropy and variance in real time; the feature aggregation module is used to aggregate multi-scale temporal features, spatial dependency features, and behavior pattern features through a gated feature fusion unit to obtain contextual features; that is, using a learnable weight gating mechanism to dynamically weight and fuse the three-channel output features; The abnormal feature fusion layer is used to perform cross-modal fusion of the contextual features through a hierarchical attention mechanism and a spatiotemporal compression module to obtain a comprehensive behavioral representation. The hierarchical attention mechanism uses a multi-level attention network. The first level of attention allocates feature importance in the temporal dimension (such as weighting the period of sudden traffic increase), and the second level of attention strengthens the association of abnormal nodes in the spatial dimension (such as C&C communication nodes). The spatiotemporal compression module uses an adaptive pooling layer to compress variable-length sequences into fixed-dimensional behavioral fingerprint vectors while preserving the statistical distribution characteristics of time series features (such as the skewness / kurtosis of traffic distribution). The anomaly prediction output layer is used to map the comprehensive behavior representation to the abnormal behavior probability distribution and classify and output the transmission behavior identification results. The anomaly prediction output layer is constructed based on a classifier and a confidence calibration module. The classifier adopts a network structure with a fully connected layer and a GELU activation function. The main branch predicts normal / abnormal, and the auxiliary branch outputs the anomaly type (such as DDoS / port scanning). The confidence calibration module adopts a temperature scaling layer to adjust the output probability distribution through a learnable temperature parameter and solve the confidence offset problem in class imbalance scenarios. The behavior optimization function of the abnormal transmission behavior recognition model is set to use Lookahead optimizer and RAdam optimizer; learning rate: cosine annealing scheduling (initial value 0.001); gradient clipping threshold: 1.0; Setting the behavior loss function of the abnormal transmission behavior recognition model to adopt classification cross entropy loss; Obtain a large amount of historical network transmission data flow including at least source IP address, destination IP address, port, communication protocol, packet size, transmission interval and identification bit distribution; Each of the historical network transmission data streams is preprocessed, including at least data cleaning, feature extraction and conversion, data dimensionality reduction, data bucketing, and data formatting; data cleaning includes at least removing duplicate data (checking whether there are duplicate records in the data set; if there are duplicate records, the duplicates can be deleted and only one record is retained), processing missing values (checking whether there are missing values in the data set; the following methods can be used to process missing values: a. delete records containing missing values; b. use the mean, median, or mode to fill missing values; c. for categorical features (such as communication protocols), fill missing values with the most common category), and filtering invalid data (checking whether the data is logically consistent, such as whether the port number is within the range of 0-65535, whether the IP address is in the correct format, etc. If invalid data is found, it can be deleted or corrected); data bucketing means that for certain features (such as packet size and transmission interval), they can be binned to simplify the data, for example, packet size can be divided into three intervals of "small", "medium", and "large", and transmission interval can be divided into three intervals of "short", "medium", and "long"; After preprocessing, each of the historical network transmission data streams is annotated with at least abnormal behaviors to construct a behavior dataset; the behavior dataset is divided into a second training set, a second validation set, and a second test set in a ratio of 8:1:1 by a stratified sampling method; The abnormal transmission behavior recognition model is trained using the second training set, behavior optimization function, and behavior loss function until the preset second early stopping condition is met. The trained abnormal transmission behavior recognition model is verified and tested using the second verification set and the second test set, respectively. The abnormal transmission behavior recognition model that passes the test is deployed to the receiving end.
[0056] In step S2, the standard encryption rule is specifically: Obtain the current date string (in the format YYYYMMDD), perform an MD5 hash calculation on the date string to obtain a hash value H1, and extract the first 8 bytes from the hash value H1 as the dynamic key K1; convert the large data to be encrypted into a UTF-8 encoded byte array B1, generate an 8-byte dynamic salt value Y1, append the dynamic salt value Y1 to the front of the byte array B1 to obtain augmented data S1, call the dynamic key K1 using the DES algorithm to encrypt the augmented data S1 to obtain ciphertext data C1, and perform Base64 encoding on the ciphertext data C1 to obtain a data encryption block; The DES algorithm is relatively weak, but the dynamic key increases variability based on the date; the salt value is added to prevent replay attacks; the steps are logically rigorous: date hash → key generation → data salt value augmentation → DES encryption → Base64 encoding, which is suitable for large data transmission in low-risk situations or low-sensitivity data processing.
[0057] The intermediate encryption rules are specifically as follows: Get the current date string, perform SHA-256 hash calculation on the date string to obtain a 32-byte hash value H2, use the first 16 bytes of the hash value H2 as the dynamic key K2, and the last 16 bytes as the dynamic key K3; convert the large data to be encrypted into a UTF-8 encoded byte array B2, divide the byte array B2 into sub-arrays B21 and B22 (if the data size is an odd number, sub-array B22 has one more byte), and use the dynamic key K2 to encrypt sub-array B21 using the AES-128 algorithm to obtain the encrypted block EA 1. Encrypt subarray B22 using the dynamic key K3 using the Blowfish algorithm to obtain an encrypted block EB1. Concatenate the encrypted blocks EA1 and EB1 to obtain concatenated data P1. Circularly shift the concatenated data P1 by 5 bits in units of bytes to obtain transformed data T1. Perform a position swap operation on every two consecutive bytes in the transformed data T1 (i.e., swap byte[i] with byte[i+1], with indexing starting at 0) to obtain ciphertext data C2. Base32 encode the ciphertext data C2 to obtain a data encryption block. SHA-256 and AES enhance hashing and encryption strength; data segmentation and two encryption algorithms provide multi-layer protection; cyclic shift and byte swapping add obfuscation; step logic: date hashing → key segmentation → data segmentation → parallel encryption → concatenation → shift transformation → byte swapping; suitable for large data transmission in medium-risk situations or medium-sensitive data, such as configuration files.
[0058] The advanced encryption rules are specifically as follows: Get the current date string, perform SHA-512 hash calculation on the date string to obtain a 64-byte hash value H3, use the first 32 bytes of the hash value H3 as the dynamic key K4, the middle 16 bytes as the dynamic key K5, and the last 16 bytes as the dynamic salt value Y2; convert the large data to be encrypted into a UTF-8 encoded byte array B3, divide the byte array B3 into sub-arrays B31, B32, and B33 (if the size is not divisible, sub-array B33 is slightly larger), use the AES-256 algorithm to call the dynamic key K4 to encrypt the sub-array B31 to obtain the encrypted block EA2, and use the 3DES algorithm to call the dynamic key K5 Subarray B32 is encrypted to obtain an encrypted block EB2. A byte-by-byte XOR operation is performed on the subarray B32 and the dynamic salt value Y2 to obtain an encrypted block EC2. The encrypted blocks EA2, EB2, and EC2 are concatenated to obtain concatenated data P2. The upper 4 bits and the lower 4 bits of each byte in the concatenated data P2 are swapped (i.e., bits [7-4] are swapped with bits [3-0]) to obtain transformed data Q. The transformed data Q is divided into 16-byte blocks. The 16-byte blocks are reordered in a preset reverse order (e.g., block 1 → block n, which is reversed to block n → block 1) to obtain ciphertext data C3. The ciphertext data C3 is Base16 encoded to obtain a data encryption block.
[0059] SHA-512 and AES-256 provide high-strength hashing and encryption; triple block processing (AES, 3DES, XOR salt) increases complexity; bit-level obfuscation and block permutation add deep obfuscation and diffusion; step logic: date hash → key salt split → data split → parallel encryption and XOR → splicing → bit obfuscation → block permutation; suitable for large data transmission in high-risk situations or highly sensitive data such as passwords or identity information.
[0060] Convert the big data to be encrypted into a UTF-8 encoded byte array, that is, convert structured data, semi-structured data or unstructured data into a byte array first to ensure universality.
[0061] The step S3 is specifically as follows: The transmission end obtains big data to be transmitted, where the data type of the big data is structured data, semi-structured data or unstructured data; the big data is segmented based on a preset file size to obtain a number of data blocks, and each of the data blocks is preprocessed, including at least data cleaning and data standardization.
[0062] The step S4 is specifically as follows: The transmission end obtains real-time security monitoring data including at least network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data, performs preprocessing on each of the real-time security monitoring data including at least data cleaning and data standardization through a streaming computing engine, and inputs the preprocessed real-time security monitoring data into a deployed data transmission risk identification model. The data transmission risk identification model performs parallel inference on multiple GPUs to output a transmission risk identification result containing risk items and risk levels; parsing the transmission risk identification result, and matching the standard encryption rule from the encryption rule set when the risk level is low risk; matching the intermediate encryption rule from the encryption rule set when the risk level is medium risk; and matching the advanced encryption rule from the encryption rule set when the risk level is high risk; Based on the matching encryption rules, each data block is encrypted in parallel to obtain a corresponding data encryption block, each data encryption block is named based on the current date string and the data block number, and each data encryption block is compressed using the DEFLATE algorithm to obtain a data compression block; the encryption rule is a standard encryption rule, an intermediate encryption rule or an advanced encryption rule.
[0063] The DEFLATE algorithm is a lossless data compression algorithm that combines the LZ77 algorithm and Huffman coding; LZ77 compression finds repeated strings in the input data and uses pointers and lengths to represent repeated patterns; Huffman coding encodes the dictionary matches generated by LZ77 compression to further compress the data.
[0064] The step S5 is specifically as follows: The transmitting end calculates the data fingerprint of each of the data compression blocks through the SM3 algorithm, constructs a Merkle tree based on the data fingerprint of each of the data compression blocks, uploads the root hash value of the Merkle tree to the blockchain, waits for the root hash value to be passed by the blockchain consensus, obtains the blockchain transaction information feedback including at least the transaction hash, block height and block hash, and sends the blockchain transaction information and each data compression block to the receiving end in sequence through a secure communication protocol.
[0065] The basic data for data fingerprint calculation includes the content of the data compression block and the file name.
[0066] A Merkle Tree, also known as a hash tree, is a tree-like data structure based on a hash function. It aggregates the hash values of a large amount of data layer by layer through a hierarchical hashing method, ultimately generating a unique root hash value (Merkle Root), thereby efficiently verifying the integrity and consistency of the data.
[0067] The step S6 is specifically as follows: The receiving end receives each of the transmitted data compression blocks in real time, and during the receiving process, collects a real-time network transmission data stream including at least the source IP address, destination IP address, port, communication protocol, packet size, transmission interval, and identification bit distribution, performs preprocessing on each of the real-time network transmission data streams through a streaming computing engine, including at least data cleaning and data standardization, and inputs each of the preprocessed real-time network transmission data streams into a deployed abnormal transmission behavior recognition model. The abnormal transmission behavior recognition model performs parallel inference on multiple GPUs to output transmission behavior recognition results; The receiving end analyzes the transmission behavior identification result. When the transmission behavior identification result carries abnormal transmission behavior, the receiving end stops receiving the data compression block and pushes an early warning notification to a pre-associated management terminal for security defense.
[0068] The step S7 is specifically as follows: The receiving end obtains a root hash value from the blockchain through the blockchain transaction information, performs integrity verification on each data compression block based on the root hash value, decompresses each data compression block through the DEFLATE algorithm to obtain a data encryption block, matches encryption rules from a preset encryption rule set based on the encoding format of the data encryption block, decrypts and verifies each data encryption block based on the matched encryption rule and the date character string carried in the file name of the data encryption block to obtain a data block, and splices each data block based on the data block number carried in the file name of the data encryption block to obtain big data, thereby completing the transmission of big data.
[0069] The step S8 is specifically as follows: The receiving end records in real time a big data reception log including at least real-time security monitoring data, reception time, and blockchain transaction information, encrypts the big data reception log into an encrypted log using the advanced encryption rules in the encryption rule set, and uploads the encrypted log to the blockchain; constructs an incremental data set using the big data reception log, annotates the incremental data set for abnormal behavior, and iteratively optimizes the deployed abnormal transmission behavior identification model.
[0070] The big data receiving log also includes data block feature information, transmission behavior details, system resource status, identity authentication information, and threat intelligence related data; the data block feature information includes: a. Data block check value: hash value of each received data block (such as SM3 value), which is used to compare with the original fingerprint of the transmission end; b. Compression / encryption metadata: compression rate, encryption rule type (standard / intermediate / advanced), dynamic key generation parameters (such as date string); c. Data block integrity status: verification result (success / failure) and failure reason (such as hash mismatch); the transmission behavior details include: a. Network layer indicators: transmission delay, packet loss rate, bandwidth occupancy rate; b. Abnormal behavior mark: output of abnormal transmission behavior recognition model The original risk score and specific features that trigger the warning (such as abnormal port access frequency); c. Defense action record: operations triggered by security defense (such as pausing reception and requesting data block retransmission); the system resource status includes: a. Receiving end load: CPU / memory usage, disk I / O speed; b. Decryption / decompression performance: single data block processing time, number of parallel tasks; the identity authentication information includes: a. Digital certificate fingerprint of the transmitting / receiving end; b. TLS / SSL protocol version and key exchange algorithm of the communication session; the threat intelligence-related data includes: a. Real-time threat intelligence matching results (such as whether the source IP is in the known malicious IP library); b. CVE number related to the received data in the vulnerability scan results.
[0071] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for secure transmission of big data combining blockchain and machine learning, characterized by: The steps include: Step S1: Create a data transmission risk identification model, train the data transmission risk identification model, and then deploy it to the transmission end; Creating an abnormal transmission behavior recognition model, training the abnormal transmission behavior recognition model, and deploying it to a receiving end; Step S2: setting an encryption rule set including standard encryption rules, intermediate encryption rules, and advanced encryption rules, and pre-setting the encryption rule set into the transmitting end and the receiving end; Step S3: The transmission end obtains big data to be transmitted, the data type of which is structured data, semi-structured data or unstructured data, and performs data sharding and preprocessing on the big data to obtain a plurality of data blocks; Step S4: The transmission end obtains real-time security monitoring data, inputs the real-time security monitoring data into a deployed data transmission risk identification model to obtain a transmission risk identification result, matches a corresponding encryption rule from a preset encryption rule set based on the transmission risk identification result, encrypts each data block based on the matched encryption rule to obtain a data encryption block, and compresses each data encryption block to obtain a data compression block; the encryption rule is a standard encryption rule, an intermediate encryption rule, or an advanced encryption rule; Step S5: The transmitting end calculates the data fingerprint of each of the data compression blocks, constructs a Merkle tree based on each of the data fingerprints, uploads the root hash value of the Merkle tree to the blockchain, obtains blockchain transaction information, and sends the blockchain transaction information and each of the data compression blocks to the receiving end in sequence; Step S6: The receiving end receives the transmitted blockchain transaction information and each data compression block in real time, and performs security defense through the deployed abnormal transmission behavior recognition model during the receiving process; Step S7: After the receiving end performs integrity verification on each data compression block using the blockchain transaction information, it decompresses each data compression block to obtain a data encryption block, decrypts the data encryption block using the preset encryption rule set to obtain a data block, and concatenates each data block to obtain big data, thereby completing the transmission of the big data. Step S8: The receiving end records the big data reception log in real time, encrypts the big data reception log into an encrypted log using the encryption rule set, and uploads it to the blockchain, and iteratively optimizes the deployed abnormal transmission behavior recognition model using the big data reception log.
2. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: In step S1, creating a data transmission risk identification model and training the data transmission risk identification model and deploying it to the transmission end are specifically as follows: Create a data transmission risk identification model based on the data input preprocessing layer, risk feature extraction layer, risk feature fusion layer and risk prediction output layer; The data input preprocessing layer is constructed based on a network traffic preprocessing module, a system log preprocessing module, a network device status preprocessing module, a user behavior preprocessing module, a vulnerability scanning preprocessing module, and a threat intelligence preprocessing module; the network traffic preprocessing module, the system log preprocessing module, the network device status preprocessing module, the user behavior preprocessing module, the vulnerability scanning preprocessing module, and the threat intelligence preprocessing module are all constructed based on standardized units and embedded units; The network traffic preprocessing module is used to standardize and preliminarily embed network traffic data to obtain a network traffic embedding vector; The system log preprocessing module is used to standardize and preliminarily embed system log data to obtain a system log embedding vector; the network device status preprocessing module is used to standardize and preliminarily embed network device status data to obtain a network device status embedding vector; the user behavior preprocessing module is used to standardize and preliminarily embed user behavior data to obtain a user behavior embedding vector; the vulnerability scanning preprocessing module is used to standardize and preliminarily embed vulnerability scanning data to obtain a vulnerability scanning embedding vector; the threat intelligence preprocessing module is used to standardize and preliminarily embed threat intelligence data to obtain a threat intelligence embedding vector; The risk feature extraction layer is constructed based on a network traffic feature extraction module, a system log feature extraction module, a network device status feature extraction module, a user behavior feature extraction module, a vulnerability scanning feature extraction module, and a threat intelligence feature extraction module; the network traffic feature extraction module, the system log feature extraction module, the network device status feature extraction module, the user behavior feature extraction module, the vulnerability scanning feature extraction module, and the threat intelligence feature extraction module are all constructed based on a convolutional neural network unit, a gated recurrent unit, and a splicing unit; The network traffic feature extraction module is used to extract network traffic risk features from the network traffic embedding vector; the system log feature extraction module is used to extract system log risk features from the system log embedding vector; the network device status feature extraction module is used to extract network device status risk features from the network device status embedding vector; the user behavior feature extraction module is used to extract user behavior risk features from the user behavior embedding vector; the vulnerability scanning feature extraction module is used to extract vulnerability scanning risk features from the vulnerability scanning embedding vector; the threat intelligence feature extraction module is used to extract threat intelligence risk features from the threat intelligence embedding vector; The risk feature fusion layer is used to fuse network traffic risk features, system log risk features, network device status risk features, user behavior risk features, vulnerability scanning risk features, and threat intelligence risk features through a multi-head self-attention mechanism unit to obtain a risk fusion feature; The risk prediction output layer is constructed based on the risk item identification module, the risk level classification module and the result output module; The risk item identification module is used to infer the risk fusion features to obtain the risk item probability distribution; the risk level classification module is used to infer the risk fusion features to obtain the risk level probability distribution; the result output module is used to output the transmission risk identification result carrying the risk item and risk level based on the risk item probability distribution and the risk level probability distribution; the risk level is low risk, medium risk or high risk; Setting the risk optimization function of the data transmission risk identification model to adopt Adam optimizer; The risk loss function of the data transmission risk identification model is set as: ; in, represents the loss value of the risk loss function; Represents the risk term cross entropy loss, using classification cross entropy loss; represents the risk level cross entropy loss, using classification cross entropy loss; Both represent weight coefficients; Obtain a large amount of historical security monitoring data, including at least network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data; The network traffic data is cleaned at least by removing noise data and duplicate data, including data standardization of unified time format and standardized numerical fields, including feature extraction of session features and traffic features; the system log data is cleaned at least by formatting logs and removing irrelevant logs, including data standardization of unified log levels and unified time formats, including feature extraction of event features and text features; the network device status data is cleaned at least by removing abnormal status data and duplicate data, including data standardization of unified device identification and standardized numerical fields, including feature extraction of performance features and status features; the user behavior data is cleaned at least by formatting logs and removing irrelevant logs, including data standardization of unified log levels and unified time formats, including feature extraction of event features and text features. The data cleaning includes at least removing invalid behavior data and removing duplicate data, including data standardization of unified user identification and unified time format, and extracting behavioral features and time features; the vulnerability scanning data includes at least removing invalid vulnerability data and removing duplicate data, including data standardization of unified vulnerability level and unified time format, and extracting vulnerability features and risk features; the threat intelligence data includes at least removing invalid intelligence data and removing duplicate data, including data standardization of unified threat type and unified time format, and extracting threat features and credibility features, so as to complete the preprocessing of each historical security monitoring data; Constructing a risk data set after labeling each of the preprocessed historical safety monitoring data, including at least risk items and risk levels; dividing the risk data set into a first training set, a first validation set, and a first test set in a ratio of 8:1:1 by a stratified sampling method; The data transmission risk identification model is trained using the first training set, risk optimization function, and risk loss function until the preset first early stopping condition is met. The trained data transmission risk identification model is verified and tested using the first validation set and the first test set, respectively. The data transmission risk identification model that passes the test is deployed to the transmission end.
3. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: In step S1, creating an abnormal transmission behavior recognition model and training the abnormal transmission behavior recognition model and deploying it to the receiving end are specifically as follows: An abnormal transmission behavior recognition model is created based on the abnormal feature extraction layer, the abnormal feature fusion layer and the abnormal prediction output layer; The abnormal feature extraction layer is constructed based on a multi-scale temporal channel, a spatial dependency channel, a behavioral pattern channel, and a feature aggregation module; the multi-scale temporal channel is used to extract multi-scale temporal features from the input network transmission data stream through a bidirectional gated recurrent unit and a dilated convolutional layer; the spatial dependency channel is used to extract spatial dependency features from the input network transmission data stream through a graph convolutional network unit; the behavioral pattern channel is used to extract behavioral pattern features from the input network transmission data stream through a self-attention mechanism and a statistical feature extractor; The feature aggregation module is used to aggregate multi-scale temporal features, spatial dependency features, and behavioral pattern features through a gated feature fusion unit to obtain contextual features; The abnormal feature fusion layer is used to perform cross-modal fusion of the context features through a hierarchical attention mechanism and a spatiotemporal compression module to obtain a comprehensive behavior representation; The anomaly prediction output layer is used to map the comprehensive behavior representation to the abnormal behavior probability distribution and classify and output the transmission behavior recognition results; Setting the behavior optimization function of the abnormal transmission behavior identification model to adopt Lookahead optimizer and RAdam optimizer; Setting the behavior loss function of the abnormal transmission behavior recognition model to adopt classification cross entropy loss; Obtain a large amount of historical network transmission data flow including at least source IP address, destination IP address, port, communication protocol, packet size, transmission interval and identification bit distribution; Performing preprocessing on each of the historical network transmission data streams, including at least data cleaning, feature extraction and conversion, data dimensionality reduction, data bucketing, and data formatting; Constructing a behavior data set after marking each of the pre-processed historical network transmission data streams at least including abnormal behaviors; Dividing the behavioral dataset into a second training set, a second validation set, and a second test set in a ratio of 8:1:1 by a stratified sampling method; The abnormal transmission behavior recognition model is trained using the second training set, behavior optimization function, and behavior loss function until the preset second early stopping condition is met. The trained abnormal transmission behavior recognition model is verified and tested using the second verification set and the second test set, respectively. The abnormal transmission behavior recognition model that passes the test is deployed to the receiving end.
4. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: In step S2, the standard encryption rule is specifically: Obtain the current date string, perform MD5 hash calculation on the date string to obtain a hash value H1, extract the first 8 bytes from the hash value H1 as the dynamic key K1; convert the large data to be encrypted into a UTF-8 encoded byte array B1, generate an 8-byte dynamic salt value Y1, append the dynamic salt value Y1 to the front of the byte array B1 to obtain augmented data S1, call the dynamic key K1 using the DES algorithm to encrypt the augmented data S1 to obtain ciphertext data C1, and perform Base64 encoding on the ciphertext data C1 to obtain a data encryption block; The intermediate encryption rules are specifically as follows: Obtain the current date string, perform a SHA-256 hash calculation on the date string to obtain a 32-byte hash value H2, use the first 16 bytes of the hash value H2 as the dynamic key K2, and the last 16 bytes as the dynamic key K3; convert the large data to be encrypted into a UTF-8 encoded byte array B2, divide the byte array B2 into sub-arrays B21 and B22, call the dynamic key K2 using the AES-128 algorithm to encrypt the sub-array B21 to obtain an encrypted block EA1, call the dynamic key K3 using the Blowfish algorithm to encrypt the sub-array B22 to obtain an encrypted block EB1, concatenate the encrypted blocks EA1 and EB1 to obtain concatenated data P1, cyclically shift the concatenated data P1 by 5 bits in units of bytes to obtain transformed data T1, perform a position swap operation on every two consecutive bytes in the transformed data T1 to obtain ciphertext data C2, and perform Base32 encoding on the ciphertext data C2 to obtain a data encryption block; The advanced encryption rules are specifically as follows: Get the current date string, perform SHA-512 hash calculation on the date string to obtain a 64-byte hash value H3, use the first 32 bytes of the hash value H3 as the dynamic key K4, the middle 16 bytes as the dynamic key K5, and the last 16 bytes as the dynamic salt value Y2; convert the large data to be encrypted into a UTF-8 encoded byte array B3, divide the byte array B3 into sub-arrays B31, B32, and B33, use the dynamic key K4 to encrypt the sub-array B31 through the AES-256 algorithm to obtain the encrypted block EA2, and use the 3DES algorithm to call the encrypted block EA2. The subarray B32 is encrypted using the dynamic key K5 to obtain an encrypted block EB2. A byte-by-byte XOR operation is performed on the subarray B32 and the dynamic salt value Y2 to obtain an encrypted block EC2. The encrypted blocks EA2, EB2, and EC2 are concatenated to obtain concatenated data P2. The upper 4 bits and the lower 4 bits of each byte in the concatenated data P2 are swapped to obtain transformed data Q. The transformed data Q is divided into 16-byte blocks. Each of the 16-byte blocks is reordered in a preset reverse order to obtain ciphertext data C3. The ciphertext data C3 is Base16 encoded to obtain a data encryption block.
5. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: The step S3 is specifically as follows: The transmission end obtains big data to be transmitted, where the data type of the big data is structured data, semi-structured data or unstructured data; the big data is segmented based on a preset file size to obtain a number of data blocks, and each of the data blocks is preprocessed, including at least data cleaning and data standardization.
6. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: The step S4 is specifically as follows: The transmission end obtains real-time security monitoring data including at least network traffic data, system log data, network device status data, user behavior data, vulnerability scanning data, and threat intelligence data, performs preprocessing on each of the real-time security monitoring data including at least data cleaning and data standardization through a streaming computing engine, and inputs the preprocessed real-time security monitoring data into a deployed data transmission risk identification model. The data transmission risk identification model performs parallel inference on multiple GPUs to output a transmission risk identification result containing risk items and risk levels; parsing the transmission risk identification result, and matching the standard encryption rule from the encryption rule set when the risk level is low risk; matching the intermediate encryption rule from the encryption rule set when the risk level is medium risk; and matching the advanced encryption rule from the encryption rule set when the risk level is high risk; Based on the matching encryption rules, each data block is encrypted in parallel to obtain a corresponding data encryption block, each data encryption block is named based on the current date string and the data block number, and each data encryption block is compressed using the DEFLATE algorithm to obtain a data compression block; the encryption rule is a standard encryption rule, an intermediate encryption rule or an advanced encryption rule.
7. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: The step S5 is specifically as follows: The transmitting end calculates the data fingerprint of each of the data compression blocks through the SM3 algorithm, constructs a Merkle tree based on the data fingerprint of each of the data compression blocks, uploads the root hash value of the Merkle tree to the blockchain, waits for the root hash value to be passed by the blockchain consensus, obtains the blockchain transaction information feedback including at least the transaction hash, block height and block hash, and sends the blockchain transaction information and each data compression block to the receiving end in sequence through a secure communication protocol.
8. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: The step S6 is specifically as follows: The receiving end receives each of the transmitted data compression blocks in real time, and during the receiving process, collects a real-time network transmission data stream including at least the source IP address, destination IP address, port, communication protocol, packet size, transmission interval, and identification bit distribution, performs preprocessing on each of the real-time network transmission data streams through a streaming computing engine, including at least data cleaning and data standardization, and inputs each of the preprocessed real-time network transmission data streams into a deployed abnormal transmission behavior recognition model. The abnormal transmission behavior recognition model performs parallel inference on multiple GPUs to output transmission behavior recognition results; The receiving end analyzes the transmission behavior identification result. When the transmission behavior identification result carries abnormal transmission behavior, the receiving end stops receiving the data compression block and pushes an early warning notification to a pre-associated management terminal for security defense.
9. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: The step S7 is specifically as follows: The receiving end obtains a root hash value from the blockchain through the blockchain transaction information, performs integrity verification on each data compression block based on the root hash value, decompresses each data compression block through the DEFLATE algorithm to obtain a data encryption block, matches encryption rules from a preset encryption rule set based on the encoding format of the data encryption block, decrypts and verifies each data encryption block based on the matched encryption rule and the date character string carried in the file name of the data encryption block to obtain a data block, and splices each data block based on the data block number carried in the file name of the data encryption block to obtain big data, thereby completing the transmission of big data.
10. The method for secure transmission of big data combining blockchain and machine learning as claimed in claim 1, characterized in that: The step S8 is specifically as follows: The receiving end records in real time a big data reception log including at least real-time security monitoring data, reception time, and blockchain transaction information, encrypts the big data reception log into an encrypted log using the advanced encryption rules in the encryption rule set, and uploads the encrypted log to the blockchain; An incremental data set is constructed by using the big data reception log, abnormal behaviors are annotated on the incremental data set, and then the deployed abnormal transmission behavior recognition model is iteratively optimized.
Citation Information
Patent Citations
Data transceiving method, sending end and receiving end
CN112671745A
Data transmission method and device, storage medium and electronic equipment
CN114363888A
Malicious traffic identification method and system based on data enhancement and feature fusion
CN116318928A
File encryption storage method, file decryption method and file encryption and decryption storage system
CN117708854A
Big data control system based on artificial intelligence and block chain
CN118378670A
Cited By
Intelligent terminal identity authentication and data security management method, system and device based on block chain, and medium
CN121000383A
Pipeline construction process based on digital-intelligent collaborative scheduling
CN121051775A
Power distribution terminal key management method and system based on trusted computing
CN121690575A