Data communication method and system based on information security
By adopting identity authentication, data encryption, real-time intrusion detection and automatic defense mechanisms in the data communication system, combining machine learning and reinforcement learning technology, dynamically adjusting defense strategies and resource allocation, the problems of insufficient coordination and flexibility of intrusion detection accuracy and defense mechanisms in the existing technology are solved, and more efficient network attack defense and data communication security guarantees are achieved.
Patent Information
- Application Number
- CN202510192160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing data communication technologies have shortcomings in intrusion detection accuracy and coordination and flexibility of defense mechanisms, making it difficult to effectively deal with complex and diverse cyber attacks.
The data communication method based on information security is adopted, through identity authentication, data encryption, real-time intrusion detection and automatic defense mechanisms, combined with machine learning and reinforcement learning technology, defense strategies and resource allocation are dynamically adjusted to achieve comprehensive monitoring and precise defense of network traffic.
It improves the accuracy and efficiency of intrusion detection, enhances the coordination and flexibility of defense mechanisms, can more effectively deal with complex and diverse network attacks, and ensures the security and reliability of data communication.
Smart Images

Figure CN119675999B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data communication, and in particular to a data communication method and system based on information security. Background Art
[0002] In the digital age, data communication technology is the key to the fast and accurate transmission of information, involving communication protocols, encryption algorithms, network architecture and security mechanisms. As the basis of Internet communication, the TCP / IP protocol family ensures reliable data transmission and addressing routing; encryption technologies such as AES (symmetric encryption) and RSA (asymmetric encryption) ensure the confidentiality and integrity of data transmission; network architecture has evolved from traditional LAN (local area network) and WAN (wide area network) to cloud computing and SDN (software defined network), improving the flexibility and resource utilization of the network; security mechanisms such as firewalls, IDS (intrusion detection system) and IPS (intrusion prevention system) protect network security and prevent unauthorized access and intrusion. These technologies together constitute the framework of modern data communication, ensuring the efficient and secure transmission of information.
[0003] Although existing data communication technologies have made remarkable achievements in ensuring information transmission, they still face a series of severe challenges and problems as network attack methods become increasingly complex and diverse.
[0004] 1.1 Accuracy and efficiency of intrusion detection
[0005] Traditional intrusion detection systems (IDS) face challenges in terms of accuracy and efficiency, which are mainly reflected in the following aspects: First, the static detection mode based on predefined rules is difficult to deal with new attacks, especially zero-day attacks, because their features have not yet been included in the feature library of IDS and are easily missed. Secondly, IDS mainly focuses on conventional features in data packet feature extraction, lacks in-depth analysis of application layer protocols and comprehensive perception of network situations, resulting in a high false alarm rate. Finally, IDS lacks real-time and adaptive capabilities, and may experience delays during peak hours or when facing large-scale network attacks. It is difficult to quickly adapt to the new network environment, reducing the efficiency and accuracy of intrusion detection.
[0006] 2. Issues of coordination and flexibility of defense mechanisms
[0007] Network security devices such as firewalls, IDS, IPS, etc. usually operate independently in the defense system, lacking effective coordination, resulting in defense loopholes. For example, after the IDS detects an intrusion, the firewall may not be able to adjust the strategy in time to effectively block it, or the linkage between the IPS and IDS is not tight enough, resulting in a delayed response to the attack. In addition, existing defense strategies are often based on fixed rules and configurations, which are difficult to flexibly adjust according to real-time attack situations. For new attacks or rapidly changing attack patterns, traditional defense strategies may not be able to respond in a timely and effective manner. In terms of isolation technology, most existing isolation technologies use coarse-grained methods, such as isolating the entire network segment or device, which may have a significant impact on normal business operations. In cloud computing and virtualization environments, this coarse-grained isolation method cannot meet the needs of fine isolation of different applications or services. Summary of the invention
[0008] The main purpose of the present invention is to provide a data communication method and system based on information security, which can effectively solve the problems mentioned in the background technology.
[0009] To achieve the above object, the technical solution adopted by the present invention is:
[0010] A data communication method based on information security comprises the following steps:
[0011] S1. Preparation
[0012] S1.1. The communicating parties perform identity authentication. The initiator sends a digital certificate and session identifier. After the receiver verifies the legitimacy of the certificate, it encrypts the symmetric key with the initiator's public key and sends it back to complete the key exchange.
[0013] S2, transmission stage
[0014] S2.1. The sender uses a symmetric key and AES algorithm to encrypt and group data. Prior to this, the sender first uses a hash function to calculate the original data, generates a unique hash value and attaches it. The sender then adds header information to each group of encrypted data and encapsulates the data, which is then transmitted over a secure network connection.
[0015] S3: Intrusion Detection and Prevention Phase
[0016] S3.1, IDS monitors network traffic in real time, extracts data packet features and compares them with normal models, and marks suspicious data packets.
[0017] S3.2. The machine learning model analyzes suspicious data packets to determine whether they are intrusion behaviors.
[0018] S3.3. After the intrusion is confirmed, the automatic defense mechanism is triggered, including blocking the attack source, notifying both communicating parties, preparing data backup and recovery, and dynamically adjusting the defense strategy.
[0019] S4, Receiving and Verifying Phase
[0020] S4.1. The receiver decapsulates the data, verifies and reassembles it according to the sequence number, and notifies the sender to retransmit if there is any problem.
[0021] S4.2. Decrypt the data with the symmetric key, calculate the hash value and compare it with the hash value attached by the sender during data transmission to verify integrity.
[0022] S5: Ending and cleaning phase
[0023] S5.1. After both parties confirm the communication is complete, they release and clean up resources.
[0024] S5.2. The system records key events in the communication process and generates logs, which are regularly analyzed to optimize the system and strengthen defense, and to update intrusion detection models.
[0025] Preferably, the communication initiator in S1 first sends an identity authentication request to the receiver, including its own digital certificate and a randomly generated session identifier; after receiving the request, the receiver verifies the legitimacy of the initiator's digital certificate through the certificate authority. If the certificate is legitimate, the receiver generates its own symmetric encryption key and encrypts the key with the initiator's public key, and then sends the encrypted key, its own digital certificate and confirmation information of the session identifier back to the initiator; the initiator uses its own private key to decrypt and obtain the symmetric encryption key. At this time, the two parties complete the key exchange, and subsequent communications will use the symmetric key for encryption.
[0026] Preferably, the S2 is divided into data encryption and data transmission, wherein:
[0027] Data encryption: The sender groups the data to be transmitted, and encrypts each group of data using a symmetric encryption key and the AES encryption algorithm. At the same time, a unique serial number is generated for each group of data, which is used by the receiver to reassemble the data in sequence. The encrypted data is encapsulated and a data packet header is added. The packet header contains the total length of the data, the length of each group of data, and the serial number information, so that the receiver can parse and process the data.
[0028] Data transmission: The sender sends the encapsulated data to the receiver through a secure network connection. During the transmission process, the network transport layer is responsible for reliable data transmission.
[0029] Preferably, the real-time monitoring of network traffic described in S3.1 is divided into: dynamic adaptive IDS deployment and optimization, deep data packet feature extraction and situational awareness, and adaptive model update for real-time anomaly detection, specifically:
[0030] The steps for dynamic adaptive IDS deployment and optimization include:
[0031] Intelligent node selection: Using network traffic analysis and topology awareness technology, the importance of nodes and the degree of potential risk are calculated through the degree centrality in the graph-based node centrality algorithm, and key monitoring points are dynamically selected; the node degree centrality is calculated using Python's NetworkX library, and the calculation formula is: ,in Representation Node The degree centrality of Is a node The degree, is the total number of nodes in the network.
[0032] Elastic resource allocation: IDS is deployed using Docker in containerization technology, and elastic resource allocation is achieved with the help of the container orchestration platform Kubernetes. The expansion and contraction mechanism is automatically triggered based on network traffic fluctuations and preset monitoring indicators. For example, when the network traffic rate exceeds 1,000 packets per second or the CPU usage exceeds 80%, Kubernetes automatically increases or decreases the number of IDS container copies and resource allocation.
[0033] The steps of deep data packet feature extraction and context awareness include:
[0034] Multi-dimensional feature fusion: Extract request parameter semantic features, URL structure features, and page access sequence features from application layer protocol packets, while considering network traffic time series features; use the TF-IDF algorithm in natural language processing technology to extract text features from HTTP request parameters and URLs, and use the ARIMA model in the time series analysis algorithm to model the traffic time series and extract trend features; for time series analysis, in the ARIMA model, if the time series data To build a model, first perform a stationarity test (such as an ADF test). If the data is not stationary, perform a difference process to make it stationary. Then determine the parameters of the ARIMA (p, d, q) model, where p is the number of autoregressive terms, q is the number of moving average terms, and d is the number of differences. This can be determined by observing the autocorrelation function (ACF) and partial autocorrelation function (PACF) graphs or using automatic parameter selection methods (such as AIC and BIC criteria). Finally, fit the ARIMA model using the training data and extract the coefficients and residuals of the model as features.
[0035] Contextual information association: Establish user behavior portraits and business process model libraries, and extract and analyze features based on network topology, user behavior patterns, and business logic context information. Use the K-Means algorithm in the clustering algorithm to cluster user login time and location to determine normal activity patterns, use directed graphs to represent business processes and determine normal process paths and interaction patterns, and compare current network activities with predefined models in real-time monitoring to determine whether there are any abnormalities.
[0036] The steps of adaptive model updating for real-time anomaly detection include:
[0037] Online learning and incremental update: The incremental ID3 algorithm in the online learning algorithm is used. When a new data packet arrives, the decision tree model is updated according to its characteristics, and the model performance is regularly evaluated and the parameters are adjusted. When new data arrives, the decision tree branches are traversed from the root node according to the sample characteristics. If the classification result of the leaf node is inconsistent with the actual label of the sample, the decision tree structure and node division are adjusted according to the new sample characteristics and labels.
[0038] Model adaptation based on feedback mechanism: Establish a linkage feedback mechanism with the firewall to adjust the normal network behavior model according to the attack characteristics and sources; handle false positives and missed positives at the same time, and adjust the feature extraction method or model threshold by analyzing the log database; after the firewall blocks the attack, it will feed back the information to the IDS, and the IDS will adjust the monitoring sensitivity and threshold of the network behavior of the relevant IP segment accordingly; if it is found that a certain feature combination often leads to false positives, adjust the processing method of this feature in the model.
[0039] Preferably, the machine learning-based intrusion analysis described in S3.2 is subdivided into an integrated machine learning model architecture and anomaly detection based on a generative adversarial network (GAN) and intrusion analysis strategy optimization driven by reinforcement learning, specifically:
[0040] The integrated machine learning model architecture includes:
[0041] Multi-model fusion and collaborative decision-making: Build a framework that integrates decision trees, support vector machines, neural networks, and random forests, and make final decisions through weighted fusion strategies; for example, input data packet features into different models at the same time, and each model outputs a prediction result. Weights are assigned based on indicators such as accuracy and recall rate in historical detections. If the comprehensive anomaly score exceeds the preset threshold, the data packet is judged to be abnormal.
[0042] Automatic feature selection and optimization: Use the recursive feature elimination (RFE) method in the feature selection algorithm to automatically screen feature subsets, regularly re-evaluate feature importance, and adjust feature combinations according to changes in network environment and attack patterns; use the RFE algorithm to gradually eliminate unimportant features, initially train the model with all features, then delete the least important features based on the model's feature importance assessment, retrain with the remaining features and repeat the process, regularly check feature contribution, and if a feature ranks low in importance in multiple consecutive evaluations, remove it. At the same time, pay attention to new potential features and add them to the test when appropriate.
[0043] Anomaly detection based on Generative Adversarial Networks (GANs) includes:
[0044] Abnormal sample generation and expansion: Construct GAN to generate synthetic abnormal samples similar to real attack samples to expand the training data set. The generator receives random noise vectors to generate fake abnormal samples. The discriminator distinguishes between real and generated samples. Through training, the generator generates more realistic samples. Construct a generator containing and the discriminator The loss function of the generator is , the loss function of the discriminator is , by maximizing the true sample score (expected ) and minimize the score of generated samples (expected ) to optimize. During training, the discriminator strives to improve the accuracy of judgment on real samples (making Close to 1), reducing the misjudgment of generated samples (making Close to 0), so as to promote the generator to improve the quality of generated samples. In the calculation process, the calculation is based on the logarithmic form, and the parameters of the discriminator are continuously adjusted so that the discriminator can more accurately distinguish between real samples and generated samples; Express expectations, is a real data sample, is the actual data distribution, is random noise, is a random noise distribution, Represents random noise According to the random noise distribution For sampling, This means that for random noise According to random noise distribution The expectation of sampling, Represents real data samples From the real data distribution The sample obtained from Represents the real data sample According to the real data distribution expectations for sampling; Representation Generator Receive random noise As input, and output false anomaly samples, Representation Discriminator For samples generated by the generator A judgment value is given. Represents the discriminator for the real data sample The discrimination results are obtained by using real attack samples to train GAN. After reaching convergence, the generator is used to generate new abnormal samples and mix them with a small amount of real samples for training the machine learning model.
[0045] Unsupervised anomaly detection and GAN improvement: Unsupervised anomaly detection is performed based on the GAN architecture, and the improved WassersteinGAN (WGAN) is used to improve training stability and the quality of generated samples. In WGAN, the loss function of the discriminator becomes In the calculation of the WGAN discriminator loss function, the logarithmic form is abandoned and the approximate calculation of the Wasserstein distance is adopted; this calculation method makes the training of the discriminator more stable and can effectively avoid the problem of gradient disappearance in traditional GAN; in the same formula Express expectations, is a real data sample, is the actual data distribution, is random noise, is a random noise distribution, Represents random noise According to the random noise distribution For sampling, This means that for random noise According to random noise distribution expectations for sampling; Represents real data samples From the real data distribution The sample obtained from Represents the real data sample According to the real data distribution expectations for sampling; Representation Generator Receive random noise As input, and output false anomaly samples, Representation Discriminator For samples generated by the generator A judgment value is given. Represents the discriminator for the real data sample The discrimination result; at the same time, the gradient penalty term is introduced ,in Express Find the gradient, that is, calculate the function about The vector of partial derivatives of It is a discriminator In the sample The output at Denotes the L2 norm. When training the discriminator, in addition to calculating the above loss function, the gradient of the discriminator is also penalized, where is a hyperparameter that controls the intensity of the gradient penalty. (in is a random number, ) is a randomly sampled point between the real data sample and the generated data sample. The gradient norm (such as L2 norm) of the discriminator at this point is calculated and compared with a preset threshold (such as 1) to perform gradient penalty to optimize the training process of WGAN.
[0046] Reinforcement learning-driven intrusion analysis strategy optimization includes:
[0047] Intelligent detection strategy learning: Apply the deep Q network (DQN) in the reinforcement learning algorithm to allow IDS to autonomously learn the optimal intrusion detection strategy; define the state space including network traffic characteristics, system status, historical attack information and the action space including detection algorithm selection, feature extraction method, and threshold adjustment strategy, build DQN to approximate the Q value function with a neural network, and the intelligent agent interacts with the environment to learn, select actions according to the state and adjust the strategy according to the reward; construct a neural network with an input layer, a hidden layer, and an output layer as DQN, input the state vector and output the Q value corresponding to each action, the intelligent agent selects an action from the DQN according to the current state and executes it in the environment, the environment returns the reward and the next state, and the DQN updates the parameters based on the sample training in the experience playback buffer; the Q value function calculation formula is ,in It's the state. It's action. are the parameters of the neural network; the parameters of the DQN are updated by minimizing the mean squared error (MSE) between the predicted Q-values and the target Q-values, which are typically estimated using a target network, a regularly updated copy of the DQN used to stabilize training.
[0048] Adaptive threat response and policy adjustment: Combining reinforcement learning and real-time threat intelligence, according to the new attack types and methods, the detection strategy and defense measures are automatically adjusted through the reinforcement learning algorithm, such as increasing the monitoring weight of relevant features, adjusting the threshold, etc., and optimizing the strategy according to the detection results and defense effect. For example, it connects with the threat intelligence source to obtain information, triggers policy adjustment when the IDS detects that the network activity matches the threat intelligence, automatically adjusts the traffic monitoring threshold and frequency for new DDoS attacks, and increases the learning rate of relevant features in the deep learning model. After the policy is adjusted, the agent is given positive or negative rewards according to the effect to optimize the strategy.
[0049] Preferably, the automatic defense mechanism described in S3.3 includes: intelligent blocking and isolation strategy, collaborative defense and emergency response mechanism, and self-learning and optimization of defense strategy, specifically:
[0050] Intelligent blocking and isolation strategies include:
[0051] Dynamic blocking based on risk assessment: Introduce a risk assessment model to comprehensively consider the attack type, intensity, target importance and overall network security situation to calculate the risk value, and adopt different blocking strategies according to the risk value; Risk value The calculation formula is ,in is the risk weight of the attack type, is the attack strength evaluation value, is the target importance level, It is the evaluation value of the overall network security situation. , , , are weight coefficients of corresponding factors, which are determined through analysis of historical data and experiments; for example, the risk weight of DDoS attack is set to 0.8 (higher risk), the risk weight of port scanning attack is set to 0.4 (medium risk), the risk weight of malware propagation attack is set to 0.6 (higher risk), etc.; if the network traffic suddenly increases to more than 5 times the normal traffic, the attack intensity is assessed as high and set to 0.8, the traffic increase is between 2-5 times, the assessment is medium and set to 0.5, the traffic increase is between 1-2 times, the assessment is low and set to 0.2; the importance level of core servers is set to 0.9 (very important), important databases are set to 0.8, and critical business applications are set to 0.7, etc.; if multiple devices issue alarms at the same time, or there are multiple high-risk security vulnerabilities that have not been repaired, the overall security situation of the network is assessed as poor and set to 0.7, if only a small number of devices issue alarms, and the number of security vulnerabilities is small and the risk is low, then it is assessed as good and set to 0.3. When the risk value is higher than the preset high-risk threshold (such as 0.6), take immediate and strict blocking measures, such as directly discarding all packets from the attack source and setting long-term blocking rules on network boundary devices (such as firewalls), such as blocking for 24 hours. If the risk value is in the medium-risk range (such as 0.3-0.6), flexibly handle it according to the specific circumstances of the attack. For example, for suspected port scanning behavior, you can temporarily limit the speed of the relevant ports and continue to monitor the development of the attack behavior. If the risk value is lower than the low-risk threshold (such as 0.3), you can take milder measures, such as issuing an alarm and recording relevant information, and increasing the frequency of monitoring the source.
[0052] Application of adaptive isolation and micro-isolation technology: Using software-defined network (SDN) technology, combined with network traffic analysis and security device alarms, to identify and isolate attacked areas or devices in real time; further apply micro-isolation technology to subdivide the network into smaller security areas, use SDN controllers to issue flow table rules based on detected anomalies to cut off the traffic between the attacked network segment or device and other parts, define network access rules for each application or service and apply them to the corresponding container or Pod to achieve micro-isolation; for example, cut off the traffic between the virtual local area network (VLAN) where the attacked virtual machine is located and other VLANs, and modify the flow table of the SDN switch to prevent data packets from being transmitted between different VLANs. Define a network policy to allow the Pod named "web-server" to communicate with the Pod named "database" on a specific port (such as port 3306 for database access), while prohibiting other unauthorized Pods from accessing the "database" Pod.
[0053] The coordinated defense and emergency response mechanism includes:
[0054] Collaborative defense across devices and systems: Establish a deep collaboration mechanism between IDS and firewalls based on the Syslog protocol. When IDS detects an intrusion, it provides detailed attack feature information to the firewall. The firewall dynamically adjusts the filtering rules based on this information. Through the unified security management platform, the equipment is centrally managed and configured, and the operation status and attack events are monitored in real time to analyze the security situation and make decisions. For example, the Syslog protocol is used to send the IDS alarm information to the firewall. The firewall dynamically adjusts the filtering rules based on the attack source IP address range, attack traffic characteristics and other information provided by the IDS, such as adding access control rules for the attack source IP address on the firewall to block all traffic from these addresses, or filtering specific types of data packets based on the characteristics of the attack traffic. The management platform displays the security situation through a visual interface and uniformly configures collaborative defense rules. Administrators can view the IDS detection log, firewall rule settings and traffic statistics on the management platform, uniformly configure cross-device collaborative defense rules, and specify the linkage measures that the IDS and firewall should take under different types of attacks.
[0055] Automated emergency response and plan execution: Develop an automated emergency response plan library covering various common attack types and network security incidents, and use automated scripts and workflow engines to automatically execute plans, including data backup (using professional tools to automatically back up regularly and trigger incremental backups during attacks), service switching (achieved through DNS redirection or load balancer configuration), notification of relevant personnel (automatic notification via email, SMS or instant messaging tools, including attack information and measures taken), and starting traffic cleaning services (cooperating with professional providers to automatically start and configure policy parameters through API interfaces), etc. During the emergency response process, operation logs and event progress are recorded in real time to facilitate review analysis and optimize the plan process.
[0056] Self-learning and optimization of defense strategies
[0057] Strategy improvement based on experience feedback: Establish a defense strategy effectiveness evaluation and feedback system, collect defense measures execution data by deploying monitoring modules on IDS and other security devices, store them in the database, use the policy gradient algorithm in the machine learning algorithm to optimize the defense strategy according to the reward signal, regularly evaluate and update the strategy and apply it to the actual system for testing and verification; for example, use network performance monitoring tools to collect network performance indicators, associate IDS alarms with actual attacks to analyze false positives and negatives, use the policy network to input the network state feature vector to output the probability distribution of defense measures, calculate the reward value according to the defense effect and update the policy network parameters to adjust the strategy selection tendency.
[0058] Strategy verification and optimization based on simulation environment: Build a network security simulation environment, configure attack tools to simulate different types of attack scenarios, implement multiple defense strategies for attack design and test them in the simulation environment, collect attack success rate, defense success rate, false alarm rate, missed alarm rate, network performance indicators, and system resource utilization evaluation data, optimize defense strategies based on the results and test them repeatedly, update the verified effective strategies to the actual IDS system, and establish strategy version management and rollback mechanisms. For example, create a virtual network with multiple devices and network connections, use attack tools to simulate attacks of different intensities, frequencies, and methods, apply defense strategies to network devices to observe the effects, obtain and analyze various performance indicator data by writing scripts, improve strategies based on problems and test again, and first pilot the deployment of optimized strategies in actual applications, then gradually promote them and regularly evaluate and adjust them.
[0059] Preferably, after receiving the data in S4.1, the receiver first decapsulates the data, extracts the information in the data packet header, and sequentially verifies and reassembles the received encrypted data group according to the serial number to ensure the integrity and sequentiality of the data; if data loss or sequence error is found, the receiver notifies the sender to resend the corresponding data group.
[0060] In S4.2, the reorganized encrypted data is decrypted using the previously exchanged symmetric encryption key to obtain the original data; the decrypted data is integrity verified, the hash value of the data is calculated, and compared with the hash value attached by the sender during data transmission. If the two hash values are consistent, it means that the data has not been tampered with during transmission and the integrity is guaranteed; if they are inconsistent, the receiver discards the data and requests the sender to resend it.
[0061] Preferably, in S5.1, when the data transmission is completed and both parties confirm that the data is correct, the communicating parties send a communication end signal. After receiving the other party's end signal, both parties perform final confirmation and cleanup to ensure that all resources for this communication are correctly released, including closing the encrypted connection, clearing temporary data and keys.
[0062] The communication system in S5.2 records key events and operations in the entire communication process, including identity authentication, data encryption and decryption, intrusion detection and defense events, and generates detailed system logs; regularly analyzes the system logs to summarize the security status and potential problems in the communication process, and at the same time, uses log data to update and optimize the intrusion detection model.
[0063] A data communication system based on information security, used to implement the above-mentioned data communication method based on information security, comprising:
[0064] Identity authentication module: ensures the legitimacy of the identities of both communicating parties, including certificate management and key exchange functions.
[0065] Data encryption and transmission module: ensure the confidentiality and integrity of data during transmission.
[0066] Intrusion detection and prevention module: monitor network security status in real time, analyze and prevent intrusion behaviors.
[0067] Data receiving and verification module: The receiver decapsulates, reassembles, decrypts and verifies the integrity of the transmitted data.
[0068] Ending and cleaning module: releases and cleans up resources after the communication is completed to prevent resource leakage.
[0069] Preferably, the intrusion detection and defense module realizes comprehensive monitoring and analysis of network traffic through dynamic and adaptive IDS deployment, deep data packet feature extraction and real-time anomaly detection model update; the intrusion detection and defense module also includes an intrusion analysis submodule and an automatic defense submodule, wherein the intrusion analysis submodule uses an integrated machine learning model, a generative adversarial network and reinforcement learning technology to accurately determine whether there is an intrusion behavior; after detecting an intrusion, the automatic defense submodule takes intelligent blocking and isolation, collaborative defense and emergency response measures to ensure system security.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] 1. In terms of real-time monitoring of network traffic, through intelligent node selection and elastic resource allocation, the system can dynamically adjust monitoring points and resource allocation according to network traffic and topology. Deep packet feature extraction combines multi-dimensional feature fusion and contextual information association to improve the ability to identify complex attacks. The adaptive model update of real-time anomaly detection uses online learning algorithms and model adaptation based on feedback mechanisms to quickly respond to new network environments and attack features, improving detection accuracy and efficiency.
[0072] 2. In terms of intrusion analysis based on machine learning, multi-model fusion and collaborative decision-making improve detection accuracy, and automatic feature selection optimization improves model efficiency. GAN generates abnormal samples to expand the data set, and WGAN improves training and detection capabilities. Reinforcement learning enables IDS to learn autonomously and adjust strategies intelligently. Combined with real-time intelligence, detection and defense measures are adaptively adjusted, such as responding to DDoS attacks and optimizing strategies based on the results. Overall, the detection and response capabilities for unknown attacks are enhanced, data communication security is guaranteed, and intrusion analysis is made smarter, more efficient and accurate.
[0073] 3. In terms of automatic defense mechanism, intelligent blocking and isolation strategies are based on dynamic blocking based on risk assessment, combined with SDN and micro-isolation technology for precise prevention and control. Collaborative defense improves defense efficiency through IDS and firewall collaboration and unified platform management. The automated emergency response plan library enables rapid response and record optimization. Defense strategies are continuously improved through experience feedback and simulation environment optimization and improvement, continuously improving defense effects and adaptability, and building an overall efficient and intelligent automatic defense system to effectively ensure data communication security and reduce network security risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 The present invention is a schematic diagram of the steps of a data communication method based on information security. DETAILED DESCRIPTION
[0075] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0076] like Figure 1 As shown, a data communication method based on information security includes the following five steps:
[0077] 1. Preparation
[0078] The communicating parties perform identity authentication. The initiator sends a digital certificate and session identifier. After the receiver verifies the legitimacy of the certificate, it encrypts the symmetric key with the initiator's public key and sends it back to complete the key exchange.
[0079] 2. Transmission phase
[0080] The sender uses a symmetric key and the AES algorithm to encrypt the data and group it. Prior to this, the sender first uses a hash function to calculate the original data, generates a unique hash value and attaches it. The sender then adds header information to each group of encrypted data and encapsulates the data for transmission over a secure network connection.
[0081] 3. Intrusion detection and prevention stage
[0082] IDS monitors network traffic in real time, extracts data packet features and compares them with normal models, and marks suspicious data packets; machine learning models analyze suspicious data packets to determine whether they are intrusions; after confirming the intrusion, it triggers the automatic defense mechanism, including blocking the attack source, notifying both communicating parties, preparing data backup and recovery, and dynamically adjusting the defense strategy.
[0083] 4. Receiving and Verifying Phase
[0084] The receiver decapsulates the data, verifies and reassembles it according to the sequence number, and notifies the sender to retransmit if there is a problem. It decrypts the data with the symmetric key, calculates the hash value, and compares it with the hash value attached by the sender during data transmission to verify integrity.
[0085] 5. Closing and Cleaning Phase
[0086] After both parties of the communication confirm that the communication is complete, resources are released and cleaned up; the system records key events in the communication process to generate logs, conducts regular analysis to optimize the system and strengthen defense, and updates intrusion detection models.
[0087] In practical application, the present invention is fully described below with a simulated embodiment:
[0088] 1. Hypothetical Data Communication Implementation Scenario
[0089] In a large enterprise with multiple branches, data communication between the headquarters and branches is required frequently, including financial report transmission, business data synchronization, video conferencing and other types of data interaction. The enterprise network architecture is complex, covering local area networks, wide area networks and cloud computing resources, and faces various network security threats from both inside and outside.
[0090] 1. Preparation
[0091] Identity Authentication
[0092] When a server at the headquarters needs to send important financial data to a client at a branch office, the server first sends an identity authentication request to the client, including its own digital certificate (the certificate is issued by an authoritative certificate authority to ensure its legitimacy and credibility) and a randomly generated session identifier.
[0093] After receiving the request, the client verifies the legitimacy of the server's digital certificate by connecting to the online verification system of the certificate authority. If the certificate is legitimate, the client generates its own symmetric encryption key (for example, a 256-bit key required for the AES algorithm) and encrypts the key using the server's public key (obtained from the server's digital certificate). The client then sends the encrypted key back to the server along with its own digital certificate and confirmation of the session identifier.
[0094] The server uses its own private key to decrypt and obtain the symmetric encryption key. At this point, the two parties have completed the key exchange, and subsequent communications will use the symmetric key for encryption.
[0095] (II) Transmission phase
[0096] Data encryption
[0097] The server groups the financial data files to be transmitted according to a certain size, and encrypts each group of data using the previously exchanged symmetric encryption key and AES encryption algorithm. At the same time, a unique serial number is generated for each group of data, which is used by the client to reorganize the data in sequence. The encrypted data is encapsulated and a data packet header is added. The packet header contains information such as the total length of the data, the length of each group of data, and the serial number. For example, the packet header may indicate that the entire data file has 100 groups of data, each group of data is 1024 bytes long, and the current data packet is the 5th group, etc.
[0098] Data Transfer
[0099] The server sends the encapsulated data to the client of the branch office through a secure virtual private network (VPN) connection established within the enterprise. During the transmission process, the network transport layer protocol (such as TCP) is responsible for reliable data transmission, ensuring that the data arrives in order and is not lost or damaged. If data is lost or erroneous, the transport layer will automatically request retransmission.
[0100] 3. Intrusion Detection and Prevention Phase
[0101] Real-time monitoring of network traffic
[0102] Dynamic and adaptive IDS deployment and optimization
[0103] Using the traffic monitoring system and topology management tools in the enterprise network, the importance and potential risk level of each network node are calculated based on the degree centrality in the graph node centrality algorithm. For example, the degree centrality of each switch, router, and server node in the enterprise network is calculated through Python's NetworkX library. If the degree centrality of a core switch is high (that is, there are more devices and network links connected to the switch), and the recent network traffic passing through the node has fluctuated greatly, there may be potential risks, then the IDS monitoring focus will be tilted towards this node. When the network traffic rate exceeds 5,000 packets per second or the CPU usage exceeds 70% (preset monitoring indicators), Kubernetes automatically increases the number of IDS container copies near the node and allocates more resources to improve monitoring capabilities.
[0104] Deep packet feature extraction and context awareness
[0105] For application layer protocol packets, especially HTTP protocol packets involving financial data transmission, request parameter semantic features, URL structure features, and page access sequence features are extracted. For example, analyze whether the parameters in the financial report upload request conform to the normal format and business logic, whether the URL points to a legitimate financial system interface, etc. At the same time, the TF-IDF algorithm in natural language processing technology is used to extract text features from HTTP request parameters and URLs, and the ARIMA model in the time series analysis algorithm is used to model the network traffic time series to extract trends and other features. In the application of the ARIMA model, the traffic time series is first tested for stationarity (such as ADF test). If the data is not stationary, differential processing is performed to make it stationary. Then, the parameters of the ARIMA (p, d, q) model are determined by observing the autocorrelation function (ACF) and partial autocorrelation function (PACF) graphs, such as determining p=2, d=1, q=3. Finally, the ARIMA model is fitted using the training data, and the coefficients and residuals of the model are extracted as features.
[0106] Establish user behavior portraits and business process model libraries for enterprise employees, and extract and analyze features based on contextual information such as network topology, user behavior patterns, and business logic. Use the K-Means algorithm in the clustering algorithm to cluster the login time and location of employees to determine normal activity patterns. For example, it is found that most employees log in to the financial system from the company's internal IP address during working hours on weekdays. If a large number of login attempts from unfamiliar IP addresses appear in a certain period of time and the operation behavior is abnormal (such as frequent downloading of large amounts of data), it will be marked as suspicious behavior. Use a directed graph to represent the business process of financial data transmission, determine the normal process path and interaction mode, and compare the current network activity with the predefined model in real-time monitoring to determine whether it is abnormal.
[0107] Adaptive model updating for real-time anomaly detection
[0108] Adopt the incremental ID3 algorithm in the online learning algorithm. When a new network data packet arrives at the IDS, the decision tree model is updated according to its characteristics. For example, when new data arrives, traverse along the decision tree branch starting from the root node according to the sample characteristics. If the classification result of the leaf node is inconsistent with the actual label of the sample (whether it is a normal data packet or an abnormal data packet), adjust the decision tree structure and node division according to the new sample characteristics and labels. Evaluate the model performance regularly (such as every hour), such as accuracy, recall rate and other indicators, and adjust the model parameters according to the evaluation results.
[0109] Establish a linkage feedback mechanism with the firewall. When the firewall blocks a suspected attack, it will feed back the attack characteristics and source information to the IDS. For example, if the firewall finds that a large number of connection requests from a certain IP segment meet the characteristics of port scanning and blocks them, it will feed back the information to the IDS. Based on this, the IDS adjusts the monitoring sensitivity and threshold of the network behavior of the relevant IP segment, such as more stringent review of subsequent connection requests for the IP segment. If it is found that a certain feature combination (such as a specific combination of packet size and protocol type) often leads to false alarms, adjust the way the feature is handled in the model, such as reducing its priority in the decision tree or adjusting the feature weight.
[0110] Intrusion Analysis Based on Machine Learning
[0111] Integrated machine learning model architecture
[0112] A framework integrating decision trees, support vector machines, neural networks, and random forests is constructed for intrusion analysis. When a data packet needs to be detected, its features are input into different models at the same time. For example, for a suspicious network data packet, its features such as packet length, protocol type, source IP address, and destination IP address are extracted and input into the decision tree, support vector machine, neural network, and random forest models for analysis. Each model outputs a prediction result, and a weight is assigned based on its accuracy and recall in historical detection. Assume that the decision tree has an accuracy of 85% and a recall of 80% in past detections; the support vector machine has an accuracy of 88% and a recall of 82%; the neural network has an accuracy of 90% and a recall of 85%; the random forest has an accuracy of 87% and a recall of 83%. Weights can be assigned to each model based on these indicators, such as a neural network weight of 0.35, a support vector machine weight of 0.3, a random forest weight of 0.25, and a decision tree weight of 0.1. The output results of each model are combined to calculate the anomaly score. If the anomaly score exceeds the preset threshold (such as 0.7), the data packet is judged to be abnormal.
[0113] Anomaly Detection Based on Generative Adversarial Networks (GAN)
[0114] GAN is constructed to generate synthetic abnormal samples similar to real attack samples to expand the training data set. The generator receives random noise vectors to generate fake abnormal samples, and the discriminator distinguishes between real and generated samples. GAN is trained using real attack samples collected by the enterprise in the past (such as various types of malware attacks, DDoS attacks, etc.). After multiple iterations of training, the generator can generate more realistic abnormal samples. After reaching convergence, the generator generates new abnormal samples and mixes them with a small number of real samples to train the machine learning model, improving the model's ability to recognize unknown attack patterns.
[0115] Improved WassersteinGAN (WGAN) is used to improve training stability and the quality of generated samples. In WGAN, the loss function of the discriminator is no longer in logarithmic form, but is approximated by the Wasserstein distance. At the same time, a gradient penalty term is introduced. When training the discriminator, in addition to calculating the above loss function, the gradient of the discriminator is also penalized. For example, the intensity hyperparameter of the gradient penalty is set to 0.1, a point is randomly sampled between the real data sample and the generated data sample, the gradient norm of the discriminator at this point (such as the L2 norm) is calculated, and the gradient penalty is compared with the preset threshold (such as 1), the training process of WGAN is optimized, the generated anomaly samples are more consistent with the distribution characteristics of the real attack samples, and the accuracy of anomaly detection is improved.
[0116] Reinforcement learning driven intrusion analysis strategy optimization
[0117] The deep Q network (DQN) is applied to allow IDS to autonomously learn the optimal intrusion detection strategy. Define a state space including network traffic characteristics (such as traffic size, fluctuation), system status (such as CPU usage, memory occupancy), historical attack information (attack type, time, frequency, etc.) and an action space including detection algorithm selection, feature extraction method, threshold adjustment strategy, etc. Construct a neural network with input layer, hidden layer and output layer as DQN, input state vector and output Q value corresponding to each action. For example, when network traffic suddenly increases and system CPU usage increases, the agent selects an action from DQN according to the current state, such as adjusting the frequency of feature extraction to a higher frequency, or selecting a more complex detection algorithm. The agent performs the selected action in the environment, and the environment returns the reward and the next state. If an attack is detected and successfully blocked, a positive reward (such as +10) is given; if a false positive or false negative occurs, a negative reward (such as -5) is given. The DQN is trained to update parameters based on the samples stored in the experience replay buffer, and the parameters of the DQN are updated by minimizing the mean square error (MSE) between the predicted Q value and the target Q value. The target Q-value is typically estimated using a target network, which is a regularly updated copy of the DQN (e.g., every other week) to stabilize training.
[0118] Combined with real-time threat intelligence, connect with professional threat intelligence source providers to obtain information. When the IDS detects that network activities match threat intelligence, such as finding that the behavior pattern of a certain IP address is similar to the characteristics of a new type of DDoS attack recently reported, it automatically adjusts the detection strategy and defense measures. For example, automatically adjust the traffic monitoring threshold, increase the monitoring frequency of the IP address to once per second, and increase the learning rate of relevant features (such as traffic peaks, packet size distribution, etc.) in the deep learning model to improve the detection ability of this type of attack. After the strategy is adjusted, give the agent positive or negative rewards to optimize the strategy based on the detection results and defense effects. If the attack is successfully blocked, increase the probability of the agent choosing this strategy in similar situations; if the attack still causes a certain impact, adjust the strategy or explore other possible strategies.
[0119] Automatic defense mechanism
[0120] Intelligent blocking and isolation strategies
[0121] The risk assessment model is introduced to calculate the risk value by comprehensively considering factors such as attack type, intensity, target importance, and overall network security situation. For example, for a suspected DDoS attack, the risk value is calculated based on the attack type (DDoS attack risk weight is set to 0.8), intensity (network traffic suddenly increases to more than 10 times the normal traffic, the attack intensity is assessed as high, set to 0.9), target importance (the target is the core business server, the importance level is set to 0.9), and the overall network security situation (multiple devices sound alarms at the same time, and there are multiple security vulnerabilities that have not been fixed, the overall network security situation is assessed as poor, set to 0.7). The risk value calculation formula is: When the risk value is higher than the preset high risk threshold (such as 1.5), take immediate and strict blocking measures, such as directly discarding all packets from the attack source and setting long-term blocking rules on network edge devices (such as firewalls), such as blocking for 72 hours. At the same time, notify the network administrator for further investigation and processing.
[0122] Using software-defined networking (SDN) technology, combined with network traffic analysis and security device alarms, the attacked area or device can be identified in real time. When a server is detected to be under attack, the SDN controller issues flow table rules based on the detected abnormal situation to cut off the traffic between the network segment where the server is located and other parts. For example, the traffic between the virtual local area network (VLAN) where the attacked server is located and other VLANs is cut off, and the flow table of the SDN switch is modified to prevent data packets from being transmitted between different VLANs. Micro-isolation technology is further applied to define network access rules for each application or service and apply them to the corresponding container or Pod to achieve micro-isolation. For example, for the financial system application of an enterprise, a network policy is defined to allow only clients from a specific IP segment to access specific ports of the financial system (such as port 8080 for financial data query), while prohibiting other unauthorized access.
[0123] Collaborative defense and emergency response mechanism
[0124] Establish a deep collaboration mechanism between IDS and firewall based on Syslog protocol. When IDS detects intrusion, it sends detailed attack feature information (such as attack source IP address, attack type, attack time, etc.) to the firewall through Syslog protocol. The firewall dynamically adjusts the filtering rules based on this information, such as adding access control rules for attack source IP addresses on the firewall to block all traffic from these addresses, or filtering specific types of data packets based on the characteristics of attack traffic (such as specific protocol type and port number combination).
[0125] All security devices in the enterprise network can be centrally managed and configured through a unified security management platform, and the operating status and attack events can be monitored in real time. The management platform displays the security situation through a visual interface, such as network traffic diagrams, attack event distribution maps, etc., so that administrators can intuitively understand the network security status. Administrators can view information such as IDS detection logs, firewall rule settings and traffic statistics on the management platform, and uniformly configure cross-device collaborative defense rules. For example, specify the linkage measures that IDS and firewalls should take when encountering large-scale DDoS attacks, such as IDS sending the attack source IP list to the firewall in a timely manner, the firewall automatically starting the traffic cleaning function and adjusting the filtering rules, and notifying the network administrator for emergency handling.
[0126] Develop an automated emergency response plan library covering various common attack types and network security incidents. For example, for DDoS attacks, the plan includes automatically starting the traffic cleaning service (cooperating with professional traffic cleaning service providers to automatically start and configure policy parameters such as cleaning thresholds and traffic forwarding rules through API interfaces); for data leakage incidents, the plan includes automatically stopping related data transmission services, performing data backup (using professional backup tools to automatically backup regularly, and triggering incremental backups during attacks to back up data to off-site storage devices), and notifying relevant personnel (automatically notifying the company's security manager, technical team members, etc. through emails, text messages, or instant messaging tools, including attack information and measures taken). During the emergency response process, record the operation log and event progress in real time to facilitate post-event review and analysis and optimize the plan process.
[0127] Self-learning and optimization of defense strategies
[0128] Establish a defense strategy effectiveness evaluation and feedback system, and deploy monitoring modules on IDS and other security devices (such as firewalls, servers, etc.) to collect data on the execution of defense measures. For example, use network performance monitoring tools to collect network performance indicators (such as network delay, packet loss rate, etc.), and associate IDS alarms with actual attacks to analyze false positives and negatives. Input the network status feature vector (including network traffic, attack frequency, device load, etc.) into the policy network, and output the probability distribution of defense measures. Calculate the reward value based on the defense effect. For example, if an attack is successfully blocked and has little impact on network performance, a higher reward value (such as +8) is given; if a false alarm occurs and normal business is affected, a lower reward value (such as -3) is given. Update the policy network parameters based on the reward value and adjust the policy selection tendency so that the system can choose a better defense strategy when facing similar situations in the future.
[0129] Build a network security simulation environment, use Mininet tools to build a virtual network topology, and simulate various structures and devices of the enterprise network. Configure attack tools (such as Metasploit) to simulate different types of attack scenarios, such as simulating DDoS attacks of different intensities, various types of malware propagation attacks, etc. Design and implement a variety of defense strategies for different types of attacks, and test them in a simulation environment. Collect evaluation data such as attack success rate, defense success rate, false alarm rate, missed alarm rate, network performance indicators, and system resource utilization, optimize defense strategies based on the results, and test them repeatedly. For example, if a certain defense strategy is found to have a low defense success rate against a specific type of attack in a simulation environment, adjust the strategy parameters (such as adjusting firewall rules, IDS detection thresholds, etc.) or improve the strategy logic (such as optimizing feature extraction algorithms, adjusting machine learning model parameters, etc.) and then test it again. Update the strategies that have been verified to be effective in the simulation environment to the actual IDS system, and establish a strategy version management and rollback mechanism. In actual applications, first pilot the deployment of optimized strategies in some branches or non-critical business systems, observe their effects, and conduct evaluations. If the effect is good, it will be gradually extended to the entire enterprise network; if any problems occur, it can be rolled back to the previous stable version in time to ensure the safe and stable operation of the enterprise network.
[0130] 4. Receiving and Verifying
[0131] Data reception and decapsulation
[0132] After receiving the data, the client of the branch office first decapsulates the data, extracts the information in the data packet header, and sequentially verifies and reassembles the received encrypted data group according to the sequence number to ensure the integrity and sequence of the data. If data is found to be missing or in the wrong sequence, the client notifies the server to resend the corresponding data group. For example, if the client finds that the data packet sequence number is interrupted or out of order, it immediately sends a request to the server to resend the missing or erroneous data packet.
[0133] Data decryption and integrity verification
[0134] The client uses the previously exchanged symmetric encryption key to decrypt the reorganized encrypted data to obtain the original financial data. Because the sender first uses the hash function to calculate the original data before using the symmetric key and AES algorithm to encrypt and group the data, and generates a unique hash value and attaches it, the integrity of the decrypted data is verified, and the hash value of the data is calculated (for example, using hash algorithms such as MD5 or SHA-256) and compared with the hash value attached by the server during data transmission. If the two hash values are consistent, it means that the data has not been tampered with during transmission and the integrity is guaranteed; if they are inconsistent, the client discards the data and requests the server to resend it.
[0135] 5. Closing and Cleaning Phase
[0136] Resource release and cleanup
[0137] When the data transmission is completed and both parties confirm that the data is correct, the two communicating parties send a communication end signal. After receiving the other party's end signal, both parties perform final confirmation and cleanup to ensure that all resources for this communication are correctly released. The server closes the encrypted connection with the client, clears the temporary data and keys generated during this communication, and releases related memory resources. The client also performs the same operations, including closing the connection, clearing the temporary data cached locally, etc., to prevent resource leakage and prepare for the next communication.
[0138] The communication system records key events and operations in the entire communication process, including the time, method, and results of identity authentication, the process and related parameters of data encryption and decryption, intrusion detection and defense events (such as suspicious behaviors detected, intrusion types, defense measures taken, etc.), and generates detailed system logs.
[0139] Analyze system logs regularly (e.g., weekly) to summarize security situations and potential problems in the communication process. For example, by analyzing logs, it is found that the number of network connection requests from a specific area has increased abnormally in a certain period of time recently, and most of them are marked as suspicious behaviors by IDS, which may pose a potential risk of attack. At the same time, use log data to update and optimize intrusion detection models. If it is found that certain features appear frequently in multiple false alarm events, it may be necessary to adjust the weights or processing methods of these features in the model; if new attack patterns or features are found, update the rules and parameters of the model in a timely manner to improve the security of the system and the accuracy of detection.
[0140] 2. Implementation Effects and Advantages
[0141] 1. Improved security
[0142] Strict identity authentication prevents illegal access and impersonation, ensures the security of key exchange and encrypted communications, and reduces the risk of data theft and tampering.
[0143] Data encryption ensures transmission confidentiality, and the AES algorithm and symmetric key exchange provide reliable protection.
[0144] Intrusion detection and defense multi-dimensional monitoring and analysis, intelligent node selection and deep feature extraction to improve detection accuracy, machine learning to enhance detection of complex unknown attacks, automatic defense mechanism to accurately respond to intrusions, risk assessment and SDN and other technologies to reduce the impact of attacks.
[0145] 2. Communication efficiency guarantee
[0146] Data grouping and sequence number management ensures the integrity of the transmission order, reduces retransmissions, improves efficiency, and adapts to network instability.
[0147] The intrusion detection module dynamically adapts without affecting communication efficiency, and intelligently adjusts resources and model updates to ensure smooth communication.
[0148] Automatic defense coordination and emergency response quickly handle incidents and reduce disruptions, such as ensuring bandwidth and business continuity during DDoS attacks.
[0149] 3. System adaptability and scalability
[0150] The system architecture adapts to complex networks and diverse business requirements, and operates effectively in different environments and data communication types.
[0151] It has strong scalability, and as the enterprise grows and upgrades, new equipment and systems can be easily included in the protection, machine learning models can be updated and optimized, and the modular design does not affect stability.
[0152] 4. Convenience of management and maintenance
[0153] The unified security management platform provides centralized management functions, and its visual interface helps administrators understand network security trends, monitor device operation status and various security events. Through this platform, administrators can configure security rules in a unified manner, reducing management complexity.
[0154] The defense strategy has the ability to self-learn and optimize, which can effectively reduce manual intervention. By building a simulation environment, various defense strategies can be verified and optimized. At the same time, the system will record key events in the communication process to generate logs. Analyzing the system logs will help discover potential problems and solve them in a timely manner, thereby improving the maintenance efficiency of the system.
[0155] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A data communication method based on information security, characterized in that: The following steps are involved: S1. Preparation S1.
1. The communicating parties perform identity authentication. The initiator sends a digital certificate and session identifier. After the receiver verifies the legitimacy of the certificate, it encrypts the symmetric key with the initiator's public key and sends it back to complete the key exchange. S2, transmission stage S2.
1. The sender uses a symmetric key and the AES algorithm to encrypt and group data. Before that, the sender first uses a hash function to calculate the original data, generates a unique hash value and attaches it. Then the sender adds header information to each group of encrypted data and encapsulates the data, which is then transmitted over a secure network connection. S3: Intrusion Detection and Prevention Phase S3.
1. IDS monitors network traffic in real time, extracts data packet features and compares them with normal models, and marks suspicious data packets. Real-time monitoring of network traffic is divided into: dynamic adaptive IDS deployment and optimization, deep data packet feature extraction and situational awareness, and adaptive model update for real-time anomaly detection. Specifically: The steps for dynamic adaptive IDS deployment and optimization include: Intelligent node selection: Using network traffic analysis and topology awareness technology, the importance of nodes and the degree of potential risk are calculated through the degree centrality in the graph-based node centrality algorithm, and key monitoring points are dynamically selected; the node degree centrality is calculated using Python's NetworkX library, and the calculation formula is: ,in Representation Node The degree centrality of Is a node The degree, is the total number of nodes in the network; Elastic resource allocation: IDS is deployed using Docker in containerization technology, and elastic resource allocation is achieved with the help of Kubernetes, a container orchestration platform; the expansion and contraction mechanism is automatically triggered based on network traffic fluctuations and preset monitoring indicators; The steps of deep data packet feature extraction and context awareness include: Multi-dimensional feature fusion: Extract request parameter semantic features, URL structure features, and page access sequence features from application layer protocol packets, while considering network traffic time series features; use the TF-IDF algorithm in natural language processing technology to extract text features from HTTP request parameters and URLs, and use the ARIMA model in the time series analysis algorithm to model the traffic time series and extract trend features; Contextual information association: Establish user behavior portraits and business process model libraries, and extract and analyze features based on network topology, user behavior patterns, and business logic context information; cluster user login times and locations using the K-Means algorithm in the clustering algorithm to determine normal activity patterns, use directed graphs to represent business processes and determine normal process paths and interaction patterns, and compare current network activities with predefined models in real-time monitoring to determine whether they are abnormal; The steps of adaptive model updating for real-time anomaly detection include: Online learning and incremental update: Adopt the incremental ID3 algorithm in the online learning algorithm. When a new data packet arrives, update the decision tree model according to its characteristics, regularly evaluate the model performance and adjust the parameters. When new data arrives, traverse along the decision tree branches starting from the root node according to the sample characteristics. If the classification result of the leaf node is inconsistent with the actual label of the sample, adjust the decision tree structure and node division according to the new sample characteristics and labels. Model adaptation based on feedback mechanism: Establish a linkage feedback mechanism with the firewall to adjust the normal network behavior model according to the attack characteristics and sources; handle false positives and false negatives at the same time, and adjust the feature extraction method or model threshold by analyzing the log database; after the firewall blocks the attack, it will feed back the information to the IDS, and the IDS will adjust the monitoring sensitivity and threshold of the network behavior of the relevant IP segment accordingly; if it is found that a certain feature combination often leads to false positives, adjust the processing method of the feature in the model; S3.
2. The machine learning model analyzes suspicious data packets to determine whether they are intrusion behaviors. The intrusion analysis based on machine learning is subdivided into integrated machine learning model architecture and anomaly detection based on generative adversarial network (GAN) and intrusion analysis strategy optimization driven by reinforcement learning. Specifically: The integrated machine learning model architecture includes: Multi-model fusion and collaborative decision-making: Build a framework that integrates decision trees, support vector machines, neural networks, and random forests, and make final decisions through weighted fusion strategies; Automatic feature selection and optimization: Use the recursive feature elimination (RFE) method in the feature selection algorithm to automatically screen feature subsets, regularly re-evaluate feature importance, and adjust feature combinations based on changes in network environment and attack patterns; Anomaly detection based on Generative Adversarial Networks (GANs) includes: Abnormal sample generation and expansion: Construct GAN to generate synthetic abnormal samples similar to real attack samples to expand the training data set. The generator receives random noise vectors to generate fake abnormal samples. The discriminator distinguishes between real and generated samples. Through training, the generator generates more realistic samples. Construct a generator containing and the discriminator The loss function of the generator is , the loss function of the discriminator is , in the calculation process, the logarithmic form is relied upon for calculation; Express expectations, is a real data sample, is the actual data distribution, is random noise, is a random noise distribution, Represents random noise According to the random noise distribution For sampling, This means that for random noise According to random noise distribution The expectation of sampling, Represents real data samples From the real data distribution The sample obtained from Represents the real data sample According to the real data distribution expectations for sampling; Representation Generator Receive random noise As input, and output false anomaly samples, Representation Discriminator For samples generated by the generator A judgment value is given. Represents the discriminator for the real data sample The discrimination results; use real attack samples to train GAN, and after reaching convergence, use the generator to generate new abnormal samples and mix them with a small amount of real samples to train the machine learning model; Unsupervised anomaly detection and GAN improvement: Unsupervised anomaly detection is performed based on the GAN architecture, and the improved WassersteinGAN (WGAN) is used to improve training stability and the quality of generated samples. In WGAN, the loss function of the discriminator becomes In the calculation of the WGAN discriminator loss function, the logarithmic form is abandoned and the approximate calculation of the Wasserstein distance is adopted. Express expectations, is a real data sample, is the actual data distribution, is random noise, is a random noise distribution, Represents random noise According to the random noise distribution For sampling, This means that for random noise According to random noise distribution expectations for sampling; Represents real data samples From the real data distribution The sample obtained from Represents the real data sample According to the real data distribution expectations for sampling; Representation Generator Receive random noise As input, and output false anomaly samples, Representation Discriminator For samples generated by the generator A judgment value is given. Represents the discriminator for the real data sample The judgment result of Reinforcement learning-driven intrusion analysis strategy optimization includes: Intelligent detection strategy learning: Apply the deep Q network (DQN) in the reinforcement learning algorithm to allow IDS to autonomously learn the optimal intrusion detection strategy; define the state space including network traffic characteristics, system status, historical attack information and the action space including detection algorithm selection, feature extraction method, and threshold adjustment strategy, build DQN to approximate the Q value function with a neural network, and the intelligent agent interacts with the environment to learn, select actions according to the state and adjust the strategy according to the reward; construct a neural network with an input layer, a hidden layer, and an output layer as DQN, input the state vector and output the Q value corresponding to each action, the intelligent agent selects an action from the DQN according to the current state and executes it in the environment, the environment returns the reward and the next state, and the DQN updates the parameters based on the sample training in the experience playback buffer; the Q value function calculation formula is ,in It's the state. It's action. are the parameters of the neural network; the parameters of the DQN are updated by minimizing the mean squared error (MSE) between the predicted Q-values and the target Q-values, which are usually estimated using a target network, which is a regularly updated copy of the DQN used to stabilize training; Adaptive threat response and policy adjustment: Combining reinforcement learning and real-time threat intelligence, the system automatically adjusts detection strategies and defense measures based on new attack types and methods through reinforcement learning algorithms, connects with threat intelligence sources to obtain information, and triggers policy adjustments when IDS detects that network activities match threat intelligence. It automatically adjusts traffic monitoring thresholds and frequencies for new DDoS attacks, and increases the learning rate of related features in the deep learning model. After the policy is adjusted, the intelligent agent is given positive or negative rewards based on the effect to optimize the policy. S3.
3. After confirming the intrusion, trigger the automatic defense mechanism, including blocking the attack source, notifying both parties in communication, preparing data backup and recovery, and dynamically adjusting the defense strategy; S4, Receiving and Verifying Phase S4.
1. The receiver decapsulates the data, verifies and reassembles it according to the sequence number, and notifies the sender to retransmit if there is any problem; S4.
2. Decrypt the data using the symmetric key, calculate the hash value and compare it with the hash value that the sender attached when transmitting the data to verify the integrity; S5: Ending and cleaning phase S5.
1. After both parties of the communication confirm that the communication is complete, release and clean up the resources; S5.
2. The system records key events in the communication process and generates logs, which are regularly analyzed to optimize the system and strengthen defense, and to update intrusion detection models.
2. A data communication method based on information security according to claim 1, characterized in that: In S1, the communication initiator first sends an identity authentication request to the receiver, including its own digital certificate and a randomly generated session identifier; after receiving the request, the receiver verifies the legitimacy of the initiator's digital certificate through the certificate authority. If the certificate is legitimate, the receiver generates its own symmetric encryption key and encrypts the key with the initiator's public key, and then sends the encrypted key, its own digital certificate and confirmation information of the session identifier back to the initiator; the initiator uses its own private key to decrypt and obtain the symmetric encryption key. At this time, the two parties complete the key exchange, and subsequent communications will use the symmetric key for encryption.
3. A data communication method based on information security according to claim 1, characterized in that: The S2 is divided into data encryption and data transmission, where: Data encryption: The sender groups the data to be transmitted, and encrypts each group of data using a symmetric encryption key and the AES encryption algorithm. At the same time, a unique serial number is generated for each group of data, which is used by the receiver to reassemble the data in sequence. The encrypted data is encapsulated and a data packet header is added, which contains the total length of the data, the length of each group of data, and the serial number information. Data transmission: The sender sends the encapsulated data to the receiver through a secure network connection. During the transmission process, the network transport layer is responsible for reliable data transmission.
4. A data communication method based on information security according to claim 1, characterized in that: The automatic defense mechanism described in S3.3 includes: intelligent blocking and isolation strategy, collaborative defense and emergency response mechanism, and self-learning and optimization of defense strategy, specifically: Intelligent blocking and isolation strategies include: Dynamic blocking based on risk assessment: Introduce a risk assessment model to comprehensively consider the attack type, intensity, target importance and overall network security situation to calculate the risk value, and adopt different blocking strategies according to the risk value; Risk value The calculation formula is ,in is the risk weight of the attack type, is the attack strength evaluation value, is the target importance level, It is the evaluation value of the overall network security situation. , , , are the weight coefficients of the corresponding factors, which are determined through analysis of historical data and experiments; Adaptive isolation and micro-isolation technology application: Utilize software-defined network (SDN) technology, combined with network traffic analysis and security device alarms, to identify and isolate attacked areas or devices in real time; further apply micro-isolation technology to subdivide the network into smaller security areas, use SDN controllers to issue flow table rules based on detected anomalies to cut off traffic between the attacked network segment or device and other parts, define network access rules for each application or service and apply them to the corresponding container or Pod to achieve micro-isolation; The coordinated defense and emergency response mechanism includes: Collaborative defense across devices and systems: Establish a deep collaboration mechanism between IDS and firewalls based on the Syslog protocol. When IDS detects an intrusion, it provides detailed attack feature information to the firewall. The firewall dynamically adjusts the filtering rules based on this information. Through a unified security management platform, devices are centrally managed and configured, and the operating status and attack events are monitored in real time to conduct security situation analysis and decision-making. Automated emergency response and plan execution: Develop an automated emergency response plan library covering various common attack types and network security incidents, and use automated scripts and workflow engines to automatically execute plans, including data backup, service switching, notification of relevant personnel, and initiation of traffic cleaning services. During the emergency response process, real-time records of operation logs and event progress are recorded to facilitate review analysis and optimization of the plan process; Self-learning and optimization of defense strategies Strategy improvement based on experience feedback: Establish a defense strategy effect evaluation and feedback system, deploy monitoring modules on IDS and other security devices to collect defense measures execution data, store it in the database, use the policy gradient algorithm in the machine learning algorithm to optimize the defense strategy based on the reward signal, regularly evaluate and update the strategy and apply it to the actual system for testing and verification; Strategy verification and optimization based on simulation environment: Build a network security simulation environment, configure attack tools to simulate different types of attack scenarios, implement multiple defense strategies for attack designs and test them in a simulation environment, collect attack success rate, defense success rate, false alarm rate, missed alarm rate, network performance indicators, and system resource utilization evaluation data, optimize defense strategies based on the results and test them repeatedly, update the verified effective strategies to the actual IDS system, and establish strategy version management and rollback mechanisms.
5. The data communication method based on information security according to claim 1, characterized in that: After receiving the data in S4.1, the receiver first decapsulates the data, extracts the information in the data packet header, and sequentially verifies and reassembles the received encrypted data group according to the sequence number to ensure the integrity and sequentiality of the data; If data is found to be missing or in the wrong order, the receiver notifies the sender to resend the corresponding data group; In S4.2, the reorganized encrypted data is decrypted using the previously exchanged symmetric encryption key to obtain the original data; the decrypted data is integrity verified, the hash value of the data is calculated, and compared with the hash value attached by the sender during data transmission. If the two hash values are consistent, it means that the data has not been tampered with during transmission and the integrity is guaranteed; if they are inconsistent, the receiver discards the data and requests the sender to resend it.
6. A data communication method based on information security according to claim 1, characterized in that: In S5.1, when the data transmission is completed and both parties confirm that the data is correct, the communicating parties send a communication end signal. After receiving the end signal from the other party, both parties perform final confirmation and cleanup to ensure that all resources for this communication are correctly released, including closing the encrypted connection, clearing temporary data and keys; The communication system in S5.2 records key events and operations in the entire communication process, including identity authentication, data encryption and decryption, intrusion detection and defense events, and generates detailed system logs; regularly analyzes the system logs to summarize the security status and potential problems in the communication process, and at the same time, uses log data to update and optimize the intrusion detection model.
7. A data communication system based on information security, characterized in that: A data communication method based on information security for implementing any one of claims 1 to 6, comprising: Identity authentication module: ensures the legitimacy of the identities of both parties in communication, including certificate management and key exchange functions; Data encryption and transmission module: ensure the confidentiality and integrity of data during transmission; Intrusion detection and prevention module: monitor network security status in real time, analyze and prevent intrusion behaviors; Data receiving and verification module: The receiver decapsulates, reassembles, decrypts and verifies the integrity of the transmitted data; Ending and cleaning module: releases and cleans up resources after the communication is completed to prevent resource leakage.
8. The data communication system based on information security according to claim 7, characterized in that: The intrusion detection and defense module realizes comprehensive monitoring and analysis of network traffic through dynamic and adaptive IDS deployment, deep data packet feature extraction and real-time anomaly detection model update; the intrusion detection and defense module also includes an intrusion analysis submodule and an automatic defense submodule, wherein the intrusion analysis submodule uses an integrated machine learning model, a generative adversarial network and reinforcement learning technology to accurately determine whether there is an intrusion behavior; after detecting an intrusion, the automatic defense submodule takes intelligent blocking and isolation, collaborative defense and emergency response measures to ensure system security.
Citation Information
Patent Citations
Intrusion detection system and method based on intelligent network
CN118413406A
Substation network security defense system based on artificial intelligence
CN119276602A