Real-time network intrusion prevention method based on multi-source data fusion

By integrating multi-source data and analyzing machine learning models, the problem of misjudgment or missed judgment caused by the uniform time window in existing technologies has been solved, enabling efficient detection and rapid response to network intrusion behavior.

CN120474771BActive Publication Date: 2025-10-28HEFEI TANOVO INFORMATION SECURITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510601239.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-10-28
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In existing technologies, the unified time window analysis method may overlook abnormal login behavior within a second-level time scale, leading to misjudgment or missed judgment in network intrusion prevention.

Method used

A real-time network intrusion prevention method that integrates multi-source data is adopted. By dividing time windows with different time granularities and combining network traffic, system logs and application logs, the method uses granularity matching and time synchronization characteristics to analyze abnormal behavior using a machine learning model and triggers corresponding response measures according to the severity.

Benefits of technology

It improves the accuracy and real-time performance of abnormal behavior detection, ensures flexible defense strategies, enables rapid response to potential security threats, and reduces cybersecurity risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474771B_ABST
    Figure CN120474771B_ABST
Patent Text Reader

Abstract

This invention discloses a real-time network intrusion prevention method based on multi-source data fusion, specifically relating to the field of network defense technology. By determining the sampling frequency based on the time granularity differences of different data sources and dividing flexible time windows, this method extracts network behavior features such as granularity matching degree and time synchronization degree, and combines them with machine learning models for analysis to accurately identify abnormal behaviors in the network. After identifying abnormal behaviors, the system triggers corresponding alarm mechanisms based on the severity of the behavior and takes corresponding defensive measures, thereby achieving effective detection and response to complex network intrusions and significantly improving network security protection capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network defense technology, and specifically to a real-time network intrusion prevention method based on multi-source data fusion. Background Technology

[0002] Network intrusion prevention refers to the use of a series of technical means and strategies to prevent unauthorized access or attacks, ensuring the security and integrity of computer network systems. These defensive measures can include firewalls, intrusion detection systems, intrusion prevention systems, encryption technologies, authentication, vulnerability scanning, etc. The goal of network intrusion prevention is to promptly detect and respond to potential security threats, prevent data leaks, system damage, or service interruptions, thereby ensuring the secure operation of the network environment.

[0003] The existing technology has the following shortcomings:

[0004] In existing technologies, time series analysis methods can combine data from different time periods to discover progressively abnormal attack patterns. However, when analyzing login behavior on a sensitive system, system logs may record user login information every second, while network traffic may be tallied every minute. In such cases, using a uniform time window (e.g., 1 minute) for summary analysis may overlook abnormal login behavior occurring on a second-level time scale, leading to misjudgments or missed detections of attacks. Summary of the Invention

[0005] The purpose of this invention is to provide a real-time network intrusion prevention method based on multi-source data fusion to address the shortcomings of the prior art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a real-time network intrusion prevention method based on multi-source data fusion, comprising:

[0007] Real-time network behavior data is collected from multiple data sources, including network traffic data, system log data, and application log data.

[0008] Based on the time granularity differences of each data source, the sampling frequency of each data source is determined, and several time windows are divided according to the sampling frequency.

[0009] Feature extraction is performed on network behavior data within each time window, and the features include granularity matching degree and time synchronization degree.

[0010] Based on the extracted network behavior features, machine learning models are used to analyze the features and identify abnormal behaviors in the network.

[0011] When abnormal behavior is detected, an alarm mechanism is triggered, which then triggers corresponding response measures based on the severity of the abnormal behavior.

[0012] Preferably, the sampling frequency of each data source is determined based on the time granularity differences of each data source. Specifically, the sampling frequency of network traffic data is sampled every second or every millisecond; the sampling frequency of system log data is recorded every second or every minute; and the sampling frequency of application log data is recorded every second or every minute.

[0013] Preferably, the time window division includes:

[0014] Analyze the sampling frequency of each data source to determine the time granularity of the data source;

[0015] Adjust the time window to ensure that the time windows of all data sources are aligned. If the sampling frequencies of different data sources are different, use the smallest time granularity as the unified time window, and aggregate high-frequency data sources or interpolate low-frequency data sources.

[0016] A sliding window mechanism is adopted to ensure that network behavior data can be updated in real time within each time window, and the sliding step size is set to the minimum sampling frequency of the data source.

[0017] Preferably, feature extraction is performed on the network behavior data within each time window. The features include granularity matching degree and time synchronization degree. The granularity matching degree is obtained by defining data source X: network traffic data; data source Y: system log data; and converting these two data sources into discrete time series. Where n and m are the lengths of the two data sources, respectively, for each pair of time points and Calculate the probability of them occurring simultaneously, i.e.: ;in, This represents the number of times x and y appear simultaneously in the data, where n and m are the lengths of the time series X and Y, respectively. The marginal probability distributions p(x) and p(y) are calculated; the marginal probability distribution p(x) is the probability of each value in X appearing, calculated as follows: The marginal probability distribution p(y) is the probability of each value in Y occurring, and it is calculated as follows: Summing the joint probability p(x,y) and marginal probabilities p(x) and p(y) for each pair of x and y yields the mutual information of X and Y. The granularity matching degree is obtained by standardizing the mutual information value. EP stands for granularity matching degree, where max(I(X,Y)) is the maximum mutual information value calculated among all time series pairings.

[0018] The preferred method for obtaining time synchronization is as follows: Define data source X: network traffic data; data source Y: system log data, and convert these two data sources into discrete time series. Where n and m are the lengths of the two data sources, respectively; calculate the cross-correlation function between the two time series X and Y. And find the maximum cross-correlation value: During the calculation, τ gradually increases from 0 until it reaches the limit of the sequence length. The maximum cross-correlation value corresponds to the optimal time synchronization. After calculating the cross-correlation values ​​under all lags τ, the lag value that maximizes the cross-correlation value is selected. ;in, It is the lag value that maximizes the cross-correlation function; the corresponding maximum cross-correlation value is: This indicates that X and Y are lagging. The optimal time alignment or synchronization is found at a certain point, and the degree of time synchronization is expressed as the standardized result of the maximum cross-correlation value, as shown in the formula: In the formula, SX represents the time synchronization degree. It is the maximum cross-correlation value.

[0019] Preferably, based on the extracted network behavior features, a machine learning model is used to analyze the features and identify abnormal behaviors in the network, specifically:

[0020] Granularity matching degree and time synchronization degree are converted into comprehensive feature vectors. The comprehensive feature vectors are used as input to the machine learning model. The machine learning model uses the prediction of network abnormal behavior analysis value labels for each set of comprehensive feature vectors as the prediction objective and minimizes the sum of prediction errors for all network abnormal behavior analysis value labels as the training objective. The machine learning model is trained until the sum of prediction errors converges and the model training stops. The network abnormal behavior analysis value is determined based on the model output. The machine learning model is a multinomial regression model.

[0021] Preferably, the obtained network abnormal behavior analysis value is compared with a preset threshold. If the network abnormal behavior analysis value is greater than or equal to the preset threshold, it is determined to be abnormal behavior and response measures are taken; if the network abnormal behavior analysis value is less than the preset threshold, it is determined to be normal behavior and no further response is taken.

[0022] Preferably, when abnormal behavior is detected, the alarm mechanism triggers corresponding response measures based on the severity of the abnormal behavior, specifically including:

[0023] Based on the comparison between the network abnormal behavior analysis values ​​and the adjusted preset thresholds, the abnormal behaviors are classified into minor abnormalities, moderate abnormalities, and severe abnormalities.

[0024] Minor anomaly response measures: Generate and store alarm logs, send a notification to the network administrator to remind them to conduct further manual checks;

[0025] Moderate anomaly response measures: Record detailed abnormal behavior logs, automatically restrict access permissions for abnormal IP addresses, notify administrators, and strengthen monitoring;

[0026] Severe anomaly response measures: Automatically isolate the attack source, activate the intrusion prevention system to counteract it, notify the administrator, and urgently initiate backup and recovery operations.

[0027] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0028] 1. This invention effectively avoids the misjudgment or missed judgment problems caused by a uniform time window in existing technologies by flexibly dividing time windows with different time granularities. By combining network traffic data, system log data, and application log data, and utilizing features such as granularity matching degree and time synchronization degree, a machine learning model is used to analyze and predict abnormal network behavior, improving the accuracy and real-time performance of abnormal behavior detection. This method, by introducing a sliding window mechanism and multi-source data fusion, not only improves data synchronization but also enables real-time response to potential security threats.

[0029] 2. This invention automatically triggers different levels of alarm mechanisms by comparing predicted abnormal behavior analysis values ​​with preset thresholds, ensuring a flexible defense strategy. Minor anomalies are manually checked and alerted; moderate anomalies automatically restrict access and enhance monitoring; and severe anomalies automatically isolate the attack source and initiate countermeasures, ensuring rapid response and handling of network attacks, significantly improving network protection capabilities and reducing network security risks. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0031] Figure 1 This is a mind map of the method of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0033] For examples, see Figure 1As shown in this embodiment, the real-time network intrusion prevention method based on multi-source data fusion includes:

[0034] Real-time network behavior data is collected from multiple data sources, including network traffic data, system log data, and application log data.

[0035] Based on the time granularity differences of each data source, the sampling frequency of each data source is determined, and several time windows are divided according to the sampling frequency.

[0036] Feature extraction is performed on network behavior data within each time window, and the features include granularity matching degree and time synchronization degree.

[0037] Based on the extracted network behavior features, machine learning models are used to analyze the features and identify abnormal behaviors in the network.

[0038] When abnormal behavior is detected, an alarm mechanism is triggered, which then triggers corresponding response measures based on the severity of the abnormal behavior.

[0039] Real-time network behavior data is collected from multiple data sources, including network traffic data, system log data, and application log data, specifically:

[0040] Network traffic data refers to data packets transmitted over a computer network and related information about network connections. This data includes the amount of data transmitted, packet size, source and destination IP addresses, transmission protocols (such as TCP and UDP), port numbers, and communication direction.

[0041] Network traffic data can be captured using traffic acquisition tools (such as Wireshark and Tcpdump) or network monitoring devices (such as Intrusion Detection Systems (IDS) and Intrusion Prevention Systems (IPS). These devices monitor packet transmission in the network in real time and record detailed information for each packet.

[0042] System log data refers to event information recorded by the operating system and server during operation. This log data includes system-level events such as user logins, changes in operation permissions, system errors, service startup or shutdown, and hardware status.

[0043] System logs are typically generated and stored automatically by the operating system, such as Syslog in Linux and Event Viewer in Windows. Administrators can also configure the logging system to export important log events to a centralized log management system (such as ELK Stack or Splunk).

[0044] Application log data refers to log files generated during application runtime. These logs contain information such as the application's running status, user actions, and error reports. Application logs can record detailed information about the application's business processes, user behavior, API calls, and more.

[0045] Application logs are typically embedded by developers within applications (e.g., using tools like Log4j or Logback). These logs can be stored on the local file system or collected and managed using log aggregation tools such as Fluentd or Graylog.

[0046] By collecting these three types of data sources (network traffic data, system log data, and application log data), comprehensive information support can be provided for network intrusion prevention.

[0047] Network traffic data provides traffic-level insights, helping to identify abnormal traffic and external attacks. System log data reveals abnormal behavior at the operating system level, such as unauthorized logins and privilege escalation. Application log data provides detailed behavioral information at the application layer, helping to identify business-related attacks, such as SQL injection and API abuse. Integrating and analyzing this data enables multi-faceted monitoring and detection of network intrusions, improving the accuracy and real-time performance of intrusion prevention.

[0048] Different data sources (network traffic data, system log data, application log data) have different time granularities. The sampling frequency for each data source needs to be determined, specifically including:

[0049] Network traffic data is typically collected in packets, and the sampling frequency depends on the configuration settings of traffic analysis tools (such as Wireshark and Tcpdump) or network monitoring equipment (such as IDS / IPS systems). Based on network traffic characteristics, sampling is usually done in seconds or milliseconds. To ensure real-time performance, the sampling frequency of network traffic data is generally set to 1 Hz or 1000 Hz. For example, if a traffic acquisition tool captures the metadata of each packet (such as source IP, destination IP, port number, protocol, etc.), then the number of packets collected per second is the network traffic sampling frequency.

[0050] System logs are typically event logs automatically generated by the operating system, and are usually recorded once each time an event is triggered. The sampling frequency of system logs is based on the frequency of events recorded in the log. System logs are usually recorded according to the time of event occurrence, and the sampling frequency is not fixed, but it is generally recorded once per second or per minute (e.g., security events, login events, system errors, etc.). For example, whenever a user attempts to log in, the system generates a login event, and the sampling frequency of this event depends on the user's login frequency, typically per second, per minute, or per hour.

[0051] Application log data consists of runtime status logs within an application, typically recorded by developers using logging frameworks such as Log4j and Logback. The sampling frequency of application logs is usually determined by the application's operational frequency. For example, an application might log once for each user request, API call, or specific business process trigger, with a sampling frequency typically per second or per minute. For instance, a web application's login logs might record each user's login request per second or per minute.

[0052] Based on the sampling frequency of each data source, time windows are divided to ensure data consistency and synchronization. The specific details of time window division include:

[0053] The size of the time window should be based on the sampling frequency of each data source. If the data source has a high sampling frequency (e.g., network traffic is sampled once per second), the time window should be shorter; if the data source has a low sampling frequency (e.g., system logs are recorded once per minute), the time window can be longer.

[0054] Common time window sizes: High-frequency data sources (e.g., network traffic): Time window size is typically between 1 second and 1 minute. Medium-frequency data sources (e.g., system logs): Time window size is typically between 1 minute and 5 minutes. Low-frequency data sources (e.g., application logs): Time window size is typically between 5 minutes and 1 hour.

[0055] To ensure that time windows from different data sources are synchronized as much as possible, the smallest time granularity can be selected as the unified time window standard. For example, if network traffic data is sampled in seconds, while system logs and application logs are sampled in minutes, every minute can be chosen as the unified time window size. For network traffic data sampled in seconds, summaries (such as total traffic, total number of packets, number of connections, etc.) can be performed every minute. For system logs and application logs sampled in minutes, data in minutes can be used directly.

[0056] To achieve real-time analysis, a sliding window mechanism can be used. After each time window ends, the window slides forward and updates with new data. The sliding step size is typically set to the minimum sampling frequency of the data source to ensure timely data updates for each time window. For example, if network traffic data is sampled every second, then data collected every second will be added to the current window, and the window will slide once per second; if system log data is sampled every minute, then data collected every minute will be added to the current window, and the window will slide once per minute.

[0057] After determining the time windows for each data source, it is necessary to synchronize the data within different time windows. This involves mapping data from different data sources to the same time window and summarizing or interpolating data at different time scales to ensure that multi-source data can be analyzed within a unified time window. For low-frequency data, interpolation techniques (such as linear interpolation or time series interpolation) can be used to adjust it to the same time scale as high-frequency data for data fusion and analysis.

[0058] In real-time network intrusion prevention methods that integrate multi-source data, granularity matching is an important indicator for measuring the degree of difference in time granularity between different data sources. Higher granularity matching indicates better synchronization of the data sources within the time window, leading to higher reliability and accuracy of the analysis results.

[0059] Define data source X: network traffic data (sampled per second); data source Y: system log data (sampled per minute). Convert these two data sources into discrete time series. Here, n and m represent the lengths of the two data sources, respectively. For data sources with different sampling frequencies, it is often necessary to align their time points. For example, network traffic data may be sampled once per second, while system log data may be sampled once per minute. System log data should be interpolated or aggregated to ensure it is aligned with the network traffic data on the same time scale. For example, the minute-by-minute data from the system logs can be aggregated to the second-by-second data, or the network traffic data can be summarized by minute (e.g., the total traffic per minute).

[0060] For each pair of time points and To calculate the probability of their simultaneous occurrence, we can statistically analyze each pair ( This is achieved by determining the frequency of occurrence, which is then converted into probability. ;in, This represents the number of times x and y appear simultaneously in the data, where n and m are the lengths of the time series X and Y, respectively. The marginal probability distributions p(x) and p(y) are calculated; the marginal probability distribution p(x) is the probability of each value in X appearing, calculated as follows: The marginal probability distribution p(y) is the probability of each value in Y occurring, and it is calculated as follows: The mutual information of X and Y is obtained by summing the joint probability p(x,y) and marginal probabilities p(x) and p(y) for each pair of x and y, expressed as: A higher mutual information value I(X,Y) indicates a better match in time granularity between the two data sources. The granularity matching degree is calculated by standardizing the mutual information value. EP stands for Granularity Match, where max(I(X,Y)) is the maximum mutual information value calculated among all possible time series pairings. The closer the standardized mutual information value is to 1, the better the time granularity of the two data sources matches.

[0061] Time synchronization reflects the consistency of time alignment between different data sources and measures whether their time windows are synchronized.

[0062] Define data source X: network traffic data (sampled per second); data source Y: system log data (sampled per minute). Convert these two data sources into discrete time series. , where n and m are the lengths of the two data sources, respectively.

[0063] Calculate the cross-correlation function between two time series X and Y And find the maximum cross-correlation value: During the calculation, τ gradually increases from 0 until it reaches the limit of the sequence length. It should be noted that the lag τ can be negative (indicating that X leads Y) or positive (indicating that X lags Y).

[0064] The maximum cross-correlation value corresponds to optimal time synchronization. After calculating the cross-correlation values ​​for all lags τ, select the lag value that maximizes the cross-correlation value: ;in, This is the lag value that maximizes the cross-correlation function. The corresponding maximum cross-correlation value is: This indicates that X and Y are lagging. This provides optimal time alignment or synchronization. Time synchronization is typically expressed as the normalized result of the maximum cross-correlation value. To map it to a range of 0 to 1, the formula is: In the formula, SX represents the time synchronization degree. This is the maximum cross-correlation value. The closer the time synchronization value is to 1, the better the time synchronization between the two data sources and the more accurate the time window alignment. The closer the value is to 0, the worse the time synchronization.

[0065] Based on the extracted network behavior features, a machine learning model is used to analyze these features and identify abnormal behaviors in the network, specifically:

[0066] Convert the granularity matching degree and the time synchronization degree into a comprehensive feature vector, and use the comprehensive feature vector as the input of a machine learning model. The machine learning model takes predicting the network anomaly behavior analysis value label for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of prediction errors for all network anomaly behavior analysis value labels as the training target, and trains the machine learning model until the sum of prediction errors reaches convergence and then stops the model training. Determine the network anomaly behavior analysis value according to the model output result, where the machine learning model is a polynomial regression model.

[0067] Compare the obtained network anomaly behavior analysis value with a preset threshold. If the network anomaly behavior analysis value is greater than or equal to the preset threshold, it is determined as an abnormal behavior, and response measures (such as alarm, automatic protection, etc.) are taken; if the network anomaly behavior analysis value is less than the preset threshold, it is determined as a normal behavior and no further response is made.

[0068] When an abnormal behavior is identified, trigger an alarm mechanism, and the alarm mechanism triggers corresponding response measures according to the severity of the abnormal behavior.

[0069] Input the network anomaly behavior analysis value (such as \(\hat{y}\)) predicted by the trained machine learning model (such as a polynomial regression model).

[0070] Output If the network anomaly behavior analysis value predicted by the model is greater than or equal to the preset threshold \(T\), the network behavior is considered abnormal; otherwise, the network behavior is considered normal.

[0071] When an abnormal behavior is identified, the alarm mechanism triggers corresponding response measures according to the severity of the abnormal behavior. Specifically, it includes:

[0072] Based on the comparison between the network anomaly behavior analysis value and the adjusted preset threshold, determine the severity of the anomaly. According to the size of the predicted network anomaly behavior analysis value, it can be divided into different severity levels:

[0073] Minor anomaly: The network anomaly behavior analysis value is slightly higher than the preset threshold (for example: \(\hat{y}\geq T\) and \(\hat{y}<T + \delta\)), indicating a lower degree of anomaly, which may be a false alarm or a minor anomaly.

[0074] Moderate anomaly: The network anomaly behavior analysis value is close to or slightly higher than the preset threshold (for example: \(\hat{y}\geq T+\delta\) and \(\hat{y}<T + 2\delta\)), indicating that there is a certain threat to the system and further monitoring is required.

[0075] Severe anomaly: The network anomaly behavior analysis value is much higher than the preset threshold (for example: \(\hat{y}\geq T + 2\delta\)), indicating an obvious intrusion behavior or attack activity, and an emergency response is required.

[0076] Here, δ is a flexible threshold interval used to distinguish between mild, moderate, and severe anomalies.

[0077] Depending on the severity of the abnormal behavior, the alarm mechanism will trigger different response measures.

[0078] Minor anomaly response measures (low-level alert):

[0079] Log alerts: Generate and store logs of minor anomalies for later analysis and review.

[0080] Notify Administrator: Send a minor warning to the network administrator via email, SMS, or system notification.

[0081] No automatic defense: Since the risk of minor anomalies is low, automatic defense measures may not be taken, but continuous monitoring is required.

[0082] Moderate anomaly response measures (medium-level alert):

[0083] Log alerts: Log detailed information for moderate anomalies and mark them as events of concern.

[0084] Automatic access restriction: If the anomaly originates from a specific IP address or device, consider limiting the access frequency of that IP or suspending its access to prevent the potential spread of attacks.

[0085] Notify Administrator: Send a warning to the network administrator via SMS, email, or system notification, requesting manual verification of the anomaly.

[0086] Enhance monitoring: Strengthen monitoring of relevant network traffic, logs, and system behavior, and conduct security checks on potentially affected systems.

[0087] Severe Anomaly Response Measures (High-Level Alert):

[0088] Immediately log alerts: Generate and record detailed logs, and mark events as high-risk.

[0089] Automatically isolate attack sources: Automatically block the IP address or device of the attack source through firewalls or intrusion prevention systems (IDS / IPS) and cut off the network connection with the attack source.

[0090] Activate Intrusion Prevention System: Trigger intrusion prevention system or firewall rules to restrict or block malicious traffic.

[0091] Activate automatic countermeasures: Based on the preset automated emergency response plan, activate defensive measures (such as blocking attack paths, reconfiguring the system, etc.).

[0092] Notify the administrator: Notify the network administrator via high-priority email, SMS, telephone, or other means, requesting timely action.

[0093] Trigger backup and restore: Initiate backup and restore operations according to the organization's security policy to ensure the integrity of critical system data.

[0094] Even minor anomalies should be monitored to ensure no potential risks emerge. For moderate and severe anomalies, log analysis and backtracking should be performed after response to identify the attack source, attack method, and potential vulnerabilities. For moderate and severe anomalies, network administrators should manually intervene based on the detailed alarm information to further analyze and confirm the specific circumstances of the anomaly.

[0095] To improve the effectiveness of the alarm mechanism, the following optimization measures can be taken:

[0096] Reduce false positives: Reduce false positives for minor anomalies by adjusting the threshold δ and the model's sensitivity. Dynamically adjust the response strategy based on changes in network activity. For example, when multiple minor anomalies are triggered, the response level can be dynamically increased to moderate or severe anomalies. Continuously improve the system's accuracy in predicting anomalous behavior by utilizing historical data feedback and model optimization.

[0097] In this application, when abnormal network behavior is identified, the alarm mechanism triggers corresponding response measures based on the severity of the abnormal behavior. Minor anomalies are typically only logged and alerted; moderate anomalies automatically restrict some access and enhance monitoring; severe anomalies initiate automatic isolation of the attack source, activate countermeasures, and urgently notify the administrator. This response mechanism ensures a rapid and effective response to network security threats, maximizing the security of the network environment.

[0098] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0099] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A real-time network intrusion prevention method based on multi-source data fusion, characterized in that: include: Real-time network behavior data is collected from multiple data sources, including network traffic data, system log data, and application log data. Based on the time granularity differences of each data source, the sampling frequency of each data source is determined, and several time windows are divided according to the sampling frequency. Feature extraction is performed on network behavior data within each time window, and the features include granularity matching degree and time synchronization degree. The method for obtaining the granularity matching degree is as follows: Define data source X: network traffic data; data source Y: system log data; convert these two data sources into discrete time series. Where n and m are the lengths of the two data sources, respectively, for each pair of time points and Calculate the probability of them occurring simultaneously, i.e.: ;in, This represents the number of times x and y appear simultaneously in the data, where n and m are the lengths of the time series X and Y, respectively. The marginal probability distributions p(x) and p(y) are calculated; the marginal probability distribution p(x) is the probability of each value in X appearing, calculated as follows: The marginal probability distribution p(y) is the probability of each value in Y occurring, and it is calculated as follows: Summing the joint probability p(x,y) and marginal probabilities p(x) and p(y) for each pair of x and y yields the mutual information of X and Y. The granularity matching degree is obtained by standardizing the mutual information value. ;EP stands for granularity matching degree, where max(I(X,Y)) is the maximum mutual information value calculated among all time series pairings; The method for obtaining time synchronization is as follows: Define data source X: network traffic data; data source Y: system log data, and convert these two data sources into discrete time series. Where n and m are the lengths of the two data sources, respectively; calculate the cross-correlation function between the two time series X and Y. And find the maximum cross-correlation value: During the calculation, τ gradually increases from 0 until it reaches the limit of the sequence length. The maximum cross-correlation value corresponds to the optimal time synchronization. After calculating the cross-correlation values ​​under all lags τ, the lag value that maximizes the cross-correlation value is selected. ;in, It is the lag value that maximizes the cross-correlation function; the corresponding maximum cross-correlation value is: This indicates that X and Y are lagging. The optimal time alignment or synchronization is found at a certain point, and the degree of time synchronization is expressed as the standardized result of the maximum cross-correlation value, as shown in the formula: In the formula, SX represents the time synchronization degree. It is the maximum cross-correlation value; Based on the extracted network behavior features, machine learning models are used to analyze the features and identify abnormal behaviors in the network. When abnormal behavior is detected, an alarm mechanism is triggered, which then triggers corresponding response measures based on the severity of the abnormal behavior.

2. The real-time network intrusion prevention method based on multi-source data fusion according to claim 1, characterized in that: Based on the time granularity differences of each data source, the sampling frequency of each data source is determined, specifically including: network traffic data is sampled at a frequency of every second or every millisecond; system log data is sampled at a frequency of every second or every minute; and application log data is sampled at a frequency of every second or every minute.

3. The real-time network intrusion prevention method based on multi-source data fusion according to claim 2, characterized in that: The time window division includes: Analyze the sampling frequency of each data source to determine the time granularity of the data source; Adjust the time window to ensure that the time windows of all data sources are aligned. If the sampling frequencies of different data sources are different, use the smallest time granularity as the unified time window, and aggregate high-frequency data sources or interpolate low-frequency data sources. A sliding window mechanism is adopted to ensure that network behavior data can be updated in real time within each time window, and the sliding step size is set to the minimum sampling frequency of the data source.

4. The real-time network intrusion prevention method based on multi-source data fusion according to claim 1, characterized in that: Based on the extracted network behavior features, a machine learning model is used to analyze these features and identify abnormal behaviors in the network, specifically: Granularity matching degree and time synchronization degree are converted into comprehensive feature vectors. The comprehensive feature vectors are used as input to the machine learning model. The machine learning model uses the prediction of network abnormal behavior analysis value labels for each set of comprehensive feature vectors as the prediction objective and minimizes the sum of prediction errors for all network abnormal behavior analysis value labels as the training objective. The machine learning model is trained until the sum of prediction errors converges and the model training stops. The network abnormal behavior analysis value is determined based on the model output. The machine learning model is a multinomial regression model.

5. The real-time network intrusion prevention method based on multi-source data fusion according to claim 4, characterized in that: The obtained network abnormal behavior analysis value is compared with a preset threshold. If the network abnormal behavior analysis value is greater than or equal to the preset threshold, it is determined to be abnormal behavior and response measures are taken. If the network abnormal behavior analysis value is less than the preset threshold, it is judged as normal behavior and no further response is taken.

6. The real-time network intrusion prevention method based on multi-source data fusion according to claim 5, characterized in that: When abnormal behavior is detected, the alarm mechanism triggers corresponding response measures based on the severity of the abnormal behavior, specifically including: Based on the comparison between the network abnormal behavior analysis values ​​and the adjusted preset thresholds, the abnormal behaviors are classified into minor abnormalities, moderate abnormalities, and severe abnormalities. Minor anomaly response measures: Generate and store alarm logs, send a notification to the network administrator to remind them to conduct further manual checks; Moderate anomaly response measures: Record detailed abnormal behavior logs, automatically restrict access permissions for abnormal IP addresses, notify administrators, and strengthen monitoring; Severe anomaly response measures: Automatically isolate the attack source, activate the intrusion prevention system to counteract it, notify the administrator, and urgently initiate backup and recovery operations.

Citation Information

Patent Citations

  • Method, system and device for monitoring network information security vulnerabilities and storage medium

    CN119691753A

  • Systems and methods for correlating probability models with non-homogenous time dependencies to generate time-specific data processing predictions

    US20230305904A1