Data outbound-oriented low-overhead multi-source risk information acquisition method and system

Through dynamic risk entropy and improved random forest algorithms, key collection points are screened, combined with inverse probability weight correction, the problems of resource waste and insufficient risk assessment in data outbound scenarios are solved, efficient and low-overhead risk information collection and compliance are achieved, and the accuracy and efficiency of cross-border data monitoring are improved.

CN120542908AActive Publication Date: 2025-08-26积至(海南)信息技术有限公司

Patent Information

Application Number
CN202510603004.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-26
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the data outbound scenario, the existing technology has problems such as redundant data collection leading to resource waste, insufficient flexibility of static risk assessment models, high risk information misreport rate, and lack of sample distribution deviation accumulation and compliance dynamic adaptation, which is difficult to meet the needs of efficient monitoring in cross-border high-bandwidth scenarios.

Method used

Dynamic risk entropy is used to integrate business criticality and traffic fluctuation to select key acquisition points, combine the improved random forest algorithm to perform multi-dimensional feature classification, and dynamically adjust model parameters through inverse probability weight correction and closed-loop feedback mechanism to achieve adaptive acquisition strategy regulation.

Benefits of technology

Significantly reduce redundant data collection, improve risk identification accuracy and resource efficiency, ensure compliance, and form a dynamic adaptive risk information collection process of "intelligent screening-graded acquisition-queue optimization-bias correction".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542908A_ABST
    Figure CN120542908A_ABST
Patent Text Reader

Abstract

The invention discloses a data outbound-oriented low-overhead multi-source risk information acquisition method and system in the technical field of data cross-border. The method comprises the following steps: S1, intelligent screening of main business risk acquisition points: based on dynamic risk entropy, business criticality and flow fluctuation are taken as screening key acquisition points; s2, linkage of communication behavior classification and risk assessment: capturing key acquisition point communication metadata based on an improved random forest algorithm, and realizing abnormal behavior detection and classification through a cooperative mechanism of feature extraction, model training, dynamic classification and feedback optimization; s3, dynamically regulating and controlling an acquisition strategy and a transmission queue, namely dynamically regulating the acquisition granularity and the transmission priority according to the risk level and the queue load, and balancing the resource overhead and the monitoring requirement; and S4, sample bias correction and closed-loop optimization: correcting sample distribution deviation through inverse probability weight, and feeding back a false alarm rate and a load state to acquisition point selection and model parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-border data, and in particular to a low-overhead multi-source risk information collection method and system for data outbound transmission. Background Art

[0002] With the acceleration of global data cross-border flows, data outbound flows face severe security and compliance challenges. Traditional technical solutions present many challenges in large-scale data monitoring:

[0003] Static data collection mechanisms lead to resource waste. Existing methods often use fixed collection points or uniform sampling strategies, failing to dynamically adjust the collection scope based on node risk. Consequently, large amounts of low-risk data are redundantly collected, resulting in wasted storage and bandwidth resources. For example, full monitoring of non-critical nodes such as log servers consumes over 60% of the transmission queue capacity, while the actual risk contribution rate is less than 5%. This extensive data collection model is unable to support the efficient monitoring needs of cross-border, high-bandwidth scenarios.

[0004] Secondly, the lack of flexibility and precision in risk assessment models has become a core bottleneck restricting risk identification. Traditional solutions rely on rule engines or simple threshold judgments, which can only identify known risk patterns and have a high rate of missing new attacks. Classification methods based on static models such as logistic regression and SVM have partially improved automation capabilities, but are limited by the single feature dimension and cannot capture high-level risk characteristics such as time fluctuations and address dispersion. Furthermore, model updates rely on manual intervention, making it difficult to respond to network traffic fluctuations in real time, resulting in a significant increase in false positives as business complexity increases.

[0005] Finally, the accumulation of sample bias and the lack of dynamic compliance adaptation further weaken the long-term effectiveness of the monitoring system. Existing collection strategies do not consider the authenticity of data distribution. After long-term operation, training samples deviate significantly from the distribution of full business traffic, and model performance decays over time. At the same time, static strategies are difficult to meet the dynamic requirements of multinational regulations. For example, GDPR requires that the granularity of personal data collection be ≤256 bytes, while traditional solutions use a fixed granularity, which poses both compliance risks and wastes resources. These defects result in the existing system's scalability, real-time performance, and compliance in cross-border scenarios being unable to meet actual needs. A new solution that integrates dynamic weighting, closed-loop optimization, and adaptive regulation is urgently needed.

[0006] Based on the above background research, the main similar patents include: "Data Outbound Security Compliance Control Method and System Based on Dynamic Risk Assessment" (CN116187766A), "A Data Outbound Compliance Path Recommendation Method and System Based on Classification and Grading" (CN118587069A), "A Data Outbound Compliance Control Method and System" (CN118484838A), "A Method and Device for Outbound Data Security Management" (CN118611894A), etc.

[0007] The defects of the above prior art are:

[0008] First, redundant data collection leads to wasted resources and the lack of flexibility of static risk assessment models. Traditional methods use fixed collection points or uniform sampling strategies, resulting in redundant monitoring of a large number of low-risk nodes, which consumes a large amount of transmission resources while insufficiently contributing to the actual risk. Furthermore, static models based on rule engines or simple threshold judgments cannot capture new attacks in dynamic network environments, have a high risk information underreporting rate, and are unable to respond to traffic fluctuations in real time.

[0009] Second, regarding the problems of accumulated sample distribution deviations in risk information collection and the lack of dynamic compliance adaptation, existing technologies cause distortion of training data due to long-term biased collection and cannot meet the real-time requirements of data export regulations in multiple countries. Summary of the Invention

[0010] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and in the abstract and title of the present invention to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.

[0011] Therefore, the purpose of the present invention is to provide a low-overhead, high-accuracy adaptive data outbound multi-source risk information collection method. By integrating business criticality and traffic fluctuations through dynamic risk entropy to screen key collection points, the improved random forest algorithm is combined to perform multi-dimensional feature classification of communication metadata, and resource adaptive collection strategy control is achieved based on the exponential decay-enhancement model and priority queue. The inverse probability weight correction and closed-loop feedback mechanism are further introduced to dynamically adjust the model parameters to adapt to the real-time requirements of regulations such as GDPR, thereby realizing a closed-loop dynamic collection method of "intelligent screening-tiered collection-queue optimization-deviation correction". While ensuring compliance, it can significantly improve the efficiency and accuracy of large-scale cross-border data risk monitoring.

[0012] To solve the above technical problems, the present invention provides a low-overhead multi-source risk information collection method for data export, which adopts the following technical solution: comprising the following steps:

[0013] Step S1: Intelligent screening of main business risk collection points: based on dynamic risk entropy, business criticality and traffic fluctuation as the key collection points for screening;

[0014] Step S2: Linking communication behavior classification and risk assessment: Based on an improved random forest algorithm, communication metadata from key collection points is captured. Through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization, abnormal behavior detection and classification are achieved.

[0015] Step S3: Dynamically adjust the collection strategy and transmission queue: Dynamically adjust the collection granularity and transmission priority according to the risk level and queue load to balance resource consumption and monitoring needs;

[0016] Step S4: Sample bias correction and closed-loop optimization: Correct the sample distribution bias through inverse probability weighting, and feed back the false alarm rate and load status to the collection point selection and model parameters.

[0017] Optionally, step S1 further includes: calculating a comprehensive risk value based on historical risk events, business dependencies, and real-time traffic fluctuations of data outbound business nodes; and selecting nodes with top P% of comprehensive risk values ​​as key collection points.

[0018] Optionally, step S1 specifically includes:

[0019] Step 1.1: Collect risk events for the business node over the past D days. Risk events are pre-defined events based on business data attributes. Calculate the risk entropy value to measure uncertainty.

[0020] Step 1.2: Consider the criticality of the business node in the entire business process;

[0021] Step 1.3: Introduce the real-time traffic peak ratio to identify abnormal nodes with sudden traffic increases;

[0022] Step 1.4: Generate a comprehensive risk score for the node by fusing risk entropy, business criticality, and traffic fluctuation indicators through dynamic weighting.

[0023] Optionally, the comprehensive risk score of the node ω i The calculation is as follows:

[0024] ω i =α·H i +β·C i +γ·V i (1)

[0025] Wherein, i represents the i-th business node; is the node risk entropy value, p k is the probability of occurrence of historical risk event K; N represents the total number of historical risk events that have occurred; C i∈[0,1], business critical system, determined by business process dependency, C i =1 indicates a core node; represents the flow fluctuation coefficient; α, β, γ are dynamic weight coefficients, satisfying α+β+γ=1, with initial values ​​of α=0.5, β=0.3, γ=0.2.

[0026] Optionally, step S2 specifically includes:

[0027] Rapid metadata feature extraction: Capture network communication raw data packets at key collection points, including IP addresses, port numbers, protocol headers, timestamps, and payload lengths, capture communication metadata in real time, and extract four-dimensional features:

[0028] Time fluctuation: Calculate the coefficient of variation of communication intervals to identify burst traffic;

[0029] Address dispersion: Count the distribution of destination addresses and discover unconventional cross-border routes;

[0030] Protocol compliance: marking the use of encryption protocols;

[0031] Payload characteristics: Analyze the packet length distribution and detect covert channels;

[0032] Define the eigenvector: F = <Δt, L, D, E>,

[0033] in, is the communication interval fluctuation coefficient, where σ t is the standard deviation of the communication time interval, μ t is the mean; is the risk level, MTU is the maximum transmission unit of the network; is the destination address concentration, f is the frequency of the destination address; is a binary variable used to determine encryption compliance.

[0034] Optionally, step S2 further includes: using an improved random forest model to classify communication behaviors into three categories: normal, suspicious, and illegal. The classification process of a single tree is as follows:

[0035] 1) Extract N samples from the training set, where N is the same size as the original dataset, to generate a differentiated training subset;

[0036] 2) Then perform random feature selection. When each node splits, randomly select 2 candidate features from the 4 features to reduce the risk of overfitting.

[0037] 3) Then, when making a node split decision, the optimal split point is selected based on the Gini impurity minimization criterion. The calculation formula is as follows:

[0038]

[0039] Where t is the current node to be split, p(i\t) is the proportion of samples of category i in node t;

[0040] The splitting goal is to maximize the reduction of Gini impurity:

[0041]

[0042] Among them, N is the total number of samples in the current node to be split, t left , t right are the left and right child nodes generated by splitting; N left is the left child node after splitting (t left ) contains the number of samples; N right is the right child node after splitting (t right ) the number of samples included;

[0043] 4) Repeat the splitting until the termination condition is met to generate a complete decision tree;

[0044] According to the above process, T decision trees are pre-trained in parallel, and the output is a set of decision trees and a set of feature importance weights;

[0045] For the real-time captured traffic, each tree starts from the root node, selects a split path based on the feature value, and finally reaches the leaf node to output the category label. Then, the prediction results of all trees are counted, and the category with the highest vote is output as the final category. The calculation formula is as follows:

[0046]

[0047] Among them, h t (F) is the prediction result of the t-th tree, I(·) is the indicator function, and its value is 1 when the prediction is category c, otherwise it is 0.

[0048] Optionally, step S2 further includes:

[0049] Based on the risk behavior classification results and system resource status, the risks are mapped to the following levels:

[0050]

[0051] Among them, N risk Number of high-risk behaviors in the current cycle; N max The maximum number of historical risks within the window period; Q c is the current transmission queue occupancy rate; Q max is the queue capacity upper limit; α1 and α2 are weight coefficients, and α1+α2=1; R max is the maximum value of the historical risk index; R minis the minimum value of the historical risk index, when R max =R min , L=3.

[0052] Optionally, step S3 specifically includes:

[0053] Step 3.1: Based on the risk level L (L∈{1,2,3}) output in step S2, the exponential decay enhancement model is used to dynamically adjust the acquisition parameters. This model uses an exponential function to synchronously adjust the acquisition granularity decay and sampling frequency enhancement to achieve a nonlinear balance between risk level and resource usage. The formula is as follows:

[0054] g=g base ·e -k(L-1) ,f=f base ·e k(L-1) (7)

[0055] Among them, g is the current node collection granularity; g base is the basic collection granularity, which can be set according to the business data type; f is the sampling frequency of the current node; f base is the basic sampling frequency, which can be adjusted according to the system processing capacity and business real-time requirements; k is the adjustment coefficient, which controls the rate of granularity attenuation and frequency enhancement;

[0056] Step 3.2: During step 3.1, monitor the transmission queue occupancy rate Q in real time. c , when Q c / Q max When the value is greater than 0.8, the load balancing rule is triggered: the collection granularity of non-critical data is automatically relaxed by 1.5 times, and its queue priority is lowered;

[0057] Step 3.3: Implement queue control based on priority index to achieve hierarchical transmission. The calculation formula is as follows:

[0058]

[0059] Among them, R t is the real-time risk coefficient, S d For data sensitivity, high-priority data with a PI ≥ 0.8 enters the real-time queue and is directly transmitted to the risk control engine using zero-copy technology to ensure end-to-end latency ≤ 50ms; ordinary data with a PI ≤ 0.8 is stored in the batch queue and processed asynchronously with a delay of ≤ 5min after compression; low-value data with a PI < 0.5 triggers a dynamic discard strategy: if Q c >90%, old data is eliminated according to the LRU algorithm.

[0060] Optionally, step S4 specifically includes:

[0061] Step 4.1: Compare the actual collection distribution of each collection node i with the real traffic distribution and calculate the inverse probability weight. The calculation formula is as follows:

[0062]

[0063] Among them, x i is the feature vector of the i-th sample; π obs (x i ) is the actual acquisition probability, indicating that the feature is x i The probability of a sample being selected under the current collection strategy can be obtained by statistical calculation of the collection log; true (x i ) is the true distribution probability, indicating that the feature is x i The probability of natural occurrence of samples in the full business traffic can be calculated through offline full mirror data; θ i is the weight adjustment factor, when π obs (x i )<π true (x i ), it means that insufficient samples of this type have been collected;

[0064] Step 4.2: Based on the ratio of false alarm rate to missed alarm rate, adjust the node weights of step S1 in reverse. When the false alarm rate is high, increase the historical risk entropy weight α to suppress over-mining. When the missed alarm rate is high, reduce the business criticality weight β to strengthen node coverage. The calculation formula is as follows:

[0065]

[0066] The false alarm rate / missing alarm rate can be calculated based on the statistical value of the classification result.

[0067] A low-overhead, multi-source risk information collection system for data export, used to implement the above method, includes:

[0068] Intelligent screening unit, based on dynamic risk entropy, business criticality and traffic fluctuation as key collection points for screening;

[0069] The linkage unit, based on an improved random forest algorithm, captures communication metadata from key collection points and implements abnormal behavior detection and classification through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization.

[0070] Dynamic control unit dynamically adjusts the collection granularity and transmission priority according to the risk level and queue load, balancing resource consumption and monitoring needs;

[0071] The correction and optimization unit corrects the sample distribution deviation through inverse probability weighting, and feeds back the false alarm rate and load status to the collection point selection and model parameters.

[0072] In summary, the present invention has at least one of the following beneficial effects:

[0073] 1. The present invention first selects key collection points based on the dynamic risk entropy, business criticality, and traffic fluctuations of data outbound business nodes to reduce redundant data collection. It then captures communication metadata from key nodes and, based on an improved random forest algorithm, achieves efficient and low-overhead abnormal behavior detection and classification through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization. It then dynamically adjusts the collection granularity and transmission priority based on the risk level and queue load to balance resource overhead and monitoring needs. Finally, it corrects sample distribution deviations through inverse probability weighting, and feeds back false alarm rates, load status, and other factors into collection point selection and model parameters, forming an adaptive risk information collection and processing flow of "intelligent screening - hierarchical collection - queue optimization - deviation correction."

[0074] 2. The present invention calculates the comprehensive risk value based on historical risk events of business node data outflow, business dependencies and real-time traffic fluctuations, and selects the nodes with the top P% comprehensive risk value as key collection points to ensure that most historical risk events are covered while reducing a large amount of redundant data collection.

[0075] 3. This invention captures communication metadata from key nodes and then, based on an improved random forest algorithm, achieves efficient and low-overhead detection and classification of abnormal behavior through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization. First, four-dimensional features are extracted from the communication metadata: time fluctuation, address dispersion, protocol compliance, and payload characteristics. Using the improved random forest model, communication behavior is classified into three risk categories: normal business, potential risk, and abnormal behavior. Risks are then mapped to the following levels based on the risk behavior classification results and system resource status.

[0076] 4. The collection strategy and transmission queue of the present invention are dynamically regulated to achieve precise control through real-time linkage between risk level and resource status. Based on the risk level output in stage 2, an exponential decay enhancement model is preferably used to dynamically adjust the collection parameters. The model synchronously adjusts the collection granularity (attenuation) and sampling frequency (enhancement) through an exponential function to achieve nonlinear balanced collection of risk level and resource occupancy, and then queue control is performed based on the priority index to achieve hierarchical transmission.

[0077] 5. This method achieves sample distribution correction by comparing the actual collection distribution of each collection node with the true traffic distribution and calculating the inverse probability weight. The node risk weights of Phase 1 are then adjusted inversely based on the ratio of the false alarm rate to the missed alarm rate. When the false alarm rate is high, the historical risk entropy weight is increased to suppress over-collection. When the missed alarm rate is high, the business-criticality weight is reduced to strengthen node coverage. This establishes a feedback iterative mechanism, ultimately forming a dynamic adaptive processing flow of "collection-correction-optimization" to continuously improve system accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0079] Figure 1 This is a flow chart of a low-overhead multi-source risk information collection method for data export in the present invention;

[0080] Figure 2 This is a flow chart of the classification process of a single tree in step S3 of the present invention;

[0081] Figure 3 This is a principle block diagram of a low-overhead multi-source risk information collection system for data export in the present invention. DETAILED DESCRIPTION

[0082] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making creative work shall fall within the scope of protection of the present invention.

[0083] In the description of the present invention, it should be noted that the terms "upper," "lower," "inner," "outer," "top / bottom," and the like, indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0084] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "installed," "provided with," "mounted / connected," and "connected" should be understood in a broad sense. For example, "connected" can mean a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be internal communication between two components. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention in specific circumstances.

[0085] Example 1

[0086] The present invention provides a low-overhead multi-source risk information collection method for data outbound risk monitoring. It integrates business criticality, historical risk events, and real-time traffic fluctuations through dynamic risk entropy, screens the top P% high-risk nodes for deployment of probes, reduces redundant data collection, and significantly reduces resource waste. An improved random forest algorithm is used to improve the accuracy of risk information classification based on multi-dimensional features such as time fluctuations and address dispersion of communication metadata, solving the defect of static models with a high underreporting rate of new attack risk information. At the same time, an inverse probability weight correction and sliding window mechanism are introduced to dynamically correct sample distribution deviations, and node risk weights are adjusted through a false alarm rate / missing alarm rate feedback closed loop to ensure the long-term stability of the model and regulatory adaptability, thereby realizing a dynamic and adaptive risk information collection process of "intelligent screening-tiered collection-queue optimization-deviation correction", which improves risk identification accuracy, resource efficiency, and compliance coverage in data cross-border scenarios.

[0087] Reference Figure 1 The present invention discloses a low-overhead multi-source risk information collection method for data export. The specific processing flow includes the following steps:

[0088] Step S1: Intelligent screening of main business risk collection points: Based on dynamic risk entropy, business criticality and traffic fluctuations as the key collection points for screening, redundant data collection is reduced; Business node data outbound risk quantification: Based on the historical risk events, business dependencies and real-time traffic fluctuations of the data outbound business nodes, the comprehensive risk value is calculated.

[0089] First, risk events from business nodes over the past D days are collected. Risk events are predefined based on business data attributes (e.g., data leaks, unencrypted private data, unencrypted data transmission), and risk entropy is calculated to measure uncertainty. Second, the criticality of the business node within the overall business process is considered (e.g., payment gateways are critical nodes, log servers are non-critical nodes). Finally, the real-time traffic peak ratio is introduced to identify abnormal nodes with sudden traffic increases. A comprehensive risk score for the node is generated by combining these three indicators using dynamic weights (risk entropy, business criticality, and traffic fluctuations).

[0090] The comprehensive risk score ω of the node i The calculation is as follows:

[0091] ω i =α·H i +β·C i +γ·V i (1)

[0092] Wherein, i represents the i-th business node; is the node risk entropy value, p k is the probability of occurrence of historical risk event K; N represents the total number of historical risk events that have occurred; C i ∈[0,1], business critical system, determined by business process dependency, C i =1 indicates a core node; represents the flow fluctuation coefficient; α, β, γ are dynamic weight coefficients, satisfying α+β+γ=1, and the preferred initial values ​​are α=0.5, β=0.3, γ=0.2.

[0093] Dynamic deployment decision of data collection probes: select nodes ranked in the top P% of comprehensive risk values ​​as key collection points to ensure coverage of most historical risk events while reducing a large amount of redundant data collection.

[0094] Step S2: Communication behavior classification and risk assessment linkage: Based on the improved random forest algorithm, the communication metadata of key collection points is captured, and through the collaborative mechanism of feature extraction, model training, dynamic classification and feedback optimization, efficient and low-cost abnormal behavior detection and classification are achieved; the specific processing flow is as follows Figure 2 As shown:

[0095] Rapid metadata feature extraction: Capture network communication raw data packets at key collection points, including IP addresses, port numbers, protocol headers, timestamps, and payload lengths, capture communication metadata in real time, and extract four-dimensional features:

[0096] Time fluctuation: Calculate the coefficient of variation of communication intervals to identify burst traffic (such as DDoS attacks);

[0097] Address dispersion: Counting the distribution of destination addresses to identify unconventional cross-border routes (such as illegal data outflow);

[0098] Protocol compliance: marking the use of encryption protocols (e.g., unencrypted transmission triggers high risk);

[0099] Payload characteristics: Analyze the data packet length distribution and detect covert channels (such as using small packet lengths to transmit sensitive data).

[0100] Define the eigenvector: F = <Δt, L, D, E>,

[0101] in, is the communication interval fluctuation coefficient, where σ t is the standard deviation of the communication time interval, μ t is the mean; is the risk level,(MTU is the maximum transmission unit of the network); is the destination address concentration, f is the frequency of the destination address; is a binary variable used to determine encryption compliance.

[0102] Classification and risk assessment linkage: Using an improved random forest model, communication behaviors are classified into three categories: normal, suspicious, and illegal. The classification process for a single tree is as follows:

[0103] 1) Extract N samples from the training set, where N is the same size as the original dataset, to generate a differentiated training subset;

[0104] 2) Then perform random feature selection. When each node splits, randomly select 2 candidate features from the 4 features to reduce the risk of overfitting.

[0105] 3) Then, when making a node split decision, the optimal split point is selected based on the Gini impurity minimization criterion. The calculation formula is as follows:

[0106]

[0107] Where t is the current node to be split, p(i\t) is the proportion of samples of category i in node t;

[0108] The splitting goal is to maximize the reduction of Gini impurity:

[0109]

[0110] Among them, N is the total number of samples in the current node to be split, t left , t right are the left and right child nodes generated by splitting; N left is the left child node after splitting (t left ) contains the number of samples; N right is the right child node after splitting (t right ) the number of samples included;

[0111] 5) Repeat the split until the termination condition is met. Repeat the split until the termination condition is met (for example: the number of node samples is less than 5 or the tree depth is greater than 10 or when N left When / N exceeds the preset range (such as <0.1 or >0.9), the current split is terminated to avoid overfitting, thereby generating a complete decision tree.

[0112] According to the above process, T decision trees are pre-trained in parallel, and the output is a set of decision trees and a set of feature importance weights;

[0113] For the real-time captured traffic, each tree starts from the root node, selects a split path based on the feature value, and finally reaches the leaf node to output the category label. Then, the prediction results of all trees are counted, and the category with the highest vote is output as the final category. The calculation formula is as follows:

[0114]

[0115] Among them, h t (F) is the prediction result of the t-th tree, I(·) is the indicator function, and its value is 1 when the prediction is category c, otherwise it is 0.

[0116] In the present invention, data outbound risks are preferably classified into normal business, potential risks, and abnormal behavior. To further manage and control risks, the risks are mapped to the following levels based on the risk behavior classification results and system resource status:

[0117]

[0118]

[0119] Among them, N risk Number of high-risk behaviors in the current cycle; N max The maximum number of historical risks within the window period; Q c is the current transmission queue occupancy rate; Q max is the queue capacity upper limit; α1 and α2 are weight coefficients, and α1+α2=1; R max is the maximum value of the historical risk index; R min It is the minimum value of the historical risk index.

[0120] When R max =R min , L=3.

[0121] Step S3: Dynamically adjust the collection strategy and transmission queue: Dynamically adjust the collection granularity and transmission priority according to the risk level and queue load to balance resource consumption and monitoring needs;

[0122] Specifically include:

[0123] Step 3.1: Based on the risk level L (L∈{1,2,3}) output in step S2, the exponential decay enhancement model is used to dynamically adjust the acquisition parameters. This model uses an exponential function to synchronously adjust the acquisition granularity decay and sampling frequency enhancement to achieve a nonlinear balance between risk level and resource usage. The formula is as follows:

[0124] g=g base ·e-k(L-1) ,f=f base ·e k(L-1) (7)

[0125] Among them, g is the current node collection granularity; g base is the basic collection granularity, which can be set according to the business data type; f is the sampling frequency of the current node; f base is the basic sampling frequency, which can be adjusted according to the system processing capacity and business real-time requirements; k is the adjustment coefficient, which controls the rate of granularity attenuation and frequency enhancement;

[0126] When L=1, the basic particle size g base = 256 bytes and low frequency f base =0.5Hz to perform baseline acquisition, capturing only metadata such as protocol type and address; when L=2, press g=g base ·e -k The granularity is reduced to approximately 128 bytes, while the sampling rate is increased to approximately 2 Hz, and sensitive fields (such as personal data identifiers) are fully recorded. If L = 3, the granularity is further reduced to approximately 64 bytes and the sampling rate is increased to approximately 10 Hz, and full traffic mirroring and real-time compliance rule matching are simultaneously enabled.

[0127] Step 3.2: During step 3.1, monitor the transmission queue occupancy rate Q in real time. c , when Q c / Q max When the value is >0.8, the load balancing rule is triggered: the collection granularity of non-critical data is automatically relaxed by 1.5 times, and its queue priority is lowered to avoid resource overload.

[0128] Step 3.3: Implement queue control based on priority index to achieve hierarchical transmission. The calculation formula is as follows:

[0129]

[0130] Among them, R t is the real-time risk coefficient, S d For data sensitivity, high-priority data with a PI ≥ 0.8 enters the real-time queue and is directly transmitted to the risk control engine using zero-copy technology to ensure end-to-end latency ≤ 50ms; ordinary data with a PI ≤ 0.8 is stored in the batch queue and processed asynchronously with a delay of ≤ 5min after compression; low-value data with a PI < 0.5 triggers a dynamic discard strategy: if Q c >90%, old data is eliminated according to the LRU algorithm.

[0131] Step S4: Sample bias correction and closed-loop optimization: Correct the sample distribution bias through inverse probability weighting, and feed back the false alarm rate and load status to the collection point selection and model parameters.

[0132] Specifically include:

[0133] Step 4.1: Compare the actual collection distribution of each collection node i with the real traffic distribution and calculate the inverse probability weight. The calculation formula is as follows:

[0134]

[0135] Among them, x i is the feature vector of the i-th sample; π obs (x i ) is the actual acquisition probability, indicating that the feature is x i The probability of a sample being selected under the current collection strategy can be obtained by statistical calculation of the collection log; true (x i ) is the true distribution probability, indicating that the feature is x i The probability of natural occurrence of samples in the full business traffic can be calculated through offline full mirror data; θ i is the weight adjustment factor, when π obs (x i )<π true (x i ), indicating insufficient sample collection for that type. For example, if log data accounts for 40% of actual traffic but only 20% is collected, the log samples are weighted twice to compensate for sampling bias. A sliding window (window size = minimum 1000 records or 5% of the total sample size) is used to dynamically update weights to avoid long-term bias accumulation.

[0136] Step 4.2: Based on the ratio of false alarm rate to missed alarm rate, adjust the node weights of step S1 in reverse. When the false alarm rate is high, increase the historical risk entropy weight α to suppress over-mining. When the missed alarm rate is high, reduce the business criticality weight β to strengthen node coverage. The calculation formula is as follows:

[0137]

[0138] The false alarm rate / missing alarm rate can be calculated based on the statistical value of the classification result.

[0139] This will establish a feedback and iteration mechanism, ultimately forming a dynamic adaptive processing flow of "acquisition-correction-optimization" to continuously improve system accuracy and efficiency.

[0140] Example 2

[0141] Based on the same concept as the first embodiment above, a low-overhead multi-source risk information collection system for data export is also provided. The system includes:

[0142] Intelligent screening unit, based on dynamic risk entropy, business criticality and traffic fluctuation as key collection points for screening;

[0143] The linkage unit, based on an improved random forest algorithm, captures communication metadata from key collection points and implements abnormal behavior detection and classification through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization.

[0144] Dynamic control unit dynamically adjusts the collection granularity and transmission priority according to the risk level and queue load, balancing resource consumption and monitoring needs;

[0145] The correction and optimization unit corrects the sample distribution deviation through inverse probability weighting, and feeds back the false alarm rate and load status to the collection point selection and model parameters.

[0146] The above are all preferred embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the scope of protection of the present invention.

Claims

1. A low-overhead multi-source risk information collection method for data export, characterized by: The following steps are involved: Step S1: Intelligent screening of main business risk collection points: based on dynamic risk entropy, business criticality and traffic fluctuation as the key collection points for screening; Step S2: Linking communication behavior classification and risk assessment: Based on an improved random forest algorithm, communication metadata from key collection points is captured. Through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization, abnormal behavior detection and classification are achieved. Step S3: Dynamically adjust the collection strategy and transmission queue: Dynamically adjust the collection granularity and transmission priority according to the risk level and queue load to balance resource consumption and monitoring needs; Step S4: Sample bias correction and closed-loop optimization: Correct the sample distribution bias through inverse probability weighting, and feed back the false alarm rate and load status to the collection point selection and model parameters.

2. The low-overhead multi-source risk information collection method for data export according to claim 1, characterized in that: The step S1 further includes: calculating a comprehensive risk value based on historical risk events, business dependencies, and real-time traffic fluctuations of data outbound business nodes; and selecting nodes with a top P% ranking in comprehensive risk value as key collection points.

3. The low-overhead multi-source risk information collection method for data export according to claim 2, characterized in that: The step S1 specifically includes: Step 1.1: Collect risk events for the business node over the past D days. Risk events are pre-defined events based on business data attributes. Calculate the risk entropy value to measure uncertainty. Step 1.2: Consider the criticality of the business node in the entire business process; Step 1.3: Introduce the real-time traffic peak ratio to identify abnormal nodes with sudden traffic increases; Step 1.4: Generate a comprehensive risk score for the node by fusing risk entropy, business criticality, and traffic fluctuation indicators through dynamic weighting.

4. The low-overhead multi-source risk information collection method for data export according to claim 3, characterized in that: The comprehensive risk score ω of the node i The calculation is as follows: oh i =α·H i +β·C i +γ·V i (1) Wherein, i represents the i-th business node; is the node risk entropy value, p k is the probability of occurrence of historical risk event K; N represents the total number of historical risk events that have occurred; C i ∈[0,1], business critical system, determined by business process dependency, C i =1 indicates a core node; represents the flow fluctuation coefficient; α, β, γ are dynamic weight coefficients, satisfying α+β+γ=1, with initial values ​​of α=0.5, β=0.3, γ=0.

2.

5. The low-overhead multi-source risk information collection method for data export according to claim 4, characterized in that: The step S2 specifically include: Rapid metadata feature extraction: Capture network communication raw data packets at key collection points, including IP addresses, port numbers, protocol headers, timestamps, and payload lengths, capture communication metadata in real time, and extract four-dimensional features: time Fluctuation: Calculate the coefficient of variation of communication intervals to identify burst traffic; Address dispersion: Count the distribution of destination addresses and discover unconventional cross-border routes; Protocol compliance: marking the use of encryption protocols; Payload characteristics: Analyze the packet length distribution and detect covert channels; Define the eigenvector: F = <Δt, L, D, E>, in, is the communication interval fluctuation coefficient, where σ t is the standard deviation of the communication time interval, μ t is the mean; is the risk level, MTU is the maximum transmission unit of the network; is the destination address concentration, f is the frequency of the destination address; is a binary variable used to determine encryption compliance.

6. The low-overhead multi-source risk information collection method for data export according to claim 5, characterized in that: The step S2 specifically includes: using the improved random forest model to classify communication behaviors into three categories: normal, suspicious, and illegal. The classification process of a single tree is as follows: 1) Extract N samples from the training set, where N is the same size as the original dataset, to generate a differentiated training subset; 2) Then perform random feature selection. When each node splits, randomly select 2 candidate features from the 4 features to reduce the risk of overfitting. 3) Then, when making a node split decision, the optimal split point is selected based on the Gini impurity minimization criterion. The calculation formula is as follows: Where t is the current node to be split, p(i\t) is the proportion of samples of category i in node t; The splitting goal is to maximize the reduction of Gini impurity: Among them, N is the total number of samples in the current node to be split, t left , t righ t are the left and right child nodes generated by splitting; N left is the left child node after splitting (t left ) contains the number of samples; N righ t is the right child node after splitting (t righ t ) the number of samples included; 4) Repeat the splitting until the termination condition is met to generate a complete decision tree; According to the above process, T decision trees are pre-trained in parallel, and the output is a set of decision trees and a set of feature importance weights; For the real-time captured traffic, each tree starts from the root node, selects a split path based on the feature value, and finally reaches the leaf node to output the category label. Then, the prediction results of all trees are counted, and the category with the highest vote is output as the final category. The calculation formula is as follows: Among them, h t (F) is the prediction result of the t-th tree, I(·) is the indicator function, and its value is 1 when the prediction is category c, otherwise it is 0.

7. The low-overhead multi-source risk information collection method for data export according to claim 6, characterized in that: The step S2 specifically further includes: Based on the risk behavior classification results and system resource status, the risks are mapped to the following levels: Among them, N risk Number of high-risk behaviors in the current cycle; N max The maximum number of historical risks within the window period; Q c is the current transmission queue occupancy rate; Q max is the queue capacity upper limit; α1 and α2 are weight coefficients, and α1+α2=1; R max is the maximum value of the historical risk index; R min is the minimum value of the historical risk index, when R max =R min , L=3.

8. The low-overhead multi-source risk information collection method for data export according to claim 7, characterized in that: The step S3 specifically includes: Step 3.1: Based on the risk level L (L∈{1,2,3}) output in step S2, the exponential decay enhancement model is used to dynamically adjust the acquisition parameters. This model uses an exponential function to synchronously adjust the acquisition granularity decay and sampling frequency enhancement to achieve a nonlinear balance between risk level and resource usage. The formula is as follows: g=g base ·e -k(L-1) ,f=f base ·e k(L-1) (7) Among them, g is the current node collection granularity; g base is the basic collection granularity, which can be set according to the business data type; f is the sampling frequency of the current node; f base is the basic sampling frequency, which can be adjusted according to the system processing capacity and business real-time requirements; k is the adjustment coefficient, which controls the rate of granularity attenuation and frequency enhancement; Step 3.2: During step 3.1, monitor the transmission queue occupancy rate Q in real time. c , when Q c / Q max When the value is greater than 0.8, the load balancing rule is triggered: the collection granularity of non-critical data is automatically relaxed by 1.5 times, and its queue priority is lowered; Step 3.3: Implement queue control based on priority index to achieve hierarchical transmission. The calculation formula is as follows: Among them, R t is the real-time risk coefficient, S d For data sensitivity, high-priority data with a PI ≥ 0.8 enters the real-time queue and is directly transmitted to the risk control engine using zero-copy technology to ensure end-to-end latency ≤ 50ms; ordinary data with a PI ≤ 0.8 is stored in the batch queue and processed asynchronously with a delay of ≤ 5min after compression; low-value data with a PI < 0.5 triggers a dynamic discard strategy: if Q c >90%, old data is eliminated according to the LRU algorithm.

9. The low-overhead multi-source risk information collection method for data export according to claim 8, characterized in that: The step S4 specifically includes: Step 4.1: Compare the actual collection distribution of each collection node i with the real traffic distribution and calculate the inverse probability weight. The calculation formula is as follows: Among them, x i is the feature vector of the i-th sample; π obs (x i ) is the actual acquisition probability, indicating that the feature is x i The probability of a sample being selected under the current collection strategy can be obtained by statistical calculation of the collection log; true (x i ) is the true distribution probability, indicating that the feature is x i The probability of natural occurrence of samples in the full business traffic can be calculated through offline full mirror data; θ i is the weight adjustment factor, when π obs (x i )<π true (x i ), it means that insufficient samples of this type have been collected; Step 4.2: Based on the ratio of false alarm rate to missed alarm rate, adjust the node weights of step S1 in reverse. When the false alarm rate is high, increase the historical risk entropy weight α to suppress over-mining. When the missed alarm rate is high, reduce the business criticality weight β to strengthen node coverage. The calculation formula is as follows: The false alarm rate / missing alarm rate can be calculated based on the statistical value of the classification result.

10. A low-overhead multi-source risk information collection system for data export, used to implement the method described in any one of claims 1 to 9, characterized in that: include: Intelligent screening unit, based on dynamic risk entropy, business criticality and traffic fluctuation as key collection points for screening; The linkage unit, based on an improved random forest algorithm, captures communication metadata from key collection points and implements abnormal behavior detection and classification through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization. Dynamic control unit dynamically adjusts the collection granularity and transmission priority according to the risk level and queue load, balancing resource consumption and monitoring needs; The correction and optimization unit corrects the sample distribution deviation through inverse probability weighting, and feeds back the false alarm rate and load status to the collection point selection and model parameters.

Citation Information

Patent Citations

  • Data cross-border compliance management and control method and device, computer equipment and storage medium

    CN114760149A

  • Data outbound security compliance management and control method and system based on dynamic risk assessment

    CN116187766A

  • Internet online sales data intelligent screening management system

    CN118503544A

  • Exit data security management method and device

    CN118611894A

  • Cross-border data circulation supervision method, platform, equipment and medium

    CN119048025A

Cited By

  • Smart park safety operation situation real-time monitoring system based on unified sensing base

    CN121585739A

  • Financial network isolation control method and system based on service awareness

    CN121967089A