A low-overhead multi-source risk information collection method and system for data export

By dynamically selecting key collection points, using an improved random forest algorithm, and applying inverse probability weight correction, the problems of resource waste and insufficient risk assessment in data export scenarios are solved, achieving efficient and low-cost risk information collection and compliance adaptation.

CN120542908BActive Publication Date: 2025-11-25积至(海南)信息技术有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510603004.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-11-25
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as redundant data collection leading to resource waste, insufficient flexibility of static risk assessment models, high rate of missed risk information, accumulation of sample distribution bias, and lack of dynamic adaptation to compliance in cross-border data export scenarios, making it difficult to meet the needs of efficient monitoring in cross-border high-bandwidth scenarios.

Method used

The method employs dynamic risk entropy fusion to select key collection points based on business criticality and traffic fluctuations, combines an improved random forest algorithm for multi-dimensional feature classification, and achieves adaptive resource regulation through an exponential decay-enhancement model and priority queue. It also introduces an inverse probability weight correction mechanism to form a closed-loop dynamic collection method.

Benefits of technology

It achieves low-cost, high-accuracy data outbound risk monitoring, reduces redundant data collection, improves the accuracy and compliance of risk identification, dynamically adapts to changes in cross-border traffic, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542908B_ABST
    Figure CN120542908B_ABST
Patent Text Reader

Abstract

The application discloses a kind of low overhead multi-source risk information acquisition method and system for data export in the technical field of data cross-border, comprising the following steps: step S1: subject business risk acquisition point intelligent screening: based on dynamic risk entropy, business criticality and flow fluctuation as screening key acquisition point;Step S2: linkage behavior classification and risk assessment linkage: based on improved random forest algorithm, capture key acquisition point communication metadata, through the collaborative mechanism of feature extraction, model training, dynamic classification and feedback optimization, realize abnormal behavior detection and classification;Step S3: acquisition strategy and transmission queue dynamic regulation: according to risk level and queue load, dynamically adjust acquisition granularity and transmission priority, balance resource overhead and monitoring demand;Step S4: sample bias correction and closed-loop optimization: correct sample distribution deviation by inverse probability weight, and false positive rate, load state is fed back to acquisition point selection and model parameters.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data cross-border, and in particular to a low-overhead multi-source risk information collection method and system for data export. BACKGROUND

[0002] With the acceleration of global data cross-border flow, data export scenarios face serious security and compliance challenges. Traditional technical solutions have many challenges in large-scale data monitoring:

[0003] Static collection mechanism leads to resource waste. Existing methods mostly use fixed collection points or uniform sampling strategies, which fail to dynamically adjust the collection range according to node risk, resulting in a large amount of low-risk data being redundantly collected, causing storage and bandwidth resource waste. For example, full monitoring of non-critical nodes such as log servers occupies more than 60% of the transmission queue capacity, while the actual risk contribution rate is less than 5%. This extensive collection mode is difficult to support efficient monitoring needs in cross-border high-bandwidth scenarios.

[0004] Secondly, the flexibility and accuracy of the risk assessment model are the core bottleneck restricting risk identification. Traditional solutions rely on rule engines or simple threshold judgments, which can only identify known risk patterns and have a high false negative rate for new attacks. Classification methods based on static models such as logistic regression and SVM, although partially improve the automation capability, are limited by single feature dimension, and cannot capture high-order risk features such as time fluctuations and address dispersion. In addition, model updates rely on manual intervention, making it difficult to respond to network traffic fluctuations in real time, resulting in a significant increase in false positives as business complexity increases.

[0005] Finally, the accumulation of sample bias and the lack of dynamic adaptation of compliance further weaken the long-term effectiveness of the monitoring system. Existing collection strategies do not consider the authenticity of data distribution, and the training sample is severely biased from the business full flow distribution after long-term operation, and the model performance decays over time. At the same time, static strategies cannot meet the dynamic requirements of multi-country regulations, such as GDPR requiring personal data collection granularity ≤256 bytes, while traditional solutions use fixed granularity, which not only has compliance risks but also causes resource waste. These defects result in the existing system being unable to meet the actual demand in terms of scalability, real-time performance and compliance in cross-border scenarios, and a new solution that integrates dynamic weights, closed-loop optimization and adaptive control is urgently needed.

[0006] Based on the background research above, the main similar patents are: "Data outbound security compliance control method and system based on dynamic risk assessment" (CN116187766A), "Data outbound compliance path recommendation method and system based on classification and grading" (CN118587069A), "Data outbound compliance control method and system" (CN118484838A), "Method and device for outbound data security management" (CN118611894A) and the like.

[0007] Defects of the prior art described above:

[0008] First, the resource waste caused by redundant data collection and the lack of flexibility of static risk assessment model. The traditional method uses fixed collection points or uniform sampling strategy, which leads to a large number of low-risk nodes being monitored redundantly, occupying a large amount of transmission resources, while the actual risk contribution rate is insufficient. At the same time, the static model based on rule engine or simple threshold judgment cannot capture new attacks in dynamic network environment, and the risk information missing report rate is high and cannot respond to traffic fluctuations in real time;

[0009] Second, to solve the problems of sample distribution deviation accumulation and lack of compliance dynamic adaptation, the existing technology causes distortion of training data due to long-term bias collection, and cannot meet the real-time requirements of multi-country data outbound regulations. SUMMARY

[0010] This section aims to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the specification to avoid obscuring the purpose of this section, abstract and title, and such simplifications or omissions cannot be used to limit the scope of the present application.

[0011] Therefore, the purpose of the present application is to provide a low-cost, high-accuracy adaptive data outbound multi-source risk information collection method. By dynamically fusing risk entropy with traffic fluctuations to screen key collection points, combining an improved random forest algorithm for multi-dimensional feature classification of communication metadata, and based on an exponential decay-enhancement model and priority queue to realize resource-adaptive collection strategy regulation. Further introduce inverse probability weight correction and closed-loop feedback mechanism, dynamically adjust the model parameters to adapt to the real-time requirements of GDPR and other regulations, so as to realize the "intelligent screening-classified collection-queue optimization-deviation correction" closed-loop dynamic collection method, which can significantly improve the efficiency and accuracy of large-scale cross-border data risk monitoring while ensuring compliance.

[0012] To solve the above technical problems, the present application provides a low-cost multi-source risk information collection method for data outbound, which adopts the following technical scheme: comprising the following steps:

[0013] Step S1: Intelligent screening of main business risk collection points: based on dynamic risk entropy, business criticality, and traffic fluctuation as screening key collection points;

[0014] Step S2: Linkage of communication behavior classification and risk assessment: based on the improved random forest algorithm, the communication metadata of key collection points are captured, and through the collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization, abnormal behavior detection and classification are realized;

[0015] Step S3: Dynamic regulation of collection strategy and transmission queue: according to the risk level and queue load, the collection granularity and transmission priority are dynamically adjusted to balance resource consumption and monitoring demand;

[0016] Step S4: Sample bias correction and closed-loop optimization: correct sample distribution bias through inverse probability weighting, and feed false positive rate and load state to collection point selection and model parameters.

[0017] Optionally, the step S1 further comprises: calculating a comprehensive risk value based on historical risk events, business dependency relationships, and real-time traffic fluctuations of data outbound business nodes; and selecting the top P% of nodes with the highest comprehensive risk value as key collection points.

[0018] Optionally, the step S1 specifically comprises:

[0019] Step 1.1: Collect risk events of business nodes in the past D days, risk events are events defined in advance according to business data attributes, and calculate risk entropy value to measure uncertainty;

[0020] Step 1.2: Combine the criticality of the business node in the entire business process;

[0021] Step 1.3: Introduce real-time traffic peak proportion to identify abnormal nodes with traffic surge;

[0022] Step 1.4: Integrate risk entropy, business criticality, and traffic fluctuation through dynamic weight to generate a comprehensive risk score of the node.

[0023] Optionally, the comprehensive risk score ω i is calculated as follows:

[0024] ω i = α·H i + β·C i + γ·V i (1)

[0025] Where i represents the i-th business node; is the node risk entropy value, p k is the occurrence probability of historical risk event K; N represents the total amount of historical risk events; C i∈[0, 1], business critical system, determined by business process dependency, C i = 1 represents a core node; represents the traffic fluctuation coefficient; α, β, γ are dynamic weight coefficients, satisfying α + β + γ = 1, and the initial values are α = 0.5, β = 0.3, and γ = 0.2.

[0026] Optionally, the step S2 specifically comprises:

[0027] Metadata feature rapid extraction: capturing network communication raw data packets at key collection points, including IP address, port number, protocol header, timestamp, and payload length, real-time capturing of communication metadata, and extracting four-dimensional features:

[0028] Time fluctuation: calculating the coefficient of variation of communication interval to identify burst traffic;

[0029] Address dispersion: statistics of destination address distribution to find irregular cross-border paths;

[0030] Protocol compliance: marking encrypted protocol usage;

[0031] Payload feature: analyzing packet length distribution to detect hidden channels;

[0032] Definition of feature vector: F = <Δt, L, D, E>,

[0033] Wherein, is the communication interval fluctuation coefficient, wherein σ t is the standard deviation of communication time interval, μ t is the mean; is the risk level, and MTU is the network maximum transmission unit; is the destination address concentration, and f is the destination address occurrence frequency; is a binary variable, used to judge the encryption compliance.

[0034] Optionally, the step S2 specifically further comprises: through the improved random forest model, the communication behavior is divided into normal, suspicious and illegal three categories, and the classification processing flow of a single tree is:

[0035] 1) Extract N samples from the training set, N is consistent with the size of the original data set, and generate a differentiated training subset;

[0036] 2) Then randomly select features, and randomly select 2 candidate features from 4 features when each node is split, to reduce the risk of overfitting;

[0037] 3) Then make a node splitting decision, select the optimal splitting point based on the Gini impurity minimization criterion, and the calculation formula is as follows:

[0038]

[0039] Wherein, t is the current node to be split, p(i\t) is the proportion of samples of class i in node t;

[0040] The splitting target is to maximize the Gini impurity reduction:

[0041]

[0042] Wherein, N is the total number of samples in the current node to be split, t left , t right are the left and right child nodes generated by splitting respectively; N left is the number of samples contained in the left child node (t left ) after splitting; N right is the number of samples contained in the right child node (t right ) after splitting;

[0043] 4) Repeat the splitting until the termination condition is met, and generate a complete decision tree;

[0044] According to the above processing flow, T decision trees are pre-trained in parallel, and the output results are a decision tree set and a feature importance weight set;

[0045] For real-time captured traffic, from the root node of each tree, the splitting path is selected according to the feature value, and finally the class label is output at the leaf node. Then the prediction results of all trees are counted, and the highest-vote class is output as the final class. The calculation formula is as follows:

[0046]

[0047] Wherein, h t (F) is the prediction result of the tth tree, I(·) is an indicator function, which is 1 when the prediction is class c, and 0 otherwise.

[0048] Optionally, the step S2 further comprises:

[0049] According to the risk behavior classification result and the system resource state, the risk is mapped as follows:

[0050]

[0051] Wherein, N risk is the number of high-risk behaviors in the current period; N max is the maximum number of risks in the history within the window period; Q c is the current transmission queue occupancy rate; Q max is the queue capacity upper limit; α1 and α2 are weight coefficients, and α1+α2=1; R max is the maximum value of the historical risk index; R minThe minimum value of the historical risk index is when R max = R min , L = 3.

[0052] Optionally, the step S3 specifically comprises:

[0053] Step 3.1: Based on the risk level L (L e {1, 2, 3}) output by step S2, a exponential decay enhancement model is used to dynamically adjust the collection parameters, which synchronously adjusts the collection granularity decay and the sampling frequency enhancement through an exponential function, realizes the nonlinear balance of risk level and resource occupation, and the formula is as follows:

[0054] g = g base · e -k(L-1) , f = f base · e k(L-1) (7)

[0055] Wherein, g is the current node collection granularity; g base is the basic collection granularity, which can be set according to the business data type; f is the current node sampling frequency; f base is the basic sampling frequency, which can be adjusted according to the system processing capacity and business real-time requirement; k is the adjustment coefficient, which controls the rate of granularity decay and frequency enhancement;

[0056] Step 3.2: During the execution of step 3.1, the transmission queue occupancy rate Q c is monitored in real time, and when Q c / Q max > 0.8, the load balancing rule is triggered: the collection granularity of non-critical data is relaxed by 1.5 times, and the queue priority is reduced;

[0057] Step 3.3: Based on the priority index, the queue control is realized to realize hierarchical transmission, and the calculation formula is as follows:

[0058]

[0059] Wherein, R t is the real-time risk coefficient, S d is the data sensitivity, high priority data PI > 0.8 enters the real-time queue, uses zero copy technology to directly transmit to the risk control engine, and ensures that the end-to-end delay is less than or equal to 50ms; ordinary data 0.5 <= PI < 0.8 is stored in the batch queue, and after compression, it is processed asynchronously with a delay of less than or equal to 5min; low value data PI < 0.5 triggers a dynamic discard strategy: if Q c > 90% for 3 consecutive periods, the old data is eliminated according to the LRU algorithm.

[0060] Optionally, the step S4 specifically comprises:

[0061] Step 4.1: Calculate the inverse probability weight by comparing the actual collection distribution of each collection node i with the real traffic distribution, the calculation formula is as follows:

[0062]

[0063] wherein x i is the feature vector of the i-th sample; π obs (x i ) is the actual collection probability, indicating the probability of the sample with feature x i being selected under the current collection strategy, which can be calculated by collecting logs; π true (x i ) is the real distribution probability, indicating the natural occurrence probability of the sample with feature x i in the full traffic, which can be calculated by offline full-mirror data; θ i is the weight adjustment factor, when π obs (x i ) < π true (x i ), it indicates that the collection of this type of sample is insufficient.

[0064] Step 4.2: According to the false positive rate and false negative rate ratio, adjust the node weight of step S1 in reverse, when the false positive rate is high, increase the historical risk entropy weight α, and suppress over-sampling, when the false negative rate is high, reduce the weight β of business criticality, and strengthen the node coverage, the calculation formula is as follows:

[0065]

[0066] wherein the false positive rate / false negative rate can be calculated based on the statistical value of the classification result.

[0067] A low-overhead multi-source risk information collection system for data export, for realizing the above method, comprising:

[0068] Intelligent screening unit, based on dynamic risk entropy, business criticality and traffic fluctuation as the screening key collection point;

[0069] Linkage unit, based on improved random forest algorithm, capturing key collection point communication metadata, through the cooperative mechanism of feature extraction, model training, dynamic classification and feedback optimization, realizing abnormal behavior detection and classification;

[0070] Dynamic regulation unit, dynamically adjusting the collection granularity and transmission priority according to the risk level and queue load, balancing resource overhead and monitoring demand;

[0071] Correction and optimization unit, correcting sample distribution deviation by inverse probability weight, and feeding back false positive rate and load state to collection point selection and model parameters.

[0072] In summary, the present application includes at least one of the following benefits:

[0073] 1、The present application first selects key collection points based on the dynamic risk entropy of data outbound service nodes, service criticality and traffic fluctuations, reducing redundant data collection; then captures key node communication metadata, and through the collaborative mechanism of feature extraction, model training, dynamic classification and feedback optimization based on the improved random forest algorithm, realizes efficient and low-cost anomaly behavior detection and classification; further adjusts the collection granularity and transmission priority according to the risk level and queue load, balances resource overhead and monitoring demand; finally corrects the sample distribution deviation through inverse probability weight, and feeds back the false positive rate, load state, etc. to the collection point selection and model parameters, forming an adaptive risk information collection and processing process of "intelligent screening - hierarchical collection - queue optimization - deviation correction".

[0074] 2、The present application calculates the comprehensive risk value based on the historical risk events, business dependency relationship and real-time traffic fluctuations of the data outbound of the service nodes, selects the nodes with the top P% of comprehensive risk value as the key collection points, ensures that most of the historical risk events are covered, and at the same time reduces a large amount of redundant data collection.

[0075] 3、The present application captures key node communication metadata, and then based on the improved random forest algorithm through the collaborative mechanism of feature extraction, model training, dynamic classification and feedback optimization, realizes efficient and low-cost anomaly behavior detection and classification. First, four-dimensional features are extracted from the communication metadata: time fluctuation, address dispersion, protocol compliance, and load characteristics. Through the improved random forest model, the communication behavior is divided into three risk types: normal business, potential risk and abnormal behavior, and according to the risk behavior classification result and system resource state, the risk is mapped as follows.

[0076] 4、The present application dynamically adjusts the collection strategy and transmission queue, realizes accurate control through real-time linkage of risk level and resource state, optimally adopts an exponential decay enhancement model to dynamically adjust the collection parameters based on the risk level output in stage 2, the model synchronously adjusts the collection granularity (decay) and sampling frequency (enhancement) through the exponential function, realizes nonlinear balanced collection of risk level and resource occupation, and then performs queue control based on the priority index to realize hierarchical transmission.

[0077] 5、The application realizes sample distribution correction by calculating inverse probability weight through comparing the actual collection distribution of each collection node with the real traffic distribution. Then, according to the false alarm rate and the false negative rate ratio, the node risk weight of stage 1 is adjusted reversely. When the false alarm rate is high, the historical risk entropy weight is improved, and the over-collection is inhibited. When the false negative rate is high, the weight of business criticality is reduced, and the node coverage is strengthened, so as to realize the establishment of a feedback iteration mechanism, and finally form a dynamic adaptive processing flow of "collection-correction-optimization", continuously improving the system accuracy and efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0079] Figure 1 A flow chart of a low-overhead multi-source risk information collection method for data outbound of the present application;

[0080] Figure 2 A classification processing flow chart of a single tree of step S3 of the present application;

[0081] Figure 3 A principle block diagram of a low-overhead multi-source risk information collection system for data outbound of the present application. DETAILED DESCRIPTION

[0082] The technical solutions of the embodiments of the present application will be described clearly and completely in the following by combining the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0083] In the description of the present application, it should be noted that the terms "upper", "lower", "inner", "outer", "top / bottom end" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

[0084] In the description of the present application, it should be pointed out that, unless otherwise explicitly specified and limited, the terms "mounting", "provided with", "sleeved / connected", "connected" and the like should be understood broadly, for example, "connected" can be fixedly connected, can also be detachably connected, or integrally connected, can be mechanically connected, can also be electrically connected, can be directly connected, can also be indirectly connected through an intermediate medium, and can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0085] Embodiment one

[0086] The low-overhead multi-source risk information collection method for data outbound risk monitoring provided by the present application reduces resource waste by screening the first P% high-risk node to deploy a probe through dynamic risk entropy fusion of business criticality, historical risk events and real-time traffic fluctuations, and reduces redundant data collection. The improved random forest algorithm is used to improve the accuracy of risk information classification for the multi-dimensional features such as time fluctuation and address dispersion of communication metadata, solving the defect of high false negative rate of static models for new attack risk information. At the same time, the inverse probability weight correction and sliding window mechanism are introduced to dynamically correct the sample distribution deviation, and the false positive rate / miss rate feedback closed loop is used to adjust the node risk weight, ensuring the long-term stability and regulatory adaptability of the model, so as to realize the dynamic adaptive risk information collection process of "intelligent screening - hierarchical collection - queue optimization - deviation correction", and improve the risk identification accuracy, resource efficiency and compliance coverage in the data cross-border scenario.

[0087] Reference Figure 1 The present application discloses a low-overhead multi-source risk information collection method for data outbound, and the specific processing flow includes the following steps:

[0088] Step S1: Intelligent screening of main business risk collection points: based on dynamic risk entropy, business criticality and traffic fluctuation as the key screening points to reduce redundant data collection; risk quantification of data outbound business nodes: based on the historical risk events, business dependency relationship and real-time traffic fluctuation of data outbound business nodes, the comprehensive risk value is calculated.

[0089] Firstly, the risk events of the business node in the past D days are collected, the risk events are events (such as data leakage, privacy data without desensitization, unencrypted data transmission, etc.) defined according to the business data attributes, and the risk entropy value is calculated to measure the uncertainty; secondly, the criticality of the business node in the whole business process (such as payment gateway as a critical node, log server as a non-critical node) is combined; finally, the real-time traffic peak ratio is introduced to identify abnormal nodes with traffic surge. Through dynamic weight (risk entropy, business criticality, traffic fluctuation) fusion of the three indexes, the comprehensive risk score of the node is generated.

[0090] The comprehensive risk score ω of the node i is calculated as follows:

[0091] ω i = α · H i + β · C i + γ · V i (1)

[0092] where i represents the i-th business node; is the node risk entropy value, p k is the occurrence probability of historical risk events K; N represents the total amount of historical risk events; C i ∈ [0, 1], the business criticality system, determined by the degree of dependence of the business process, C i = 1 represents a core node; represents the traffic fluctuation coefficient; α, β, γ are dynamic weight coefficients, satisfying α + β + γ = 1, and the initial values are preferably α = 0.5, β = 0.3, and γ = 0.2.

[0093] Data collection probe dynamic deployment decision: select the nodes with the top P% comprehensive risk values as key collection points to ensure that most historical risk events are covered while reducing a large amount of redundant data collection.

[0094] Step S2: Linkage behavior classification and risk assessment linkage: based on the improved random forest algorithm, capture the communication metadata of the key collection points, and through the cooperative mechanism of feature extraction, model training, dynamic classification, and feedback optimization, realize efficient and low-cost abnormal behavior detection and classification; the specific processing flow is as shown in Figure 2

[0095] Metadata feature rapid extraction: capture network communication raw data packets at key collection points, including IP address, port number, protocol header, timestamp, and payload length, capture communication metadata in real time, and extract four-dimensional features:

[0096] Time fluctuation: calculate the coefficient of variation of communication interval to identify sudden traffic (such as DDoS attack);

[0097] Address dispersion: statistics of destination address distribution to find irregular cross-border paths (such as data violation export);

[0098] Protocol compliance: mark the use of encrypted protocols (such as non-encrypted transmission triggering high risk);

[0099] Payload feature: analyze the length distribution of data packets to detect hidden channels (such as using small packet length to transmit sensitive data).

[0100] Define the feature vector: F = <Δt, L, D, E>, ​

[0101] wherein, is the coefficient of interval fluctuation of communication, wherein σ t is the standard deviation of communication time interval, μ t is the mean value; is the risk level, (MTU is the maximum transmission unit of network); is the destination address concentration, f is the frequency of occurrence of destination address; is a binary variable, used to determine the encryption compliance.

[0102] Classification and risk assessment linkage: through the improved random forest model, the communication behavior is divided into normal, suspicious and illegal three categories, and the classification processing flow of a single tree is as follows:

[0103] 1) Extract N samples from the training set, N is consistent with the size of the original data set, and generate a differentiated training subset;

[0104] 2) Then perform random feature selection, and when each node is split, randomly select 2 candidate features from 4 features to reduce the risk of overfitting;

[0105] 3) Then make a node splitting decision, select the optimal splitting point based on the Gini impurity minimization criterion, and the calculation formula is as follows:

[0106]

[0107] wherein, t is the current node to be split, p(i\t) is the proportion of samples of class i in node t;

[0108] The splitting target is to maximize the Gini impurity drop:

[0109]

[0110] wherein, N is the total number of samples in the current node to be split, t left , t right are the left and right child nodes generated by splitting respectively; N left is the number of samples contained in the left child node (t left ) after splitting; N right is the number of samples contained in the right child node (t right ) after splitting;

[0111] 5) Repeat the splitting until the termination condition is met (such as: the number of node samples <5 or the tree depth >10 or when N left / N exceeds the preset range (such as <0.1 or >0.9)), end the current splitting to avoid overfitting, and thus generate a complete decision tree.

[0112] According to the above processing flow, the T decision trees are pre-trained in parallel, and the output results are a decision tree set and a feature importance weight set;

[0113] For real-time captured traffic, from the root node of each tree, a split path is selected according to the feature value, finally reaching the leaf node to output the category label, and then the prediction results of all trees are counted, and the highest-vote category is output as the final category, and the calculation formula is as follows:

[0114]

[0115] Wherein, h t (F) is the prediction result of the tth tree, I(·) is an indicator function, and the value is 1 when the prediction is class c, otherwise 0.

[0116] In the present application, the data outbound risk classification is preferably divided into: normal business, potential risk, abnormal behavior; in order to further accurately control the risk, according to the risk behavior classification result, the system resource state is mapped as follows:

[0117]

[0118]

[0119] Wherein, N risk The number of high-risk behaviors in the current period; N max The maximum number of risks in the history in the window period; Q c The current transmission queue occupancy rate; Q max The queue capacity upper limit; α1 and α2 are weight coefficients, and α1+α2=1; R max The maximum value of the historical risk index; R min The minimum value of the historical risk index.

[0120] When R max =R min , L=3.

[0121] Step S3: acquisition strategy and transmission queue dynamic control: according to the risk level and queue load, dynamically adjust the acquisition granularity and transmission priority, balance resource consumption and monitoring demand;

[0122] Specifically, it includes:

[0123] Step 3.1: based on the risk level L (L∈{1,2,3}) output by step S2, the acquisition parameters are dynamically adjusted by using an exponential decay enhancement model, which synchronously adjusts the acquisition granularity decay and sampling frequency enhancement through an exponential function, realizes the nonlinear balance of risk level and resource occupation, and the formula is as follows:

[0124] g=g base ·e-k(L-1) f = f base ·e k(L-1) (7)

[0125] Where g is the granularity of data collection at the current node; g base The basic data collection granularity can be set according to the data type of the business; f is the sampling frequency of the current node; f base The base sampling frequency can be adjusted according to the system's processing capacity and the real-time requirements of the business; k is an adjustment coefficient that controls the rate of granular attenuation and frequency enhancement.

[0126] When L=1, the basic particle size is g base =256 bytes and low frequency f base =0.5Hz performs baseline acquisition, capturing only metadata such as protocol type and address; when L=2, press g=g base ·e -k The granularity is reduced to approximately 128 bytes, while the sampling rate is increased to approximately 2Hz, and all sensitive fields (such as personal data identifiers) are recorded. If L=3, the granularity is further reduced to approximately 64 bytes and the sampling rate is increased to approximately 10Hz, while full traffic mirroring and real-time compliance rule matching are enabled simultaneously.

[0127] Step 3.2: During the execution of step 3.1, monitor the transmission queue occupancy rate Q in real time. c When Q c / Q max When the value is greater than 0.8, the load balancing rule is triggered: the collection granularity of non-critical data is automatically widened by 1.5 times and its queue priority is reduced to avoid resource overload.

[0128] Step 3.3: Implement hierarchical transmission by using queue control based on priority index. The calculation formula is as follows:

[0129]

[0130] Among them, R t To measure the real-time risk coefficient, S d For data sensitivity, high-priority data with a PI ≥ 0.8 enters the real-time queue and is directly transmitted to the risk control engine using zero-copy technology, ensuring an end-to-end latency of ≤ 50ms. Ordinary data with a PI ≤ 0.5 and a PI < 0.8 is stored in a batch queue, compressed, and processed asynchronously with a latency of ≤ 5 minutes. Low-value data with a PI < 0.5 triggers a dynamic discard strategy: if PI < 0.5 for three consecutive periods... c If the threshold is >90%, old data will be evicted using the LRU algorithm.

[0131] Step S4: Sample bias correction and closed-loop optimization: Correct sample distribution bias through inverse probability weights, and feed back false alarm rate and load status to the selection of collection points and model parameters.

[0132] Specifically comprising:

[0133] Step 4.1: Calculate the inverse probability weight by comparing the actual collection distribution of each collection node i with the true traffic distribution, and the calculation formula is as follows:

[0134]

[0135] Wherein, x i is the feature vector of the i-th sample; π obs (x i ) is the actual collection probability, which represents the probability that the sample with feature x i is selected under the current collection strategy, which can be calculated by collecting logs; π true (x i ) is the true distribution probability, which represents the natural occurrence probability of the sample with feature x i in the full traffic, which can be calculated by offline full-mirror data; θ i is the weight adjustment factor, which represents that the collection of this type of sample is insufficient when π obs (x i ) < π true (x i ); for example, if log data accounts for 40% of real traffic but only 20% is collected, then log samples are weighted by 2 times to compensate for sampling bias. A sliding window (window size = minimum 1000 or 5% of total sample size) is used to dynamically update the weight to avoid long-term bias accumulation.

[0136] Step 4.2: According to the false positive rate and false negative rate ratio, adjust the node weight of step S1 in reverse, when the false positive rate is high, increase the historical risk entropy weight α, and suppress over-sampling, when the false negative rate is high, reduce the weight β of business criticality, and strengthen the node coverage, and the calculation formula is as follows:

[0137]

[0138] Wherein, the false positive rate / false negative rate can be calculated based on the statistical value of the classification result.

[0139] Thus, a feedback iteration mechanism is established, and a "collection-correction-optimization" dynamic adaptive processing flow is finally formed, continuously improving system accuracy and efficiency.

[0140] Embodiment two

[0141] Based on the same idea as embodiment one, a low-overhead multi-source risk information collection system for data export is also included, the system comprises:

[0142] An intelligent screening unit based on dynamic risk entropy, business criticality and traffic fluctuation as a screening key collection point;

[0143] The linkage unit captures key collection point communication metadata based on an improved random forest algorithm, realizes abnormal behavior detection and classification through the cooperative mechanism of feature extraction, model training, dynamic classification and feedback optimization;

[0144] The dynamic regulation unit dynamically adjusts the collection granularity and transmission priority according to the risk level and queue load, balances the resource overhead and monitoring demand;

[0145] The correction and optimization unit corrects the sample distribution deviation through inverse probability weight, and feeds back the false alarm rate and load state to the collection point selection and model parameters.

[0146] The above are preferred embodiments of the present application, not limited to the protection scope of the present application, therefore: any equivalent changes made according to the structure, shape, principle of the present application should be covered within the protection scope of the present application.

Claims

1. A low-overhead, multi-source risk information collection method for cross-border data transfer, characterized in that: Includes the following steps: Step S1: Intelligent screening of risk collection points for main business: Screening key collection points based on dynamic risk entropy, business criticality and traffic fluctuations; Step S2: Linking Communication Behavior Classification and Risk Assessment: Based on the improved random forest algorithm, capture key data collection points' communication metadata, and achieve abnormal behavior detection and classification through a collaborative mechanism of feature extraction, model training, dynamic classification, and feedback optimization. Step S3: Dynamic adjustment of data acquisition strategy and transmission queue: Dynamically adjust the data acquisition granularity and transmission priority according to the risk level and queue load to balance resource consumption and monitoring needs; Step S4: Sample bias correction and closed-loop optimization: Correct sample distribution bias by inverse probability weights, and feed back false alarm rate and load status to the selection of collection points and model parameters; Step S1 further includes: calculating a comprehensive risk value based on historical risk events, business dependencies, and real-time traffic fluctuations of data outbound business nodes; and selecting the nodes with the highest comprehensive risk values ​​(P%) as key data collection points. Step S3 specifically includes: Step 3.1: Risk level based on the output of step S2 An exponential decay enhancement model is used to dynamically adjust the acquisition parameters. This model synchronously adjusts the attenuation of acquisition granularity and the enhancement of sampling frequency through an exponential function, achieving a nonlinear balance between risk level and resource consumption. The formula is as follows: (7); in, Collect granularity for the current node; The basic data collection granularity can be set according to the data type of the business. The sampling frequency for the current node; The base sampling frequency can be adjusted according to the system's processing capacity and the real-time requirements of the business. The adjustment coefficient controls the rate of particle size attenuation and frequency enhancement. Step 3.2: During the execution of Step 3.1, monitor the transmission queue occupancy rate in real time. ,when When this happens, the load balancing rule is triggered: the collection granularity of non-critical data is automatically widened by 1.5 times, and its queue priority is reduced; Step 3.3: Implement hierarchical transmission by using queue control based on priority index. The calculation formula is as follows: (8); in, For real-time risk coefficient, For data sensitivity, high-priority data with a PI ≥ 0.8 enters the real-time queue and is directly transmitted to the risk control engine using zero-copy technology, ensuring an end-to-end latency of ≤ 50ms. Ordinary data with a PI ≤ 0.5 and a PI < 0.8 is stored in a batch queue, compressed, and processed asynchronously with a latency of ≤ 5 minutes. Low-value data with a PI < 0.5 triggers a dynamic discard strategy: if this occurs for three consecutive periods... Old data is evicted using the LRU algorithm; Step S4 specifically includes: Step 4.1: By comparing each data collection node The inverse probability weight is calculated based on the actual data collection distribution and the real traffic distribution, using the following formula: (9); in, Let i be the feature vector of the i-th sample; The actual sampling probability represents the feature. The probability of a sample being selected under the current collection strategy can be obtained through statistical calculation of the collection logs; Let be the true distribution probability, representing the feature as The probability of a sample naturally occurring in the entire business traffic can be calculated using offline full-volume mirror data; As the weight adjustment factor, when When the time is insufficient, it indicates that the sample collection is inadequate. Step 4.2: Based on the ratio of false positive rate to false negative rate, adjust the node weights in step S1 in reverse. When the false positive rate is high, increase the weight of historical risk entropy. To curb over-sampling, when the false negative rate is high, reduce the weight of business criticality. To enhance node coverage, the calculation formula is as follows: ; ; The false alarm rate / false negative rate can be calculated based on the statistical values ​​of the classification results.

2. The method for collecting low-overhead, multi-source risk information for cross-border data transfer as described in claim 1, characterized in that: Step S1 specifically includes: Step 1.1: Collect past T data from the service node d Risk events are predefined events based on business data attributes. Risk entropy values ​​are calculated to measure uncertainty. Step 1.2: Consider the criticality of each business node within the overall business process; Step 1.3: Introduce real-time traffic peak percentage to identify abnormal nodes with sudden traffic surges; Step 1.4: Generate a comprehensive risk score for the node by dynamically weighting and integrating three indicators: risk entropy, business criticality, and traffic fluctuation.

3. The method for collecting low-overhead, multi-source risk information for cross-border data transfer as described in claim 2, characterized in that: The node's comprehensive risk score The calculation is as follows: (1); in, Indicates the first Each business node; , is the node risk entropy value. Let K be the probability of occurrence of a historical risk event; N represents the total number of historical risk events that have occurred. Business criticality is determined by the degree of dependence on business processes. Indicates the core node; , representing the flow fluctuation coefficient; , , For dynamic weighting coefficients, satisfying The initial value is , , .

4. The method for collecting low-overhead, multi-source risk information for cross-border data transfer as described in claim 3, characterized in that: The specific steps S2 are as follows include: Rapid metadata feature extraction: Capture raw network communication data packets at key collection points, including IP address, port number, protocol header, timestamp, and payload length, and extract four-dimensional features in real time. time Fluctuations: Calculate the coefficient of variation of communication intervals to identify bursts of traffic; Address dispersion: Statistical analysis of destination address distribution to identify unconventional cross-border routes; Protocol compliance: Mark the usage of encryption protocols; Payload characteristics: Analyze data packet length distribution to detect covert channels; Define the eigenvector: , in, , where is the communication interval fluctuation coefficient, The standard deviation of the communication time interval. The mean; , which represents the risk level; MTU is the maximum transmission unit of a network. , representing the concentration of destination addresses. Frequency of occurrence of the destination address; , is a binary variable used to determine encryption compliance.

5. A low-overhead, multi-source risk information collection method for cross-border data transfer as described in claim 4, characterized in that: Step S2 further includes: using an improved random forest model, classifying communication behaviors into three categories: normal, suspicious, and illegal. The classification process for a single tree is as follows: 1) Extract N samples from the training set, where N is the same size as the original dataset, to generate a differential training subset; 2) Then, feature selection is performed randomly. When splitting at each node, two candidate features are randomly selected from the four features to reduce the risk of overfitting. 3) Then, node splitting decisions are made, and the optimal splitting point is selected based on the Gini impurity minimization criterion. The calculation formula is as follows: (2); Where t is the current node to be split. This represents the proportion of samples of category i in node t; The splitting objective is to maximize Decrease in impurity: (3); Where N is the total number of samples in the current node to be split. , These are the left and right child nodes generated from the split; The left child node after splitting The number of samples included; The right child node after splitting The number of samples included; 4) Repeat the splitting until the termination condition is met to generate a complete decision tree; Based on the above processing flow, T decision trees are pre-trained in parallel, and the output is a set of decision trees and a set of feature importance weights; For the real-time captured traffic, starting from the root node of each tree, a splitting path is selected based on the feature value, eventually reaching the leaf node to output the class label. Then, the prediction results of all trees are counted, and the class with the highest number of votes is output as the final class. The calculation formula is as follows: (4); in, Let t be the prediction result for the t-th tree. This is an indicator function; its value is 1 if the predicted class is c, and 0 otherwise.

6. The method for collecting low-overhead, multi-source risk information for cross-border data transfer as described in claim 5, characterized in that: Step S2 specifically also includes: Based on the risk behavior classification results and system resource status, the risks are mapped to the following levels: (5); (6); in, The number of high-risk behaviors in the current cycle; The highest number of historical risks within the window period; This represents the current transmission queue occupancy rate. This represents the maximum queue capacity. and These are the weighting coefficients, and ; This represents the highest historical risk index. When the historical risk index reaches its minimum value, At that time, L=3.

7. A low-overhead, multi-source risk information collection system for cross-border data transfer, used to implement the method described in any one of claims 1-6, characterized in that: include: The intelligent filtering unit filters key collection points based on dynamic risk entropy, business criticality, and traffic fluctuations. The linkage unit, based on an improved random forest algorithm, captures communication metadata of key collection points and achieves abnormal behavior detection and classification through a collaborative mechanism of feature extraction, model training, dynamic classification and feedback optimization. The dynamic control unit dynamically adjusts the data acquisition granularity and transmission priority according to the risk level and queue load, balancing resource consumption and monitoring needs. The correction and optimization unit corrects sample distribution bias through inverse probability weights and feeds back false alarm rate and load status to the selection of collection points and model parameters.

Citation Information

Patent Citations

  • Data outbound compliance management and control method and system

    CN118484838A

  • Data outbound compliance path recommendation method and system based on classification and grading

    CN118587069A

  • Exit data security management method and device

    CN118611894A

  • Data cross-border compliance management and control method and device, computer equipment and storage medium

    CN114760149A

  • Data outbound security compliance management and control method and system based on dynamic risk assessment

    CN116187766A