Data leakage risk assessment system

Through endpoint log traceability and traffic metadata analysis, the risk of IoT privacy leakage is detected without relying on data packet content, which solves the detection difficulties of existing technology under encrypted traffic and improves the leakage detection capability of encrypted traffic.

CN120217434AInactive Publication Date: 2025-06-27BEIJING YUNXIN INTERNET TECHNOLOGY SERVICE CO LTD

Patent Information

Application Number
CN202510303756.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology cannot parse data content when facing encrypted traffic, resulting in no way to start with potential data leakage, increasing the risk of IoT privacy leakage.

Method used

Through endpoint log traceability, traffic metadata analysis and protocol behavior mapping, it does not rely on packet content to detect IoT privacy leakage risks from endpoint interaction and traffic patterns.

Benefits of technology

It has relatively improved the ability to detect leakage under encrypted traffic, reduced the misjudgment caused by single feature detection, and improved the traceability of data leakage incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217434A_ABST
    Figure CN120217434A_ABST
Patent Text Reader

Abstract

The invention discloses a data leakage risk assessment system, and particularly relates to the field of Internet of Things privacy leakage, and the system comprises a preprocessing module, an acquisition module, and an integration module. The preprocessing module executes traceability analysis and integrity verification on the endpoint log by constructing a preprocessing list of hierarchical JSON and verification Hash of the endpoint log, and the preprocessing module is used for identifying integrity abnormity of privacy leakage data of the Internet of Things; and the acquisition module executes encryption identifier extraction and protocol type mapping of an endpoint log through segmented sampling and parallel protocol analysis of network traffic metadata, and the acquisition module is used for acquiring an abnormal traffic mode of privacy leakage of the Internet of Things. Through endpoint log tracing, flow metadata analysis and protocol behavior mapping, the privacy leakage risk of the Internet of Things is detected from endpoint interaction and a flow mode without depending on the content of a data packet, and the leakage detection capability under encrypted flow is relatively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things (IoT) privacy leakage, and more specifically, to a data leakage risk assessment system. Background Art

[0002] IoT privacy leakage refers to the leakage or abuse of user privacy data caused by factors such as unauthorized access, data hijacking, and protocol vulnerabilities during the data collection, transmission, or storage process of IoT devices. With the large-scale deployment of IoT devices, scenarios such as encrypted traffic, remote communication, and edge computing have increased the complexity of IoT privacy leakage; IoT privacy protection can prevent privacy information from being exposed or illegally exploited in an unauthorized environment by means of endpoint security, encryption mechanisms, access control, traffic monitoring, etc., and execute data leakage prevention strategies.

[0003] In the detection of IoT privacy data leakage, network traffic characteristics and protocol behaviors of network traffic are usually relied on to identify abnormal data transmissions. However, with the widespread application of network-wide encryption communication methods such as HTTPS and VPN, deep packet inspection (DPI) cannot access the data content when processing encrypted traffic, losing visibility of the content, resulting in being unable to start with potential data leakage behaviors and increasing the risk of IoT privacy leakage. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a data leakage risk assessment system. Through endpoint log tracing, traffic metadata analysis, and protocol behavior mapping, without relying on the packet content, it detects the risk of IoT privacy leakage from endpoint interactions and traffic patterns to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: A data leakage risk assessment system, including a preprocessing module, a collection module, and an integration module;

[0006] The preprocessing module performs retrospective analysis and integrity verification of endpoint logs by constructing a hierarchical JSON of endpoint logs and a preprocessing list for verifying hashes. The preprocessing module is used to identify integrity anomalies of IoT privacy leakage data;

[0007] The collection module extracts encrypted identifiers of endpoint logs and maps protocol types by segment sampling of network traffic metadata and parallel protocol parsing. The collection module is used to collect abnormal traffic patterns of IoT privacy leakage;

[0008] The integration module is used to merge and process endpoint logs and real-time stream parsing results at the backend, construct an event index, and form cross-source data standardization and field mapping. The integration module is used to perform correlation analysis on the data streams of IoT privacy leakage;

[0009] It further includes a detection module, a side-channel excavation module, and an association module;

[0010] The detection module actively detects based on three perspectives, injects a pre-abnormal mark into the target event, and adjusts the risk priority according to the pre-abnormal mark;

[0011] The side-channel excavation module, based on the pre-abnormal mark of the detection module, performs a timing comparison between the side-channel excavation and the sequence of endpoint log process calls, generates a "suspicious connection queue", and outputs a temporary feature table;

[0012] The association module is used to backtrack the intersection of side-channel features and endpoint logs through a unified event index during the association analysis phase, screen out "suspicious events", classify them as "abnormal events", and output an "abnormal event chain" to the in-depth analysis queue based on the classification results.

[0013] In a preferred embodiment, the preprocessing module encapsulates system events, process states, and user operation records into a hierarchical JSON object at fixed time intervals by the front-end user device, generates a preprocessing list according to the time stamp, and the preprocessing list includes the total amount of logs, data verification hash, and index information. The preprocessing list is submitted to the back-end for integrity verification; if the verification hash of the preprocessing list does not match the previous round of verification value stored in the back-end, a re-packaging signal is triggered;

[0014] The acquisition module is responsible for extracting traffic metadata from the front-end network monitoring process, splitting the traffic metadata into two parts: synchronous packaging and real-time stream analysis. Synchronous packaging embeds traffic data into the endpoint log storage, and real-time stream analysis is used to parse the traffic protocol, mark the detected unknown protocols, and write them into the protocol metadata field of the endpoint log, and then include them in the priority analysis queue; the traffic metadata includes port number, encryption identification bit, and protocol name;

[0015] The integration module is used to perform association matching according to the association fields and the records of real-time stream analysis after receiving the endpoint logs at the back-end; if a missing field, data deviation, or no corresponding record is identified during the matching process, it is marked as a suspected forgery event or a suspected tampering event and written into the event management index; the association fields include time identifier, device identifier, connection identifier, traffic characteristics, and process activity information.

[0016] In a preferred embodiment, the three perspectives in the detection module include the port scanning perspective, the protocol behavior anomaly perspective, and the network topology dynamic change perspective; the port scanning perspective is used to detect the opening, closing, and fluctuation of ports, the protocol behavior anomaly perspective is used to detect anomalies in protocol usage patterns and packet distributions, and the network topology dynamic change perspective is used to identify connection jumps and routing path anomalies;

[0017] After receiving the pre - anomaly marker from the detection module, the side - excavation module extracts the marked "suspicious traffic records" from the event index, and calls the side - channel analysis engine to extract features of the traffic time interval, packet size distribution, connection holding duration, and flow pattern;

[0018] The side - excavation module retrieves the endpoint logs, parses the process call sequence within the corresponding time window, and constructs a time - series feature matrix by comparing its process startup behavior, file operation behavior, remote connection behavior, and permission change behavior. The side - excavation module performs temporal matching on the traffic side - channel features and endpoint process behaviors. If feature overlap is found, "suspicious connection entries" are generated. The "suspicious connection entries" are stored in the "suspicious connection queue" through the side - excavation module and encapsulated into a temporary feature table, which includes connection identifiers, associated processes, side - channel features, and abnormal risk weights.

[0019] In a preferred embodiment, it includes a learning module and an adversarial module;

[0020] The learning module extracts endpoint features, traffic features, and side - channel features from the "abnormal event chain" output by the association module, converts them into vectorized data, and inputs the vectorized data into the generative adversarial network for training. In the model training stage of the generative adversarial network, the generator of the generative adversarial network simulates potential data leakage patterns, and the discriminator of the generative adversarial network classifies the input data. If the discriminator cannot classify the input event into the existing categories, a new abnormal category label is generated and fed back to the association module;

[0021] The adversarial module generates camouflaged traffic samples based on the generative adversarial network trained by the learning module to test the recognition ability of the current detection rules; the adversarial module uses the generator to construct abnormal data similar to the target traffic and evaluates the robustness of the existing detection logic through the discriminator;

[0022] If the recognition rate of the discriminator for the camouflaged samples is lower than the system - preset recognition rate threshold, the weight adjustment mechanism is triggered to re - optimize the feature extraction parameters and risk threshold settings, and update the detection strategy; the optimized detection rules are synchronized to the detection module, and the logs of the optimization process are stored in the evolution log of the generative adversarial network model.

[0023] In a preferred embodiment, it further includes a detection module and a response module;

[0024] The detection module is used to receive the updated detection rules from the adversarial module and the data of the "abnormal event chain" input by the association module; based on the input endpoint features, traffic features, and side - channel features, the detection module performs feature matching and risk score calculation through the detection engine, and encapsulates the score data into a risk assessment record and writes it into the risk assessment log; when the risk score exceeds the system - preset risk score threshold, the corresponding record is pushed to the warning queue;

[0025] The response module includes extracting warning event records from the warning queue, calling the front-end device to collect supplementary data, and taking protective measures such as isolation, process termination or network speed limit for endpoints suspected of data leakage based on the supplementary data; the operation data includes memory images, key directory hashes, and process status snapshots;

[0026] The processing results and response operations are fed back to the backend through the response module, providing a basis for subsequent manual audits and policy adjustments.

[0027] In a preferred embodiment, it also includes a feedback module;

[0028] The feedback module is used to receive the final decision of the abnormal events in the risk assessment report, pass the decision information to the integration module and the confrontation module through the callback interface, and update the event management index and strategy evolution log;

[0029] The feedback data is used as the input of the learning module to adjust the abnormal feature library and risk determination rules.

[0030] In a preferred embodiment, it also includes a response error processing module;

[0031] After receiving the processing result fed back by the response module, the response error processing module extracts constant parameters from the returned data. The constant parameters include the response status code, response delay, data verification hash consistency, and the matching status between the preceding exception mark and the "exception event chain";

[0032] The response error processing module is executed according to the preset judgment process: comparing whether the response status code is in the target range, checking whether the response delay exceeds the delay threshold set by the system, and verifying whether the returned data is consistent with the verification hash and index information in the preprocessing list and temporary feature table; if any judgment condition is not met, it is judged as a response error.

[0033] Technical effects and advantages of the present invention:

[0034] 1. Traditional deep packet inspection (DPI) cannot parse data content when facing encrypted communications such as HTTPS and VPN, resulting in a blind spot in traffic analysis. The present invention detects the risk of privacy leakage in the Internet of Things from endpoint interactions and traffic patterns without relying on data packet content, through endpoint log tracing, traffic metadata analysis, and protocol behavior mapping, thereby relatively improving the ability to detect leakage under encrypted traffic;

[0035] 2. The present invention adopts three-perspective detection methods: port scanning, protocol behavior anomaly detection, and network topology dynamic change analysis, and combines side channel analysis technology to evaluate data leakage risks from multiple levels of endpoint behavior, traffic characteristics, and side channel characteristics, reducing misjudgments caused by single feature detection;

[0036] 3. The correlation analysis module indexes and correlates abnormal events from different times and different sources, and generates an "abnormal event chain" for in-depth analysis. Compared with traditional single detection, the present invention can form cross-time and space behavior pattern analysis, improving the traceability ability of data leakage events;

[0037] 4. Through the learning module and the adversarial module, a generative adversarial network is used to simulate potential data leakage patterns, and the training system continuously learns new abnormal behavior patterns to cope with the adaptability of Internet of Things privacy leakage behavior. Description of the Drawings

[0038] Figure 1 It is a system module composition diagram of the present invention.

[0039] Figure 2 It is a general execution flow chart of the system of the present invention.

[0040] Figure 3 It is a preprocessing and data acquisition flow chart of the system in the present invention.

[0041] Figure 4 It is an abnormal detection and risk response flow chart of the system in the present invention.

[0042] Figure 5 It is the timing of the detection module of the present invention Figure 1 .

[0043] Figure 6 It is the timing of the detection module of the present invention Figure 2 . Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Referring to the attached Figures 1-6 description, a data leakage risk assessment system according to an embodiment of the present invention includes a preprocessing module, a collection module, and an integration module;

[0046] The preprocessing module performs retrospective analysis and integrity verification of the endpoint logs by constructing a hierarchical JSON of the endpoint logs and a preprocessing list of verification hashes. The preprocessing module is used to identify the integrity anomalies of the Internet of Things privacy leakage data;

[0047] The collection module performs the extraction of encrypted identifiers and protocol type mapping of endpoint logs through segmented sampling and parallel protocol parsing of network traffic metadata, and is used to collect abnormal traffic patterns of Internet of Things privacy leakage;

[0048] The integration module is used to merge and process endpoint logs and real-time stream parsing results at the backend, build an event index, and form cross-source data standardization and field mapping. The integration module is used to perform correlation analysis on the data stream of Internet of Things privacy leakage;

[0049] It also includes a detection module, a side-channel excavation module, and a correlation module;

[0050] The detection module actively detects based on three perspectives, injects pre-abnormal marks into suspicious target events, and adjusts the risk priority according to the pre-abnormal marks;

[0051] Based on the pre-abnormal marks of the detection module, the side-channel excavation module performs temporal comparison with the endpoint log process call sequence through side-channel excavation, generates a "suspicious connection queue", and outputs a temporary feature table;

[0052] The correlation module is used to trace back the intersection of side-channel features and endpoint logs through a unified event index during the correlation analysis phase, screen out "suspicious events" and classify them as "abnormal events", and output an "abnormal event chain" to the in-depth analysis queue based on the classification results.

[0053] The preprocessing module encapsulates system events, process status, and user operation records into a hierarchical JSON object at fixed time intervals by the front-end user device, and generates a preprocessing list according to the timestamp. The preprocessing list includes the total log volume, data verification hash, and index information. The preprocessing list is submitted to the backend for integrity verification; if the verification hash of the preprocessing list does not match the previous round of verification value stored in the backend, a re-packaging signal is triggered;

[0054] In addition, in actual applications, if a re-packaging signal is triggered, the front-end device first clears the current preprocessing cache, and then extracts system events, process status, and user operation records from the local log storage again, and reorganizes them into a hierarchical data object in chronological order; subsequently, the total log volume, data verification hash, and index information are recalculated, a new preprocessing list is generated and replaces the old version, and finally it is submitted to the backend for integrity verification again until the verification passes;

[0055] The collection module is responsible for extracting traffic metadata from the front-end network monitoring process, splitting the traffic metadata into two parts: synchronous packaging and real-time stream analysis. Synchronous packaging embeds traffic data into the endpoint log storage, and real-time stream analysis is used to parse traffic protocols, mark detected unknown protocols, and write them into the protocol metadata field of the endpoint log, and then include them in the priority analysis queue; traffic metadata includes port number, encryption identification bit, and protocol name;

[0056] The integration module is used to perform correlation matching based on the correlation fields and the records of real-time stream analysis after receiving endpoint logs at the backend; if field missing, data deviation or no corresponding record is identified during the matching process, it is marked as a suspected forgery event or a suspected tampering event and written into the event management index; the correlation fields include time identifier, device identifier, connection identifier, traffic characteristics and process activity information.

[0057] The three perspectives in the detection module include the port scanning perspective, the protocol behavior anomaly perspective and the network topology dynamic change perspective; the port scanning perspective is used to detect the opening, closing and fluctuation of ports, the protocol behavior anomaly perspective is used to detect the anomalies in protocol usage patterns and packet distributions, and the network topology dynamic change perspective is used to identify sudden connection jumps and routing path anomalies;

[0058] The detection module defines the port scanning anomaly degree D based on the port scanning perspective P ; defines the protocol behavior anomaly degree D based on the protocol behavior anomaly perspective Q ; defines the network topology anomaly degree D based on the network topology dynamic change perspective T ;

[0059]

[0060] Where P o is the port status indicator actually detected by the front-end device; P r represents the expected port status indicator obtained based on historical data or security baselines; |P o -P r | represents the absolute deviation between the actual and the expected; κ P is the smoothing constant for port scanning, which is used to prevent the denominator from being zero in practical applications and its value is a positive real number; tanh -1 (x) is the inverse hyperbolic tangent function, which is used to non-linearly amplify the normalized deviation and enhance the sensitivity to small differences;

[0061] Where Q o is the protocol usage indicator actually detected by the front-end device; Q r is the expected protocol usage indicator; |Q o -Q r | is the absolute difference between the actual and the expected protocol behaviors; κ Q is the smoothing constant for protocol behavior, κ Q is a positive real number, which is used to normalize the difference; arcsinh(x) is the inverse hyperbolic sine function, which is used to non-linearly map the normalized difference to highlight the degree of anomaly;

[0062] Where T o is the network topology indicator actually detected; T rDenote the reference topology metrics set based on the normal network structure; |T o -T r | denotes the absolute deviation between the observation and the expectation; is the network topology normalization constant, is a positive real number; ln(x) is the natural logarithm function, which is used in practical applications to perform non-linear scaling on the squared normalization deviation, amplifying the influence of larger outliers;

[0063] Based on D P 、D Q 、D T Calculate the comprehensive anomaly risk score R det ;

[0064]

[0065] where α P ,α Q 、α T are the weight coefficients from the perspectives of port, protocol, and network topology respectively; κ PQ is the joint normalization constant of port and protocol, a positive real number; is the normalization constant of network topology in the comprehensive score; exp(x) is the natural exponential function, with the sum of two terms as the exponential input, making the overall risk score have non-linear response characteristics; ln(1 + ·) represents logarithmic compression of the exponential result, and finally outputs R det as the comprehensive anomaly risk value;

[0066] After receiving the pre-anomaly mark from the detection module, the side-digging module extracts the marked "suspicious traffic records" from the event index, and calls the side-channel analysis engine to extract features of the traffic time interval, packet size distribution, connection holding duration, and flow direction pattern;

[0067] The side-digging module retrieves the endpoint logs, parses the process call sequence within the corresponding time window, and constructs a time series feature matrix by comparing its process startup behavior, file operation behavior, remote connection behavior, and permission change behavior. The side-digging module performs temporal matching on the traffic side-channel features and the endpoint process behavior. If feature overlap is found, "suspicious connection entries" are generated. The "suspicious connection entries" are stored in the "suspicious connection queue" through the side-digging module and encapsulated into a temporary feature table. The temporary feature table includes connection identifiers, associated processes, side-channel features, and anomaly risk weights;

[0068] Regarding the above, it should be further noted that the side-digging module includes the calculation of temporal matching differences and cumulative temporal mismatch degrees;

[0069] Within the time interval [t0, t1], define the instantaneous difference between the side-channel feature S(t) and the endpoint process call sequence feature L(t) as Δ(t); additionally, integrate the instantaneous difference to obtain the overall timing mismatch degree C side , based on the overall timing mismatch degree C side define the suspicious connection risk weight W susp ;

[0070]

[0071]

[0072] where S(t) represents the side-channel feature recorded at time t; L(t) represents the process call sequence feature in the endpoint log recorded at time t; |S(t) - L(t)| represents the instantaneous absolute difference between the two features; κ side is the normalization constant of the side-channel module, κ side is used to control the difference amplitude; arctan(x) is the arctangent function, which normalizes the difference to the interval (0, π / 2) to form a smooth response; t0 and t1 are the start and end times of the timing analysis respectively; the integration result C side represents the cumulative mismatch degree between the side-channel and the endpoint process behavior within the entire time window; W susp has a numerical range between 0 and 1;

[0073] In addition, it should be noted that the lower the value of W susp , the more significant the timing matching anomaly detected by the side-channel module, indicating an increased risk of Internet of Things privacy leakage.

[0074] includes a learning module and an adversarial module;

[0075] The learning module extracts endpoint features, traffic features, and side-channel features from the "abnormal event chain" output by the association module, converts them into vectorized data, and inputs the vectorized data into a generative adversarial network for training; during the model training stage of the generative adversarial network, the generator of the generative adversarial network simulates potential data leakage patterns, and the discriminator of the generative adversarial network classifies the input data to optimize the ability to determine abnormal patterns; in addition, during the training process of the generative adversarial network, the learning module continuously compares the matching situation between new samples and existing clusters. If the discriminator cannot classify the input event into an existing category, a new abnormal category label is generated and fed back to the association module to optimize the event chain construction strategy, enabling the model to adaptively identify unknown data leakage patterns;

[0076] The learning module includes a non-linear mapping of feature vectors; assume the feature vector of the abnormal event chain output by the association module is Based on this, a non - linear mapping function is adopted to express φ(x); φ(x)=ln(1 + exp(Ax + b));

[0077] where x represents the original input feature vector of the Internet of Things privacy leakage abnormal event, and x includes endpoint, traffic and side - channel features; is the transpose operation of the vector; A is a transformation matrix of size m×n for feature reconstruction; represents that the vector b belongs to the bias term of m, represents the m - dimensional real number space; exp(·) is the exponential function, which performs exponential mapping on the input features to enhance the expression ability of non - linear features; ln(·) is the natural logarithm function, which is used to compress the dynamic range of the input value to ensure that outliers are appropriately amplified;

[0078] The loss functions of the discriminator and generator of the generative adversarial network are expressed as:

[0079] For real samples and generated samples in a batch, the discriminator loss is defined as L D ; the generator loss is defined as L G ;

[0080]

[0081] where N is the batch size; D(x) is the output probability of the discriminator for the input real sample x, and D(x) takes values between 0 and 1; G(x) is the pseudo - sample generated by the generator according to the noise vector ; Li2(·) is the dilogarithm function of the second order, and Li2(·) is used to perform non - linear weighting on the loss;

[0082] Based on the learning module, a new abnormal category determination condition is established; for the feature vector x of a single new sample, the Manhattan distance between it and the abnormal center of the j - th category is defined as d j (x); μ j,k is the central value of the j - th abnormal event chain in the k - th dimension;

[0083] If is satisfied, a new abnormal category label is generated, where τ L is the preset threshold, and K is the number of existing abnormal categories;

[0084] Based on the generative adversarial network trained by the learning module, the adversarial module actively generates camouflaged traffic samples to test the recognition ability of the current detection rules; the adversarial module uses the generator to construct abnormal data similar to the target normal traffic and evaluates the robustness of the existing detection logic through the discriminator;

[0085] If the recognition rate of the discriminator for disguised samples is lower than the preset recognition rate threshold of the system, the weight adjustment mechanism is triggered to re-optimize the feature extraction parameters and the risk threshold setting, and update the detection strategy; the optimized detection rules are synchronized to the detection module, and the logs of the optimization process are stored in the evolution log of the generative adversarial network model, which is used to support long-term policy iteration and improve the adaptability of the detection system to complex data leakage behaviors;

[0086] Calculate the similarity metric of the disguised samples based on the adversarial module; define the target normal sample vector as The similarity index S between the disguised abnormal samples generated by the generator adv ; Indicates that the input feature vector belongs to the n-dimensional real number space; erfc(·) represents the complementary error function, The role of erfc(·) is to normalize the input distance, so that when the disguised sample is close to the target sample, S adv is closer to 2, otherwise it tends to 0. The value range of S adv is usually between 0 and 2. The larger the value of S adv , the more similar the two are, and the smaller the value, the more significant the difference;

[0087] In addition, ∥y - G(z)∥1 represents the Manhattan distance between two vectors; where y k and G(z) k are the components of y and G(z) in the k-th dimension respectively; λ adv is the similarity normalization parameter, and λ adv is a positive real number; y k represents the value of the target normal sample y in the k-th dimension; G(z) k represents the value of the disguised abnormal sample generated by the generator G in the k-th dimension; n is the dimension of the feature vector used in the Internet of Things privacy leakage detection system;

[0088] If S adv < τ adv , then calculate the adjustment value ΔW of the detection parameter, and τ adv is the preset similarity threshold of the system; ΔW = λ W (τ adv - S adv ), and λ W is the adjustment rate constant; ΔW is the adjustment amount of the detection rule parameter, and the updated detection parameter based on this is W new ; W new = W old + ΔW, W oldRepresents the old value of the detection rule parameter.

[0089] It also includes a detection module and a response module;

[0090] The detection module is used to receive the updated detection rules from the adversarial module and the data of the "abnormal event chain" input by the associated module; the detection module, based on the input endpoint features, traffic features, and side-channel features, performs feature matching and risk score calculation through the detection engine, and encapsulates the score data into a risk assessment record and writes it into the risk assessment log; when the risk score exceeds the system's preset risk score threshold, the corresponding record is pushed to the warning queue to provide real-time warnings for subsequent response processes;

[0091] In the detection module, calculate the abnormality degree of each feature; calculate the abnormality degree of the endpoint, traffic, and side-channel parts in the input event chain feature vector respectively; define F E )x) as the abnormality degree of the endpoint feature, define F F (x) as the abnormality degree of the traffic feature, define F S (x) as the abnormality degree of the side-channel feature;

[0092] Through F E (x) to measure the deviation between the current endpoint state and the expected baseline, through F F )x) to evaluate whether the network traffic pattern conforms to normal communication behavior, through F S (x) to detect whether side-channel information such as packet time intervals and traffic patterns is abnormal;

[0093]

[0094] Where x E , x F , x S are the detected endpoint, traffic, and side-channel feature values respectively; x E,ref , x F,ref , x S,ref represents the reference feature value set based on historical data or security baseline; γ E , γ F , γ S represents the non-linear amplification exponent of each feature, γ E , γ F , γ S is a positive real number; κ E , κ F , κ S represents the normalization constant of each feature; arcsinh(x) represents the inverse hyperbolic sine function, and arcsinh(x) is used for non-linear mapping of the normalized deviation;

[0095] Take F E (x), FF (x), F S (x) The abnormality degrees of each part are multiplicatively combined, and a comprehensive risk score R is obtained based on the natural logarithm function score ;

[0096]

[0097] where β E , β F , β S are the weight indices of the endpoint, traffic, and side-channel characteristics respectively, which are used to reflect the importance of each characteristic in the risk of Internet of Things privacy leakage; ln(·) is the natural logarithm function, which is used to compress the dynamic range of the risk score value; R score is the finally obtained risk score; when R score ≥τ score , mark the event as high-risk and push it to the warning queue, τ score is the preset risk threshold;

[0098] The response module includes extracting warning event records from the warning queue, calling the front-end device to collect supplementary data, and taking isolation, process termination, or network speed limit protection measures for the endpoints suspected of data leakage according to the supplementary data; the operation data includes memory images, critical directory hashes, and process status snapshots;

[0099] The response module feeds back the processing results and response operations to the backend, providing a basis for subsequent manual auditing and policy adjustment.

[0100] It also includes a feedback module;

[0101] The feedback module is used to receive the final ruling of the abnormal events in the risk assessment report by humans, and transmit the ruling information to the integration module and the countermeasure module through the callback interface to update the event management index and the policy evolution log;

[0102] The feedback data is used as the input of the learning module to adjust the abnormal feature library and the risk determination rules, so as to continuously improve the detection accuracy of the system for abnormal behaviors of Internet of Things privacy leakage and realize a complete closed-loop adaptive optimization process.

[0103] It also includes a response error handling module;

[0104] After receiving the processing results feedback by the response module, the response error handling module extracts constant parameters from the returned data. The constant parameters include the response status code, response delay, data verification hash consistency, and the matching status between the pre-abnormal mark and the "abnormal event chain";

[0105] The response error handling module executes according to a preset judgment process: compare whether the response status code is within the target range, check whether the response delay exceeds the delay threshold set by the system, and verify whether the returned data is consistent with the verification hashes and index information in the preprocessing list and the temporary feature table; if any judgment condition is not met, it is determined as a response error.

[0106] When a response error is determined, a detailed response error report is generated, recording the error code, abnormal delay, data comparison result, and the deviation of the pre - exception mark; this error report is fed back to the integration module and the countermeasure module through a callback interface to trigger subsequent automatic retries or manual interventions, forming a closed - loop feedback to continuously optimize the overall response accuracy of the Internet of Things privacy leakage risk assessment.

[0107] In addition, the verification of the returned data includes the memory image, key directory hash, and process status snapshot collected from the front - end.

[0108] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data leakage risk assessment system, comprising a preprocessing module, a collection module, and an integration module; The preprocessing module performs traceability analysis and integrity verification of endpoint logs by constructing a preprocessing list of hierarchical JSON and verification hashes of endpoint logs. The preprocessing module is used to identify integrity anomalies of IoT privacy leakage data; The collection module performs encrypted identification extraction and protocol type mapping of endpoint logs through segmented sampling and parallel protocol parsing of network traffic metadata. The collection module is used to collect abnormal traffic patterns that may leak privacy in the Internet of Things. The integration module is used to merge and process endpoint logs and real-time stream analysis results at the back end, build event indexes, and form cross-source data standardization and field mapping. The integration module is used to correlate and analyze data streams for privacy leaks in the Internet of Things. It is characterized by: it also includes a detection module, a side excavation module, and a correlation module; The detection module actively detects based on three perspectives, injects pre-anomaly markers into target events, and adjusts risk priorities based on pre-anomaly markers; The side-channel mining module compares the timing of the endpoint log process call sequence based on the anomaly mark of the detection module, generates a "suspicious connection queue" and outputs a temporary feature table. The correlation module is used to trace back the intersection of side channel features and endpoint logs through a unified event index during the correlation analysis phase, screen out "suspicious events" and characterize them as "abnormal events", and output the "abnormal event chain" to the deep analysis queue based on the qualitative results.

2. A data leakage risk assessment system according to claim 1, characterized in that: The preprocessing module encapsulates system events, process status and user operation records into hierarchical JSON objects at fixed time intervals by the front-end user device, and generates a preprocessing list by timestamp. The preprocessing list includes the total log volume, data verification hash, and index information. The preprocessing list is submitted to the back-end for integrity verification; If the checksum hash of the preprocessed manifest does not match the previous round of checksums stored in the backend, a repackaging signal is triggered; The acquisition module is responsible for extracting traffic metadata from the front-end network monitoring process and splitting the traffic metadata into two parts: synchronous packaging and real-time stream analysis. Synchronous packaging embeds traffic data into endpoint log storage. Real-time stream analysis is used to parse traffic protocols, mark detected unknown protocols, and write them into the protocol metadata field of the endpoint log, which is then included in the priority analysis queue. Traffic metadata includes port number, encryption identification bit, and protocol name. The integration module is used to perform correlation matching based on the associated fields and the records of real-time stream analysis after receiving the endpoint logs at the back end; if a field is missing or the data is deviated or there is no corresponding record during the matching process, it is marked as a suspected forgery event or a suspected tampering event and written into the event management index; The associated fields include time identification, device identification, connection identification, traffic characteristics, and process activity information.

3. A data leakage risk assessment system according to claim 2, characterized in that: The three perspectives in the detection module include port scanning perspective, protocol behavior anomaly perspective and network topology dynamic change perspective; the port scanning perspective is used to detect the opening, closing and fluctuation of ports, the protocol behavior anomaly perspective is used to detect anomalies in protocol usage patterns and data packet distribution, and the network topology dynamic change perspective is used to identify connection jumps and routing path anomalies; After receiving the pre-anomaly mark from the detection module, the side-channel module extracts the marked "suspicious traffic record" from the event index and calls the side-channel analysis engine to extract features of the traffic time interval, packet size distribution, connection duration and flow pattern; The endpoint logs are retrieved through the side-mining module to analyze the process call sequence within the corresponding time window, and the process startup behavior, file operation behavior, remote connection behavior and permission change behavior are compared to construct a time series feature matrix. The side-mining module performs time-series matching on traffic side-channel features and endpoint process behaviors. If feature overlap is found, a "suspicious connection entry" is generated. The "suspicious connection entry" is stored in the "suspicious connection queue" through the side-mining module and encapsulated into a temporary feature table, which includes connection identifiers, associated processes, side-channel features, and abnormal risk weights.

4. A data leakage risk assessment system according to claim 3, characterized in that: Includes learning module and confrontation module; The learning module extracts endpoint features, traffic features, and side channel features from the "abnormal event chain" output by the association module, converts them into vectorized data, and inputs the vectorized data into the generative adversarial network for training. In the model training phase of the generative adversarial network, the generator of the generative adversarial network simulates potential data leakage patterns, and the discriminator of the generative adversarial network classifies the input data. If the discriminator cannot classify the input event into an existing category, it generates a new abnormal category label and feeds it back to the association module; The adversarial module generates disguised traffic samples based on the generative adversarial network trained by the learning module to test the recognition ability of the current detection rules. The adversarial module uses the generator to construct abnormal data similar to the target traffic and evaluates the robustness of the existing detection logic through the discriminator. If the discriminator's recognition rate for disguised samples is lower than the recognition rate threshold preset by the system, the weight adjustment mechanism is triggered, the feature extraction parameters and risk threshold settings are re-optimized, and the detection strategy is updated; the optimized detection rules are synchronized to the detection module, and the log of the optimization process is stored in the evolution log of the generative adversarial network model.

5. A data leakage risk assessment system according to claim 4, characterized in that: It also includes a detection module and a response module; The detection module is used to receive the detection rules updated by the confrontation module and the data of the "abnormal event chain" input by the association module; The detection module performs feature matching and risk score calculation based on the input endpoint features, traffic features and side channel features through the detection engine, and encapsulates the score data into risk assessment records and writes them into the risk assessment log; when the risk score exceeds the risk score threshold preset by the system, the corresponding record is pushed to the warning queue; The response module includes extracting warning event records from the warning queue, calling the front-end device to collect supplementary data, and taking protective measures such as isolation, process termination or network speed limit for endpoints suspected of data leakage based on the supplementary data; the operation data includes memory images, key directory hashes, and process status snapshots; The processing results and response operations are fed back to the backend through the response module, providing a basis for subsequent manual audits and policy adjustments.

6. A data leakage risk assessment system according to claim 5, characterized in that: It also includes a feedback module; The feedback module is used to receive the final decision of the abnormal events in the risk assessment report, pass the decision information to the integration module and the confrontation module through the callback interface, and update the event management index and strategy evolution log; The feedback data is used as the input of the learning module to adjust the abnormal feature library and risk determination rules.

7. A data leakage risk assessment system according to claim 6, characterized in that: Also includes a response error handling module; After receiving the processing result fed back by the response module, the response error processing module extracts constant parameters from the returned data. The constant parameters include the response status code, response delay, data verification hash consistency, and the matching status between the preceding exception mark and the "exception event chain"; The response error processing module is executed according to the preset judgment process: comparing whether the response status code is in the target range, checking whether the response delay exceeds the delay threshold set by the system, and verifying whether the returned data is consistent with the verification hash and index information in the preprocessing list and temporary feature table; if any judgment condition is not met, it is judged as a response error.

Citation Information

Patent Citations

  • Internet of Things security protection system and protection method thereof

    CN115118525A

  • Generative adversarial network-based power grid flow data privacy protection method and system

    CN119210801A

  • Method and device for detecting illegal external connection of terminal and storage medium

    CN119316226A

  • Medical data privacy protection method and system based on artificial intelligence

    CN119577841A

  • Data Surveillance for Privileged Assets based on Threat Streams

    US20200204574A1

Cited By

  • Detection method for disk array data leakage, product, equipment and medium

    CN120723561A

  • A method, product, device, and medium for detecting data leakage from a disk array

    CN120723561B

  • Data leakage risk identification method for light-wind storage and charging cooperative control system

    CN120874093A

  • Intelligent operation leakage analysis system and method applied to cryptographic equipment

    CN121770781A