Data flow supervision method and system

By conducting multi-dimensional analysis of logs and security data from different systems, constructing risk chains and performing matching analysis, the applicability and efficiency issues of data security supervision in existing technologies have been resolved, enabling real-time security supervision and risk identification across systems and networks.

CN120930133APending Publication Date: 2025-11-11CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510848063.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies cannot be dynamically adjusted according to the company's business needs and security strategies, resulting in insufficient applicability and effectiveness of data security supervision, especially in the cross-system and cross-network flow of structured and unstructured data, making it difficult to achieve comprehensive and real-time security supervision.

Method used

By collecting log data and security data from different systems, we analyze key risk characteristics, conduct multi-dimensional risk event correlation analysis, construct risk chains, perform risk assessment and pattern extraction, identify collaborative risk behaviors, and use a fusion model to generate related explanatory text, conduct feature and semantic matching analysis, and finally visualize and display risk information.

Benefits of technology

It enables comprehensive and real-time security monitoring of structured and unstructured data flowing across systems and networks, improves the accuracy of risk identification and response speed, reduces false alarm rate, and adapts to the security policy requirements of different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930133A_ABST
    Figure CN120930133A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a data flow supervision method and system.The method comprises the steps that log data and safety data of different systems are collected, various risk key features related to risk events are obtained through analysis, multi-dimensional risk event correlation analysis is conducted, risk links corresponding to the risk events are constructed, and the risk events are monitored; performing risk assessment calculation on each risk link, determining a risk level of each risk link, and identifying a collaborative risk behavior; and on the basis of the obtained feature matching calculation result, the obtained semantic matching analysis result and the case quality score of each candidate case obtained by matching, performing comprehensive evaluation on each candidate case matched with the to-be-identified event, and determining a final candidate case. According to the invention, the risk event identification precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data flow monitoring method and system. Background Technology

[0002] Data has become a core asset driving enterprise and social development. Its flow across systems, networks, and platforms is becoming increasingly frequent, encompassing not only the precise interaction of structured data (such as database tables) but also the widespread application of unstructured data (such as files and data transmitted via API interfaces).

[0003] However, with the continuous development of the company's business, the volume of business data has increased significantly, and the data flow process has become increasingly complex. Simultaneously, to meet the ever-growing business demands, node data centers are typically built in different physical environments for distributed storage and processing of business data. This high-frequency and complex data flow exacerbates security risks such as data leakage, tampering, and attacks. Traditional data security monitoring solutions, relying on static rules and post-event audits, struggle to address the real-time threat detection needs in dynamic data flow scenarios, revealing significant shortcomings such as low efficiency and narrow coverage. This poses a serious challenge to the protection of personal privacy, the safeguarding of corporate interests, and social stability.

[0004] Current data security supervision mainly relies on traditional implementation solutions, typically using single security protection tools such as log auditing and interface auditing. These include rule-based static auditing, which matches data access behavior patterns using preset rules; centralized log analysis, which simply aggregates scattered log data to achieve basic storage and querying; and manual tracing and response, where the security team manually analyzes logs and abnormal events and formulates handling strategies.

[0005] Currently, traditional data security oversight methods have many drawbacks in addressing the flow of business data across an enterprise. On one hand, there is a lack of unified and efficient methods for overseeing both structured data (relational database tables, etc.) and unstructured data (files, images, etc.), making it difficult to comprehensively cover various data flow scenarios. For example, database security oversight typically focuses on access control for structured data, while security oversight of files and interface data is relatively weak, leading to vulnerabilities in data security protection.

[0006] On the other hand, with the rapid growth of data and business data flow within the company, and the interaction of data from multiple data centers, traditional data security monitoring is becoming increasingly inefficient. Traditional monitoring methods often rely on manual analysis and simple rule matching, which cannot detect and handle potential security threats in a timely manner and are unable to meet the ever-increasing data security needs.

[0007] Furthermore, existing technologies lack flexible configuration and management mechanisms, making it impossible to dynamically adjust them according to different business needs (such as data sharing and data processing) and security strategies of the company, thus limiting the applicability and effectiveness of data security supervision. Summary of the Invention

[0008] This invention aims to provide a data flow monitoring method and system to address the technical problems in existing technologies, such as the inability to dynamically adjust data based on a company's business needs and security strategies regarding data sharing and processing, which limits the applicability and effectiveness of data security monitoring, and the low efficiency of data flow monitoring in the event of a surge in business data. The technical problems to be solved by this invention are achieved through the following technical solutions.

[0009] The first aspect of this invention proposes a data flow monitoring method, comprising: collecting log data and security data from different systems based on a customized collection strategy, and performing data parsing to obtain various key risk features related to risk events, including network traffic features, user behavior features, system status features, data access features, and data transmission features; performing multi-dimensional risk event correlation analysis on the collected log data, security data, and the various key risk features to construct risk links corresponding to each risk event; performing risk assessment calculations on each risk link to determine the risk level of each risk link; and extracting risk patterns based on the constructed risk links to conduct cross-system behavior analysis. The process involves several steps: First, identifying collaborative risk behaviors; second, using a fusion model to infer and generate explanatory texts related to the risk event to construct an evidence chain and assessing its confidence level. Third, performing feature matching calculations between the event to be identified and various key risk features in candidate cases to obtain feature matching results, and performing semantic matching analysis on the semantic descriptions of the event to be identified and candidate cases to obtain semantic matching analysis results. Finally, based on the obtained feature matching calculation results, semantic matching analysis results, and case quality scores of each matched candidate case, a comprehensive evaluation is conducted on each candidate case matching the event to be identified to determine the final candidate cases, and the relevant information of the event to be identified is visualized.

[0010] The second aspect of this invention proposes a data flow monitoring system that executes the data flow monitoring method described in the first aspect of this invention. The data flow monitoring system includes: a data acquisition module, which, based on a customized acquisition strategy, collects log data and security data from different systems and performs data parsing to obtain various key risk features related to risk events. These key risk features include network traffic features, user behavior features, system status features, data access features, and data transmission features; a data analysis model, used to perform multi-dimensional risk event correlation analysis on the collected log data, security data, and the aforementioned key risk features, construct risk links corresponding to each risk event, and perform risk assessment calculations on each risk link to determine the risk level of each risk link; and a data processing module, used to process the constructed risk links... The system performs risk pattern extraction and cross-system behavior correlation to identify collaborative risk behaviors. An inference processing module uses a fusion model to generate associated explanatory text related to the risk event, constructing an evidence chain and evaluating the confidence level of the constructed evidence chain. A calculation processing module performs feature matching calculations between the event to be identified and various key risk features in candidate cases, obtaining feature matching calculation results, and performs semantic matching analysis on the semantic descriptions of the event to be identified and candidate cases, obtaining semantic matching analysis results. A determination module, based on the obtained feature matching calculation results, semantic matching analysis results, and case quality scores of each matched candidate case, comprehensively evaluates each candidate case matching the event to be identified to determine the final candidate case and visualizes the relevant information of the event to be identified.

[0011] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the data flow monitoring method described in the first aspect of the present invention.

[0012] A fourth aspect of the present invention provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the data flow monitoring method described in the first aspect of the present invention.

[0013] The embodiments of the present invention have the following advantages: The data flow monitoring method provided in this invention, based on a customized collection strategy, collects log data and security data from different systems and performs data parsing to obtain various key risk features related to risk events. This method effectively collects relevant data on risk events and obtains accurate feedback on various key risk features. Multi-dimensional risk event correlation analysis is performed on the collected log data, security data, and the aforementioned key risk features to construct risk chains corresponding to each risk event. Risk assessment calculations are then performed on each risk chain to accurately determine its risk level.

[0014] Based on the constructed risk chain, risk patterns are extracted, and cross-system behavioral correlations are performed to identify collaborative risk behaviors. A fusion model is used to infer and generate associated explanatory texts related to risk events to construct an evidence chain. Feature matching calculations are performed between the event to be identified and various key risk features in candidate cases to obtain feature matching results. Semantic matching analysis is also performed between the semantic descriptions of the event to be identified and candidate cases to obtain semantic matching analysis results. A comprehensive evaluation of each candidate case matching the event to be identified is conducted to determine the final candidate cases, and the relevant information of the event to be identified is visualized. Through risk identification and analysis, by correlating security logs with multi-dimensional risk events in data, comprehensive and real-time security supervision of structured and unstructured data flowing across systems and networks in all scenarios is achieved, and risk events are accurately identified. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the steps of an embodiment of the data flow monitoring method of the present invention; Figure 2 This is a schematic diagram of the framework of an application example of the data flow supervision method of the present invention; Figure 3 This is a schematic diagram of the structure of an embodiment of the data flow monitoring system of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a computer-readable medium embodiment according to the present invention. Detailed Implementation

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0018] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0019] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0020] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0021] For ease of description, spatial relative terms such as "above," "on top of," "on the upper surface of," "above," etc., are used herein to describe the spatial positional relationship of a device or feature as shown in the figures to other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation beyond the orientation of the device as described in the figures. For example, if the device in the figures were inverted, a device described as "above" or "on top of" other devices or structures would subsequently be positioned as "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both "above" and "below." The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the spatial relative descriptions used herein will be interpreted accordingly.

[0022] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] In view of the above problems, this invention proposes a data flow supervision method and system. In terms of risk identification and analysis, this method efficiently correlates and analyzes security logs with data inputs through a modular architecture, outputting risk correlation analysis results to achieve comprehensive and real-time security supervision of structured and unstructured data flowing across systems and networks in all scenarios. Simultaneously, in terms of security incident response and handling, modular design shortens threat discovery time to eliminate regulatory blind spots, and relies on contextual correlation and machine learning models to accurately identify real threats, reducing false alarm rates and improving the accuracy of risk incident identification. Furthermore, an automated blocking mechanism quickly responds to abnormal data flow behavior, and flexible configuration and management functions are provided to adapt to the security policy requirements of different business scenarios.

[0024] It should be noted that this invention adopts a standardized and scalable process architecture. Its core processing flow covers security log collection and parsing, security risk correlation identification, and intelligent security risk suggestions. Through modular collaboration, it achieves accurate monitoring and proactive defense throughout the entire lifecycle of data flow.

[0025] Example 1 The following reference Figure 1 , Figure 2 The present invention will be described in detail below.

[0026] Figure 1 This is a flowchart illustrating an example of the data flow monitoring method of the present invention. Figure 2 This is a schematic diagram illustrating an application example of the data flow supervision method of the present invention.

[0027] like Figure 1 As shown, the data flow supervision method of the present invention includes the following steps.

[0028] In step S101, based on a customized collection strategy, log data and security data from different systems are collected and analyzed to obtain various key risk features related to risk events. These key risk features include network traffic features, user behavior features, system status features, data access features, and data transmission features.

[0029] In a specific application example, a data flow monitoring system is used in collaboration with multiple enterprise systems, which collects log data and security data from each of the multiple enterprise systems.

[0030] Specifically, based on factors such as the security requirements of different business systems within an enterprise, the update frequency of data sources, and the importance of the data, differentiated data collection strategies are formulated—that is, customized collection strategies corresponding to each enterprise system. For example, for logs from critical business systems, a collection frequency and requirements higher than a first set value are set. For logs from non-critical business systems, the collection frequency is reduced to below a second set value to minimize resource consumption. Simultaneously, the collection rules also consider data security and compliance requirements, ensuring that the collected data complies with relevant external and internal regulations.

[0031] Figure 2 This is a schematic diagram of the framework of an application example of the data management optimization method of the present invention.

[0032] from Figure 2 As can be seen from the example, the application of this invention includes a central node and multiple child nodes (e.g., child node 1, child node 2, and child node 3) that can communicate with (transmit data) the central node. The central node, as the core control center of the entire system (e.g., a data management system or data management platform), possesses powerful computing and storage capabilities, adopts a high-performance server cluster architecture, and can be flexibly expanded according to the data volume and business needs of various enterprises. After deploying multiple high-performance servers to form a central node cluster in a suitable location within the system, big data processing frameworks such as Hadoop and Spark, as well as related software for data management, analysis, and monitoring, are installed and configured. Simultaneously, network configuration and security settings are performed on the backend servers to achieve centralized management of each child node. This allows the central node to receive, store, and process international traffic data, log data, and other data from child node 1; domestic traffic data, log data, and other data from child node 2; and targeted traffic data, log data, and other data from child node 3. This ensures centralized data management and efficient processing, guaranteeing the stable operation of the central node and data security. The multiple child nodes are distributed across different business systems or network areas within the enterprise.

[0033] Optionally, the customized collection strategy includes a dynamically updated collection frequency and a data source update frequency. For example, the dynamically updated collection frequency includes: when the collection target is log data from a critical business system, adjusting the collection frequency to a value higher than a first set value (e.g., 30%); when the collection target is log data from a non-critical business system, adjusting the collection frequency to a value less than or equal to a second set value (e.g., 20%).

[0034] For critical business data in key business systems, this specifically includes business data related to the purchase of domestic traffic, international traffic, and targeted traffic on the client side. The critical business data also includes business data related to the purchase of domestic traffic, international traffic, and targeted traffic on the operations management side. For customized data collection strategies, such as a data collection delay assessment value less than a preset value (e.g., 2 seconds), data is collected immediately when critical business data is generated, reducing the number of polling iterations.

[0035] It should be noted that, in other embodiments, the customized data acquisition strategy may also include the amount of data to be acquired, information on changes in data format, and system load indicators (specifically including system CPU utilization, memory usage, network bandwidth, etc.). The above are merely optional examples and should not be construed as limiting the present invention.

[0036] Specifically, various log data include host logs, VPN logs, bastion host logs, database logs, and interface logs. Security data includes asset information, alarm data, and event data. Host logs contain information such as host operating status and operational behaviors, including user logins, file access, and process startups, to detect host-level risk activities and potential risks. VPN logs include virtual private network usage information, including connection time, user identity, and accessed internal resources, enabling monitoring of remote access security. Bastion host logs audit operational operations, including operational instructions, operation time, and operation objects, to ensure traceability of operational operations. Database logs contain records of database queries, modifications, and deletions to prevent data leaks and tampering. Interface logs include information on inter-system interface calls, including call time, call parameters, and return results, helping to analyze interface security and stability. Security data includes asset information, such as device model, configuration, and location. Alarm data includes alerts issued by security devices or systems when abnormal situations are detected, promptly reflecting potential security risks. Event data provides a detailed description of security incidents, including the time, location, and scope of impact of the incident.

[0037] In one specific implementation, a distributed collection probe is used to selectively crawl data from diverse data sources such as network and security logs, host logs, and API logs. The distributed collection probe possesses high concurrency processing capabilities and fault tolerance mechanisms, enabling efficient data collection from multiple data sources simultaneously. During the crawling process, incremental synchronization and breakpoint resumption technologies are employed to ensure data integrity. The incremental synchronization technology records the last modification time or version number of the data, acquiring only newly added or changed data, significantly reducing data transmission volume and improving collection efficiency. The breakpoint resumption technology automatically records the collection progress in abnormal situations such as network interruptions or device failures, resuming collection from the breakpoint after network recovery or device repair, avoiding data loss and duplicate collection. The collection results are encapsulated in a standardized message queue in real time and pushed to subsequent data parsing processes. For example, the message queue adopts a first-in, first-out (FIFO) principle to ensure data sequence and timeliness. The message queue has buffering and load balancing functions to cope with the pressure of peak data collection periods and ensure data transmission stability. During the push process, the data is encrypted to prevent data theft or tampering during transmission.

[0038] Next, the collected log data and security data are analyzed to obtain various key risk features related to risk events. These key risk features include network traffic features, user behavior features, system status features, data access features, and data transmission features.

[0039] Specifically, multi-dimensional data processing is performed through data standardization and cleaning engines to remove noise, duplicate values, and erroneous values ​​from various log data and security data, thereby improving data quality.

[0040] The data is parsed. For example, the VPN logs are parsed to record the precise time (accurate to milliseconds), the IP address used for login (including IPv4 and IPv6 formats), the specific resources accessed (such as server name, database name, file path, etc.), user identification (such as username, employee number), and login method (such as password login, certificate login). The bastion host logs are parsed to record the complete content of operation commands, the precise time of command execution, the operator's identity information, the assets operated on (such as server IP, port), and the operation result (success, failure, and reason for failure). The host logs are parsed to record system login logs (login time, logged-in user, login source IP, login method), process startup logs (process name, process ID, startup time, startup user), and file operation logs (operation type such as create, delete, modify, file path, operation time, and operator user). The interface logs are parsed to record the precise time of interface calls, call parameters (including parameter name, parameter value, parameter type), call results (return status code, return data), and the caller's IP address. Analyze the precise startup and shutdown times, system exception information (error codes, error descriptions, and locations of exceptions) and system resource usage (CPU usage, memory usage, and disk I / O) recorded in each system log.

[0041] For example, a data lineage awareness engine can be used to automatically parse and classify heterogeneous data collected from different enterprise systems. By analyzing the source, flow, and transformation process of the data, a data lineage graph is constructed. During the parsing process, the data is divided into different categories based on its format, structure, and semantic features, such as log data, security data, asset data, and alarm data. Further subdivision is then performed, such as dividing log data into system logs, application logs, and security logs, providing a clear data structure and classification basis for subsequent data processing and analysis.

[0042] For example, the feature engineering module is invoked synchronously to extract key risk features from multi-dimensional data processing data using statistical feature extraction and binning techniques. Representative and discriminative features are extracted based on different analysis scenarios and needs. For instance, in intrusion detection scenarios, network traffic features (traffic size, direction, protocol type, etc.), user behavior features (login time, login location, operation frequency, etc.), and system status features (CPU utilization, memory usage, disk I / O, etc.) are extracted; in data leakage detection scenarios, data access features (access time, access frequency, access permissions, etc.) and data transmission features (transmission protocol, transmission volume, transmission destination, etc.) are extracted. These scenario-based feature extractions provide high-quality feature input for subsequent risk analysis models, improving the accuracy and efficiency of security analysis. Thus, network traffic features, user behavior features, system status features, data access features, and data transmission features (i.e., key risk features) are obtained.

[0043] In another implementation, to dynamically optimize the customized data acquisition strategy, continuously monitor data source characteristics and system load, and dynamically adjust one or more parameters in the acquisition strategy, such as the acquisition frequency, using reinforcement learning algorithms such as policy gradient methods and Actor-Critic algorithms. By collecting information in real time on the data source's update frequency, data volume, data format changes, and system load indicators such as CPU utilization, memory usage, and network bandwidth, a data source characteristic and system load model is constructed. The reinforcement learning algorithm automatically adjusts the acquisition strategy, such as acquisition frequency, acquisition depth, and acquisition method, based on the data acquisition results and system performance indicators fed back by the model.

[0044] For example, when an increased update frequency is detected in a certain data source, the collection frequency of that data source is automatically increased. When the system load is too high, the collection frequency is reduced to below a specified value, or the collection of non-critical data sources is paused to ensure the stable operation of the entire system.

[0045] It should be noted that the above is only a preferred example and should not be construed as a limitation of the present invention.

[0046] Next, in step S102, a multi-dimensional risk event correlation analysis is performed on the collected log data, security data, and various risk key features to construct risk links corresponding to each risk event, and risk assessment calculations are performed on each risk link to determine the risk level of each risk link.

[0047] Specifically, multi-dimensional risk event correlation analysis is conducted on the collected log data, security data, and key risk characteristics. This multi-dimensional risk event correlation analysis includes time-based correlation analysis, subject-based correlation analysis, object-based correlation analysis, operation-based correlation analysis, and multi-dimensional correlation result fusion processing. Thus, correlation analysis is performed on events from multiple dimensions such as time, subject, object, and operation to determine the potential connections between events in each dimension, providing basic data and preliminary correlation clues for subsequent risk analysis.

[0048] For time-based correlation analysis, different time windows can be set based on the actual business scenario (e.g., office scenario, customer traffic purchase scenario) and security requirements, such as 1 minute, 5 minutes, 15 minutes, 30 minutes, etc. For example, for a transaction system, a shorter time window (e.g., 1 minute) can be set to capture rapidly occurring risky behaviors; while for an office scenario, a longer time window (e.g., 30 minutes) can be set.

[0049] A risk sequence pattern library is built to determine whether events to be processed are correlated in a temporal dimension. For example, "port scan → vulnerability exploitation → privilege escalation → data theft". Regular expressions or specialized sequence pattern matching algorithms are used to find combinations of events that satisfy the risk sequence pattern within a time window. For example, if a port scan event targeting a server is detected within 10 minutes and is subsequently determined to be a risky event, then a temporal correlation exists. Conversely, if a port scan event targeting a server is detected within 10 minutes and is subsequently determined to be a non-risky event, then no temporal correlation exists.

[0050] For object-dimensional correlation analysis, user and account information is extracted from event data (including risk events) and integrated with the user identity management system to obtain detailed information about the entities, such as user roles, departments, and permission scopes. Activities of the same or related entities are correlated. For example, if the same user performs abnormal logins and accesses sensitive data on different systems within a short period, these activities are correlated to analyze potential risks. Correlation rules can be set; for instance, if the same user accesses sensitive data on different systems more than a threshold (e.g., 5 times) within one hour, object correlation analysis is triggered.

[0051] The operation types in the events to be processed are categorized, such as login, access, modification, deletion, upload, and download. The operation types in the events to be processed are then identified.

[0052] For operation-level correlation analysis, an operation sequence rule base is constructed, and correlation analysis is performed using a rule matching algorithm. The operation sequence rule base includes specific operation sequence patterns (sequence patterns containing risky operations) and / or patterns conforming to attack chains. It is determined whether the current operation is a specific operation sequence pattern (sequence pattern containing risky operations) or conforms to an attack chain pattern, or a combination of a specific operation sequence pattern (sequence pattern containing risky operations) and a pattern conforming to an attack chain. For example, if an operation involving obtaining permissions (such as unauthorized access) is performed first, followed by a data download operation, this operation sequence is determined to pose a security risk.

[0053] For multi-dimensional correlation result fusion, the correlation results generated from the analysis of each dimension are merged to synthesize the evidence from each dimension and form a preliminary correlation network. For example, if two pending events are found to be similar in the time dimension, are the same user operation in the subject dimension, target the same data asset in the object dimension, and the operation sequence matches an attack pattern in the operation dimension, then these correlation results are integrated to form a more comprehensive correlation network. Weighted voting or Bayesian networks can be used for correlation result fusion.

[0054] After establishing multi-dimensional risk event correlations, the risk event chain is constructed based on preliminary correlation clues. This step, based on the identified related events, determines the initial point of the attack, analyzes each intermediate event, and connects them according to chronological order and causal relationships to form a complete risk path.

[0055] Further constructing the risk chain includes the following steps: Step S201: Identify the starting point of the risky link.

[0056] The analysis specifically examines the characteristics of the events to be processed, including their type, severity, frequency, and source. For example, when the event to be processed is a Category I risk event, the initial point of the risk chain is the Category I risk event. Examples of Category I risk events include logging in outside of working hours, logging in from a different location, and successfully logging in after multiple failed attempts. When the event to be processed is a Category II risk event, the initial point of the risk chain is the Category II risk event, i.e., an access control event, such as unauthorized access or privilege escalation. When the event to be processed is a Category III risk event, the initial point of the risk chain is the Category III risk event (i.e., a Category III risk point or attack point), such as a vulnerability scanning event.

[0057] Optionally, model matching can also be used to identify the initial point of a risk chain. Specifically, a risk starting point model is constructed based on historical attack data and expert experience. This risk starting point model includes the characteristics and rules of various risk starting points. The constructed risk starting point model is used to match risk events in the risk chain. When the characteristics of a risk event in the risk chain match the risk characteristics and rules in the risk starting point model, the risk event in the risk chain is marked as a risk point or initial point.

[0058] Furthermore, based on known attack chain models, events in the intermediate stages of each risk chain are analyzed and identified. Pattern matching and rule engines are used to find events that match the attack chain model. Known attack chain models include APT attack chains (information gathering → vulnerability exploitation → privilege escalation → lateral movement → data theft) and web attack chains (SQL injection → privilege acquisition → data tampering), etc.

[0059] Furthermore, this includes contextual analysis of risk events within the risk chain. Specifically, it combines the contextual information of risk events within the risk chain with network topology, system configuration, and business logic to further analyze events in intermediate stages. For example, by analyzing network traffic data, it can be determined that after gaining initial privileges, attackers move laterally through specific network paths to access sensitive data in other systems.

[0060] Next, based on the temporal and logical relationships of each event in the risk chain, the risk event serving as the initial point and each intermediate event are connected in chronological order and causal relationship to form a complete risk path, i.e., constructing the risk path. A directed acyclic graph (DAG) is used to represent the risk path, where the starting point represents the initial point (i.e., the risk event), intermediate nodes represent intermediate joint events (containing the risk event and other events), and edges represent the temporal and causal relationships between two adjacent nodes.

[0061] To collect more accurate and complete information, based on attack pattern knowledge and contextual information, we infer intermediate steps or information that exist but are not directly observed. For example, according to known attack patterns, attackers typically perform information gathering operations after gaining access, but if this information gathering operation is not found in the risk chain of the current event, we infer that there is an information gathering stage or that information has been lost.

[0062] In one specific implementation, the context information (or context data) includes asset topology relationships, risk intelligence information, and historical alarm processing records. Specifically, the asset topology relationships include physical connections between assets (e.g., port connections between servers and switches, link connections between switches and routers), logical associations (e.g., associations between business systems and the servers and databases supporting those systems), the name of the business system to which the asset belongs, the importance level of the business system (e.g., core business, important business, general business), and information about the department to which the asset belongs and its head.

[0063] Specifically, risk intelligence information includes detailed information about risky IP addresses (such as the geographical location of the IP address), the ISP to which they belong, the time when the risky behavior was first discovered, the type of risky behavior (such as scanning, attack, network control), and a severity score. Risk intelligence information also includes registration information of risky domains (registrant, registered email, registration time), resolved IP addresses, associated risky software information (risky software name, version, and harm characteristics), access paths of risky URLs, and associated risky page content characteristics (such as contained risky scripts and phishing information).

[0064] Furthermore, historical alert handling records include detailed information about historical alert events, detailed steps of the handling measures (such as the patch version installed and the configuration parameters modified), handling results (successfully eliminated the risk, partially mitigated the risk, or not eliminated the risk), and the identity information of the handling personnel. Detailed information about historical alert events includes the alert time, alert type (such as vulnerability exploitation or attack risk), and specific information about the affected assets (asset name and IP address).

[0065] Specifically, further data analysis and investigation are needed to fill in any missing links or information. For example, checking system logs, network traffic data, and security device alarm information can help find evidence related to information collection. Machine learning algorithms, such as anomaly detection algorithms, can be used to discover overlooked abnormal behaviors or operations.

[0066] Next, the integrity of the risk link is assessed and a risk score is given. The risk link integrity assessment evaluates the completeness and rationality of the constructed link, checking whether the risk link contains necessary intermediate links and whether there are logical jumps or breakpoints. Specifically, this includes determining assessment criteria and rules, extracting link information, and performing integrity checks.

[0067] To determine the evaluation criteria and rules, necessary intermediate steps should be identified. Based on attack chain models (such as APT attack chains) and industry security standards, the intermediate steps that must be included in different types of attack chains should be determined. For example, for attack chains involving database operations, they should typically include steps such as vulnerability exploitation (targeting database vulnerabilities) and privilege escalation (gaining database administrator privileges).

[0068] Determine the logical order and dependencies between each step to avoid logical jumps or breakpoints. For example, the privilege escalation step must be performed after initial privileges are obtained, and the lateral movement step usually follows privilege escalation.

[0069] For link information extraction, link data is collected, specifically extracting relevant information from the constructed risk event links, including the event type, occurrence time, involved subjects, and objects at each stage. For example, a risk link may contain the following information: initial login event (time: 2024-01-01 10:00, subject: hacker account, object: a server) and abnormal file access event (time: 2024-01-01 10:05, subject: hacker account, object: sensitive files on the server).

[0070] The extracted information is organized according to time sequence and causal relationships to form a clear link structure diagram, which facilitates subsequent analysis.

[0071] Integrity checks refer to checks on the completeness of each step in a risk chain, comparing it against assessment standards and rules to ensure that all necessary intermediate steps are included. If one or more critical steps are missing, the risk chain is deemed incomplete. For example, in the risk chain example above, if the vulnerability exploitation step is missing and the user is directly redirected from initial login to file access, then the risk chain is incomplete.

[0072] Examine whether the logical order and dependencies between each step in the risk chain conform to the rules. For example, check if there is a situation where data theft occurs before privilege escalation; if so, a logical problem is considered to exist. If not, no logical problem is considered to exist.

[0073] Next, the constructed risk link is scored, which is used to characterize the degree of potential harm.

[0074] Specifically, the following scoring factors were determined, along with their respective weights in the calculation: severity of the event, scope of impact, intent of the risk party (e.g., attacker), and asset sensitivity. Based on the determined scoring factors and their respective weights, a weighted scoring model is used to calculate the risk score for the current risk link using the following expression: R 当前 = S + B + I + A Where: R 当前This represents the risk score of the current risk chain, which is a weighted assessment of the current risk chain based on the severity of the event, the scope of impact, the intent of the risker, and the sensitivity of the asset; S represents the severity score of the risk events included in the current risk chain. This indicates the severity weight of the risk events included in the current risk chain; B represents the weight of the impact range of the current risk link; B represents the score of the impact range of the current risk link. I represents the intention weight of the risker; I represents the intention score of the risker. A represents the asset sensitivity weight of the asset involved in the current risk link; A represents the asset sensitivity score of the asset involved in the current risk link.

[0075] Specifically, the severity score of the risk events included in the current risk chain is calculated using the following expression: S = ; Where S represents the severity score of the risk events contained in the current risk link; A quantitative score for the risk consequences of the i-th risk event; This represents the consequence weight of the risk consequences arising from the i-th risk event; λ represents the time decay coefficient; Time_Decay represents the number of days from the date of occurrence of the risk event contained in the current risk link to the specified time.

[0076] The impact scope score of the risk event in the current risk link is calculated using the following expression: Region_Multiplier; Wherein, B represents the impact range score of the current risk link; Affected_Users represents the number of affected users involved in the current risk link; Total_User represents the total number of users involved in the current risk link; Affected_Systems represents the number of affected systems involved in the current risk link; Region_Multiplier represents the regional impact coefficient involved in the current risk link.

[0077] The intention score of the risk-taker is calculated using the following expression: I = ; Where I represents the risker's intention score; The weight of the j-th intention of the risker is represented by j, which is a positive integer, specifically 1, 2, ..., m, where m represents the number of types of intentions of the risker. The presence of intent is indicated by 1 and 0, respectively, representing the presence and absence of intent; Preparation_Level represents the attacker's level of preparedness for the attack.

[0078] The asset sensitivity score of the assets involved in the current risk chain is calculated using the following expression: A = ; Where A represents the asset sensitivity score of the assets involved in the current risk chain; The value score is assigned to the k-th asset class, where k is a positive integer, specifically 1, 2, ..., p, and p represents the number of asset classes. This indicates the criticality coefficient of the assets involved in the current risk chain; This indicates the asset exposure status of the assets involved in the current risk link, specifically represented by 1 or 0 to indicate asset exposure status and resource non-exposure status, respectively.

[0079] Furthermore, based on the classification range to which the calculated risk score belongs, the current risk link is determined to be either Level 1, Level 2, Level 3, or Level 4 risk.

[0080] For example, if the calculated risk score is between 0 and 2, then the current risk link is determined to be Level 1 risk, i.e., low risk.

[0081] For example, if the calculated risk score is between 2 and 4, then the current risk link is determined to be a level 2 risk, i.e., a medium risk.

[0082] For example, if the calculated risk score is between 4 and 7, then the current risk link is determined to be Level 3 risk, i.e., high risk.

[0083] For example, if the calculated risk score is greater than 7, then the current risk link is determined to be Level 4 risk, which is extremely high risk.

[0084] It should be noted that the above is only an optional example and should not be construed as a limitation of the present invention.

[0085] Next, in step S103, risk patterns are extracted based on the constructed risk links, and cross-system behavior associations are performed to identify collaborative risk behaviors.

[0086] Specifically, based on the constructed risk chain, the process proceeds to the risk pattern extraction stage to extract risk patterns.

[0087] It should be noted that in this example, by analyzing historical cases and summarizing patterns from the constructed risk chain, the key characteristics, triggering conditions, and specific sequences of risk patterns are defined. Risk pattern extraction is an abstraction and summary of attack behavior, enabling the rapid identification of risk events similar to the risk patterns.

[0088] Specifically, this involves collecting risk events and handling cases for a company over a historical period of time, from the current time back to a certain point in the past. This includes information such as risk event descriptions, attack methods, scope of impact, and handling measures.

[0089] Data mining techniques such as association rule mining and cluster analysis are used to identify risk patterns and characteristics from data collected over historical periods. For example, association rule mining can reveal a strong correlation between "multiple failed login attempts and abnormal IP addresses" and "hacking attacks."

[0090] Define the key risk characteristics of each risk model, such as the attacker's behavioral characteristics (e.g., attack frequency, attack time, attack tools, etc.), the characteristics of the affected system (e.g., system type, operating system version, application software version, etc.), and the attack time characteristics (e.g., peak attack period, attack duration, etc.).

[0091] For triggering conditions and specific sequences, determine the triggering conditions and specific sequences for risk patterns. For example, when a large number of abnormal login requests from the same IP address are detected, the detection of a DDoS attack pattern is triggered.

[0092] Specific sequences include sending risky emails disguised as normal emails, emails containing risky links or attachments, and users being infected with risky software after clicking on the links or downloading the attachments.

[0093] For example, the extracted risk patterns can be validated and optimized by incorporating expert experience. The accuracy and effectiveness of the risk patterns can be assessed based on actual risk events and handling experience. For instance, if the triggering conditions of a certain pattern cause the false positive rate to exceed the false positive threshold, the triggering conditions of that risk pattern will be adjusted and optimized.

[0094] It should be noted that in this example, cross-validation and ROC curve analysis are used to evaluate pattern performance. Actual data is used for testing to verify the performance of risky patterns. By continuously adjusting and optimizing the patterns, the accuracy and recall of risky pattern identification are improved. The above is only an optional example and should not be construed as limiting the invention.

[0095] Next, the extracted risk patterns are named and categorized to build a risk pattern library, which facilitates management and retrieval. For example, risk patterns can be divided into multiple categories such as network attack patterns, data breach patterns, privilege abuse patterns, and malware infection patterns. Furthermore, the risk pattern library supports adding, deleting, modifying, and querying risk patterns. The library is regularly updated, adding new risk patterns and deleting outdated or invalid ones. Version control technology can be used to record and manage changes to the pattern library.

[0096] Based on a constructed risk pattern library, new risk events similar to known risk patterns can be quickly identified. When a new event occurs, it is matched against existing risk patterns in the library, specifically calculating the similarity between the new event and existing risk patterns to determine the existing risk patterns similar to the new event. Similarity is calculated using methods such as string matching algorithms and vector space models. If the calculated similarity exceeds a preset threshold, the existing risk pattern similar to the new event is determined.

[0097] It should be noted that the string matching algorithm is suitable for risk pattern matching based on text features, such as risky URLs and specific commands. Specifically, it calculates whether the feature string of the new event completely or partially matches the strings of existing risk patterns in the risk pattern library to determine existing risk patterns similar to the new event. In other implementations, similarity can also be calculated using methods such as cosine similarity, Levenshtein distance algorithm, Jaro-Winkler distance, vector space model, and feature vectors. The above are only illustrative examples and should not be construed as limiting the invention.

[0098] Next, cross-system behavior correlation is performed to identify collaborative risk behaviors.

[0099] For enterprise systems with multiple subsystems, risky operations may be carried out across different systems to achieve attack objectives. Cross-system behavior correlation uses methods such as user identity mapping, session tracking analysis, and data flow tracking to link behaviors scattered across various subsystems and identify coordinated attack behaviors.

[0100] For user identity mapping, user identity information, including usernames, accounts, passwords, and digital certificates, is collected from various subsystems (or multiple systems). User identity mapping rules are established to associate user accounts across different subsystems. For example, mapping can be done using unique identifiers such as usernames, email addresses, and mobile phone numbers. Hash algorithms can be used to encrypt user identity information to protect user privacy. A unique identifier is generated for each user session, recording the session's start time, end time, and operation content to track each user session. Specifically, user session activities across different subsystems (or different systems) are tracked and associated. Multiple sessions of the same user are associated using unique identifiers to analyze cross-system behavioral patterns. For example, if a user's session switching frequency across different subsystems (or different systems) is found to be higher than the statistical average, it is identified as risky behavior.

[0101] By constructing activity sequences of the same user in different subsystems (or different systems), information such as the time, type, and goal of the activities can be recorded.

[0102] For example, association rule matching algorithms can be used to find relationships between activities across systems. For instance, if a user first downloads a file in the office system and then imports data into the business system, these activities can be linked together to analyze for potential risks.

[0103] Generate unique identifiers, or data tags, for each data point, recording its source, destination, transmission time, and content. Analyze the flow paths and behaviors of data across different systems. Track the data flow process using data identifiers to identify risky data transfer behaviors. For example, sensitive data might be transferred from a database system to an external network without authorization or approval.

[0104] Next, behavioral patterns scattered across multiple systems are analyzed to identify coordinated attack behaviors. For example, an attacker might gather information in one system, exploit vulnerabilities in another, and steal data in a third. This coordinated attack behavior can be identified through cross-system behavioral correlation analysis. Coordinated attack behaviors are identified by collecting correlative evidence that supports their identification (such as network traffic data, system logs, and security device alarm information).

[0105] Preferably, a comprehensive risk assessment is conducted for cross-system behaviors. Risk scores are calculated based on factors such as the degree of correlation, scope of impact, and potential hazards of the cross-system behaviors. The risk assessment is performed using a combination of the Analytic Hierarchy Process (AHP) and fuzzy comprehensive evaluation.

[0106] Specifically, the degree of correlation includes data interaction frequency (measuring the frequency of data exchange between systems, such as daily or weekly interaction counts), system coupling (reflecting the degree of interdependence between systems, such as the impact of a system failure on other systems), and interface complexity (assessing the design and implementation complexity of cross-system interfaces). The scope of impact includes the number of systems involved (counting the number of systems involved in cross-system activities), business criticality (judging the importance of the involved systems in the business process), and user coverage (determining the number of users affected by cross-system activities). Potential hazards include data risks (including risks of data leakage, tampering, and loss), business continuity risks (assessing the impact of activity failures on the normal operation of business processes), and compliance risks (checking whether the activity complies with relevant regulations).

[0107] The weights are determined using the analytic hierarchy process (AHP). Pairwise comparisons are made of the importance of each factor at the same level relative to a factor at the level above it. For example, a 1-9 scale is used for scoring, constructing an m x m judgment matrix A = ,in This represents the importance ratio of the i-th factor to the j-th factor, where i, j, and n are all positive integers. In this example, the number of rows and columns of the judgment matrix A are equal, and m is used to represent both.

[0108] Calculate the product of the elements in each row of the judgment matrix: = ; in, To determine the product of the elements in the i-th row of matrix A, where i = 1, 2, ..., m; This represents the importance ratio of the i-th factor to the j-th factor, where j = 1, 2, ..., m; calculate The nth root: = , i = 1, 2, ..., m; For vectors After normalization, the weight vector W is obtained, W= ,in, = , i = 1, 2, ..., m; j is a positive integer, specifically 1, 2, ..., N.

[0109] Next, calculate the largest eigenvalue of the judgment matrix. = , where A×W is the product of the judgment matrix A and the weight vector W, n represents the order, and j is a positive integer, specifically 1, 2, ..., N.

[0110] The consistency index (CI) is used to calculate the importance of factors at the same level relative to a factor at the next higher level. .

[0111] Find the average random consistency index RI (different RI values ​​correspond to different orders n), and calculate the consistency ratio CR = When the calculated CR is less than 0.1, the constructed judgment matrix passes the consistency test and the calculated weight vector is valid; when the calculated CR is greater than or equal to 0.1, the constructed judgment matrix fails the consistency test and the calculated weight vector is invalid, and the judgment matrix needs to be readjusted.

[0112] In addition, the fuzzy comprehensive evaluation method is used to calculate the risk score. This includes determining the evaluation set, establishing a fuzzy relation matrix, performing fuzzy comprehensive evaluation, and calculating the risk score for cross-system behavior.

[0113] For a given evaluation set, let the evaluation set V = { }, such as V = {low risk, relatively low risk, medium risk, high risk, extremely high risk}, and assign corresponding scores. v1, v2, ..., v s These represent different indicators.

[0114] To establish the fuzzy relation matrix, each indicator is scored relative to the membership degree of each level in the evaluation set, resulting in a membership degree vector for each indicator. The membership degree vectors of all indicators form the fuzzy relation matrix R = ,in, This represents the membership degree of the i'th indicator to the j'th evaluation level.

[0115] For fuzzy comprehensive evaluation, the evaluation of each indicator layer is performed to obtain the evaluation results of the criterion layer. = ∘ ,in, The weight vector of the indicator layer. Let ∘ be the fuzzy relation matrix of the index layer, and ∘ be the fuzzy composition operator (using a weighted average operator). = ).

[0116] A comprehensive evaluation of the criterion layer yields the evaluation result B = Z∘ of the target layer. Where Z is the weight vector of the criterion layer, and R is the weight vector of the criterion layer. The resulting fuzzy relation matrix.

[0117] Compare the evaluation result B with the score vector of the evaluation set. Multiplying these together yields a risk score X for cross-system behavior: X = B × ; = ; Where X represents the risk score for cross-system behavior; B represents the evaluation result of the target layer obtained by comprehensively evaluating the criteria layer; Represents the score vector of the evaluation set; Let v be the value of the j'-th element in the evaluation result B of the target layer, where j' is less than N', and both j' and N' are positive integers. j' Let j' represent the j-th evaluation set.

[0118] The risk level of cross-system activities is determined based on the risk score X, and corresponding risk response measures are formulated. For example, high-risk activities require immediate corrective action, while low-risk activities can be maintained under routine monitoring.

[0119] It should be noted that the above is only a preferred example and should not be construed as a limitation of the present invention.

[0120] Next, in step S104, model inference is used to generate associated explanatory text related to the risk event in order to construct a chain of evidence and assess the association confidence of the constructed chain of evidence.

[0121] Specifically, the process involves using a model to infer and generate explanatory text related to the risk event. For example, an existing large language model can be used to generate explanatory text associated with the risk event. The reasoning process of the large language model is recorded, including the knowledge used, the reasoning steps, and the intermediate results. The explanatory text includes the log recording time of the risk event, network traffic data, and device alarm information.

[0122] Next, collect the generated related explanatory texts and relevant evidence, such as risk event logs, system configurations, network traffic data, and security device alarm information.

[0123] Relevant evidence is organized in a logical order to form a complete chain of evidence. For example, blockchain technology can be used to store and verify the formed chain of evidence, ensuring its authenticity and immutability.

[0124] Further evaluate the association confidence of the constructed chain of evidence (e.g., the current chain of evidence). Specifically, based on factors such as data noise and incomplete knowledge, use Bayesian methods to calculate the association confidence of the constructed chain of evidence.

[0125] Let H be a hypothesis (such as the inference that the association is true), and E be the observed evidence. According to Bayes' theorem, the posterior probability P(H|E) (i.e., the confidence level of hypothesis H given evidence E) is calculated as follows:

[0126] P(H|E)= ; Where P(H) is the prior probability, specifically representing the probability that hypothesis H is true before the evidence E is observed; P(E|H) is the likelihood function, specifically representing the probability that the evidence E is observed given that hypothesis H is true.

[0127] P(E) is the marginal probability of evidence E, which can be calculated using the law of total probability: P(E) = ,in, Let i represent the i-th possible hypothesis, where i is a positive integer.

[0128] The calculated confidence level of the current evidence chain is compared with a specified threshold. When the calculated confidence level is less than or equal to the specified threshold (e.g., 95%), the association analysis result of the current evidence chain is modified and optimized. When the calculated confidence level is greater than the specified threshold, the association analysis result of the current evidence chain is stored. By calculating the confidence level to modify and optimize the association analysis relationship of each evidence chain, evidence chains with higher confidence and more reliable evaluation of inference associations can be obtained.

[0129] It should be noted that the above is only a preferred example and should not be construed as a limitation of the present invention.

[0130] Next, in step S105, feature matching calculations are performed on the event to be identified and various risk key features in the candidate cases to obtain feature matching calculation results, and semantic matching analysis is performed on the semantic descriptions of the event to be identified and the candidate cases to obtain semantic matching analysis results.

[0131] Specifically, the event to be identified (or the event to be processed) is matched with various key risk features in the candidate cases to obtain the feature matching calculation results.

[0132] By identifying key characteristics of the event to be identified, such as the risk type (e.g., data breach, malicious software infection), affected assets (e.g., servers, databases), and attack methods (e.g., SQL injection, DDoS attacks), potentially relevant candidate cases are selected from the case library. For example, if the event to be identified is a data breach risk, the affected asset is a database, and the attack method is SQL injection, then candidate cases involving database data breaches and SQL injection attacks are selected from the case library. The candidate cases may be one or more.

[0133] Specifically, cosine similarity and other vector similarity algorithms are used to calculate the degree of matching between the event to be identified and each candidate case in various dimensions of features (such as risk type, affected assets, degree of harm, and response measures). Each feature is converted into a vector representation, and the cosine similarity between the vectors of the event to be identified and each candidate case in each dimension of features (such as risk type, affected assets, degree of harm, and response measures) is calculated. When the calculated cosine similarity values ​​of at least two dimensions of features in the current candidate case are greater than a set value (e.g., 95%), the event to be identified is determined to be similar to the current candidate case. When the calculated cosine similarity values ​​of at least two dimensions of features in the current candidate case are less than or equal to the set value (e.g., 95%), the event to be identified is determined to be dissimilar to the current candidate case, and the search for similar candidate cases continues until a similar candidate case is found, or until a specified number of similar candidate cases (e.g., two, three, or more) are found.

[0134] Next, natural language processing techniques, such as the BERT model, are used to analyze the semantic similarity between the description of the event to be identified and the descriptions of candidate cases. The descriptions of both events are converted into vector representations, and the similarity between the description vectors of the event and the candidate cases is calculated to obtain the semantic matching analysis results. Simultaneously, factors such as keywords and semantic relationships in the descriptions are considered to improve the accuracy of semantic matching.

[0135] It should be noted that the above is only a preferred example and should not be construed as a limitation of the present invention.

[0136] Next, in step S106, based on the obtained feature matching calculation results, the obtained semantic matching analysis results, and the case quality scores of each candidate case, a comprehensive evaluation is performed on each candidate case that matches the event to be identified, so as to determine the final candidate case and visualize the relevant information of the event to be identified.

[0137] Specifically, based on the obtained feature matching calculation results, the obtained semantic matching analysis results, and the case quality scores of each candidate case obtained by matching, a comprehensive evaluation is performed on each candidate case that matches the event to be identified.

[0138] Based on the obtained feature matching calculation results, semantic matching analysis results, and case quality scores of each candidate case (such as the handling effect, applicability, and update time of the candidate cases), a comprehensive score and ranking are performed on each candidate case that matches the event to be identified. Specifically, factors such as the historical handling effect value, user score, timeliness parameter, and whether it is a recent case are calculated for each candidate case.

[0139] The following expression is used to calculate the overall score of the current candidate cases that match the event to be identified: Q = ; Where: Q represents the comprehensive score of the current candidate cases that match the event to be identified; This represents the historical handling effect value of the current candidate case; This indicates the user rating of the current candidate case; This indicates the timeliness parameter of the current candidate case. = ,in, Indicates the timeliness coefficient. Indicates the number of days to update; This represents the applicability parameter of the current candidate case; α, β, γ, and δ are the first dynamic weight, second dynamic weight, third dynamic weight, and fourth dynamic weight corresponding to the historical processing effect value, user rating, timeliness parameter, and adaptability parameter, respectively. For example, μ=0.01 specifically means that it decays to 37% after 100 days.

[0140] For the above dynamic weights to satisfy α+β+γ+δ=1, they can be adjusted according to user needs. For example, when more attention is paid to the effect, the first dynamic weight α will be increased.

[0141] For emergency scenarios (emphasizing the effectiveness and timeliness of handling), α=0.4, β=0.2, γ=0.3, δ=0.1.

[0142] For scenarios where the treatment effect and evaluation need to be balanced, α=0.3, β=0.3, γ=0.2, δ=0.2.

[0143] Specifically, the historical treatment effect value of the current candidate case is calculated using the following expression: = ; in, P represents the historical treatment effect value of the current candidate case; i Z represents the success rate of matching the pending event with the i-th current candidate case, where i is a positive integer and n represents the number of all historical candidate cases that have been matched; i λ represents the i-th resource efficiency ratio, which is specifically the ratio of recovered losses to invested resources; λ is the time decay coefficient. This represents the time difference between the event to be processed and the i-th candidate case. In this example, λ is 0.05, specifically representing a decay to 50% over 30 days.

[0144] The user rating for the current candidate case is calculated using the following expression: = P平均评分 (1 - ) log(G+1); in, P represents the user rating of the current candidate case; 平均评分 This represents the average rating based on the user's historical evaluation scores; This represents the variance of a user's historical ratings; G represents the number of historical rating cases, expressed as log(G+1), which uses a logarithmic function to avoid excessive weighting due to linear growth in the number of ratings. The larger the variance, the lower the reliability of the rating.

[0145] The applicability parameters for the current candidate case are calculated using the following expression: = ; in, Indicates the applicability parameters of the current candidate case; This represents the j-th feature of the event to be identified; Let j represent the j-th feature in the current case, where j and m are positive integers, m represents the number of features contained in the event to be identified, and j is less than m.

[0146] Specifically, when the overall score of the current candidate case exceeds the similarity threshold (e.g., 0.7), a specified number (e.g., two) of the candidate cases with high similarity are selected as reference cases, thus determining the final candidate cases, and the relevant information of the event to be identified is visualized. If no matching candidate case that meets the conditions is found, the matching conditions are adjusted, such as lowering the similarity threshold or changing the feature weights, and the candidate case matching is re-executed.

[0147] For example, choosing appropriate visualization tools, such as Gephi or D3.js, can graphically display relevant information about the risk events to be identified (including related networks, related event information, related networks, risk links, correlation strength, and inference conclusions). Interactive methods such as node zooming, dragging, filtering, and detailed viewing can be used to view the displayed information or data. Key information, such as event type, occurrence time, and risk score, should be labeled in the visualization charts to help users quickly understand the attack process.

[0148] Preferably, the correlation analysis results are exported to common file formats such as Excel, PDF, and CSV. A detailed analysis report is also generated, including an overview of the correlation events, risk assessment, recommended measures, and supporting evidence. Report templates can be customized to meet user needs.

[0149] It should be noted that in other implementations, a user-friendly query interface is also designed, allowing users to query historical analysis results based on conditions such as time range, event type, correlation strength, and risk score. Trend analysis is performed on the historical analysis results to display the changing trends of security risks, providing a reference for enterprise security decisions. The above is only described as a preferred example and should not be construed as limiting the present invention.

[0150] In one alternative implementation, the existing BERT model is used to perform text vectorization representation of the alarm description, converting each word in the alarm description into a 768-dimensional vector. For example, for the alarm description "SQL injection attempt detected from IP192.168.1.100", key semantics (such as "SQL injection attempt") can be identified and converted into corresponding vector representations to obtain risk text features.

[0151] By using entity recognition technology, attack characteristics can be extracted, such as injection points (e.g., the name of an input box in a web form, or a URL parameter name) and payload fragments (e.g., the specific content and length of "'OR '1'='1").

[0152] Perform time-series correlation analysis to construct a timeline of risk events and obtain time-series correlation characteristics. Arrange risk events that generate alarm information for the same asset or related assets at different times in chronological order to form an event timeline. For example, for a certain server, a port scan alarm first appears (time: 2024-01-01 10:00:00), followed by a vulnerability detection alarm (time: 2024-01-01 10:05:00), and finally an attack payload delivery alarm (time: 2024-01-01 10:10:00).

[0153] For example, by using the LSTM-CRF model to analyze the event timeline, common attack chain patterns can be identified, such as port scanning → vulnerability detection → attack payload delivery.

[0154] The LSTM-CRF model learns the temporal relationships and characteristics between different attack stages by learning from a large amount of historical attack data labeled with risk types. For example, the probability of a vulnerability detection alarm appearing within a certain time range (such as within 5 minutes) after a port scan alarm is relatively high.

[0155] By setting the sliding window size to 5 minutes, the frequency of similar alarms within the window (e.g., the number of times a port scan alarm occurs within 5 minutes) and the correlation between cross-device alarms (e.g., the number of times two devices simultaneously generate the same type of alarm within 5 minutes) are statistically analyzed to obtain the sliding window's characteristics. Furthermore,

[0156] For example, asset relationship graph embeddings can be generated based on GraphSAGE, representing the asset topology as graph data, where nodes represent assets and edges represent the connections between assets. The GraphSAGE algorithm is used to learn the graph data, generating a 128-dimensional graph embedding vector for each asset to reflect its position in the topology and its relationships with other assets, such as its distance from the core business system and the number of connected assets. This yields enhanced topology features.

[0157] Based on the asset relationship graph embedding vectors, critical paths are identified, such as the access chain from VPN / bastion host to host server to database. Simultaneously, combined with asset value tags (e.g., database server weight = 0.9, ordinary server weight = 0.5), an attack impact surface score is calculated. For example, if the attack path involves a high-weight database server, and this server is connected to many other important assets, the attack impact surface score will be higher.

[0158] Based on real attack logs, adversarial examples are generated. The real attack logs are modified by randomly adjusting the timestamp distribution (randomly shifting the timestamps within a certain range, such as ±10 minutes), making them no longer regular. Key fields are obfuscated, and IP addresses are partially replaced (e.g., replacing 192.168.1.100 with 192.168.1.101) or encrypted (using a symmetric encryption algorithm to encrypt the IP addresses). By generating adversarial examples, the robustness of the model against new and variant attacks is improved.

[0159] A comprehensive feature is formed based on the risk text features, temporal correlation features, and topological structure enhancement features.

[0160] In another optional implementation, a risk assessment model is constructed using a multimodal fusion network algorithm, a BERT training model, and a retrieval enhancement generation algorithm. The risk assessment model comprises an input layer, an intermediate processing layer, and an output layer.

[0161] First, a multimodal fusion network algorithm is used to fuse the original multimodal data (such as the original data collected by different devices and sensors mentioned above), extracting high-level features such as risk text features, temporal correlation features, and topology enhancement features from the data of different modalities. A gated attention mechanism is used for feature selection, dynamically assigning different feature weights to each feature according to the characteristics of the risk event. For example, for risk events related to high-risk assets, the feature weight of topology enhancement features is emphasized (e.g., the feature weight is adjusted to 0.6), because topological relationships can more directly affect the degree of risk. For alerts involving new attack methods, the feature weight of risk text features is emphasized (e.g., the feature weight is adjusted to 0.5). Next, the risk text features (e.g., a 768-dimensional feature vector), temporal correlation features (e.g., a 64-dimensional feature vector), and topology enhancement features (e.g., a 128-dimensional feature vector) are concatenated to form a comprehensive feature vector of a specified dimension (specifically 800 to 1000 dimensions, e.g., 960 dimensions) as the input data for the input layer.

[0162] In the intermediate processing layer, a BERT pre-trained language model based on the Transformer architecture is used to train the input data from the input layer, resulting in more accurate risk types and treatment priorities. The BERT model internally consists of multiple identical Transformer encoder layers stacked on top of each other. First, all parameters of the BERT layers are frozen, and only the top classification head (usually a fully connected layer) is trained. A large learning rate (e.g., 0.001) is initially used to enable the model to quickly learn the basic features of the input risk event classification. Then, all BERT parameters are unfrozen, and end-to-end fine-tuning is performed on all BERT layer parameters using a smaller learning rate (e.g., 0.0001) to improve the model's generalization ability.

[0163] During end-to-end fine-tuning training, a weighted cross-entropy loss function is used, with a class imbalance coefficient set (e.g., α = 0.8) to address the issue of imbalanced sample sizes across different risk types. Simultaneously, a priority ranking loss is incorporated to optimize the accuracy of the top K (i.e., Top-K) results (e.g., the top three). position (Accuracy) ensures that the model's predictions of handling priorities are more accurate. Through the above training, the model semantically encodes historical handling cases and alarm details, constructing a vector database, where text is converted into 768-dimensional vectors during encoding.

[0164] At the output layer, a retrieval-enhanced generation algorithm is first used to query the vector database generated by the risk assessment model in real time. Based on the risk text features, temporal correlation features, and topological structure enhancement features of the event to be processed, the algorithm recalls the top five (e.g., Top-5) similar candidate cases (i.e., historical cases or historical handling cases), along with related alarm details and assessment results, from the vector database. This outputs the top five similar candidate cases, related alarm details, and assessment results. Regular expressions are used to restrict the output format, such as using priority levels P0 to P3, to ensure that the generated assessment results can be correctly parsed and processed by the program. Simultaneously, an output length limit is set to avoid generating excessively long results.

[0165] Specifically, by inputting the events to be processed into the aforementioned risk assessment model, the results of classification and priority of handling are obtained. Among them, the event classification output is different risk types, such as data leakage, risky software infection, and hacking risk; the handling priority output is a specific priority level, such as P0 (highest priority, must be handled immediately), P1 (high priority, must be handled within 1 hour), P2 (medium priority, must be handled within 24 hours), and P3 (low priority, can be handled later).

[0166] Furthermore, the relevant information of the events to be processed (including the type of risk event, the priority of handling, and the judgment suggestions for generating risk events) will be visualized.

[0167] Optionally, a collaborative decision-making mechanism between the aforementioned risk assessment model and large-scale model technology can be established, including cascading trigger logic. For routine alarms (accounting for 95% of traffic), the aforementioned model processes them and quickly provides assessment results. For complex or low-confidence alarms (confidence less than 0.6), the large-scale model (using advanced models such as GPT-4) is triggered for in-depth analysis, providing a deeper understanding and analysis of the identified risk events. The output of the risk assessment model is used as the baseline result, and the suggestions from the large-scale model are used as a correction reference. When the results of the two conflict, a manual arbitration mechanism is initiated, with security experts making the final judgment. Security experts can make accurate decisions based on their experience and professional knowledge, comprehensively considering various factors.

[0168] In the verification and optimization stage of the risk assessment model output, an adversarial test set can be constructed through offline evaluation. 20% of the log samples are selected from historical data and obfuscated by inserting attack fragments (such as inserting a simulated SQL injection attack code into normal logs) and modifying key fields (such as replacing normal access IP addresses with risky IP addresses).

[0169] The key metrics of the evaluation results were validated, including the F1-score for risk event classification. For known attack types, the F1-score should be greater than 0.89; for novel risk types, the F1-score should be greater than 0.75, to measure the model's ability to classify different types of attacks. The F1-score is the harmonic mean of precision and recall, which comprehensively reflects the model's classification performance.

[0170] Optionally, online optimization can be performed during model operation. A shadow mode can be used, where the new risk assessment model and the old risk assessment rules run in parallel for a period of time (e.g., one week), comparing their alarm assessment efficiency. For example, the average assessment time can be calculated, requiring the new model's average assessment time to be 30% shorter than the old rule. Simultaneously, the accuracy and recall of both models are compared to ensure the new risk assessment model performs better. Human-participated feedback learning is used, sampling the results of human review (5% sampling rate) as feedback signals for the model's reinforcement learning. The algorithm optimizes the inference path of the large model, improving the accuracy of the assessment results. The above are merely optional examples and should not be construed as limiting the invention.

[0171] In the risk mitigation recommendation generation stage, key elements such as the scope of impact (e.g., the number of affected assets and business systems), severity (e.g., the amount of data loss and business interruption time caused by a data breach), and potential sources (e.g., the attacker's IP address and attack methods) are automatically identified and extracted. For example, by analyzing the correlation between asset information and business systems in alarm logs, the number of affected assets and business systems can be determined; the severity of the attack methods and the scope of impact can be assessed; and the source of the attacker can be determined through IP reputation databases and attack signature analysis.

[0172] The identified risk events are assessed from multiple dimensions, including technology (such as the difficulty of exploiting vulnerabilities and the stealth of attacks), business (such as the impact on business continuity and the impact on customer trust), and compliance (such as whether relevant laws and standards are violated), and a comprehensive risk score is given.

[0173] The comprehensive risk score for the risk event to be identified is calculated using the following expression: R = T + Y + C Where R represents the comprehensive risk score of the risk event to be identified; T represents the risk score in the technology dimension; Y represents the risk score in the business dimension; and C represents the risk score in the compliance dimension.

[0174] Because there is a certain synergistic effect between the difficulty of vulnerability exploitation and the stealth of attacks (for example, highly stealthy attacks may be more difficult to exploit, but once successfully exploited, their impact is greater), the following expression (specifically a non-linear scoring formula) is used to calculate the technical dimension risk score of the risk event to be identified: T = ɑ×( )+ ; Where T represents the technical risk score, specifically expressed by quantitative parameters of vulnerability exploitation difficulty and attack concealment; and These represent the normalized scores for vulnerability exploitation difficulty and attack concealment, respectively; α is the total weight coefficient for the technical dimension. The parameters representing the control of the synergistic effect intensity can be adjusted according to the actual situation; This represents the offset, used to adjust the base value of the technical dimension risk score.

[0175] For the normalized scores of vulnerability exploitation difficulty and attack concealment, for example, low, medium and high are mapped to values ​​between 0 and 1, such as low = 0.2, medium = 0.5 and high = 0.8.

[0176] Because there is a dynamic relationship between business continuity and customer trust (e.g., the longer the business interruption, the faster customer trust declines), the following expression is used to calculate the business dimension score of the risk event to be identified: Y= ( ) +

[0177] Where Y represents the business dimension score of the risk event to be identified; and δ represents the first and second scores for the impact on business continuity and customer trust, respectively; δ represents the total weight coefficient of the business dimension; λ represents the parameter that controls the rate of decay of the impact on customer trust; and threshold represents the adjustment threshold, indicating the point at which customer trust begins to decline significantly. This indicates the offset.

[0178] The first and second scores for the impact on business continuity and customer trust are normalized scores, for example, scores based on the duration of business interruption or the number of customer complaints.

[0179] Because the severity of compliance is not linear (e.g., the gap between minor and serious violations is much larger than the gap between full compliance and minor violations), a non-linear scoring formula based on the severity of the violation is used to calculate: C = ζ (1 - ); Where C represents the compliance dimension risk score of the risk event to be identified; Compliance_level represents the initial compliance score; ζ represents the total weight coefficient of the compliance dimension (or the maximum possible deduction); k represents the parameter that controls the steepness of the scoring curve; and μ represents the threshold of the compliance score, used to distinguish between minor and serious violations.

[0180] For the initial compliance score, for example, the initial compliance score obtained based on the preliminary judgment is one of the following three: full compliance, represented by 0; minor violation, represented by 1; and serious violation, represented by 2.

[0181] Furthermore, by comparing the similarity between current alerts and historical cases in terms of attack type, affected assets, and attack methods, relevant experience can be identified. The similarity calculation uses a cosine similarity algorithm to calculate the similarity value between the feature vectors of the current alert and historical cases. When the similarity value exceeds a certain threshold (e.g., 0.8), the two are considered to have high similarity, and the handling experience of historical cases can be referenced. False alarm characteristics in alerts are analyzed, such as normal business operations being mistakenly identified as attack behavior. By establishing a false alarm feature library, alerts are filtered to reduce the number of false alarms. The false alarm feature library can include common false alarm scenarios (such as network traffic generated by automatic system updates, logs generated by employees accessing business systems normally) and corresponding characteristics (such as traffic patterns, access time, access frequency, etc.).

[0182] The risk assessment system described above generates recommendations for identifying risk events, such as confirmed risk (clearly existing security threat), suspected risk (potential security threat, requiring further confirmation), and false alarm risk (no security threat). The criteria for judgment include feature matching degree, model confidence level, and historical case references.

[0183] For confirmed risk events, provide initial response and handling recommendations, such as isolating affected systems (isolation methods, post-isolation operations), blocking attack sources (IP blocking methods, blocking duration), and collecting attack evidence (evidence content, storage location). Generate the basis and explanations for each recommendation to improve its credibility and understandability. For example, for the recommendation to isolate affected systems, explain that a system vulnerability has created a risk or alert, and the attacker has already gained partial access to the system; to prevent further spread of the attack, the system needs to be isolated immediately.

[0184] Furthermore, the accompanying drawings are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes shown in the drawings do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0185] Compared with existing technologies, the data flow monitoring method provided in this invention, based on a customized collection strategy, collects log data and security data from different systems and performs data parsing to obtain various key risk features related to risk events. This effectively collects relevant data on risk events and obtains accurate feedback on various key risk features. Multi-dimensional risk event correlation analysis is performed on the collected log data, security data, and the aforementioned key risk features to construct risk links corresponding to each risk event. Risk assessment calculations are performed on each risk link to accurately determine its risk level. Based on the constructed risk links, risk patterns are extracted, and cross-system behavioral correlations are performed to identify collaborative risk behaviors. A fusion model is used to infer and generate associated explanatory text related to the risk event to construct an evidence chain. Feature matching calculations are performed between the event to be identified and various key risk features in candidate cases to obtain feature matching results. Semantic matching analysis is also performed between the semantic descriptions of the event to be identified and candidate cases to obtain semantic matching analysis results. A comprehensive evaluation is conducted on each candidate case matching the event to be identified to determine the final candidate case, and the relevant information of the event to be identified is visualized. In terms of risk identification and analysis, by correlating security logs with multi-dimensional risk events in data, comprehensive and real-time security supervision of structured and unstructured data flowing across systems and networks in all scenarios can be achieved, and risk events can be accurately identified.

[0186] Example 2 The following is a system embodiment of the present invention, which implements the data flow supervision method embodiment of embodiment 1. For details not disclosed in the system embodiment of the present invention, please refer to the method embodiment of the present invention.

[0187] Figure 3 This is a schematic diagram of an example of the data flow monitoring system of the present invention.

[0188] Reference Figure 3 The second aspect of this disclosure provides a data flow monitoring system that employs the data flow monitoring method described in the first aspect of this invention.

[0189] Specifically, the data flow monitoring system 100 includes a data acquisition module 10, a data analysis module 20, a data processing module 30, an inference processing module 40, a calculation processing module 50, and a determination module 60.

[0190] In one specific implementation, the data acquisition module 10, based on a customized acquisition strategy, collects log data and security data from different systems and performs data parsing to obtain various key risk features related to risk events. These key risk features include network traffic features, user behavior features, system status features, data access features, and data transmission features. The data analysis model 20 performs multi-dimensional risk event correlation analysis on the collected log data, security data, and the aforementioned key risk features, constructs risk links corresponding to each risk event, and performs risk assessment calculations on each risk link to determine its risk level. The data processing module 30 extracts risk patterns based on the constructed risk links and performs cross-system behavior correlation to identify collaborative risk behaviors. The inference processing module 40 uses a fusion model to infer and generate related explanatory text associated with the risk event to construct an evidence chain and assess the correlation confidence of the constructed evidence chain. The calculation processing module 50 performs feature matching calculations between the event to be identified and various key risk features in candidate cases to obtain feature matching calculation results, and performs semantic matching analysis on the semantic descriptions of the event to be identified and candidate cases to obtain semantic matching analysis results. The determination module 60 comprehensively evaluates each candidate case that matches the event to be identified based on the obtained feature matching calculation results, the obtained semantic matching analysis results, and the case quality score of each candidate case obtained by matching, so as to determine the final candidate case and visualize the relevant information of the event to be identified.

[0191] According to the optional implementation method, risk assessment calculations are performed on each risk link, including: determining the following scoring factors and determining the calculation weight of each scoring factor: event severity, scope of impact, risker's intent, and asset sensitivity.

[0192] Based on the determined scoring factors and their respective weights, a weighted scoring model is used to calculate the risk score for the current risk link using the following expression: R 当前 = S + B + I + A Where: R 当前 This represents the risk score of the current risk chain, which is a weighted assessment of the current risk chain based on the severity of the event, the scope of impact, the intent of the risker, and the sensitivity of the asset; S represents the severity score of the risk events included in the current risk chain. This indicates the severity weight of the risk events included in the current risk chain; B represents the weight of the impact range of the current risk link; B represents the score of the impact range of the current risk link. I represents the intention weight of the risker; I represents the intention score of the risker. A represents the asset sensitivity weight of the asset involved in the current risk link; A represents the asset sensitivity score of the asset involved in the current risk link.

[0193] Based on the risk score, the current risk link is determined to be either Level 1, Level 2, Level 3, or Level 4 risk.

[0194] According to an optional implementation, the severity score of the risk events included in the current risk link is calculated using the following expression: S = ×

[0195] Where S represents the severity score of the risk events contained in the current risk link; A quantitative score for the risk consequences of the i-th risk event; This represents the consequence weight of the risk consequences arising from the i-th risk event; λ represents the time decay coefficient; Time_Decay represents the number of days from the date of occurrence of the risk event contained in the current risk link to the specified time.

[0196] The impact scope score of the risk event in the current risk link is calculated using the following expression: Region_Multiplier Wherein, B represents the impact range score of the current risk link; Affected_Users represents the number of affected users involved in the current risk link; Total_User represents the total number of users involved in the current risk link; Affected_Systems represents the number of affected systems involved in the current risk link; Region_Multiplier represents the regional impact coefficient involved in the current risk link.

[0197] According to the optional implementation method, the intention score of the risker is calculated using the following expression: I =

[0198] Where I represents the risker's intention score; The weight of the j-th intention of the risker is represented by j, which is a positive integer, specifically 1, 2, ..., m, where m represents the number of types of intentions of the risker. The presence of intent is indicated by 1 and 0, respectively, representing the presence and absence of intent; Preparation_Level represents the attacker's level of preparedness for the attack.

[0199] The asset sensitivity score of the assets involved in the current risk chain is calculated using the following expression: A =

[0200] Where A represents the asset sensitivity score of the assets involved in the current risk chain; The value score is assigned to the k-th asset class, where k is a positive integer, specifically 1, 2, ..., p, and p represents the number of asset classes. This indicates the criticality coefficient of the assets involved in the current risk chain; This indicates the asset exposure status of the assets involved in the current risk link, specifically represented by 1 or 0 to indicate asset exposure status and resource non-exposure status, respectively.

[0201] According to an optional implementation, the candidate cases that match the event to be identified are comprehensively evaluated to determine the final candidate cases.

[0202] The following expression is used to calculate the overall score of the current candidate cases that match the event to be identified: Q =

[0203] Where Q represents the comprehensive score of the current candidate cases that match the event to be identified; This represents the historical handling effect value of the current candidate case; This indicates the user rating of the current candidate case; This indicates the timeliness parameter of the current candidate case. = ,in, Indicates the timeliness coefficient. Indicates the number of days to update; The applicability parameters of the current candidate case are represented by α, β, γ, and δ, which are the first dynamic weight, second dynamic weight, third dynamic weight, and fourth dynamic weight corresponding to the historical processing effect value, user rating, timeliness parameter, and adaptability parameter, respectively.

[0204] According to the optional implementation method, by identifying the risk type, affected assets, and attack methods in the event to be identified, potentially relevant candidate cases are selected from the case library. Specifically, this includes calculating the similarity with the current candidate cases. When the calculated comprehensive score of the current candidate cases exceeds the similarity threshold, a specified number of candidate cases are selected as reference cases.

[0205] According to the optional implementation method, the following related explanatory texts and related evidence generated are collected and organized in logical order to form a complete chain of evidence: risk event logs, system configuration, network traffic data, and security device alarm information; the association confidence of the constructed chain of evidence is further evaluated.

[0206] According to an optional implementation, the method further includes: formulating customized collection strategies corresponding to each enterprise system based on the security requirements of different business systems of the enterprise, the update frequency of the data source, and the importance of the data; the customized collection strategy includes a dynamically updated collection frequency and a data source update frequency.

[0207] It should be noted that the data flow supervision method executed by the data flow supervision system in Embodiment 2 is the same as that in Embodiment 1, therefore, the description of the same parts is omitted.

[0208] Compared with existing technologies, this invention, based on a customized collection strategy, collects log data and security data from different systems and performs data analysis to obtain various key risk features related to risk events. This effectively collects relevant data on risk events and provides accurate feedback on various key risk features. Multi-dimensional risk event correlation analysis is performed on the collected log data, security data, and the aforementioned key risk features to construct risk links corresponding to each risk event. Risk assessment calculations are then performed on each risk link to accurately determine its risk level. Based on the constructed risk links, risk patterns are extracted, and cross-system behavioral correlations are established to identify collaborative risk behaviors. A fusion model is used to infer and generate associated explanatory text related to the risk event to construct an evidence chain. Feature matching calculations are performed between the event to be identified and various key risk features in candidate cases to obtain feature matching results. Semantic matching analysis is also performed between the semantic descriptions of the event to be identified and candidate cases to obtain semantic matching analysis results. A comprehensive evaluation of each candidate case matching the event to be identified is conducted to determine the final candidate cases, and the relevant information of the event to be identified is visualized. In terms of risk identification and analysis, by correlating security logs with multi-dimensional risk events in data, comprehensive and real-time security supervision of structured and unstructured data flowing across systems and networks in all scenarios can be achieved, and risk events can be accurately identified.

[0209] Figure 4This is a schematic diagram of an embodiment of an electronic device according to the present invention.

[0210] like Figure 4 As shown, the electronic device is embodied in the form of a general-purpose computing device. There can be one or more processors working collaboratively. This invention also does not preclude distributed processing, meaning that processors can be distributed across different physical devices. The electronic device of this invention is not limited to a single entity, but can also be the sum of multiple physical devices.

[0211] The memory stores a computer-executable program, typically machine-readable code. The computer-readable program can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some steps of the method.

[0212] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and may also be non-volatile memory, such as read-only memory (ROM).

[0213] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with external devices. The I / O interface can represent one or more of several bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0214] It should be understood that Figure 4 The electronic device shown is merely one example of the present invention, and the electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as displays, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. Any electronic device capable of executing a computer-readable program in memory to implement the method of the present invention or at least some steps of the method can be considered as an electronic device covered by the present invention.

[0215] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software, or by combining software with necessary hardware. Therefore, as... Figure 5 As shown, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) or on a network, and includes several commands to cause a computing device (such as a personal computer, server, or network device, etc.) to execute the above-described method according to the embodiments of the present invention.

[0216] The software product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0217] The computer-readable storage medium may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting programs for use by or in connection with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0218] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0219] The aforementioned computer-readable medium carries one or more programs, which, when executed by a device, enable the computer-readable medium to implement the data interaction method of this disclosure.

[0220] Those skilled in the art will understand that the above modules can be distributed in the device as described in the embodiments, or they can be modified to be uniquely different from one or more devices in this embodiment. The modules in the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0221] Through the description of the above embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or on a network, including several commands to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of the present invention.

[0222] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0223] In the detailed description above, reference has been made to the accompanying drawings, which form part of this document. In the drawings, similar symbols typically identify similar parts unless the context otherwise indicates otherwise. The illustrated embodiments described in the detailed specification, drawings, and claims are not intended to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0224] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for regulating data flow, characterized in that, The method includes: Based on a customized collection strategy, log data and security data from different systems are collected and analyzed to obtain various key risk features related to risk events. These key risk features include network traffic features, user behavior features, system status features, data access features, and data transmission features. Multi-dimensional risk event correlation analysis is performed on the collected log data, security data, and key risk features of various types to construct risk links corresponding to each risk event, and risk assessment calculations are performed on each risk link to determine the risk level of each risk link. Based on the constructed risk links, risk patterns are extracted and cross-system behavior is correlated to identify collaborative risk behaviors; The fusion model is used to infer and generate related explanatory texts associated with risk events in order to construct a chain of evidence and assess the association confidence of the constructed chain of evidence. The event to be identified is matched with various key risk features in the candidate cases to obtain the feature matching calculation results. The semantic matching analysis is performed on the semantic descriptions of the event to be identified and the candidate cases to obtain the semantic matching analysis results. Based on the obtained feature matching calculation results, semantic matching analysis results, and case quality scores of each candidate case, a comprehensive evaluation is performed on each candidate case that matches the event to be identified, so as to determine the final candidate case and visualize the relevant information of the event to be identified.

2. The data flow supervision method according to claim 1, characterized in that, The risk assessment calculation for each risk link includes: The following scoring factors were identified, and their respective weights were determined: severity of the event, scope of impact, intent of the risk stakeholder, and asset sensitivity. Based on the determined scoring factors and their respective weights, a weighted scoring model is used to calculate the risk score for the current risk link using the following expression: R 当前 = S + B + I + A; Where: R 当前 This represents the risk score of the current risk chain, which is a weighted assessment of the current risk chain based on the severity of the event, the scope of impact, the intent of the risker, and the sensitivity of the asset; S represents the severity score of the risk events included in the current risk chain. This indicates the severity weight of the risk events included in the current risk chain; B represents the weight of the impact range of the current risk link; B represents the score of the impact range of the current risk link. I represents the intention weight of the risker; I represents the intention score of the risker. A represents the asset sensitivity weight of the assets involved in the current risk link; A represents the asset sensitivity score of the assets involved in the current risk link. Based on the risk score, the current risk link is determined to be either Level 1, Level 2, Level 3, or Level 4 risk.

3. The data flow supervision method according to claim 2, characterized in that, Further includes: The severity score of the risk events included in the current risk chain is calculated using the following expression: S = × ; Where S represents the severity score of the risk events contained in the current risk link; A quantitative score for the risk consequences of the i-th risk event; This represents the consequence weight of the risk consequences arising from the i-th risk event; λ represents the time decay coefficient; Time_Decay represents the number of days from the occurrence date of the risk event included in the current risk link to the specified time; and / or The impact scope score of the risk event in the current risk link is calculated using the following expression: Region_Multiplier; Wherein, B represents the impact range score of the current risk link; Affected_Users represents the number of affected users involved in the current risk link; Total_User represents the total number of users involved in the current risk link; Affected_Systems represents the number of affected systems involved in the current risk link; Region_Multiplier represents the regional impact coefficient involved in the current risk link.

4. The data flow supervision method according to claim 2, characterized in that, Further includes: The intention score of the risk-taker is calculated using the following expression: I = ; Where I represents the risker's intention score; The weight of the j-th intention of the risker is represented by j, which is a positive integer, specifically 1, 2, ..., m, where m represents the number of types of intentions of the risker. The presence of intent is indicated by 1 and 0, representing the presence and absence of intent, respectively; Preparation_Level represents the attacker's level of preparedness; and / or The asset sensitivity score of the assets involved in the current risk chain is calculated using the following expression: A = ; Where A represents the asset sensitivity score of the assets involved in the current risk chain; The value score is assigned to the k-th asset class, where k is a positive integer, specifically 1, 2, ..., p, and p represents the number of asset classes. This indicates the criticality coefficient of the assets involved in the current risk chain; This indicates the asset exposure status of the assets involved in the current risk link, specifically represented by 1 or 0 to indicate asset exposure status and resource non-exposure status, respectively.

5. The data flow supervision method according to claim 1, characterized in that, The process of comprehensively evaluating each candidate case that matches the event to be identified to determine the final candidate case includes: The following expression is used to calculate the overall score of the current candidate cases that match the event to be identified: Q = ; Where Q represents the comprehensive score of the current candidate cases that match the event to be identified; This represents the historical handling effect value of the current candidate case; This indicates the user rating of the current candidate case; This indicates the timeliness parameter of the current candidate case. = ,in, Indicates the timeliness coefficient. Indicates the number of days updated; The applicability parameters of the current candidate case are represented by α, β, γ, and δ, which are the first dynamic weight, second dynamic weight, third dynamic weight, and fourth dynamic weight corresponding to the historical processing effect value, user rating, timeliness parameter, and adaptability parameter, respectively.

6. The data flow supervision method according to claim 1, characterized in that, Further includes: By identifying the risk type, affected assets, and attack methods in the event to be identified, potentially relevant candidate cases are selected from the case library. Specifically, this includes calculating the similarity with the current candidate cases. When the overall score of the current candidate case exceeds the similarity threshold, a specified number of candidate cases are selected as reference cases.

7. The data flow supervision method according to claim 1, characterized in that, Further includes: Collect the following related explanatory texts and evidence generated, and organize them in a logical order to form a complete chain of evidence: risk event logs, system configuration, network traffic data, and security device alarm information; Further assess the confidence level of the association of the constructed chain of evidence.

8. The data flow supervision method according to claim 1, characterized in that, Further includes: Based on the security requirements of different business systems of an enterprise, the update frequency of data sources, and the importance of data, formulate customized data collection strategies corresponding to each enterprise system. The customized acquisition strategy includes dynamically updated acquisition frequency and data source update frequency.

9. A data flow monitoring system, characterized in that, The data flow supervision method according to any one of claims 1 to 8 is implemented, wherein the data flow supervision system comprises: The data acquisition module, based on a customized acquisition strategy, collects log data and security data from different systems and performs data analysis to obtain various key risk features related to risk events. These key risk features include network traffic features, user behavior features, system status features, data access features, and data transmission features. The data analysis model is used to perform multi-dimensional risk event correlation analysis on the collected log data, security data, and key risk features, construct risk links corresponding to each risk event, and perform risk assessment calculations on each risk link to determine the risk level of each risk link. The data processing module is used to extract risk patterns and perform cross-system behavior correlation based on the constructed risk links in order to identify collaborative risk behaviors. The inference processing module is used to generate related explanatory texts associated with risk events using a fusion model to construct a chain of evidence and assess the confidence level of the constructed chain of evidence. The calculation and processing module is used to perform feature matching calculations between the event to be identified and various risk key features in the candidate cases to obtain the feature matching calculation results, and to perform semantic matching analysis between the semantic descriptions of the event to be identified and the candidate cases to obtain the semantic matching analysis results. The determination module, based on the obtained feature matching calculation results, semantic matching analysis results, and case quality scores of each matched candidate case, comprehensively evaluates each candidate case that matches the event to be identified, in order to determine the final candidate case, and visualizes the relevant information of the event to be identified.

10. The data flow monitoring system according to claim 9, characterized in that, The risk assessment calculation for each risk link includes: The following scoring factors were identified, and their respective weights were determined: severity of the event, scope of impact, intent of the risk stakeholder, and asset sensitivity. Based on the determined scoring factors and their respective weights, a weighted scoring model is used to calculate the risk score for the current risk link using the following expression: R 当前 = S + B + I + A; Where: R 当前 This represents the risk score of the current risk chain, which is a weighted assessment of the current risk chain based on the severity of the event, the scope of impact, the intent of the risker, and the sensitivity of the asset; S represents the severity score of the risk events included in the current risk chain. This indicates the severity weight of the risk events included in the current risk chain; B represents the weight of the impact range of the current risk link; B represents the score of the impact range of the current risk link. I represents the intention weight of the risker; I represents the intention score of the risker. A represents the asset sensitivity weight of the assets involved in the current risk link; A represents the asset sensitivity score of the assets involved in the current risk link. Based on the risk score, the current risk link is determined to be either Level 1, Level 2, Level 3, or Level 4 risk.

Citation Information

Cited By

  • Data security risk early warning method and system based on artificial intelligence

    CN121923947A

  • Data transaction security assessment method based on gated attention feedforward network

    CN122066518A

  • Sensitive data circulation risk assessment and detection system and method based on AI

    CN122204560A