A security alarm noise reduction method and system based on data weaving technology

By employing data weaving technology and AI-driven security alarm noise reduction methods, the problems of alarm overload and high false alarm rates in security alarm systems have been solved, enabling efficient and accurate security alarm processing and dynamic linkage response, thereby improving the overall efficiency of the security operations center.

CN120811857BActive Publication Date: 2026-03-27BEIJING HUIERTE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing security alert systems suffer from problems such as alarm overload, high false alarm rates, and low processing efficiency. Traditional methods cannot adapt to dynamically changing network attack patterns and new data sources, making it difficult for operations and maintenance personnel to effectively focus on real threats.

Method used

A security alarm noise reduction method based on data weaving technology is adopted. By constructing a unified logical data view, deploying an AI-driven data intelligence agent component system, and combining an attack chain panorama, a behavioral causal chain model, and a semantic noise reduction algorithm, intelligent access, semantic fusion, and efficient alarm processing of multi-source heterogeneous data are achieved.

Benefits of technology

It improves the accuracy and processing efficiency of security alerts, reduces false alarms and low-risk events, provides intuitive attack path analysis and dynamic linkage response strategies, and enhances the overall effectiveness of the security operations center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120811857B_ABST
    Figure CN120811857B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data noise reduction, and discloses a security alarm noise reduction method and system based on a data weaving technology, which comprises the following steps: based on a data weaving architecture, a unified logical data view is constructed; an AI-driven data intelligent agent component system is deployed, API and new data sources that change automatically are recognized and adapted, metadata, threat intelligence and infrastructure operation state information are fused; an attack chain panoramic graph and a behavior causal chain model are constructed; a semantic noise reduction algorithm, time series clustering analysis and a confidence score mechanism are used to remove duplicate alarms, identify false alarms and filter low-risk events for security alarms; and alarm deduplication, priority adjustment and linkage disposal strategy generation are carried out. The application solves the problems of alarm flooding, high false alarm rate and low processing efficiency in the existing security alarm system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data noise reduction technology, and in particular to a method and system for noise reduction of security alarms based on data weaving technology. Background Technology

[0002] With the rapid development of information technology, enterprise network environments are becoming increasingly complex, and the massive amount of alerts generated by various security devices (such as firewalls, intrusion detection systems, and endpoint protection software) is exploding. These alerts, generated by security tools from different vendors and architectures, often contain a large number of duplicates, false alarms, and low-value events, leading to "alert fatigue" among operations and maintenance personnel and making it difficult to effectively focus on real threats. Traditional methods rely on manual rules or simple threshold filtering, which cannot adapt to dynamically changing network attack patterns and constantly emerging new data sources. There is an urgent need for an innovative solution that can automatically integrate multi-source heterogeneous data, intelligently identify noise, and optimize the handling process. Summary of the Invention

[0003] The purpose of this invention is to provide a security alarm noise reduction method and system based on data weaving technology, aiming to solve at least one of the above-mentioned problems.

[0004] This invention provides a security alarm noise reduction method based on data weaving technology, comprising:

[0005] Based on the data weaving architecture, a unified logical data view is built to achieve intelligent access and semantic fusion of distributed, heterogeneous secure data sources;

[0006] Deploy an AI-driven data intelligence agent component system to automatically identify and adapt to changing APIs and new data sources, and integrate metadata, threat intelligence, and infrastructure operation status information;

[0007] Based on the fused data, a panoramic view of the attack chain, a behavioral causal chain model, and a threat intelligence relationship network are constructed.

[0008] Semantic denoising algorithms, time series clustering analysis, and confidence scoring mechanisms are used to deduplicate security alarms, identify false alarms, and filter low-risk events.

[0009] Based on the identification results, alarms are dynamically deduplicated, priority is adjusted, and a coordinated response strategy is generated.

[0010] Preferably, based on a data weaving architecture, a unified logical data view is constructed, including:

[0011] Based on the characteristics of the data weaving architecture, data mapping and transformation technologies are used to integrate distributed, heterogeneous, and secure data sources from different geographical locations, formats, and types.

[0012] For various data sources, analyze the data structure and semantic information, and establish logical mapping relationships for each data source by defining a unified data model and data dictionary;

[0013] Metadata management technology is used to collect, store and manage metadata from data sources to ensure that the meaning and relationships of various data sources can be understood and processed when building logical data views;

[0014] At the same time, a data synchronization mechanism is established to ensure that the data in the logical data view is synchronized with the original data source in real time or near real time to reflect the latest security situation.

[0015] During the construction process, a modular and scalable design approach is adopted to allow for the addition of new data sources or the adjustment of existing data sources to adapt to ever-changing security monitoring needs;

[0016] The data weaving architecture includes:

[0017] An adaptive connection mechanism is used to support access via multiple protocols, including HTTP / RESTful API, Syslog, SNMP Trap, and MQTT.

[0018] A global knowledge graph model is used to define entity relationships, attribute constraints, and business rules across data sources.

[0019] Preferably, the AI-driven data intelligence agent component system includes:

[0020] A dual-modal interactive interface is used to provide natural language query NL2SQL conversion functionality and Python / SDK custom algorithm extension functionality;

[0021] The self-evolving rule engine uses the Q-learning algorithm of reinforcement learning to dynamically optimize the weight parameters of association rules;

[0022] A causal reasoning enhancer for integrating Bayesian networks with the MITRE ATT&CK framework to label tactical phase transition probabilities.

[0023] Preferably, the attack chain panorama is modeled based on the association between the nodes of the MITRE ATT&CK framework and the actual observed attack paths, and the behavioral causal chain model uses Bayesian network inference to analyze the temporal dependencies and causal strength between events.

[0024] Preferably, the modeling process of the behavioral causal chain model includes:

[0025] The original log is abstracted as the node transition process of a finite state automaton (FSM).

[0026] Using Hidden Markov Models (HMMs) to infer state transition paths;

[0027] The knowledge distillation and compression technique is used to transform deep learning models into a set of rule expressions.

[0028] Preferably, the semantic noise reduction algorithm includes:

[0029] Spatiotemporal cube compression technology is used to map timestamps, source IPs, and destination port numbers into a three-dimensional coordinate system and perform PCA dimensionality reduction and DBSCAN density clustering.

[0030] A context-aware deduplication strategy is used to compare payload content similarity in conjunction with the SimHash locality-sensitive hashing algorithm.

[0031] The confidence score matrix is ​​used to calculate a comprehensive score by combining the completeness of the evidence chain, the frequency of historical false alarms, and the asset importance level.

[0032] Preferably, the semantic noise reduction algorithm uses an improved locality-sensitive hashing algorithm combined with natural language processing technology. It merges event clusters by calculating the Jaccard similarity matrix of alarm texts and parses verb phrases to identify different expressions that are essentially the same.

[0033] The time series clustering uses the DBSCAN density clustering algorithm to divide event clusters in both time and space dimensions, and combines the state machine model of the OPC UA protocol to determine the reasonable fluctuation range within the normal operating cycle of the equipment.

[0034] The confidence assessment model uses a gradient boosting tree classifier, and the input features include alarm source confidence weight, historical false alarm rate statistics, importance level of related assets, and current security status index.

[0035] Preferably, the coordinated response strategy includes:

[0036] Automatically generate draft firewall blocking rules;

[0037] Includes a guide to configuring whitelist exceptions for affected service IPs;

[0038] Push to the SOAR platform for automated response.

[0039] Preferably, it also includes an adaptive strategy optimization mechanism that dynamically adjusts parameter thresholds based on daily noise reduction effect feedback, and initiates a transfer learning framework to update model parameters when a new attack mode emerges.

[0040] This invention also discloses a security alarm noise reduction system based on data weaving technology, used to apply the above-mentioned security alarm noise reduction method based on data weaving technology, comprising:

[0041] The data weaving engine is configured to build a unified logical data view based on the data weaving architecture to achieve intelligent access and semantic fusion of distributed, heterogeneous secure data sources;

[0042] The access module is configured to deploy an AI-driven data intelligence agent component system that automatically identifies and adapts to changing APIs and new data sources, and integrates metadata, threat intelligence and infrastructure operation status information.

[0043] The building module is configured to construct a panoramic view of the attack chain, a behavioral causal chain model, and a threat intelligence relationship network based on the fused data;

[0044] The noise reduction module is configured to use semantic noise reduction algorithms, time series clustering analysis, and confidence scoring mechanisms to deduplicate security alarms, identify false alarms, and filter low-risk events.

[0045] The strategy adjustment module is configured to dynamically perform alarm deduplication and priority adjustment based on the identification results, and generate linkage handling strategies.

[0046] Compared with existing technologies, the advantages of this invention are that it breaks down security data silos through a data weaving architecture, utilizes AI-driven intelligent agents to achieve zero-code access to new data sources, and effectively filters invalid alarms using multi-level noise reduction algorithms. This invention possesses self-evolving capabilities, automatically optimizing the detection model as business evolves, enabling enterprises to concentrate resources on addressing real advanced threats and significantly improving the overall efficiency of security operations centers. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0048] Figure 1 This is a flowchart illustrating a security alarm noise reduction method based on data weaving technology according to the present invention.

[0049] Figure 2 This is a functional block diagram of a security alarm noise reduction system based on data weaving technology according to the present invention. Detailed Implementation

[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0051] As Figure 1 shown, the present invention provides a security alert noise reduction method based on data weaving technology, including:

[0052] Based on the data weaving architecture, a unified logical data view is constructed to achieve intelligent access and semantic fusion of distributed and heterogeneous security data sources.

[0053] Deploy an AI-driven data intelligent agent component system to automatically identify and adapt to changing APIs and new data sources, and fuse metadata, threat intelligence, and infrastructure operation status information.

[0054] Based on the fused data, construct a panoramic view of the attack chain, a behavior causality chain model, and a threat intelligence relationship network.

[0055] Adopt semantic noise reduction algorithms, time series clustering analysis, and confidence scoring mechanisms to perform duplicate alert removal, false alarm identification, and low-risk event filtering on security alerts.

[0056] Dynamically perform alert deduplication and priority adjustment according to the recognition results, and generate a linkage disposal strategy.

[0057] The present invention effectively solves the problems of alert flooding, high false alarm rate, and low processing efficiency existing in the existing security alert system. Through data weaving technology, intelligent access and semantic fusion of distributed and heterogeneous security data sources are achieved, greatly improving the availability and accuracy of data. At the same time, the AI-driven data intelligent agent component system can automatically identify and adapt to changing APIs and new data sources, ensuring the real-time nature and integrity of data. The panoramic view of the attack chain and the behavior causality chain model constructed based on the fused data provide intuitive attack paths and behavior relationships for security analysts, helping to quickly locate and analyze security events. The introduction of semantic noise reduction algorithms, time series clustering analysis, and confidence scoring mechanisms further improves the accuracy and processing efficiency of security alerts, reducing unnecessary interference and misjudgment. Finally, the dynamic alert deduplication, priority adjustment, and the generated linkage disposal strategy based on the recognition results provide strong decision-making support for security operation and maintenance personnel, effectively enhancing the overall security protection ability.

[0058] In this embodiment, before using the data, it further includes a preprocessing stage, specifically:

[0059] Spatiotemporal alignment and normalization: To address the timestamp discrepancy issue of the same security event in different systems, a dynamic time warping (DTW) algorithm is used for sequence matching; geolocation coding is implemented for network identifiers such as IP addresses and domain names to ensure comparability analysis of cross-regional events.

[0060] Feature engineering enhancement: Extract the following high-discrimination features for subsequent modeling: traffic baseline offset (compared to the historical mean ±3σ range); protocol anomaly flag combinations (TCP RST packet ratio, ICMP timeout response frequency); user behavior entropy values ​​(keyboard input speed variation coefficient, mouse movement trajectory complexity).

[0061] In some embodiments of this application, a unified logical data view is constructed based on a data weaving architecture, including:

[0062] Based on the characteristics of the data weaving architecture, data mapping and transformation technologies are used to integrate distributed, heterogeneous, and secure data sources from different geographical locations, formats, and types.

[0063] For various data sources, analyze the data structure and semantic information, and establish logical mapping relationships for each data source by defining a unified data model and data dictionary;

[0064] Metadata management technology is used to collect, store and manage metadata from data sources to ensure that the meaning and relationships of various data sources can be understood and processed when building logical data views;

[0065] At the same time, a data synchronization mechanism is established to ensure that the data in the logical data view is synchronized with the original data source in real time or near real time to reflect the latest security situation.

[0066] During the construction process, a modular and scalable design approach is adopted to allow for the addition of new data sources or the adjustment of existing data sources to adapt to ever-changing security monitoring needs;

[0067] The data weaving architecture includes: an adaptive connection mechanism to support access via multiple protocols, including HTTP / RESTful API, Syslog, SNMP Trap, and MQTT; and a global knowledge graph model to define entity relationships, attribute constraints, and business rules across data sources.

[0068] In this embodiment, the overall architecture adopts a layered decoupled design, including:

[0069] Data acquisition supports multiple protocols including HTTP / RESTful API, Syslog, SNMP Trap, and MQTT, and is compatible with structured (SQL database), semi-structured (JSON / XML logs), and unstructured data (text reports). An adaptive connection mechanism is introduced, allowing for hot-swappable expansion to adapt to new protocols.

[0070] Modeling Engine: Constructs a global knowledge graph based on ontology, defining entity relationships (e.g., IP address ↔ vulnerability CVE number), attribute constraints (CIDR block location tags), and business rules (compliance thresholds). It utilizes RDF triple storage to achieve cross-data source concept mapping, resolving ambiguities caused by terminological differences.

[0071] The orchestration pipeline adopts a DAG (Directed Acyclic Graph) workflow engine, which allows users to visually configure data processing steps, including the entire process of cleaning → standardization → association → aggregation → modeling, and supports branch condition judgment and iterative loop logic.

[0072] Understandably, the adaptive connectivity mechanism supports multiple protocol accesses, ensuring seamless connections between different data sources and improving data diversity and coverage. The global knowledge graph model defines entity relationships, attribute constraints, and business rules across data sources, providing a solid foundation for intelligent data fusion and analysis.

[0073] In this embodiment, the core objective of data weaving is to connect data sources scattered across different systems, formats, and locations in an automated and intelligent manner, forming a unified and dynamically adjustable logical view. Unlike traditional ETL (Extract-Transform-Load) models that require copying data to a central repository, data weaving only creates metadata indexes and mappings pointing to the original data. Users access the data directly, retrieving real-time updated data from the source, reducing redundant storage costs. Business departments can combine fields across systems for analysis as needed, without pre-setting a fixed model structure. For example, the finance team can independently link sales records and supply chain logs without relying on the IT team to rebuild the cube.

[0074] Data weaving is a novel data management architecture concept that optimizes the discovery and access of heterogeneous data from multiple sources. It delivers trusted data to all relevant data consumers in a flexible and business-understandable manner, enabling self-service and efficient collaboration, and achieving highly agile data delivery. A key breakthrough of data weaving is the creation of a logical data layer using data virtualization technology. This layer integrates data scattered across different systems at a single logical point, providing data users with a unified, abstract, and encapsulated logical data view. Users can query and manipulate data stored in heterogeneous data sources through this view, treating multiple heterogeneous data sources as a single homogeneous data source without needing to worry about data location, data type, or data format. This achieves unified and centralized data access and management, similar to a data platform. Its biggest difference lies in the elimination of pre-processing data transfer, in-process ETL tasks, and post-processing storage and governance (zero data transfer, maintenance-free, self-governing).

[0075] In some embodiments of this application, the AI-driven data intelligence agent component system includes: a bimodal interaction interface for providing natural language query NL2SQL conversion function and Python / SDK custom algorithm extension function; a self-evolving rule engine that uses the Q-learning algorithm of reinforcement learning to dynamically optimize the weight parameters of association rules; and a causal inference enhancer for integrating Bayesian networks and the MITRE ATT&CK framework to label tactical phase transition probabilities.

[0076] The data intelligence agent component has dynamic discovery capabilities, enabling it to monitor network topology changes in real time and automatically match and configure the communication protocols of more than 50 mainstream security products through its built-in protocol adapter library.

[0077] Specifically, the AI-driven data intelligence agent component system includes:

[0078] A dual-modal interface provides NL2SQL conversion for natural language queries, allowing users to retrieve specific threat patterns using conversational commands; it also offers a Python / SDK for advanced users to customize algorithm plugins. A built-in sandbox environment ensures the secure execution of third-party code.

[0079] The self-evolving rule engine dynamically optimizes the weight parameters of association rules based on the Q-learning algorithm of reinforcement learning. When a policy is detected to cause an increase in the false negative rate, it automatically rolls back to the best historical version and triggers the root cause tracing process.

[0080] Causal reasoning enhancer: Integrates Bayesian networks to construct attack path inference models, combines the MITRE ATT&CK framework to label tactical phase transition probabilities, and quantifies the strength of dependencies between different alerts. For example, it determines whether "illegal hacking attempt" necessarily leads to "successful privilege escalation".

[0081] Understandably, the bimodal interface not only provides NL2SQL conversion functionality for natural language queries but also supports Python / SDK custom algorithm extensions, allowing users to flexibly customize data processing and analysis logic according to actual needs. The self-evolving rule engine uses a Q-learning algorithm based on reinforcement learning to dynamically optimize the weight parameters of association rules, continuously improving analysis strategies based on historical data and real-time feedback, thereby enhancing the accuracy and efficiency of alerts. The causal inference enhancer integrates Bayesian networks and the MITRE ATT&CK framework, labeling tactical phase transition probabilities and providing security analysts with a more in-depth and comprehensive analysis of attack paths and behavioral relationships, helping to provide early warnings and prevent potential security threats.

[0082] In some embodiments of this application, the attack chain panorama is modeled based on the association between MITRE ATT&CK framework nodes and actually observed attack paths. The behavioral causal chain model uses Bayesian network inference to analyze the temporal dependencies and causal strength between events. The threat intelligence association network is used to interface with the STIX / TAXII standard interface to achieve cross-validation of IOC metrics and internal alerts.

[0083] Specifically, the construction of the attack chain panorama is achieved by deeply associating and fusing standardized attack tactics and technical nodes defined in the MITRE ATT&CK framework with specific attack paths actually monitored and captured in the network environment. This process not only integrates the structured attack knowledge system provided by the ATT&CK framework but also combines the specific attack steps, tools, processes, and goals exhibited by attackers in real network environments. This results in a dynamic and visualized panoramic view that reflects the essential laws of attacks while closely matching actual attack and defense scenarios, clearly presenting the complete link and evolutionary process of attackers from initial intrusion to achieving their attack goals. Simultaneously, the behavioral causal chain model employs advanced Bayesian network inference methods when analyzing the intrinsic connections between security events. This method can effectively quantify and analyze the temporal dependencies of different security events and the strength of their causal effects. By constructing probabilistic dependency models between event nodes, it accurately models and probabilistically infers the correlation of multi-source events in complex attack scenarios, thereby revealing the potential driving factors and evolutionary logic of attack behavior, providing strong support for tracing the root causes of attacks and predicting attack development trends. Furthermore, the threat intelligence correlation network integrated into the system's core function is to seamlessly connect to the internationally recognized STIX (Structured Threat Information Representation) / TAXII (Trusted Automatic Exchange of Indicator Information) standard interface protocol. Through this standardized interface, the system can efficiently obtain the latest IOC (Indicator of Compromise) information from authoritative external threat intelligence sources, such as malicious IP addresses, domain names, file hash values, and attack signatures, and intelligently cross-validate and correlate these external IOC indicators with the massive amounts of alerts generated by the internal security monitoring system.

[0084] Understandably, by using an attack chain panorama based on the MITRE ATT&CK framework, this invention can accurately correlate observed attack paths, providing security analysts with a clear view of attack stages and tactics. Simultaneously, the behavioral causal chain model employs Bayesian network inference to deeply analyze the temporal dependencies and causal strength between events, helping to reveal the inherent logic and development trends of attack behaviors. This method of correlation modeling and inference analysis improves the targeting and accuracy of security alerts.

[0085] In some embodiments of this application, the modeling process of the behavioral causal chain model includes:

[0086] The original log is abstracted as the node transition process of a finite state automaton (FSM).

[0087] Using Hidden Markov Models (HMMs) to infer state transition paths;

[0088] The knowledge distillation and compression technique is used to transform deep learning models into a set of rule expressions.

[0089] Specifically, the modeling process of the behavioral causal chain model includes: creating a state machine abstraction layer: abstracting the original log into a finite state automaton (FSM) node transition process, where each state represents a system behavior pattern (normal → abnormal → recovery). The most probable state transition path is inferred using a Hidden Markov Model (HMM).

[0090] Knowledge distillation compression technology: Projects high-dimensional feature vectors extracted by deep neural networks onto a low-dimensional space to generate a set of highly interpretable rule expressions, making it easier for security experts to understand the decision-making logic behind complex models.

[0091] Interactive 3D topology display: Based on WebGL rendering, the attacker's lateral movement trajectory is overlaid with GIS data showing the physical location distribution, supporting playback of the complete kill chain process along a timeline. Users can add annotations by dragging and dropping nodes to create electronic forensic report attachments.

[0092] Understandably, through the detailed modeling process of the behavioral causal chain model, this invention achieves efficient processing and accurate analysis of raw logs. Abstracting the logs into the transition process of finite state automata nodes allows for a clear presentation of the changing trends of attack behavior. Combining this with Hidden Markov Models to infer state transition paths further enhances the predictive ability for attack behavior sequences. Simultaneously, employing knowledge distillation and compression techniques transforms the complex logic of deep learning models into a concise set of rule expressions, not only reducing computational costs but also improving the interpretability and operability of the model.

[0093] In some embodiments of this application, the semantic noise reduction algorithm includes:

[0094] Spatiotemporal cube compression technology is used to map timestamps, source IPs, and destination port numbers into a three-dimensional coordinate system and perform PCA dimensionality reduction and DBSCAN density clustering.

[0095] A context-aware deduplication strategy is used to compare payload content similarity in conjunction with the SimHash locality-sensitive hashing algorithm.

[0096] The confidence score matrix is ​​used to calculate a comprehensive score by combining the completeness of the evidence chain, the frequency of historical false alarms, and the asset importance level.

[0097] Specifically, the spatiotemporal cube compression technique first maps the three key feature dimensions of network data—timestamp, source IP address, and destination port number—to the X, Y, and Z axes of a three-dimensional coordinate system, respectively, constructing a cube-shaped data structure that characterizes the spatiotemporal distribution of network behavior. Then, to reduce data dimensionality and extract key features, Principal Component Analysis (PCA) is used to perform dimensionality reduction on this three-dimensional cube data, retaining the most representative feature vectors. Based on this, density-based noisy applied spatial clustering (DBSCAN) is further applied to perform density clustering analysis on the dimensionality-reduced data points, effectively identifying dense regions (potential normal patterns or attack clusters) and sparse outliers (possible noise or anomaly candidates) in network traffic.

[0098] Context-aware deduplication strategy aims to address the redundancy and noise caused by a large amount of duplicate or highly similar payload content in network data. It introduces the SimHash locality-sensitive hashing algorithm to hash the payload content, generating hash fingerprints that maintain content similarity to a certain extent. By comparing the Hamming distance or similarity scores between the SimHash fingerprints of different payloads, the similarity of payload content can be quickly and efficiently determined. Combining network communication context information (such as session state, protocol type, request-response relationships, etc.), similar payloads are intelligently merged or marked, thereby achieving accurate deduplication and noise reduction while retaining key and valuable payload information.

[0099] The confidence score matrix, serving as the core evaluation mechanism for algorithmic decision-making, is used to comprehensively and quantitatively evaluate potential anomalies or alarms identified after the aforementioned spatiotemporal cube compression and context-aware deduplication processing. It comprehensively considers three key factors: first, the completeness of the evidence chain, i.e., the sufficiency and logical coherence of various relevant logs, traffic characteristics, behavioral indicators, and other evidence supporting the anomaly; second, historical false alarm frequency, i.e., the probability and frequency of false alarms occurring under this type of alarm or similar scenarios based on historical data statistical analysis; and third, the asset importance level, i.e., the criticality and value weight of network assets (such as servers, databases, core business systems, etc.) affected by the anomaly in the entire information system. By constructing a multi-dimensional scoring model and weight allocation mechanism, these three factors are weighted and calculated to ultimately generate a comprehensive confidence score, used to determine the true threat level of the anomaly and assist in deciding whether to trigger an alarm and the alarm level.

[0100] Understandably, through the meticulous design of semantic denoising algorithms, this invention significantly improves the purity and analysis efficiency of log data. Spatiotemporal cube compression technology effectively reduces data dimensionality while retaining key information, making subsequent processing more efficient. Combined with a context-aware deduplication strategy, redundant data is further reduced, ensuring the accuracy of the analysis. The introduction of a confidence score matrix assigns a more scientific weight to each log data entry, making security analysis more accurate and reliable.

[0101] In some embodiments of this application, the semantic noise reduction algorithm may also employ improved locality-sensitive hashing combined with natural language processing techniques. It merges event clusters by calculating the Jaccard similarity matrix of alarm texts and parses verb phrases to identify different expressions that are essentially the same. The time series clustering uses the DBSCAN density clustering algorithm to divide event clusters in both time and space dimensions, and combines the state machine model of the OPC UA protocol to determine the reasonable fluctuation range within the normal operating cycle of the device. The confidence assessment model uses a gradient boosting tree classifier, with input features including alarm source confidence weights, historical false alarm rate statistics, importance level of associated assets, and current security status index.

[0102] Specifically, semantic denoising algorithms can also employ improved Locality Sensitive Hash (LSH) technology combined with advanced Natural Language Processing (NLP) techniques. First, an improved LSH function is used to efficiently reduce the dimensionality and extract features from the alert text, mapping high-dimensional text data to a low-dimensional hash space. This allows for rapid calculation of the Jaccard similarity between different alert texts, constructing a Jaccard similarity matrix for the alert texts. Based on this similarity matrix, the system can automatically aggregate semantically similar alert events into event clusters, achieving preliminary event merging. Furthermore, the algorithm deeply integrates NLP technology, performing syntactic and semantic analysis on the alert text. It focuses on identifying and extracting verb phrases and core semantic components to accurately identify alert content that differs in expression but shares the same essential meaning. This effectively eliminates alert redundancy caused by differences in expression, significantly improving the accuracy and consistency of alert information.

[0103] In this application's embodiments, time series clustering specifically utilizes the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) density clustering algorithm. This algorithm considers not only the temporal dimension of alarm events but also integrates the spatial location information of the devices, performing comprehensive clustering analysis on alarm events in both time and space dimensions to accurately identify event clusters with close spatiotemporal correlations. Simultaneously, the system further integrates the device state machine model defined in the OPC UA (OLE for Process Control Unified Architecture) protocol to gain a deeper understanding of the various state transition logics and behavioral patterns of the device during normal operation. By comparing and analyzing the clustered event clusters with the state machine model within the device's normal operating cycle, the reasonable fluctuation range of various parameters within the normal operating cycle can be accurately determined. Alarms falling within this range are thus identified as normal fluctuations, avoiding misjudgment as abnormal events and effectively reducing the false alarm rate.

[0104] The confidence assessment model in this application employs a Gradient Boosting Decision Tree (GBDT) classifier. The input features of this model are carefully selected and constructed, including but not limited to: alarm source confidence weights, which are dynamically adjusted based on the historical reliability performance of the alarm source (such as sensors, monitoring systems, etc.); historical false alarm rate statistics, i.e., a quantitative analysis of the frequency of false alarms of this type of alarm or similar devices over a past period; the importance level of associated assets, assigned according to the criticality of the physical devices, data assets, or business systems involved in the alarm within the overall architecture; and the current security posture index, which comprehensively considers multiple factors such as the current network environment, system load, and external threat intelligence, reflecting the overall security risk level. By inputting these multi-dimensional features into the Gradient Boosting Decision Tree classifier for training and prediction, the model can accurately assess the confidence of each alarm event (i.e., the probability that the alarm is a genuine and valid abnormal event), providing a scientific basis for subsequent alarm response prioritization and automated processing. Events with output probability scores below a set threshold are marked as potential noise.

[0105] Understandably, by employing an improved Locality Sensitive Hashing (LSH) algorithm combined with Natural Language Processing (NLP) technology, this application further enhances the semantic understanding capabilities of log data. Calculating the Jaccard similarity matrix of alarm texts effectively merges event clusters, solving the problem of event identification caused by differences in wording. Simultaneously, parsing verb phrases accurately identifies alarm messages that are essentially the same but have different wording, further improving the accuracy of log data. In terms of time-series clustering, the DBSCAN density clustering algorithm in both time and space dimensions is used to divide event clusters. Combined with the state machine model of the OPC UA protocol, it can accurately determine the reasonable fluctuation range of the device within the normal operating cycle, thereby effectively distinguishing between normal operation and abnormal events. Furthermore, the confidence assessment model uses a gradient boosting tree classifier, comprehensively considering multiple dimensions such as alarm source confidence weight, historical false alarm rate statistics, importance level of associated assets, and current security posture index, providing a more scientific confidence assessment for each alarm message, making security response faster and more accurate.

[0106] In some embodiments of this application, the coordinated response strategy includes: automatically generating a draft firewall blocking rule; attaching a whitelist exception configuration guide for the affected service IPs; and pushing it to the SOAR platform for automated response.

[0107] Specifically, the coordinated response strategy includes the following three closely linked implementation steps: First, the system automatically generates a precise draft firewall blocking rule based on security analysis results. This draft must clearly define key elements such as attack characteristics, source and destination IPs / ports, and blocking actions. Second, it simultaneously generates and attaches a whitelist exception configuration guide for affected legitimate service IP addresses. The guide should detail the whitelist IP range, corresponding service ports, and configuration priorities to ensure that business continuity is not affected by the blocking rules. Finally, the aforementioned draft blocking rule and whitelist configuration guide are automatically pushed to the Security Orchestration Automation and Response (SOAR) platform. The platform then completes rule verification, approval process, and final firewall policy distribution and activation according to a preset workflow, achieving end-to-end automated response and handling of security threats.

[0108] For example, consider a scenario where a financial institution suffers a DDoS hybrid application layer attack: the WAF device reports a large number of HTTP 404 error code requests; the IDS detects an abnormal increase in TCP connection reset RST packets; and SIEM receives threat intelligence entries from dark web forums mentioning the institution's domain.

[0109] Traditional methods generate hundreds of redundant alarms, while this solution will perform the following operations:

[0110] The first step is normalization: converting the geographical locations of all relevant events into city codes in the GeoIP database;

[0111] The second step is to complete the association: supplement the missing User-Agent string information, and it was found that they all carry the same UA header characteristics;

[0112] The third step, timing alignment, revealed that the request rate exhibited a pulse-like burst pattern after resampling at the second-level granularity.

[0113] Step 4: Root cause localization: The decision tree model determines that the root cause lies in a CDN node being hijacked and acting as a reflection amplifier.

[0114] The fifth step in the response strategy is to automatically generate a draft firewall blocking rule, along with a guide to configuring whitelist exceptions for the affected service IPs.

[0115] Understandably, by automatically generating draft firewall blocking rules, this application can quickly respond to security threats and effectively prevent potential attacks. The accompanying guide to configuring whitelist exceptions for affected service IPs ensures the normal operation of critical services and reduces unnecessary interference caused by the implementation of security policies. Pushing coordinated response strategies to the SOAR platform for automated response further improves the efficiency and accuracy of security responses, enabling enterprises to quickly respond to various security incidents and ensure business continuity and data security.

[0116] In some embodiments of this application, an adaptive strategy optimization mechanism is also included, which dynamically adjusts parameter thresholds based on daily noise reduction effect feedback, and initiates a transfer learning framework to update model parameters when a new attack mode emerges.

[0117] It also includes an adaptive strategy optimization mechanism, which can collect and analyze various effect feedback data generated during the daily noise reduction process of the system in real time, such as false alarm rate, false negative rate, noise reduction efficiency, user correction records and other multi-dimensional indicators. Based on a preset evaluation model, it performs quantitative evaluation and in-depth mining of these feedback data, and then dynamically adjusts the thresholds of various key parameters related to the noise reduction algorithm in the system according to the evaluation results, such as feature extraction threshold, abnormal behavior judgment threshold, similarity matching threshold, etc., to ensure that the noise reduction system can continuously adapt to the ever-changing actual application scenarios and data characteristics, and always maintain efficient and accurate noise reduction performance. Meanwhile, this adaptive strategy optimization mechanism also possesses the ability to intelligently perceive and rapidly respond to new attack patterns. When a new attack pattern with unknown characteristics is identified in the system through the pattern recognition module or anomaly detection algorithm, a preset transfer learning framework will be automatically triggered and started. This framework can utilize existing historical attack sample data and relevant domain knowledge as prior knowledge, combined with newly collected new attack pattern sample data, to quickly update the core parameters and decision boundaries of the denoising model through parameter fine-tuning, feature transfer, or model structure adaptation. This enables the denoising model to quickly acquire the ability to identify and defend against new attack patterns without large-scale retraining, effectively improving the system's adaptive defense level and long-term robustness in complex and ever-changing network environments.

[0118] Understandably, by introducing an adaptive strategy optimization mechanism, this application can intelligently adjust parameter thresholds based on actual noise reduction feedback, thereby ensuring the continuous optimization and adaptability of the security strategy. When faced with new attack patterns, the system can automatically activate the transfer learning framework to quickly update model parameters to cope with ever-changing security threats.

[0119] This invention discloses a security alarm noise reduction system based on data weaving technology, used to apply the aforementioned security alarm noise reduction method based on data weaving technology, comprising:

[0120] The data weaving engine is configured to build a unified logical data view based on the data weaving architecture to achieve intelligent access and semantic fusion of distributed, heterogeneous secure data sources;

[0121] The access module is configured to deploy an AI-driven data intelligence agent component system that automatically identifies and adapts to changing APIs and new data sources, and integrates metadata, threat intelligence and infrastructure operation status information.

[0122] The building module is configured to construct a panoramic view of the attack chain, a behavioral causal chain model, and a threat intelligence relationship network based on the fused data;

[0123] The noise reduction module is configured to use semantic noise reduction algorithms, time series clustering analysis, and confidence scoring mechanisms to deduplicate security alarms, identify false alarms, and filter low-risk events.

[0124] The strategy adjustment module is configured to dynamically perform alarm deduplication and priority adjustment based on the identification results, and generate linkage handling strategies.

[0125] This invention utilizes a data weaving engine to integrate security data from diverse sources and formats, forming a unified logical data view that enhances data availability and analysis efficiency. The access module leverages AI technology to intelligently identify and adapt to constantly changing APIs and new data sources, ensuring data real-time performance and accuracy. The construction module, based on the fused data, comprehensively builds a panoramic view of the attack chain and a behavioral causal chain model, providing in-depth insights for security analysis. The noise reduction module, through advanced algorithms and mechanisms, effectively reduces false alarms and low-risk events, improving the accuracy and usability of security alerts. The strategy adjustment module dynamically adjusts alert processing and priorities based on the identification results, providing security operations personnel with timely and coordinated response strategies, enhancing the overall response speed and effectiveness of security protection. This series of innovative designs enables the security alert noise reduction system of this invention to demonstrate superior performance and adaptability in practical applications.

[0126] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A security alarm noise reduction method based on data weaving technology, characterized in that, Comprise: Based on the data weaving architecture, a unified logical data view is constructed to realize intelligent access and semantic fusion of distributed and heterogeneous secure data sources; Deploy AI-driven data intelligence agent component system, automatically identify and adapt to changing APIs and new data sources, integrate metadata, threat intelligence and infrastructure operation status information; Based on the fused data, build attack chain panoramic map, behavior causal chain model and threat intelligence relationship network; Use semantic noise reduction algorithm, time series clustering analysis and confidence score mechanism to remove duplicate alarms, identify false positives and filter low-risk events for security alerts; The semantic noise reduction algorithm includes: spatiotemporal cube compression technology, which is used to map timestamp, source IP and destination port number into a three-dimensional coordinate system and perform PCA dimensionality reduction and DBSCAN density clustering; Context-aware deduplication strategy, combined with SimHash local sensitive hash algorithm to compare Payload content similarity; Confidence score matrix, used to calculate the comprehensive score by integrating evidence chain integrity, historical false positive frequency and asset importance level; The semantic noise reduction algorithm uses an improved local sensitive hash combined with natural language processing technology to realize event cluster merging by calculating the Jaccard similarity matrix of the alert text, and to analyze the verb phrase to identify different expression methods of the same nature; The time series clustering uses the DBSCAN density clustering algorithm to divide the event cluster in the time-space two-dimensional space, and combines the state machine model of OPC UA protocol to determine the reasonable fluctuation range within the normal operation period of the device; The confidence evaluation model uses a gradient boosting tree classifier, and the input features include alert source credibility weight, historical false positive rate statistics, associated asset importance level and current security situation index; According to the identification result, dynamically perform alarm deduplication, priority adjustment, and generate linkage disposal strategy; Based on the data weaving architecture, a unified logical data view is constructed, including: According to the characteristics of the data weaving architecture, use data mapping and conversion technology to integrate distributed and heterogeneous secure data sources from different geographical locations, different formats and different types; For various data sources, analyze data structure and semantic information, define a unified data model and data dictionary, and establish a logical mapping relationship for each data source; Use metadata management technology to collect, store and manage the metadata of the data source to ensure that the meaning and association of various data sources can be understood and processed when constructing the logical data view; At the same time, establish a data synchronization mechanism to ensure that the data in the logical data view is synchronized with the original data source in real time or near real time to reflect the latest security situation; In the construction process, use modular and extensible design to add new data sources or adjust existing data sources to adapt to changing security monitoring needs; The data weaving architecture comprises: Adaptive connection mechanism for supporting multiple protocol access, including HTTP / RESTful API, Syslog, SNMP Trap and MQTT; A global knowledge graph model is used to define entity relationships, attribute constraints, and business rules across data sources.

2. The data weaving technology based security alert noise reduction method of claim 1, wherein, The AI-driven data intelligence agent component system includes: A dual-mode interactive interface is used to provide natural language query NL2SQL conversion functions and Python / SDK custom algorithm extension functions. A self-evolution rule engine uses a Q-learning algorithm of reinforcement learning to dynamically optimize correlation rule weight parameters. A causal reasoning enhancer is used to integrate a Bayesian network and a MITRE ATT&CK framework to label tactical phase transition probabilities.

3. The data weaving technology based security alert noise reduction method of claim 1, wherein, The attack chain panorama is based on the association modeling of MITRE ATT&CK framework nodes and actually observed attack paths, and the behavior causal chain model uses a Bayesian network to infer the temporal dependence and causal strength between events.

4. The data weaving technology based security alert noise reduction method of claim 1, wherein, The modeling process of the behavior causal chain model includes: The original log is abstracted into a finite state automaton FSM node transition process. A hidden Markov model HMM is used to infer state transition paths. A knowledge distillation compression technique is used to convert a deep learning model into a set of rule expressions.

5. The data weaving technology based security alert noise reduction method of claim 1, wherein, The linkage disposal strategy includes: Automatically generating a draft of a firewall blocking rule; Including an affected service IP whitelist exception configuration guide; Pushing to a SOAR platform for automated response.

6. The data weaving technology based security alert noise reduction method of claim 1, wherein, An adaptive strategy optimization mechanism is also included, which dynamically adjusts parameter thresholds based on daily noise reduction effect feedback and starts a transfer learning framework to update model parameters when a new attack mode appears.

7. A security alarm noise reduction system based on data weaving technology, for applying the security alarm noise reduction method based on data weaving technology as claimed in any one of claims 1-6, characterized in that, It includes: A data weaving engine is configured to build a unified logical data view based on a data weaving architecture to realize intelligent access and semantic fusion of distributed, heterogeneous secure data sources; An access module is configured to deploy an AI-driven data intelligence agent component system, automatically identify and adapt to changing APIs and new data sources, and fuse metadata, threat intelligence, and infrastructure operation state information; A construction module is configured to build an attack chain panorama, a behavior causal chain model, and a threat intelligence relationship network based on the fused data; A noise reduction module is configured to use a semantic noise reduction algorithm, time series clustering analysis, and confidence scoring mechanism to remove duplicate alerts, identify false positives, and filter low-risk events for security alerts; A strategy adjustment module is configured to dynamically perform alert deduplication, priority adjustment, and linkage disposal strategy generation based on the identification results.

Citation Information

Patent Citations

  • Alarm information intelligent noise reduction and alarm convergence system

    CN113722178A

  • Wind power plant booster station multi-source data fusion anti-misoperation locking intelligent decision and early warning method

    CN120450241A