A traffic collection system, threat analysis method, and policy generation method

By introducing large-scale intelligent models and dynamic policy engines into the traffic acquisition system, and combining semantic analysis and temporal analysis, the problems of policy rigidity and real-time performance in traditional traffic acquisition technology when facing new threats are solved, achieving accurate identification and efficient protection.

CN120785652BActive Publication Date: 2025-12-05HANGZHOU DPTECH TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511278983.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-12-05
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Traditional traffic acquisition technologies are ill-equipped to deal with emerging security risks such as advanced persistent threats (APTs) and zero-day vulnerability attacks. They suffer from rigid strategies, insufficient semantic understanding, and real-time deficiencies, making it impossible to achieve accurate capture and real-time response in dynamic and heterogeneous network environments.

Method used

A large-scale intelligent model is introduced to perform semantic and temporal analysis, combined with a threat knowledge base for comprehensive threat assessment, and a dynamic policy engine is used to generate differentiated traffic collection strategies to adapt to the actual threat situation of edge devices.

Benefits of technology

It enhances the ability to perform deep semantic analysis of traffic, enabling accurate identification and real-time response to traffic threats, reducing system resource consumption, and improving the efficiency of abnormal traffic detection and security protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120785652B_ABST
    Figure CN120785652B_ABST
Patent Text Reader

Abstract

The application provides a traffic collection system, a threat analysis method and a strategy generation method. The system comprises: a probe configured to collect original feature data of network traffic; a threat knowledge base configured to store historical threat information; a large model agent connected with the probe and the threat knowledge base, respectively, and configured to perform semantic analysis and time series analysis on the original feature data, extract multi-modal fusion features of the original feature data, obtain historical threat information associated with the multi-modal fusion features from the threat knowledge base based on the multi-modal fusion features, and generate a threat analysis result of the network traffic based on the multi-modal fusion features and the historical threat information; a dynamic strategy engine connected with the large model agent and the probe, respectively, and configured to determine a threat level of the threat analysis result, generate a traffic collection strategy corresponding to the threat level, and issue the traffic collection strategy to the probe; and the probe is further configured to collect traffic according to the traffic collection strategy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security, and in particular to a traffic collection system, a threat analysis method and a strategy generation method. BACKGROUND

[0002] With the increasing complexity of network attack means, the traditional traffic collection technology based on rules and static strategies has been difficult to cope with new security risks such as advanced persistent threats (APTs), 0day vulnerability attacks, etc. The current mainstream traffic collection scheme usually relies on a pre-defined rule set or simple statistical analysis, such as through a firewall integrated collection module or a cloud platform unified strategy issuance. Although these methods can handle known attack patterns, there are obvious detection blind spots when facing variant attacks, encrypted traffic disguises and low-rate attacks. In addition, in modern distributed network environments, traffic has the characteristics of dynamics, heterogeneity and scaling, and it is urgent to introduce intelligent traffic analysis and strategy generation mechanisms to achieve accurate capture and real-time response to abnormal traffic.

[0003] The existing technology mainly faces three problems: first, the strategy is rigid, static rules are difficult to adapt to the rapid evolution of attack means, resulting in a high false negative rate; second, semantic understanding is insufficient, traditional methods cannot deeply analyze the hidden semantic features (such as malicious code fragments in HTTP requests or DNS tunnel encoding) in traffic, making it difficult to identify potential risks in a timely manner; third, real-time defects, in the edge heterogeneous network environment, the threat of traffic of different edge devices is significantly different, and the use of a unified collection strategy can easily cause resource mismatch, and the redundant collection of low-threat devices consumes a large amount of computing resources, while the response delay of high-threat devices increases, which cannot effectively intercept transient attacks, seriously affecting the real-time protection effect. These problems are further exacerbated in the context of massive heterogeneous edge traffic, making it difficult for traditional solutions to meet the dynamic security protection needs. SUMMARY

[0004] To overcome the problems in the related art, the present application provides a traffic collection system, a threat analysis method and a strategy generation method.

[0005] According to a first aspect of an embodiment of the present application, a traffic collection system is provided, applied to an edge cloud collaborative architecture, the edge cloud collaborative architecture comprising an edge gateway, a cloud platform and an intelligent agent platform; the system comprises:

[0006] a probe deployed at the edge gateway, configured to collect raw feature data of network traffic;

[0007] a threat knowledge base deployed at the intelligent agent platform, configured to store historical threat information;

[0008] A large model agent deployed on the intelligent agent platform, in communication connection with the probe and the threat knowledge base respectively, configured to receive the original feature data, perform semantic analysis and time series analysis on the original feature data, extract multi-modal fusion features of the original feature data, and acquire historical threat information associated with the multi-modal fusion features from the threat knowledge base according to the multi-modal fusion features, and generate a threat analysis result of the network traffic based on the multi-modal fusion features and the historical threat information;

[0009] A dynamic policy engine deployed on the cloud platform, in communication connection with the large model agent and the probe respectively, configured to receive the threat analysis result, determine a threat level of the threat analysis result, generate a traffic collection policy corresponding to the threat level, and issue the traffic collection policy to the probe;

[0010] The probe is further configured to perform traffic collection according to the traffic collection policy.

[0011] According to a second aspect of the embodiments of the present application, a traffic threat analysis method is provided, applied to the large model agent of the first aspect, and the method comprises:

[0012] Receiving original feature data of network traffic;

[0013] Performing semantic analysis and time series analysis on the original feature data, and extracting multi-modal fusion features of the original feature data;

[0014] Acquiring historical threat information associated with the multi-modal fusion features from a preset threat knowledge base according to the multi-modal fusion features;

[0015] Generating a threat analysis result of the network traffic based on the multi-modal fusion features and the historical threat information.

[0016] According to a third aspect of the embodiments of the present application, a traffic collection policy generation method is provided, applied to the dynamic policy engine of the first aspect, and the method comprises:

[0017] Receiving a threat analysis result generated by the large model agent of the first aspect;

[0018] Determining a threat level of the threat analysis result, and generating a traffic collection policy corresponding to the threat level.

[0019] According to a fourth aspect of the embodiments of the present application, a large model agent is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of the second aspect when executing the computer program.

[0020] According to a fifth aspect of the embodiments of the present application, a dynamic strategy engine, a memory, a processor, and a computer program stored in the memory and executable on the processor are provided, wherein the processor implements the method of the third aspect when executing the computer program.

[0021] According to a sixth aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executable on a processor to implement the method of the second aspect or the third aspect.

[0022] The technical solutions provided by the embodiments of the present application can include the following beneficial effects:

[0023] In the embodiments of the present application, by introducing a large model agent into the traffic collection system, semantic analysis and time series analysis are carried out on real-time traffic features, and comprehensive threat analysis is carried out in combination with historical threat information in the threat knowledge base, thereby effectively improving the analysis capability of the deep semantic of the traffic, and realizing accurate identification of traffic threats. At the same time, according to the threat level of the threat analysis result, the dynamic strategy engine generates differentiated traffic collection strategies, so that the traffic collection mode can be dynamically adjusted according to the actual threat situation of the edge device, thereby effectively reducing the system resource consumption on the premise of ensuring the accuracy of threat detection, and comprehensively improving the detection efficiency of abnormal traffic and the real-time security protection effect of the system.

[0024] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which are incorporated into the specification and constitute a part of the present application, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.

[0026] Figure 1 is a structural schematic diagram of a traffic collection system according to an exemplary embodiment of the present application.

[0027] Figure 2 is an interaction logic schematic diagram of a large model agent and a threat knowledge base according to an exemplary embodiment of the present application.

[0028] Figure 3 is a flow schematic diagram of a traffic threat analysis method according to an exemplary embodiment of the present application.

[0029] Figure 4 is a flow schematic diagram of a traffic collection strategy generation method according to an exemplary embodiment of the present application.

[0030] Figure 5is a flowchart of another traffic collection strategy generation method according to an example embodiment of the present application.

[0031] Figure 6 is a structural block diagram of a large model agent 103 according to an example embodiment of the present application.

[0032] Figure 7 is a structural block diagram of a dynamic strategy engine 104 according to an example embodiment of the present application. DETAILED DESCRIPTION

[0033] The example embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements, unless the context clearly dictates otherwise. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements, unless the context clearly dictates otherwise. The following example embodiments described in the following detailed description are not meant to be all inclusive or to be the only embodiments that are consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with some aspects of the present application as detailed in the appended claims.

[0034] The terminology used in the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the present application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0035] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, a first information can be termed a second information, and similarly, a second information can also be termed a first information, without departing from the scope of the present application. As used herein, the word "if' can be construed to mean "when" or "upon" or "in response to determining" depending on the context.

[0036] With the increasing complexity of network attack means, the traditional traffic collection technology based on rules and static policies has been difficult to cope with new security risks such as advanced persistent threats (APTs), 0day vulnerability attacks, etc. The current mainstream traffic collection scheme usually relies on a pre-defined rule set or simple statistical analysis, such as through firewall integration collection module or cloud platform unified policy issuance. Although these methods can handle known attack patterns, there are obvious detection blind spots when facing variant attacks, encrypted traffic camouflage and low-rate attacks. In addition, in modern distributed network environment, traffic has the characteristics of dynamicity, heterogeneity and scaling, and intelligent traffic analysis and policy generation mechanism needs to be introduced to realize accurate capture and real-time response to abnormal traffic.

[0037] The existing technology mainly faces three problems: first, the policy is rigid, and static rules are difficult to adapt to the rapid evolution of attack means, resulting in a high false negative rate; second, the semantic understanding is insufficient, and traditional methods cannot deeply analyze the hidden semantic features (such as malicious code fragments in HTTP requests or DNS tunnel encoding) in the traffic, making it difficult to identify potential risks in time; third, the real-time performance is defective, in the edge heterogeneous network environment, the traffic threat of different edge devices is significantly different, and the use of unified collection strategy can easily cause resource mismatch, and the redundant collection of low-threat devices consumes a large amount of computing resources, while the response delay of high-threat devices increases, which cannot effectively intercept transient attacks, seriously affecting the real-time protection effect. These problems are further exacerbated in the scenario of massive heterogeneous edge traffic, making it difficult for traditional schemes to meet the dynamic security protection needs.

[0038] Based on this, in order to solve the problems in the related art, the embodiments of the present application provide a traffic collection system. The system introduces a large model agent to conduct semantic analysis and time series analysis on real-time traffic features, and combines historical threat information in the threat knowledge base for comprehensive threat analysis, effectively improving the analysis ability of the deep semantics of the traffic, and realizing accurate identification of traffic threats; at the same time, the system generates differentiated traffic collection strategies through a dynamic strategy engine according to the threat level of the threat analysis result, so that it can dynamically adjust the traffic collection mode according to the actual threat situation of the edge device, thereby effectively reducing system resource consumption on the premise of ensuring threat detection accuracy, and comprehensively improving the detection efficiency of abnormal traffic and the real-time security protection effect of the system.

[0039] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0040] Figure 1 is a structural schematic diagram of a traffic collection system according to an exemplary embodiment of the present application. As shown in Figure 1As shown, the system is applied to an edge cloud collaborative architecture, the edge cloud collaborative architecture including an edge gateway, a cloud platform and an agent platform; the system includes: a probe 101 deployed at the edge gateway, a threat knowledge base 102 deployed at the agent platform, a large model agent 103 deployed at the agent platform, and a dynamic strategy engine 104 deployed at the cloud platform; the large model agent 103 is in communication connection with the probe 101 and the threat knowledge base 102 respectively, and the dynamic strategy engine 104 is in communication connection with the large model agent 103 and the probe 101 respectively.

[0041] The edge cloud collaborative architecture in the embodiments of the present application is different from the traditional edge cloud architecture, and adopts a distributed design concept of edge gateway-cloud platform-agent platform. In actual application process, the agent platform can be directly embedded on the basis of the traditional edge gateway-cloud platform traffic collection framework to realize upgrading, without the need of large-scale modification of the hardware architecture of the existing edge gateway and cloud platform, greatly reducing the difficulty and cost of system upgrading, having the significant advantages of convenient upgrading and strong dynamic performance, and being able to well adapt to the intelligent transformation needs in different network environments.

[0042] In the traffic collection system of the embodiments of the present application, the probe 101 is mainly used for collecting original feature data of network traffic. Specifically, the probe 101 can be a lightweight micro-service module deployed at the edge gateway, having the characteristics of low resource occupation and high real-time performance, and being able to be deeply embedded in the traffic forwarding link of the gateway, realizing real-time mirror collection of network traffic passing through the gateway without significantly affecting the original data transmission performance of the gateway, realizing low-delay data collection, and being able to quickly respond to real-time traffic changes at the edge gateway, providing timely data support for subsequent threat analysis and strategy generation.

[0043] The raw feature data collected by the probe 101 can cover multi-dimensional information of network traffic, such as basic features, semantic features, and time sequence features, etc. The basic features can be the core identification of network communication, such as source IP address, destination IP address, transmission protocol type (such as TCP / UDP), port number, etc. These information is the basis for locating network connection subjects and distinguishing traffic types, and provides key clues for subsequent threat tracing. The semantic features focus on text information in the traffic, such as parameter sequences in HTTP request body, function name and parameter combination in API call path, character encoding fragments in URL, device and browser identification in User-Agent field, etc. Such features contain the business intent and potential threat signals of the traffic, such as malicious code fragments, abnormal instruction injection, etc. They are the core basis for deep analysis of the nature of the traffic. The time sequence features reflect the time dynamic characteristics of the traffic, which can include request frequency per unit time, single session duration, packet interval time distribution, connection establishment and disconnection time sequence law, etc. Such features can effectively capture low-rate attacks, periodic scanning and other time-related abnormal behaviors, making up for the shortcomings of static features in time dimension analysis. By collecting these multi-dimensional raw feature data, the probe 101 can comprehensively depict the properties and behavior patterns of network traffic, providing sufficient information basis for the accurate threat analysis of the large model agent, thereby effectively avoiding the threat detection problem caused by feature missing, and improving the coverage ability of the system to complex attack scenarios.

[0044] The threat knowledge base 102 is used to store historical threat information. The historical threat information refers to various known information related to network threats, which can provide rich known threat context for the large model agent 103, and is the core knowledge support for subsequent threat analysis, which can assist the system to accurately identify potential risks in network traffic.

[0045] The historical threat information stored in the threat knowledge base 102 has a wide range of sources and multiple dimensions, which can specifically cover three major categories of core content: first, attack patterns, such as the tactics, techniques and sub-techniques used by attackers in each stage of network attack defined based on the MITRE ATT&CK framework, providing standardized reference for quickly identifying similar attack patterns; second, historical attack cases, such as public CVE vulnerability exploitation process records, complete links of attack chains of known APT organizations, detailed feature and disposal records of various attack events, etc. Practical experience data; third, threat indicators (IoC), such as malicious IP addresses, suspicious domain names, hash values of malicious files, feature strings, etc. These information can be directly used for matching identification information. These information can come from public threat intelligence library, user historical attack defense records and third-party security vendor shared intelligence, etc., to ensure the comprehensiveness and practicality of knowledge coverage.

[0046] To achieve efficient management and correlation retrieval of massive historical threat information, the threat knowledge base 102 can use graph database technology to structure the storage of various types of historical threat information. For example, a threat knowledge graph is constructed in a "node-edge" graph structure: core elements such as attacker entities, attack methods, and vulnerability assets are taken as nodes, and logical relationships such as "exploitation relationships", "correlation characteristics", and "attack chain links" are taken as edges, forming a multi-dimensional correlated knowledge network. This storage method can intuitively display the complex internal relationships between threat information, support efficient subgraph query and knowledge reasoning, and help the large model agent 103 quickly trace associated historical threat information along the graph structure during threat analysis of the traffic, significantly improving the depth of context understanding and enhancing the accuracy of threat identification of the traffic.

[0047] In addition, the threat knowledge base 102 can also be equipped with a dynamic incremental update mechanism, which can capture the latest information from data sources in real time or at regular intervals, automatically identify new threat types, attack method variants, and new IoC indicators, and through automatic expansion and update algorithms of the knowledge graph, integrate these new data into the existing knowledge system, realizing the continuous expansion and iteration of knowledge reserves. This mechanism ensures that the knowledge base can dynamically adapt to the evolution rhythm of network threats, providing fresh historical references for real-time traffic analysis, fundamentally solving the problem of lagging response of traditional static knowledge bases to new threats, and continuously empowering the accuracy and depth of threat identification of the system.

[0048] The large model agent 103 can first be used to receive the original feature data transmitted by the probe 101, and perform semantic analysis and time series analysis on the original feature data to extract multi-modal fusion features of the original feature data.

[0049] Traditional traffic analysis techniques rely on static rules or simple statistical features, and can only achieve "feature matching" level detection, making it difficult to deal with hidden threats such as encrypted traffic disguising and malicious code variants; while the combination of semantic analysis and time series analysis in the embodiments of the present application enables the large model agent 103 to deeply analyze the internal intent and dynamic law of the traffic, realizing a qualitative change from "feature detection" to "intent understanding", and significantly improving the ability to identify new attack threats.

[0050] Specifically, in some embodiments, the large model agent 103 can include a BERT model and a Transformer network; then the large model agent 103 can use the BERT model to perform semantic analysis on the original feature data to obtain semantic features of the original feature data; and use the Transformer network to perform time series analysis on the original feature data to obtain time series features of the original feature data; finally, the semantic features and the time series features are fused to obtain multi-modal fusion features of the original feature data.

[0051] The BERT model has strong context-aware semantic understanding capability and can perform semantic deep analysis on text features such as parameter sequences in the HTTP request body, character combinations in the URL, and the User-Agent field. For example, it can convert seemingly random strings (such as URL parameters containing injection statements and malicious code fragments disguised as normal instructions) into dense vectors (Embedding) containing rich semantic information, thereby accurately capturing hidden semantic signals in traffic and identifying potential threat intentions. The Transformer network focuses on time dynamic modeling of time sequence features such as NetFlow data, request frequency, and session duration, and can effectively identify periodic patterns of low-rate attacks, time sequence mutations of abnormal connections (such as high-frequency port scanning in a short period of time), and other time-dimension abnormal patterns, making up for the shortcomings of traditional analysis methods in capturing long-period dependencies. In actual work, the BERT model and the Transformer network can work in parallel, allowing the BERT model to focus on semantic deep understanding of text features and the Transformer network to focus on dynamic modeling of time sequence features. This clear division of labor and strong targeting not only avoids the lack of adaptation of a single model to multi-modal data, but also improves analysis efficiency through parallel computing, ensuring real-time analysis in a massive heterogeneous traffic scenario. This multi-modal fusion feature extraction method lays a solid foundation for subsequent precise threat analysis combined with a threat knowledge base.

[0052] After extracting the multi-modal fusion features of the original feature data, the large model agent 103 can also be used to obtain historical threat information associated with the multi-modal fusion features from the threat knowledge base 102 based on the multi-modal fusion features, and generate a threat analysis result of the network traffic based on the multi-modal fusion features and the historical threat information.

[0053] This process achieves a "real-time analysis + historical experience" collaborative decision-making mode by associating real-time traffic features with historical threat knowledge. By introducing historical threat information as a context reference, the large model agent 103 can break through the limitations of single real-time data, deeply analyze the hidden attack intent in the traffic, for example, when the semantic features of a certain HTTP request body show abnormal parameter combinations, combined with the record of "similar parameters have been used for certain CVE exploit" in the historical threat library, the potential threat can be quickly located. At the same time, this association mechanism can greatly reduce the real-time computing load of the large model agent 103, so that it does not need to start from scratch for complete inference on each piece of traffic, but focuses directly on key threat points based on historical associated threat information, improving response speed while ensuring analysis accuracy. In addition, the introduction of historical threat information can also make the threat analysis result traceable, and security and operation personnel can quickly understand the basis for threat determination by viewing the associated historical threat information, significantly improving the credibility and explainability of subsequent strategy generation.

[0054] Specifically, in some embodiments, the large model agent 103 can first generate a retrieval identifier based on the multi-modal fusion features. Since the multi-modal fusion features are usually high-dimensional vectors, direct retrieval will result in low efficiency, so they can be converted into low-dimensional retrieval identifiers that can uniquely identify the corresponding features, thereby achieving dimension compression while preserving the uniqueness of the features, ensuring the accuracy of the retrieval, and significantly reducing the data transmission volume, laying the foundation for efficient retrieval. The form of the retrieval identifier can include the hash value of the multi-modal fusion features, the combination of key dimensions of the feature vector, or the set of semantic labels, etc.

[0055] After generating the retrieval identifier, the large model agent 103 can connect the interactive interface of the threat knowledge base 102 through tools such as the Langchain framework, and send the retrieval identifier to the threat knowledge base 102. After receiving the retrieval identifier, the threat knowledge base 102 can perform associated retrieval based on the retrieval identifier, determine historical threat information associated with the multi-modal fusion feature, and a confidence score corresponding to the historical threat information. Taking the case where the threat knowledge base 102 uses a graph database for storage as an example, it will perform a subgraph query based on the identifier in the graph database: for example, starting from the hash value, performing subgraph traversal in the knowledge graph, matching attack patterns, historical attack cases and IoC indicators associated with the current feature, and calculating and generating a confidence score according to the connection strength between nodes (such as feature overlap, historical frequency of occurrence). The confidence score is used to quantify the degree of association between historical threat information and the current multi-modal fusion feature. Finally, the threat knowledge base 102 can encapsulate the retrieved historical threat information and the corresponding confidence score into a structured data package and return it to the large model agent 103. The large model agent 103 can integrate real-time multi-modal fusion features with high-confidence historical threat information to generate threat analysis results containing threat types, risk levels, key feature bases, and other information, providing accurate input for the decision-making of the dynamic strategy engine. The interaction logic between the threat knowledge base 102 and the large model agent 103 in this embodiment is shown in the schematic diagram of Figure 2 Through this efficient retrieval and feedback mechanism, the large model agent 103 can quickly obtain historical evidence to support threat judgment, laying the foundation for generating accurate analysis results in the future.

[0056] Of course, in other scenarios, such as facing new vulnerability attacks, unknown attack method variants, or abnormal traffic patterns for the first time in a business scenario, the threat knowledge base 102 may not have included relevant historical records, resulting in the large model agent 103 being unable to obtain associated historical threat information. At this time, the large model agent 103 can rely on its pre-training capabilities and dynamic reasoning capabilities to independently perform threat analysis based on multi-modal fusion features. Specifically, a large amount of network security samples (including public vulnerability exploit data, attack chain simulation data, normal business traffic features, etc.) can be used to train the large model agent 103 in advance, enabling it to have semantic understanding, pattern recognition, and logical reasoning capabilities for unknown features. For example, when detecting abnormal session handshake timing features in encrypted traffic, the large model agent 103 can analyze the deviation of the timing features from normal encrypted protocols, combine the built-in "abnormal behavior-potential risk" reasoning logic, and determine that it may be a hidden transmission behavior of a new tunnel attack, and generate a preliminary threat analysis result (such as "detecting unknown encrypted traffic anomaly, suspected new data exfiltration attempt, risk level temporarily evaluated as medium risk").

[0057] To further improve the accuracy of the analysis results of such unknown threats, the system can also set up an artificial verification and feedback mechanism. Based on the actual network environment and business background, security operation personnel can audit and adjust the preliminary threat analysis results generated by the large model agent 103, such as correcting threat type definitions, optimizing risk level assessment, or supplementing attack path speculation. The adjusted threat analysis results can also be fed back to the threat knowledge base 102 as new historical threat information, and after structured processing, they are updated to the knowledge graph (such as adding the nodes and relationships of "new tunnel attack features -> associated behavior patterns -> damage levels"), realizing the closed-loop precipitation of knowledge. This "self-analysis + artificial optimization + knowledge update" mechanism not only ensures the timely response to unknown threats, but also continuously enhances the threat identification ability of the system by continuously accumulating practical data, so that the large model agent 103 always maintains adaptability and foresight when dealing with network threat evolution.

[0058] The dynamic strategy engine 104 is used to receive the threat analysis results output by the large model agent 103, determine the threat level of the threat analysis results, generate a traffic collection strategy corresponding to the threat level, and issue the traffic collection strategy to the probe 101 for execution. This process converts abstract threat analysis results into executable collection strategies, realizes the conversion from threat perception to active defense, and ensures that the system can take differentiated response measures for different levels of threats, thereby optimizing resource consumption while ensuring security protection effect.

[0059] Specifically, the dynamic strategy engine 104 can first perform in-depth analysis on the received threat analysis results, extract key information such as threat type, confidence score, and associated attack patterns, and determine the threat level in combination with the preset risk rating rules. For example, when the threat analysis result shows "APT attack features exist, associated historical high-risk cases and confidence score > 0.9", it is determined as a high-risk threat; if the result is "suspected XSS exploratory attack, feature matching degree is moderate and no direct associated harmful case", it is determined as a medium-risk threat; and for the case of "occasional abnormal port access, isolated features and confidence < 0.3", it is determined as a low-risk threat. This grading mechanism ensures the accuracy of strategy generation and avoids resource waste or protection omissions caused by "one-size-fits-all" collection.

[0060] After determining the threat level, the dynamic policy engine 104 can generate traffic collection policies corresponding to the threat level: For high-risk threats, in order to fully capture attack details to support source tracing and handling, a full-packet collection policy is generated, and a real-time alarm mechanism is triggered to push threat information to the security operation and maintenance platform through the interface to ensure that operation and maintenance personnel respond immediately; For medium-risk threats, in order to balance collection efficiency and information integrity, a key field collection policy is generated (such as focusing on key information such as HTTP request headers, cookie fields, and specific protocol payloads), and a periodic reporting mode is adopted (such as summarizing the collected data every 5 minutes), which reduces the pressure of real-time transmission and ensures that the threat is dynamically traceable; For low-risk threats, in order to reduce system resource consumption, a sampling collection policy is generated (such as randomly sampling traffic at a rate of 10%), and the collected data is temporarily stored in the local cache of probe 101, and only uploaded to the cloud platform when a higher risk threshold is triggered later, so as to achieve fine-grained allocation of resources and reduce the consumption of cloud platform resources. This refined and dynamic hierarchical strategy mechanism enables the system to dynamically adjust the collection granularity and response method according to the actual threat situation of traffic, achieving a precise balance between security protection and resource consumption. While ensuring the detection rate of key threats, it effectively reduces system resource consumption and comprehensively improves the efficiency of abnormal traffic detection and the real-time security protection effect.

[0061] Furthermore, the dynamic policy engine 104 can also receive user-preset custom policy rules and match threat analysis results with these rules. If a match is successful, a traffic collection policy is generated based on the custom policy rules. For example, users can customize matching rules for key attack types according to their business characteristics, such as "triggering a high-risk policy if an HTTP request contains the SQL injection keyword UNION SELECT" or "if the URL contains..." <script>标签则按中危策略采集”、"短时间内同一IP发起超过50次登录请求则判定为暴力破解并执行全量采集”等。通常情况下,自定义策略的优先级高于系统默认分级策略,能够更贴合用户的实际业务场景与防护需求,例如对核心业务系统的特定端口流量,用户可设置更严格的采集触发条件,确保关键资产得到重点防护。这种灵活性使得系统既能满足通用安全需求,又能适配个性化场景,进一步提升防御的针对性。

[0062] 策略生成后,动态策略引擎104可以通过轻量化通信协议(如MQTT或HTTP / 2)将流量采集策略下发至对应边缘网关的探针101。为适配不同探针的异构环境(如不同厂商的网关设备、不同配置的边缘节点),动态策略引擎104内部还可以设置有规则编译器,使得流量采集策略可以经规则编译器转化为统一格式的可执行指令(如JSON配置或二进制指令),再将转换后的可执行指令下发至对应边缘网关的探针101,确保探针101无论硬件型号、部署场景如何,均能准确解析并执行流量采集策略。例如,针对资源受限的边缘探针101,下发的指令可能会自动精简冗余字段,仅保留核心采集规则;而针对高性能网关探针,则可携带更详细的采集参数(如报文深度、过滤条件等)。这种灵活的下发机制确保了策略在分布式网络环境中的高效落地,最终实现"威胁等级动态适配采集策略”的智能化防御目标。

[0063] 探针101在接收到动态策略引擎104下发的流量采集策略后,则可以立即更新本地采集规则,按照流量采集策略定义的采集粒度(全报文 / 关键字段 / 抽样)与响应方式执行后续流量采集,形成"采集-分析-决策-执行”的完整闭环,实现对网络威胁的动态、智能响应。

[0064] 针对以上各实施例中的大模型智能体103,本申请实施例还提供一种流量威胁分析方法,用于展示大模型智能体103在各个实施例中所执行的步骤。

[0065] 图3是本申请根据一示例性实施例示出的一种流量威胁分析方法的流程示意图。如图3所示,该方法包括:

[0066] S301、接收网络流量的原始特征数据;

[0067] S302、对原始特征数据进行语义分析和时序分析,提取原始特征数据的多模态融合特征;

[0068] S303、根据多模态融合特征从预设的威胁知识库获取与多模态融合特征关联的历史威胁信息;

[0069] S304、基于多模态融合特征以及历史威胁信息,生成网络流量的威胁分析结果。

[0070] 上述方法中各步骤的具体实现过程详见上述系统中对应大模型智能体103的功能实现过程,同时,需要说明的是,大模型智能体103除了执行上述步骤外,上文中其他各实施例中大模型智能体103所执行的其他步骤,均适用在该方法所提供的方案中,在此不再赘述。

[0071] 针对以上各实施例中的动态策略引擎104,本申请实施例还提供一种流量采集策略生成方法,用于展示动态策略引擎104在各个实施例中所执行的步骤。

[0072] 图4是本申请根据一示例性实施例示出的一种流量采集策略生成方法的流程示意图。如图4所示,该方法包括:

[0073] S401、接收上述任一实施例中的大模型智能体生成的威胁分析结果;

[0074] S402、确定威胁分析结果的威胁等级,并生成与威胁等级对应的流量采集策略。

[0075] 基于上述系统中对动态策略引擎104的功能描述可知,动态策略引擎104生成的与威胁等级对应的流量采集策略主要包括:针对高危威胁,生成全报文采集策略,同时触发实时告警机制;针对中危威胁,生成关键字段采集策略,并采用周期上报模式;针对低危威胁,生成抽样采集策略,并将采集数据暂存于探针101的本地缓存。

[0076] 此外,动态策略引擎104还可以支持接收用户预设的自定义策略规则,并将威胁分析结果与自定义策略规则进行匹配,在匹配成功的情况下,基于自定义策略规则生成流量采集策略。同时,通常情况下,自定义策略的优先级高于系统默认分级策略。

[0077] 动态策略引擎104内部还可以设置有规则编译器,使得流量采集策略可以经规则编译器转化为统一格式的可执行指令,再将转换后的可执行指令下发至对应边缘网关的探针101。

[0078] 因此,本申请实施例还可以提供另一种流量采集策略生成方法。图5是本申请根据一示例性实施例示出的另一种流量采集策略生成方法的流程示意图。如图5所示,该方法包括:

[0079] S501、接收上述任一实施例中的大模型智能体生成的威胁分析结果;

[0080] S502、确定威胁分析结果与用户预设的自定义策略规则是否匹配成功,若匹配成功,执行步骤S503,若匹配失败,执行步骤S504;

[0081] S503、基于自定义策略规则生成流量采集策略;

[0082] S504、确定威胁分析结果的威胁等级,并生成与威胁等级对应的流量采集策略;

[0083] S505、针对高危威胁,生成全报文采集策略,同时触发实时告警机制

[0084] S506、针对中危威胁,生成关键字段采集策略,并采用周期上报模式

[0085] S507、针对低危威胁,生成抽样采集策略,并将采集数据暂存于探针的本地缓存;

[0086] S508、将流量采集策略通过规则编译器转化为统一格式的可执行指令;

[0087] S509、将转换后的可执行指令下发至探针。

[0088] 上述方法中各步骤的具体实现过程详见上述系统中对应动态策略引擎104的功能的实现过程,同时,需要说明的是,动态策略引擎104除了执行上述步骤外,上文中其他各实施例中动态策略引擎104所执行的其他步骤,均适用在该方法所提供的方案中,在此不再赘述。

[0089] 与前述流量威胁分析方法的实施例相对应,本申请实施例还提供一种大模型智能体103,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序;其中,处理器执行计算机程序时实现上述任一实施例所记载的流量威胁分析方法的步骤。

[0090] 图6是本申请根据一示例性实施例示出的一种大模型智能体103的结构示意图。如图6所示,在硬件层面,该大模型智能体103包括处理器601、内部总线602、网络接口603、内存604以及非易失性存储器605,当然还可能包括其他业务所需要的硬件。本申请的一个或多个实施例可以基于软件方式来实现,比如由处理器601从非易失性存储器605中读取对应的计算机程序到内存604中然后运行。当然,除了软件实现方式之外,本申请的一个或多个实施例并不排除其他实现方式,比如逻辑器件抑或软硬件结合的方式等等,也就是说以上处理流程的执行主体并不限定于各个逻辑单元,也可以是硬件或逻辑器件。

[0091] 与前述流量采集策略生成方法的实施例相对应,本申请实施例还提供一种动态策略引擎104,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序;其中,处理器执行计算机程序时实现上述任一实施例所记载的流量采集策略生成方法的步骤。

[0092] 图7是本申请根据一示例性实施例示出的一种动态策略引擎104的结构示意图。如图7所示,在硬件层面,该动态策略引擎104包括处理器701、内部总线702、网络接口703、内存704以及非易失性存储器705,当然还可能包括其他业务所需要的硬件。本申请的一个或多个实施例可以基于软件方式来实现,比如由处理器701从非易失性存储器705中读取对应的计算机程序到内存704中然后运行。当然,除了软件实现方式之外,本申请的一个或多个实施例并不排除其他实现方式,比如逻辑器件抑或软硬件结合的方式等等,也就是说以上处理流程的执行主体并不限定于各个逻辑单元,也可以是硬件或逻辑器件。

[0093] 示例性的,处理器包括但不限于中央处理单元(Central Processing Unit,CPU)、图形处理器(graphics processing unit,GPU)、数字信号处理器 (Digital SignalProcessor,DSP)、专用集成电路 (Application Specific Integrated Circuit,ASIC)或者现成可编程门阵列 (Field-Programmable Gate Array,FPGA)等。

[0094] 示例性的,存储器可以包括至少一种类型的存储介质,存储介质包括闪存、硬盘、多媒体卡、卡型存储器 (例如,SD或DX存储器等等)、随机访问存储器 (RAM)、静态随机访问存储器 (SRAM)、只读存储器 (ROM)、电可擦除可编程只读存储器 (EEPROM)、可编程只读存储器(PROM)、磁性存储器、磁盘、光盘等等。

[0095] 与前述方法的实施例相对应,本申请实施例还提供一种计算机可读存储介质,存储有计算机程序,计算机程序被处理器执行时实现上述任一实施例所记载的流量威胁分析方法或流量采集策略生成方法的步骤。

[0096] 与前述方法的实施例相对应,本申请实施例还提供一种计算机程序产品,包括计算机程序,该计算机程序被处理器执行时实现上述任一实施例所记载的流量威胁分析方法或流量采集策略生成方法的步骤。

[0097] 上述对本申请特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。

[0098] 本领域技术人员在考虑说明书及实践这里申请的发明后,将容易想到本申请的其它实施方案。本申请旨在涵盖本申请的任何变型、用途或者适应性变化,这些变型、用途或者适应性变化遵循本申请的一般性原理并包括本申请未申请的本技术领域中的公知常识或惯用技术手段。说明书和实施例仅被视为示例性的,本申请的真正范围和精神由上面的权利要求指出。

[0099] 应当理解的是,本申请并不局限于上面已经描述并在附图中示出的精确结构,并且可以在不脱离其范围进行各种修改和改变。本申请的范围仅由所附的权利要求来限制。

[0100] 以上所述仅为本申请的较佳实施例而已,并不用以限制本申请,凡在本申请的精神和原则之内,所做的任何修改、等同替换、改进等,均应包含在本申请保护的范围之内。< / script>

Claims

1. A flow acquisition system, characterized by, The system is applied to an edge cloud collaborative architecture, the edge cloud collaborative architecture comprises an edge gateway, a cloud platform and an agent platform; the system comprises: a probe deployed on the edge gateway, configured to collect raw feature data of network traffic; a threat knowledge base deployed on the agent platform, configured to store historical threat information; a large model agent deployed on the agent platform, in communication connection with the probe and the threat knowledge base respectively, configured to: receive the raw feature data, perform semantic analysis and time series analysis on the raw feature data, and extract multi-modal fusion features of the raw feature data; generate a retrieval identifier based on the multi-modal fusion features, and send the retrieval identifier to the threat knowledge base, so that the threat knowledge base performs associated retrieval based on the retrieval identifier, determines historical threat information associated with the multi-modal fusion features and a confidence score corresponding to the historical threat information; the confidence score is used to represent the degree of association between the historical threat information and the multi-modal fusion features; obtain the historical threat information returned by the threat knowledge base and the confidence score corresponding to the historical threat information; generate a threat analysis result of the network traffic based on the multi-modal fusion features and the historical threat information; a dynamic policy engine deployed on the cloud platform, in communication connection with the large model agent and the probe respectively, configured to receive the threat analysis result, determine a threat level of the threat analysis result, generate a traffic collection policy corresponding to the threat level, and issue the traffic collection policy to the probe; the probe is further configured to collect traffic according to the traffic collection policy.

2. The system of claim 1, wherein, The large model agent comprises a BERT model and a Transformer network; The large model agent performs semantic analysis and time series analysis on the raw feature data, and extracts multi-modal fusion features of the raw feature data, specifically comprising: The large model agent uses the BERT model to perform semantic analysis on the raw feature data, and obtains semantic features of the raw feature data; The large model agent uses the Transformer network to perform time series analysis on the raw feature data, and obtains time series features of the raw feature data; The large model agent performs feature fusion on the semantic features and the time series features, and obtains the multi-modal fusion features.

3. The system of claim 1, wherein, The dynamic policy engine generates a traffic collection policy corresponding to the threat level, specifically comprising: if the threat level is a high-risk threat, the dynamic policy engine generates a traffic collection policy of full-packet collection and real-time alarm; if the threat level is a medium-risk threat, the dynamic policy engine generates a traffic collection policy of key field collection and periodic reporting; if the threat level is a low-risk threat, the dynamic policy engine generates a traffic collection policy of sampling collection and local caching.

4. The system of claim 1, wherein, The dynamic strategy engine is also configured to receive a user preset custom policy rule, match the threat analysis result with the custom policy rule, and generate the traffic collection strategy based on the custom policy rule if the matching is successful.

5. A traffic threat analysis method characterized by, The method applied to the large model agent in the traffic collection system of any one of claims 1-4, comprising: The large model agent receives original feature data of network traffic; The large model agent performs semantic analysis and time series analysis on the original feature data to extract multi-modal fusion features of the original feature data; The large model agent acquires historical threat information associated with the multi-modal fusion features from a preset threat knowledge base according to the multi-modal fusion features; The large model agent generates a threat analysis result of the network traffic based on the multi-modal fusion features and the historical threat information.

6. A traffic collection policy generation method characterized by comprising: The method applied to the traffic collection system of any one of claims 1-4, wherein the traffic collection system comprises a dynamic strategy engine and a large model agent, and the method comprises: The dynamic strategy engine receives a threat analysis result generated by the large model agent; The dynamic strategy engine determines a threat level of the threat analysis result and generates a traffic collection strategy corresponding to the threat level.

7. A large model agent, characterized in that, A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of claim 5 when executing the computer program.

8. A dynamic policy engine characterized by, A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of claim 6 when executing the computer program.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 5-6.

Citation Information

Patent Citations

  • An acquisition strategy generation method and system based on external threats

    CN109714312A

  • Standardized processing system and method for multi-modal data acquisition and fusion of network transaction platform

    CN120336713A