An IPv6 encrypted traffic advanced persistent threat attack identification method and system

CN122457381BActive Publication Date: 2026-09-08LISHUI POWER SUPPLY COMPANY OF STATE GRID ZHEJIANG ELECTRIC POWER
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610923064.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-08
Estimated Expiration
2046-06-25

AI Technical Summary

Technical Problem

[0004]本发明针对上述不足或缺点,提供了一种IPv6加密流量高级持续威胁攻击识别方法及系统,能够解决现有技术在应对信创环境下客户端指纹多样化的现实场景时,存在握手指纹模板的识别精度与整体识别准确率难以兼顾的技术问题

Benefits of technology

[0007]采用本发明的技术方案,提供了一种IPv6加密流量高级持续威胁攻击识别方法。该方法通过特征提取、序列比对与容差判定、筛选聚焦、匹配验证与聚类分析四个核心步骤协同实现。其中,从IPv6加密链路的传输层握手报文中提取密码套件排列序列、扩展字段类型集合及密码库版本标识,为流量分析提供了源自协议协商阶段的核心特征数据;基于这些特征与预设的合法客户端变体指纹库进行序列比对与容差判定以确定套件排列吻合等级,为区分正常通信与异常流量建立了动态、可容忍的量化评估基准;依据该吻合等级筛选出低于预设合法阈值的待甄别消息并生成待甄别对象集合,实现了从海量加密流量中对可疑对象的快速聚焦与高效提纯;将待甄别对象集合中的握手指纹特征与已知的仿冒流量特征库进行匹配验证,并基于排列特征向量执行聚类分析以识别同源的隐蔽仿冒流量子集,构建了从可疑对象中精准挖掘并关联攻击组件的验证与溯源机制。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457381B_ABST
    Figure CN122457381B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network security, in particular to an IPv6 encrypted traffic advanced persistent threat attack identification method and system; the method comprises the following steps: extracting a cipher suite arrangement sequence, an extended field type set and a cipher library version identifier; performing sequence comparison and tolerance determination with a preset legal client variant fingerprint library to determine a suite arrangement coincidence level; screening out to-be-identified messages with a coincidence level lower than a preset legal threshold to generate a to-be-identified object set; and performing clustering analysis based on the arrangement feature vectors of the to-be-identified object set to identify a homologous concealed imitated traffic subset. In this way, the technical problem that the identification precision of a handshake fingerprint template and the overall identification accuracy are difficult to be considered in the prior art when dealing with the realistic scene of diversified client fingerprints in a signal creation environment is solved, and the accuracy and reliability of encrypted threat identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a method and system for identifying advanced persistent threat attacks on IPv6 encrypted traffic. Background Technology

[0002] In the field of network security, with the large-scale deployment of IPv6 (Internet Protocol version 6) and the rapid promotion of innovative information technology applications, encrypted communication based on SSL / TLS (Secure Sockets Layer / Transport Layer Security) has become the mainstream of network traffic. Advanced Persistent Threat (APT) attacks utilize encrypted channels to conceal their malicious activities, posing a serious threat to critical information infrastructure. Therefore, accurately identifying such covert attacks in encrypted traffic has become an essential capability for maintaining network security. In traditional network environments, traffic characteristics are relatively stable, and identification methods based on fixed fingerprint rule matching can meet detection needs to a certain extent. However, browsers generate various legitimate variants due to differences in cryptographic library versions, compilation options, etc. The fingerprint characteristics they present during the communication handshake process, such as the order of cipher suites and the combination of extended fields, have inherent differences, breaking the traditional assumption of a single, static fingerprint.

[0003] In the field of encrypted malicious traffic identification, some explorations have been made in existing technologies. For example, existing technology (application publication number CN116318827A) discloses a system environment identification method based on encrypted malicious traffic attacks. This method constructs a systematic fingerprint database covering various simulated real system environments and matches the attack environment fingerprint of the target traffic with the fingerprints in the database to identify the system environment in which the attack originated. However, the core of this method relies on a pre-collected, static systematic fingerprint database for precise matching. This method fails to fully consider that in actual domestic IT innovation environments, legitimate client software itself has a large number of legitimate fingerprint variants caused by version and configuration, and these variants form a dynamic range. If an overly precise fingerprint template is used for matching, these legitimate variant traffic will be misjudged as attacks, leading to a surge in false positives; if the matching conditions are relaxed to reduce false positives, it will be difficult to capture the subtle feature deviations produced when attackers imitate legitimate software, resulting in missed detections of genuine spoofing attack traffic. Therefore, existing identification methods based on precise static fingerprint matching face a technical challenge when dealing with the diverse client fingerprints in the context of the information technology innovation environment: it is difficult to balance the accuracy of the handshake fingerprint template with the overall recognition accuracy. That is, while improving the template accuracy to enhance the ability to identify spoofed traffic, it will increase the misjudgment of legitimate variant traffic; while relaxing the template to tolerate legitimate variants will weaken the ability to detect hidden spoofed traffic. Summary of the Invention

[0004] To address the aforementioned shortcomings or drawbacks, this invention provides a method and system for identifying advanced persistent threat attacks on IPv6 encrypted traffic. This method solves the technical problem that existing technologies struggle to balance the accuracy of handshake fingerprint template recognition with overall recognition accuracy when dealing with diverse client fingerprint scenarios in the context of domestic IT innovation.

[0005] This invention provides a method for identifying advanced persistent threat (APS) attacks on IPv6 encrypted traffic, comprising: extracting cipher suite permutation sequences, extended field type sets, and cipher library version identifiers from transport layer handshake messages of IPv6 encrypted links; performing sequence comparison and tolerance determination based on the cipher suite permutation sequences, extended field type sets, and cipher library version identifiers against a preset legitimate client variant fingerprint database to determine the permutation matching level; filtering out messages with matching levels below a preset legitimate threshold based on the permutation matching level to generate a set of objects to be identified; matching and verifying the handshake fingerprint features in the set of objects to be identified against a known spoofing traffic feature database, and performing cluster analysis based on the permutation feature vectors of the set of objects to be identified to identify a subset of covert spoofing traffic originating from the same source.

[0006] According to a second aspect, the present invention provides an advanced persistent threat (APPS) attack identification system for IPv6 encrypted traffic, characterized in that it includes a processor and a memory, wherein the memory stores a computer program.

[0007] This invention provides a method for identifying advanced persistent threat (APS) attacks on IPv6 encrypted traffic. The method is achieved through four core steps: feature extraction, sequence alignment and tolerance assessment, filtering and focusing, matching verification, and cluster analysis. Specifically, it extracts cipher suite permutations, extended field type sets, and cipher library version identifiers from the transport layer handshake messages of the IPv6 encrypted link, providing core feature data originating from the protocol negotiation phase for traffic analysis. Based on these features, sequence alignment and tolerance assessment are performed against a pre-defined legitimate client variant fingerprint database to determine the suite permutation matching level, establishing a dynamic and tolerable quantitative evaluation benchmark for distinguishing normal communication from abnormal traffic. Based on this matching level, messages below a pre-defined legitimate threshold are filtered out to generate a set of objects to be identified, achieving rapid focusing and efficient purification of suspicious objects from massive encrypted traffic. The handshake fingerprint features in the set of objects to be identified are matched and verified against a known spoofing traffic feature database, and cluster analysis is performed based on the permutation feature vectors to identify subsets of concealed spoofing traffic originating from the same source, constructing a verification and tracing mechanism for accurately mining and associating attack components from suspicious objects.

[0008] In this technical solution, the present invention addresses the problem described in the background art of existing methods based on static fingerprint precise matching, which struggle to balance the accuracy of handshake fingerprint template recognition with overall recognition accuracy when dealing with the diversity of client fingerprints in the context of information technology innovation. By introducing sequence comparison and tolerance judgment steps, dynamic matching levels and legal thresholds are set for normal feature offsets caused by version and compilation differences of legitimate clients. This allows the template to maintain high accuracy against standard features while accommodating the inherent differences of legitimate variants, thereby effectively avoiding the misclassification of a large number of legitimate variant traffic as attacks at the recognition front end and reducing the false positive rate. Furthermore, to address the risk of missed detection of attacker spoofing traffic that may result from accommodating legitimate variants, the present invention uses matching verification and cluster analysis steps to perform secondary matching of the selected suspicious objects with a known spoofing feature library and performs homology clustering based on subtle arrangement features. This can accurately identify and associate the real attack components and their homologous variants from traffic that deliberately imitates normal fingerprints but still has inherent deviations, effectively preventing missed detections. Therefore, the technical solution of the present invention solves the technical problem that the existing technology has difficulty in balancing the recognition accuracy of the handshake fingerprint template and the overall recognition accuracy when dealing with the diverse real-world scenarios of client fingerprints in the context of information technology innovation, thereby improving the accuracy and reliability of encryption threat identification. Attached Figure Description

[0009] Figure 1 This is a flowchart of an embodiment of the advanced persistent threat attack identification method for IPv6 encrypted traffic according to the present invention; Figure 2This is a schematic diagram of the implementation process of full-link identification and interception of covert APT attacks in IPv6 encrypted traffic based on handshake fingerprint features in one embodiment of the present invention; Figure 3 This is a schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present invention. Detailed Implementation

[0010] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0011] During the development of this invention, researchers discovered a complex and contradictory relationship between the matching accuracy of the handshake fingerprint template and the accuracy of identifying legitimate client variants through extensive experiments and data analysis: the more accurate the template, the higher the recognition rate of known attack traffic, but the risk of misjudging legitimate browser variant traffic generated under normal conditions due to different cryptographic library versions and compilation options as attacks also increases dramatically; conversely, if the template matching conditions are relaxed to tolerate legitimate variants, it becomes difficult to effectively capture the subtle feature deviations generated when attackers deliberately imitate normal handshake behavior, leading to an increased false negative rate for advanced persistent threat spoofing traffic. Based on this relationship, this invention innovatively proposes this technical solution, which utilizes multi-dimensional features such as the cipher suite arrangement sequence, extended field type set, and cryptographic library version identifier extracted from transport layer handshake messages. By dynamically comparing and judging the sequence with a preset legitimate client variant fingerprint database, and combining this with spoofing feature matching and homology clustering analysis of the selected suspicious traffic, it achieves the ability to accurately lock attack homology groups with consistent spoofing features while encompassing a massive number of legitimate variants, embodying the core concept of "dynamic inclusion first, then precise focus".

[0012] Specifically, through comparative experiments, the invention team discovered that traditional methods based on static fingerprint precise matching have technical shortcomings that cannot adapt to the highly diverse client software fingerprints in the context of the domestic IT innovation environment. Their fixed matching rules mark a large amount of normal browser traffic with feature shifts due to legitimate reasons (such as version iterations and differences in compilation configurations) as abnormal, leading to frequent false positives. Furthermore, if the rules are relaxed to reduce false positives, the inability to distinguish between "legitimate natural shifts" and "malicious deliberate imitations" causes genuine spoofing attack traffic to slip through the net. The dynamic baseline matching method based on sequence alignment and tolerance determination proposed in this invention improves compatibility with legitimate variant traffic and significantly reduces false positives. Through secondary matching verification and homogeneous clustering analysis of low-match-level traffic, it can accurately mine and associate genuine attack components from suspicious objects, ensuring a high detection rate for advanced persistent threat (APS) spoofing attacks with extremely low false positives. This solves the problem of balancing false positives and false negatives when identifying APS attacks in encrypted traffic in a diverse domestic IT innovation environment, improving the accuracy and engineering usability of encrypted threat detection.

[0013] Therefore, this invention provides a method for identifying advanced persistent threat (APS) attacks on IPv6 encrypted traffic, which can be applied to a traffic auditing system (hereinafter referred to as the "system"). This system can detect APS spoofing attacks in encrypted traffic through deep analysis and intelligent analysis of network traffic, and by running the identification method defined in this invention, thereby achieving accurate identification and interception of covert spoofing communications. Specifically, this system can be deployed in various hardware computing and execution environments, including but not limited to: network boundary firewalls, traffic probes, dedicated security devices, or cloud servers, in software, hardware, or integrated software and hardware forms.

[0014] like Figure 1 As shown, the method may include: Step S110: Extract the cipher suite arrangement sequence, extended field type set, and cipher library version identifier from the transport layer handshake message of the IPv6 encrypted link.

[0015] Among them, the cipher suite arrangement sequence refers to the list of encryption algorithm combination identifiers announced by the client to the server in order of priority during the transport layer security protocol handshake phase. This sequence is one of the core fingerprint features of the client software environment; the extension field type set refers to the set of protocol extension type identifiers declared in the handshake message to support enhanced functions, such as server name indication, application layer protocol negotiation, and other extensions; the cryptographic library version identifier refers to the version number or compilation identifier information of the underlying cryptographic library (such as OpenSSL, GmSSL) used to generate the cryptographic handshake parameters.

[0016] Specifically, the system can use a deep packet inspection engine to perform protocol parsing on captured network packets, locate the "Client Hello" message segment in the transport layer security protocol handshake message, extract the byte sequence of the "CipherSuites" field as the cipher suite permutation sequence, extract the type values ​​in the "Extensions" list to construct an extension field type set, and parse the cipher library version identifier from the relevant extensions (such as "supported_versions" or "signature_algorithms") or specific byte patterns in the message payload.

[0017] For example, the system parses the cipher suite arrangement sequence as "{0x13,0x02,0x13,0x03, 0xC0,0x2C,…}" from a single IPv6 encrypted traffic, the extended field type set as "{0x00, 0x0D, 0x10, 0x2B}", and the cipher library version identifier as "OpenSSL 3.0.2".

[0018] Step S120: Based on the cipher suite arrangement sequence, extended field type set, and cipher library version identifier, perform sequence comparison and tolerance determination with the preset legitimate client variant fingerprint library to determine the suite arrangement matching level.

[0019] Among them, the legitimate client variant fingerprint database is a pre-built database that stores the standard handshake features generated by various known legitimate client software (such as browsers of different versions and compilation configurations) during normal communication; sequence comparison refers to the process of calculating the similarity between the sequence of the cipher suite to be tested and the benchmark sequence in the fingerprint database; tolerance judgment refers to the judgment logic that allows some elements in the sequence to have positional shifts or content differences within a preset tolerance range during the comparison process; the suite arrangement matching level is a level index used to quantify the degree of matching between the feature to be tested and the legitimate benchmark, and the higher the level, the more likely it is to be legitimate traffic.

[0020] Specifically, the system can calculate the initial matching degree between the test sequence and the reference sequence using an edit distance algorithm (such as Levenshtein Distance). Then, it combines the matching of the extended field type set with the compatibility of the cryptographic library version identifier to weight and correct the initial matching degree, obtaining a comprehensive matching degree score. Finally, the comprehensive matching degree score is compared with preset high and low matching thresholds, and it is also determined whether the positional offset of the sequence elements is within the tolerance threshold. The final matching level (such as "high matching", "medium matching", "low matching") is determined by combining these two results.

[0021] For example, the system compares the sequence to be tested with the baseline sequence "Chrome 105.0.5195.127 Standard Edition" in the fingerprint database, and calculates an initial matching degree of 85%. Because the extended fields match completely and the cryptographic database version is compatible, the overall matching degree is corrected to 92%. Since 92% is higher than the high matching threshold of 90%, and the first three digits of the main kit are completely consistent, the system determines that the kit arrangement matching level of this traffic is "high matching".

[0022] Step S130: Based on the matching level of the kit arrangement, filter out the messages to be identified whose matching level is lower than the preset legal threshold, and generate a set of objects to be identified.

[0023] The preset legal threshold is a pre-defined boundary point for matching level. Traffic below this threshold is considered suspicious and needs to be further verified. The message to be verified refers to a single network session handshake message that is determined to have a matching level below the preset legal threshold in step S120. The set of objects to be verified is a structured data set composed of all messages to be verified and their associated features (such as session quintuples and timestamps).

[0024] Specifically, the system can manage the flow table system and traverse all session messages that have undergone feature extraction and level determination. The system reads the suite arrangement matching level of each message and compares it with a preset valid threshold (e.g., the lower bound of the "medium matching" level). Messages with a "low matching" level are marked as objects to be screened, and the system extracts five key identifiers (i.e., a 5-tuple) of the object: source IP (Internet Protocol) address, destination IP address, source port, destination port, and transport layer protocol, along with a timestamp, and encapsulates them into an entry for an object to be screened. Finally, all entries are added to a list to form a set of objects to be screened.

[0025] For example, the system processes 10,000 handshake messages within a monitoring period, of which 9,900 messages have a matching level of "high matching" and "medium matching", and 100 messages have a matching level of "low matching". The system marks these 100 "low matching" messages as objects to be identified, extracts their respective five-tuple information, and generates a list of objects to be identified containing 100 entries.

[0026] Step S140: Match and verify the handshake fingerprint features in the set of objects to be identified with the known spoofing traffic feature library, and perform cluster analysis based on the permutation feature vector of the set of objects to be identified to identify a subset of hidden spoofing traffic from the same source.

[0027] The spoofing traffic signature database is a database that stores distinctive handshake signature patterns left behind by known advanced persistent threat (APS) attack components when spoofing legitimate clients. Handshake fingerprint features, in this step, specifically refer to the feature combinations extracted from the target object for fine-grained comparison, which may include cipher suite permutation sequences, extended field order, and detailed parameters of specific extensions. Permutation feature vectors are numerical vectors obtained by mathematically representing cipher suite permutation sequences, and can be used to calculate the similarity distance between different sequences. Homology clustering refers to the analysis process of grouping traffic with highly similar handshake fingerprint features into the same group; these traffic streams are likely to originate from the same attack component or the same attacker. The covert spoofing traffic subset refers to the set of encrypted traffic belonging to the same APS attack activity ultimately identified through the above analysis.

[0028] Specifically, the system achieves accurate identification through a two-step process. First, matching verification: the detailed handshake fingerprint features of each object to be identified are compared item by item with records in the spoofing traffic feature database. Second, clustering analysis: the permutation sequences of cipher suites of all objects to be identified are converted into permutation feature vectors (e.g., using a bag-of-words model or an n-gram model), and then clustering algorithms (such as DBSCAN density clustering algorithm) are used to analyze these vectors, grouping objects whose vector distance is less than a preset threshold into the same cluster. Finally, the system identifies traffic that matches high similarity features in the spoofing feature database and is also grouped into the same cluster by the clustering algorithm as the same subset of covert spoofing traffic.

[0029] For example, the system processes a set of 100 objects to be identified. Through matching verification, it is found that the characteristics of 15 objects highly match the feature patterns of "Golang malware v2.3" in the spoofing library. Next, cluster analysis divides all 100 objects into 5 clusters, with the aforementioned 15 objects all concentrated in "Cluster 1," and the average cosine similarity of the permutation feature vectors of objects within this cluster is greater than 0.95. Based on this, the system determines that the traffic corresponding to "Cluster 1" constitutes a subset of homogeneous, covert spoofing traffic, and its attack source is very likely "Golang malware v2.3."

[0030] The system employing the technical solution of this embodiment can achieve its goals through four core steps: feature extraction, sequence alignment and tolerance determination, filtering and focusing, matching verification, and cluster analysis. Specifically, it extracts cipher suite permutation sequences, extended field type sets, and cipher library version identifiers from the transport layer handshake messages of IPv6 encrypted links, providing core feature data originating from the protocol negotiation phase for traffic analysis. Based on these features, sequence alignment and tolerance determination are performed against a pre-set legitimate client variant fingerprint database to determine the suite permutation matching level, establishing a dynamic and tolerable quantitative evaluation benchmark for distinguishing normal communication from abnormal traffic. Based on this matching level, messages below a pre-set legitimate threshold are filtered out to generate a set of objects to be identified, achieving rapid focusing and efficient purification of suspicious objects from massive encrypted traffic. The handshake fingerprint features in the set of objects to be identified are matched and verified against a known spoofing traffic feature database, and cluster analysis is performed based on the permutation feature vectors to identify subsets of concealed spoofing traffic originating from the same source, constructing a verification and tracing mechanism for accurately mining and associating attack components from suspicious objects.

[0031] In this technical solution, this embodiment addresses the problem described in the background art of existing methods based on static fingerprint precise matching, which struggle to balance the accuracy of handshake fingerprint template recognition with overall recognition accuracy when dealing with the diversity of client fingerprints in the context of information technology innovation. By introducing sequence comparison and tolerance judgment steps, dynamic matching levels and legal thresholds are set for normal feature offsets caused by version and compilation differences of legitimate clients. This allows the template to maintain high accuracy against standard features while accommodating the inherent differences of legitimate variants, thereby effectively avoiding the misclassification of a large number of legitimate variant traffic as attacks at the recognition front end and reducing the false positive rate. Furthermore, to address the risk of missed detection of attacker spoofing traffic that may result from accommodating legitimate variants, a matching verification and clustering analysis step is used to perform a secondary matching between the selected suspicious objects and a known spoofing feature library. Based on subtle arrangement features, homology clustering is performed, which can accurately identify and associate the real attack components and their homologous variants from those traffic that deliberately imitate normal fingerprints but still have inherent deviations, effectively preventing missed detections. Therefore, the technical solution of this embodiment solves the technical problem that the existing technology has difficulty in balancing the recognition accuracy of the handshake fingerprint template and the overall recognition accuracy when dealing with the diverse real-world scenarios of client fingerprints in the context of information technology innovation, thus improving the accuracy and reliability of encryption threat identification.

[0032] In another embodiment, Table 1 below shows a benchmark library of handshake fingerprint features for legal cryptographic browser variants.

[0033] To more intuitively illustrate the construction method of the preset range of legitimate browser variants in this solution and its specific application in the feature comparison stage, this embodiment provides a detailed explanation using real-world national cryptographic browser feature data recorded in Table 1. This benchmark library is generated through long-term tracking and extraction of TLS 1.3 handshake behavior of mainstream and officially certified national cryptographic browsers (such as QiAnXin Trusted Browser, Red Lotus Secure Browser, and TongXin UOS Built-in WebView) under specific operating system environments. Each row in the table represents an independent, legitimate client variant. The system uses this variant data as a "whitelist" benchmark for determining whether unknown traffic is abnormal. Specific feature data examples and comparison logic are as follows: 1.1 QiAnXin Trusted Browser v12.0 (UOS version) An example of the cipher suite arrangement order exhibited by this variant during the handshake phase is as follows. The system extracts this hexadecimal sequence as the main feature vector. Regarding the combination of extended fields, its common form is... Furthermore, this browser version typically links to and uses cryptographic libraries ranging from QAX-GMSSL-3.0.0 to QAX-GMSSL-3.0.8. During the sequence alignment algorithm execution, if the consecutive offset difference of the first byte of the cipher suite of the object to be examined exceeds the "offset difference threshold" recorded in Table 1 (2 in this variant), or if the combination of extended fields is not within the preset reasonable combination space, the system determines that the object has a significant deviation from the legitimate benchmark in the underlying handshake characteristics, thereby reducing its matching level or directly marking it as suspicious.

[0034] 1.2 Red Lotus Secure Browser v8.5 (Kylin ARM version) The Red Lotus browser, optimized for the domestic ARM architecture environment, features a unique combination of extended fields in its handshake fingerprint. It is worth noting that the cryptographic library version declared in this variant is identified as HLH-BabaSSL-2.1.0 to HLH-BabaSSL-2.1.5. During cross-validation, the system will focus on verifying whether the compiled version of the cryptographic library of the target object falls within this legitimate range. If an attacker attempts to impersonate this browser for communication but incorrectly spells the version declaration as BabaSSL-2.5.0 or uses a completely incorrect cryptographic suite (such as not including a Chinese cryptographic algorithm suite starting with 0xC1), the system will immediately trigger a version spoofing verification alert.

[0035] 1.3, 360 Starter Browser v11.2 (compatibility mode) and NeoKylin customized Firefox E SR.

[0036] Table 1 also covers other typical variations, such as 360 Starter Browser v11.2. This variation slightly tweaks the order of the cipher suites; an example sequence is shown below. Furthermore, its extended field combination contains specific 0x1A and 0x05 type values. The NeoKylin customized version of FirefoxESR exhibits the following behavior: The system identifies "ghost" fingerprints that are not present in the legitimate benchmark library by performing a multi-dimensional vector space comparison between the complete handshake fingerprint of the target object (including the arrangement, extended combination, and version identifier) ​​and these refined benchmarks in Table 1, thereby pinpointing potential APT attack covert channels.

[0037] In another embodiment, Table 2 below shows the false positive rate and recall evaluation matrix constructed based on the real test dataset.

[0038] To verify the effectiveness of this solution in protecting legitimate IT application innovation traffic and its accuracy in identifying APT simulated attack traffic in a real network environment, this embodiment constructs a mixed test set containing 1000 known legitimate IT application innovation browser traffic entries and 100 known APT simulated tool traffic entries. The system conducted multiple sets of comparative tests on the overall performance of the solution by adjusting the "comprehensive matching degree" threshold in the core algorithm, and the final legitimate traffic coverage (recall rate) and APT traffic false positive rate are recorded in Table 2. Specific test data and performance trade-off examples are as follows: 2.1 High-strictness matching scenarios (overall matching degree ≥ 0.90) When the system sets the overall matching threshold to 0.90, it means that only traffic whose handshake fingerprint characteristics are extremely close to the preset legal benchmark can be allowed to pass. Under this strict condition, the system's coverage rate for 1000 legitimate domestic browser traffic entries is 85%, meaning that 850 normal traffic entries are correctly identified and allowed to pass. Meanwhile, the system shows the best interception effect on 100 APT simulation tool traffic entries, with the lowest false positive rate of only 2%, indicating that a very high proportion of dangerous traffic is successfully identified.

[0039] 2.2. Scenarios with equal strictness of matching (overall matching degree ≥ 0.85) If the threshold is relaxed to 0.85, the system's coverage of legitimate traffic significantly increases to 95%, with only 50 normal traffic entries being falsely blocked. However, with the relaxation of the judgment criteria, the risk of APT attack traffic slipping through the net increases, causing the false positive rate of APT traffic to rise to 5%.

[0040] 2.3 Low-strictness matching scenarios (overall matching degree ≥ 0.80) When the threshold was further lowered to 0.80, the system aimed for the highest possible coverage of legitimate traffic, reaching 98%, with only 20 normal traffic entries affected. However, this lenient strategy sacrificed security, with the false positive rate for APT traffic soaring to 15%, allowing a large number of maliciously disguised traffic entries to pass through the screening process.

[0041] The test data above shows that the system can flexibly adjust the comprehensive matching threshold according to the security requirements of actual business scenarios. When facing production environments with extremely high availability requirements, the threshold can be appropriately lowered to reduce interference with legitimate IT innovation businesses; while during high-risk periods when advanced persistent threats (APTs) are frequent, the threshold can be increased to minimize the covert communication space for APT attacks.

[0042] In another embodiment, Table 3 below shows the performance comparison test results of this solution and the traditional static fingerprint matching method under real network traffic environment.

[0043] To visually demonstrate the technical superiority of this solution after introducing dynamic feature analysis and version spoofing verification mechanisms, this embodiment sets up a test environment simulating a real office network. The test set consists of 5000 normal domestically developed browser variant traffic entries (covering the latest version handshake data of mainstream domestic cryptographic browsers such as Qi An Xin, Red Lotus, and 360) and 200 spoofed traffic entries generated by APT simulation tools (deliberately simulating the abnormal handshake behavior of advanced threats such as Cobalt Strike). The system performs blind tests on this test set using both this solution and the traditional static fingerprint matching solution. The final detection rate and false positive rate statistics are shown in Table 3. Specific test data and performance difference analysis are as follows: 3.1 Practical Performance of this Solution (Distinguishing Legal Variations) Thanks to the system's built-in legitimate variant feature benchmark library and adaptive sequence comparison algorithm, this solution demonstrates extremely high robustness when facing massive amounts of normal domestic IT innovation business traffic. Out of 5000 normal traffic entries, the system successfully allowed 4925, achieving a legitimate traffic coverage (detection rate) of 98.5%. More importantly, the system exhibits extremely high sensitivity in identifying APT spoofing traffic, successfully blocking 196 malicious traffic entries. Furthermore, due to the introduction of cross-validation logic using cryptographic version identifiers, the system rarely misclassifies normal business traffic as a threat, keeping the false positive rate strictly controlled at an extremely low level of 0.2%, effectively ensuring the network availability of the domestic IT innovation business system.

[0044] 3.2 Limitations of Traditional Static Fingerprint Matching Methods As a comparison, traditional static fingerprint matching schemes rely on predefined fixed fingerprint templates, lacking tolerance for legitimate variants and adaptability to new camouflage techniques. While its detection rate for legitimate traffic is slightly higher (99.0%) under ideal conditions, it fails to identify deliberately fabricated "legitimate" fingerprints by attackers when facing APT attacks, resulting in a large amount of malicious traffic slipping through the net. Data shows that this scheme only intercepted 180 APT traffic entries. More critically, due to the lack of a dynamic context verification mechanism, static matching schemes are highly prone to false positives. In testing, the scheme incorrectly marked 250 variant traffic entries from legitimate domestic browsers as APT attacks, causing the false positive rate to soar to 5.0%. In actual operation and maintenance, such a high false positive rate will force security personnel to invest a significant amount of effort in manual investigation, severely impacting security operation efficiency.

[0045] In some embodiments, the step of performing sequence comparison and tolerance determination with a preset legitimate client variant fingerprint database based on the cipher suite permutation sequence, extended field type set, and cipher library version identifier includes: Obtain the base sequence corresponding to the permutation sequence of the cipher suite.

[0046] The baseline sequence refers to the order of standard cipher suites corresponding to client variants suspected to be the source of the traffic under test, obtained from the legitimate client variant fingerprint database.

[0047] Specifically, the system can quickly search the legitimate client variant fingerprint database, filter out one or several most likely candidate variants based on preliminary characteristics of the traffic under test (such as partial package headers and version identifier prefixes) or by quickly pre-matching with all baseline templates, and read the standard package arrangement order corresponding to these candidate variants as the baseline sequence. For example, based on the cryptographic library version identifier "GmSSL3.0.0" of the traffic under test, the system locks the variant "QiAnXin Trusted Browser v12.0 (UOS version)" in the fingerprint database and reads its stored baseline sequence "{0xC1,0x02,0xC1,0x01,0xC0,0x2F}".

[0048] The sequence alignment calculation is performed by comparing the cipher suite permutation sequence with the baseline sequence to obtain the initial matching degree.

[0049] The initial matching degree is a score that quantifies the overall similarity between two sequences, and it usually ranges from 0 to 1, with 1 indicating a perfect match.

[0050] Specifically, the system can calculate the initial match using an edit distance algorithm (an algorithm for calculating the degree of difference between two sequences). The algorithm measures the difference by calculating the minimum number of single-point edit operations (including insertion, deletion, and replacement) required to transform the test sequence into the reference sequence. The initial match can be expressed using a formula. Edit distance The calculation yields an edit distance, where the edit distance is the result of the algorithm, and len(SeqA) and len(SeqB) are the lengths of the test sequence and the reference sequence, respectively. For example, the test sequence is "{0xC1,0x02,0xC0,0x2F,0xC1,0x01}", with a length of 3 (in kit pairs); the reference sequence is "{0xC1,0x02,0xC1,0x01,0xC0,0x2F}", also with a length of 3. The calculated edit distance is 2 (requiring the swapping of the last two elements). Substituting into the formula... .

[0051] Based on the initial matching score, and combined with the extended field type set and the cryptographic library version identifier, the initial matching score is corrected and calculated to obtain the comprehensive matching score.

[0052] The overall matching score is the final matching score obtained by taking into account the compatibility of extended fields and the rationality of version identifiers, based on the initial matching score.

[0053] Specifically, the system can correct this using a weighted summation method. First, calculate the overlap ratio of the extended fields. (That is, the number of extension types shared by the test message and the corresponding variant of the baseline sequence, divided by the total number of extension types of the baseline variant). Next, it is determined whether the cryptographic library version identifier falls within the version range allowed by the corresponding variant of the baseline sequence, thus obtaining the version matching flag. (A match is 1, otherwise 0). Finally, according to the formula... Calculate the overall matching degree, where and These are preset weighting coefficients used to balance the impact of different features on the final score. For example, continuing from the previous example, The value is 0.333. Assuming the extended type set of the baseline variant is {0x00, 0x0d, 0x10}, and the extended type set of the message under test is {0x00, 0x0d}, then... If the version being tested is identified as "GmSSL3.0.0" within the range of versions allowed by the benchmark variant. Inside, then =1. Let ,but .

[0054] The overall matching degree is compared with a preset tolerance threshold to determine the kit arrangement matching level.

[0055] Among them, the tolerance threshold includes a comprehensive matching degree threshold used to classify the levels (such as a high matching threshold). Low matching threshold ), and a position offset tolerance threshold used to determine whether minor changes in the kit's position are acceptable. .

[0056] Specifically, the system can determine the level through multi-level judgment logic. First, the calculated comprehensive matching degree... and and Comparison. Simultaneously, calculate the average positional offset of the major suites (e.g., the first N) between the test sequence and the reference sequence. and with Comparison. If and If it is, it is judged as a "high match" level; if If it is, it is judged as "moderately consistent"; if or If so, it is judged as a "low fit" level. For example, suppose... (Position). Continuing from the previous example, The value is between 0.60 and 0.85. Calculate the average positional offset of the first three kits, assuming it to be 1.5. Since... and However, 0.733 < 0.85, so the system ultimately determined that the kit arrangement matching level of this traffic was "medium matching".

[0057] Therefore, according to the above implementation method, the system can dynamically and meticulously evaluate the degree of consistency between each encrypted traffic and the legitimate behavior pattern through multi-level and multi-feature quantitative calculation and rule judgment, providing a reliable and interpretable basis for subsequent accurate focus on suspicious traffic.

[0058] In some embodiments, the overall matching degree is compared with a preset tolerance threshold to determine the kit arrangement fit level, including: The overall matching degree is matched with preset high and low matching thresholds to make a matching decision.

[0059] Among them, the matching decision refers to the logical operation of making a preliminary classification of traffic legality based on the comparison result of the quantitative score and the preset threshold; the high matching threshold is a higher score threshold used to define traffic that is highly consistent with the legal behavior pattern; the low matching threshold is a lower score threshold used to distinguish suspicious traffic that is significantly different from the legal behavior pattern.

[0060] Specifically, the system makes decisions through comparison logic. The system reads the overall matching score calculated for the traffic to be tested and compares this score with pre-stored high-matching threshold values ​​and low-matching threshold values. Based on the comparison results, the system can generate a preliminary decision intention, such as "suspected high match," "suspected medium match," or "suspected low match." For example, setting the high-matching threshold to 0.85 and the low-matching threshold to 0.60, for a traffic to be tested with an overall matching score of 0.73, the system determines through comparison that 0.73 is greater than 0.60 but less than 0.85, thus generating a preliminary decision intention of "suspected medium match."

[0061] Tolerance determination is performed on the positional offset of sequence elements between the cipher suite permutation sequence and the reference sequence.

[0062] Among them, the sequence element position offset refers to the absolute value of the difference between the index positions of the same cipher suite encoding value in the test sequence and the reference sequence during the comparison; the tolerance judgment refers to the operation of determining whether the position offset is within the reasonable deviation range allowed by the system.

[0063] Specifically, the system can determine tolerance by iteratively calculating and comparing with a threshold. First, the system aligns and compares the sequence to be tested with the reference sequence to identify common cipher suites. For each common suite, the absolute value of the difference between its index in the sequence to be tested and its index in the reference sequence is calculated to obtain the positional offset of a single suite. Then, the average positional offset of all common suites can be calculated, or the maximum positional offset can be taken. Finally, this calculation result is compared with a preset offset tolerance threshold. If the calculation result is less than or equal to the threshold, it is determined as "tolerance passed"; otherwise, it is determined as "tolerance failed". For example, the sequence to be tested is... The baseline sequence is There are four kits: A, B, C, and D. Calculations show that the positional offsets of kits A and B are both 0, while the positional offset of kit C is... The position offset of kit D is The average position offset is If the preset offset tolerance threshold is 2 (positions), then The system determines that the sequence position offset of the flow is "tolerance-passed".

[0064] Based on the results of the matching decision and tolerance determination, the kit arrangement matching level is determined.

[0065] Among them, the suite alignment matching level is the final qualitative conclusion made on the overall consistency between traffic handshake characteristics and legitimate benchmarks after integrating evidence from both matching scores and location tolerances.

[0066] Specifically, the system can determine the final level using a decision table or conditional logic rules. The level determination rules can be designed as follows: (1) If the matching judgment is "suspected high match" and the tolerance judgment is "pass", then the final level is determined to be "high match"; (2) If the matching judgment is "suspected moderate match" and the tolerance judgment is "pass", then the final grade is determined to be "moderate match"; (3) If the matching decision is "suspected low match" or the tolerance decision is "failed", then the final level is determined to be "low match". For example, continuing from the previous two examples. For this traffic, the matching decision result is "suspected medium match" (due to the overall matching degree of 0.73), and the tolerance decision result is "passed" (due to the average offset). Based on the above rules, the system ultimately determined the suite arrangement matching level of this traffic to be "medium matching".

[0067] In summary, based on the above implementation methods, the system can form a dual verification mechanism by combining quantified feature matching degree with structured location tolerance analysis, thereby more robustly and accurately classifying the legitimacy of encrypted traffic, effectively reducing misjudgments caused by fluctuations in a single indicator or special permutation variations, and providing a more reliable classification basis for subsequent processing.

[0068] In some embodiments, based on the matching level of the suite arrangement, messages to be identified that have a matching level lower than a preset legal threshold are filtered out, and a set of objects to be identified is generated, including: Read the suite arrangement matching level of each message from the handshake fingerprint dataset.

[0069] The handshake fingerprint dataset is a structured dataset that stores the complete features of each handshake message (such as the sequence of cipher suite arrangements, the set of extended field types, and the cipher library version identifier) ​​and the calculated suite arrangement matching level after processing in steps S110 and S120.

[0070] Specifically, the system can sequentially read the "suite alignment match level" field value of each record by traversing the handshake fingerprint data table in memory or a database. This match level field is typically stored as a discrete enumeration value, such as the strings "high match," "medium match," and "low match," or the corresponding numeric codes 1, 2, and 3. For example, the system sequentially reads records from the handshake fingerprint data table. The first record has a match level field value of "high match," the second "medium match," and the third "low match."

[0071] The kit arrangement matching level is compared with a preset valid threshold.

[0072] The preset legal threshold is a decision-making boundary point, pre-set by the system administrator or based on historical traffic analysis. Messages falling below this threshold are considered suspicious and require further in-depth analysis. This threshold is typically set at the boundary between "medium match" and "low match," with the aim of filtering out traffic that significantly deviates from the normal baseline.

[0073] Specifically, the system can achieve this through simple logical comparison operations. The system compares the matching level of each message it reads with the legal threshold levels stored in the configuration. The judgment logic is: if the matching level value of a message (assuming "high matching" = 3, "medium matching" = 2, "low matching" = 1) is less than the legal threshold value, then the message meets the filtering conditions. For example, suppose the system's preset legal threshold is "medium matching" (corresponding to a value of 2). When the system reads a message with a matching level of "low matching" (corresponding to a value of 1), since 1 is less than 2, the comparison result is "below the threshold"; when it reads a message with a level of "medium matching" (value 2), 2 equals 2, and the comparison result is "not below the threshold".

[0074] Messages whose suite arrangement matching level is lower than the preset legal threshold are marked as objects to be identified, and the session identifier and source address information of the objects to be identified are extracted to generate a set of objects to be identified.

[0075] The objects to be identified are data structures that carry the core identification information of suspicious messages; the session identifier in network communication refers to the information that uniquely identifies a transport layer connection. For TCP (Transmission Control Protocol), it is usually a four-tuple (source IP address, source port, destination IP address, destination port), while for encrypted communication based on UDP (User Datagram Protocol), it may be a five-tuple (adding protocol type); the source address information here specifically refers to the network layer address of the client that initiated the encrypted handshake connection, which is a 128-bit IPv6 address in the IPv6 environment; the set of objects to be identified is a list or array of all marked objects to be identified, used for subsequent batch processing.

[0076] Specifically, the system generates the set through the following steps: First, for each message that meets the "below the threshold" condition, an empty object data structure is created to be verified. Then, the "session identifier" field (such as a TCP quadruple) and the "source address" field (IPv6 address) are extracted from the handshake fingerprint data record corresponding to the message and filled into the created object to be verified. Simultaneously, auxiliary information such as a timestamp and original fingerprint feature index can be attached to the object. Finally, this completed object is appended to a global list of objects to be verified. After all messages have been traversed and compared, this list is the final set of objects to be verified. For example, if the system processes 1000 messages, and 50 of them have a matching level below the "medium matching" threshold, the system creates objects for each of these 50 messages. Taking one record as an example, the session identifier is extracted as {source IP:2001:db8::1, source port:54322, destination IP:2001:db8::80, destination port:443}, and the source address is 2001:db8::1. The system populates this information into an object and adds this object to a list. Finally, the system generates a list containing 50 objects to be identified, i.e., the set of objects to be identified.

[0077] Therefore, according to the above implementation method, the system can efficiently and accurately filter out suspicious sessions from massive encrypted traffic based on pre-calculated fine matching levels, and organize the key network identification information of these sessions in a structured manner, providing clear and specific input targets for subsequent deep matching verification and threat clustering analysis, thereby improving the automation level and analysis efficiency of the entire detection process.

[0078] In some embodiments, the handshake fingerprint features in the set of objects to be identified are matched and verified against a known database of spoofing traffic features, and cluster analysis is performed based on the permutation feature vectors of the set of objects to be identified to identify a subset of covert spoofing traffic from the same source, including: The handshake fingerprint features in the target set are matched and verified against the spoofing traffic feature library to filter out potential spoofing communication components.

[0079] Among them, potential spoofing communication components refer to objects to be identified in the initial matching verification whose handshake fingerprint features meet the preset similarity requirements with a certain record in the spoofing traffic feature database. The network session corresponding to this object is considered to be a communication instance initiated by a certain malicious software or attack tool.

[0080] Specifically, the system can achieve screening through one-by-one comparison and threshold judgment. The system traverses each object in the set of objects to be screened, extracting feature tuples for fine-grained comparison from the detailed handshake fingerprint data associated with that object. Feature tuples may include: (1) An ordered subsequence of a sequence of cipher suites (e.g., the first 6 suites); (2) The order of the expanded field list; (3) Feature values ​​in specific extensions (such as the server_name extension). The system compares this feature tuple with the feature template of each record in the spoofing traffic feature library and calculates the feature similarity. If the feature similarity between an object and any record in the feature library exceeds the preset matching threshold, the object is marked as a "potential spoofing communication component". For example, for an object to be identified (source IP: 2001:db8:cafe::1a, destination port: 443), its handshake fingerprint feature tuple is: kit subsequence , expansion order The feature similarity is compared with the feature template recorded in the "Golang_C2_v2.1" counterfeit library, and the calculated feature similarity is 0.91. Assuming the matching threshold is 0.85, since 0.91 > 0.85, the system classifies the object as a "potential counterfeit communication component".

[0081] Cluster analysis is performed on the permutation feature vectors of potential counterfeit communication components to identify homogeneous groups that exhibit consistent permutation features.

[0082] Among them, the permutation feature vector is a fixed-length numerical vector that is converted into a permutation sequence of cipher suites through mathematical encoding (such as one-hot encoding, n-gram frequency statistics). It is used to quantify the sequence features so as to make mathematical similarity comparisons. The homogeneous group refers to a set of potential spoofing communication components that are divided by clustering algorithms and whose members have highly similar permutation feature vectors. Traffic within the same homogeneous group is likely to originate from the same attack entity or the same malware family.

[0083] Specifically, the system performs clustering analysis through the following steps: First, for all objects marked as "potential spoofing communication components," their original cipher suite permutations are converted into permutation feature vectors. Then, a suitable clustering algorithm (such as Density-Based Spatial Clustering of Applications with Noise, or DBSCAN for short) is used to analyze these feature vectors. This algorithm groups density-connected points in the vector space into clusters and filters out noise points (i.e., isolated points that cannot be assigned to any high-density region). Finally, the system outputs the clustering results; each cluster identified by the algorithm that contains at least two members (MinPts parameter) is identified as a homogeneous group. For example, suppose there are 15 potential spoofing communication components. The system inputs their respective feature vectors (e.g., 128-dimensional vectors) into the DBSCAN algorithm. The algorithm parameters are set as follows: neighborhood radius (eps) of 0.3 and minimum number of points (MinPts) of 2. After execution, the algorithm outputs three clusters (homogeneous groups): cluster A contains 7 components, cluster B contains 5 components, and cluster C contains 3 components. The average cosine similarity among members within these clusters is higher than 0.9, while the average similarity among members of different clusters is lower than 0.3.

[0084] Traffic belonging to the same source group is identified as a subset of covert spoofing traffic.

[0085] The concealed spoofing traffic subset is the final analysis result, which refers to the set of original network traffic data corresponding to all potential spoofing communication components that are determined to belong to the same source group, representing a batch of related traffic generated by the same attack activity.

[0086] Specifically, the system can determine subsets through mapping operations. After obtaining the results of the same-source groups (clusters), the system reverse-engineers all "potential spoofing communication component" objects contained within each same-source group, based on the object index used during clustering. Then, it extracts the session identifiers (such as 5-tuples) and timestamps initially used to identify the original traffic from these objects. The system packages these session identifiers and timestamps to form a structured description of a "covert spoofing traffic subset." Each subset can be associated with an attack activity ID for easy subsequent tracking and handling. For example, continuing the previous example, for the identified "same-source group A" (containing 7 components), the system uses mapping to find the original session identifiers corresponding to these 7 components, such as: {source IP:2001:db8:cafe::1a, destination port:443}, {source IP:2001:db8:cafe::2b, destination port:443}, and so on. The system aggregates these session identifiers and their respective timestamp ranges (e.g., "2023-10-26 14:00:00 to 2023-10-26 15:30:00"), generates a record, marks it as "Covered Spoofing Traffic Subset_001", and records its corresponding attack activity ID as "Campaign_Alpha".

[0087] Therefore, according to the above implementation method, the system can, based on the initial screening of suspicious traffic, further focus on high-threat targets through secondary feature matching, and automatically discover attack traffic with highly similar fingerprint characteristics using an unsupervised clustering algorithm. This not only enables the detection of individual attack instances, but also the batch discovery and attribution of related traffic under the same attack activity, thereby improving the depth and breadth of threat discovery and providing more comprehensive contextual information for network security incident response.

[0088] In some embodiments, messages with a suite arrangement matching level below a preset legal threshold are marked as objects to be screened, and the session identifier and source address information of the objects to be screened are extracted to generate a set of objects to be screened, including: The session identifier and source address information are extracted from the handshake fingerprint record corresponding to the message to be verified.

[0089] The handshake fingerprint record is a data structure that stores the complete processing result of a network handshake message, typically including the original features, the calculated matching level, and related network layer and transport layer metadata.

[0090] Specifically, the system can perform parsing through a query operation. Each handshake fingerprint record has a unique record identifier in memory or a database. The system finds the corresponding complete handshake fingerprint record based on the record identifier of the message identified as "to be verified". Then, it reads the values ​​of the "session identifier" and "source address" fields from the predefined fields of that record. The values ​​of these fields are extracted and stored during the initial message parsing (step S110). For example, the handshake fingerprint record ID corresponding to a message to be verified is FPR_1023. The system queries this record and reads that the value of its "session identifier" field is {protocol:TCP, source IP:2001:db8:85a3::8a2e:370:7334, source port:55022, destination IP:2001:db8:3c4d::15, destination port:443}, and the value of its "source address" field is 2001:db8:85a3::8a2e:370:7334.

[0091] Based on the parsed session identifier and source address information, an identifier entry for the object to be identified is constructed.

[0092] The identifier entry is a standardized data structure specifically designed for subsequent analysis processes. It encapsulates the core identifier information of a traffic to be identified and the necessary contextual metadata.

[0093] Specifically, the system can construct entries by instantiating a predefined `SuspectedObject` class or structure. The member variables of this data structure include at least: `session_id` (session identifier), `src_addr` (source address), and `detection_timestamp` (detection timestamp). The system assigns the session identifier and source address information parsed from the handshake fingerprint record to the corresponding members. Simultaneously, the system obtains the current system time and uses it as the value of `detection_timestamp` to record the moment the object was detected. For example, the system creates a new `SuspectedObject` instance. The session identifier `{TCP,2001:db8:85a3::8a2e:370:7334,55022,2001:db8:3c4d::15,443}` is assigned to the instance's `session_id` member, and the source address `2001:db8:85a3::8a2e:370:7334` is assigned to the `src_addr` member. The current system time is obtained as 2023-10-27T10:30:15Z and assigned to the detection_timestamp member. At this point, a complete identifier entry is constructed.

[0094] Add the completed identifier entries to the list to be identified, and output the list to be identified as a set of objects to be identified.

[0095] The list to be identified is a linear data structure dynamically maintained in memory, used for temporary storage and orderly management of all the identification entries constructed in this analysis cycle; the set of objects to be identified is the final output of this step, referring to the list containing all the identification entries after all the messages to be identified have been processed.

[0096] Specifically, the system adds entries using list append operations. The system maintains a global or thread-safe list object (e.g., `list` in Python, or `ArrayList` in Java). Once an entry is constructed, the system calls the `append()` or `add()` method of the list to add it to the end. After processing all messages with a match level below the threshold, this list contains all traffic identifiers requiring further analysis. The system returns this list as a result or passes it to subsequent processing modules; this list is considered the generated set of objects to be analyzed. For example, in one analysis window, the system has tagged 150 messages to be analyzed. The system constructs an entry for each message and adds it sequentially to an initially empty `ArrayList`. <suspectedobject>After processing, the ArrayList contains 150 elements. The system outputs this ArrayList object as the generated set of objects to be identified and passes it to the matching and verification module in step S140.

[0097] Therefore, according to the above implementation method, the system can automatically and structurally complete the transformation from suspicious message identification to a set of processable data objects. Through standardized data parsing, entry construction, and aggregation, it ensures that every suspicious traffic is accurately and completely identified and organized, providing a uniformly formatted and comprehensive input for subsequent in-depth threat analysis processes, and guaranteeing the smoothness of the analysis process and the traceability of the results.

[0098] In some embodiments, the handshake fingerprint features in the set of objects to be identified are matched and verified against a spoofing traffic feature library to filter out potential spoofing communication components, including: Obtain the spoofing feature records from the spoofing traffic feature library and perform feature matching calculations with the handshake fingerprint features of the object to be identified.

[0099] Feature matching calculation refers to the process of quantifying the similarity between the handshake fingerprint features to be tested and the counterfeit feature records through an algorithm. The calculation result is usually a score between 0 and 1, with a higher score indicating a higher similarity.

[0100] Specifically, the system performs the calculation through the following steps: First, for each counterfeit feature record, its feature vector is pre-extracted and stored. This vector is encoded by key fragments of the cipher suite arrangement order, a list of extension types, and hash values ​​of specific extensions. Then, the system converts the handshake fingerprint features of the object to be identified into a feature vector to be matched according to the same rules. Next, the system calculates the cosine similarity (a metric measuring the similarity between the directions of two vectors) between these two feature vectors. This calculated cosine similarity value is used as the result of this feature matching calculation. For example, the feature vector of the counterfeit feature record "CobaltStrike_Beacon_4.3" in the counterfeit feature database is... (Total 256 dimensions). The transformed feature vector of the object to be identified, "Obj_47", is: The system calculates the cosine similarity between the two vectors, obtaining a result of 0.92.

[0101] The version identification of the cryptographic database of the object being investigated is verified to determine whether the version identification falls within the scope of the spoofing feature record definition.

[0102] The spoofing range is a version number interval defined in the spoofing feature record. This interval describes the set of cryptographic library version numbers that the spoofing tool typically declares during the handshake phase to impersonate legitimate software. The version numbers may be abnormally old, abnormally advanced, or clearly inconsistent with the release patterns of legitimate software.

[0103] Specifically, the system can perform verification through interval comparison. The spoofing range stored in each spoofing feature record can be represented as a closed interval [Version_low, Version_high] or a set of multiple discontinuous intervals. The system reads the cryptographic version identifier string of the object to be verified and converts it into a comparable numerical or tuple form according to the version number parsing rules (such as "major version.minor version.revision number"). Then, it determines whether this converted version falls within any of the spoofing intervals defined by the spoofing feature record. If it does, the verification result is "passed"; otherwise, it is "failed". For example, the spoofing feature record "CobaltStrike_Beacon_4.3" defines the spoofing range as ["OpenSSL1.0.1", "OpenSSL1.0.2t"]. The cryptographic version identifier of the object to be verified, "Obj_47", is "OpenSSL1.0.2k". After analysis and comparison, "1.0.2k" is between "1.0.1" and "1.0.2t", so it is determined that the version identifier falls within the scope of spoofing, and the verification result is "passed".

[0104] Based on the results of feature matching calculation and version spoofing verification, a comprehensive judgment is made on the objects to be identified in order to filter out potential counterfeit communication components.

[0105] The comprehensive judgment is a decision-making process that makes a final binary judgment (yes or no) on whether the object to be identified is a counterfeit component based on pre-set logical rules, combined with feature matching degree and version verification results.

[0106] Specifically, the system can achieve this through a preset decision threshold and a logical "AND" operation. The system sets a decision threshold for feature matching (e.g., 0.85). For each object to be identified, if its feature matching score is greater than or equal to this decision threshold (e.g., ...), the system will determine the match. If the version spoofing verification result is "passed," the system's overall judgment for the object is "yes," thus filtering it as a potential counterfeit communication component. If either of the above two conditions is not met, the overall judgment is "no," and the object is excluded in this round of verification. For example, for the object to be identified, "Obj_47," its feature matching degree calculation result is 0.92, and the version spoofing verification result is "passed." The system's preset matching degree judgment threshold is 0.85. The verification passed, so the system's overall judgment for "Obj_47" was "yes", and the object was marked as a potential spoofing communication component.

[0107] Therefore, according to the above implementation method, the system can significantly improve the accuracy of screening by combining static fingerprint similarity (feature matching) and dynamic claim rationality (version verification). This effectively reduces misjudgments caused by coincidence of a single feature or accidental normal version number, ensuring that the screened potential counterfeit communication components have a higher threat confidence, and laying a reliable foundation for subsequent accurate clustering and source tracing analysis.

[0108] In some embodiments, the step of performing version spoofing verification on the cryptographic library version identifier of the object to be verified includes: Retrieve a list of historical cryptographic library version identifiers for known legitimate client variants.

[0109] The historical cryptographic library version identifier list is an ordered list of strings that records all cryptographic library version strings actually observed and collected from handshake messages of various legitimate client software (such as different versions of browsers) in a controlled experimental environment or long-term production environment.

[0110] Specifically, the system can obtain this list by reading a pre-built configuration file or querying a feature database. This list is typically built by security researchers before system deployment through proactive scanning, vendor collaboration, or long-term traffic log analysis, and is integrated into the system as part of a static knowledge base. For example, the system reads a list of historical cryptanalyst version identifiers from the configuration file `legitimate_versions.conf`. The first few records are: ["GmSSL3.0.0", "GmSSL 3.0.1", "GmSSL 3.0.2", "BabaSSL 2.1.0", "BabaSSL 2.1.5"], and the list contains a total of 150 different version identifiers.

[0111] Based on the historical cryptographic database version identifier list, a preset range of version spoofing identifiers is determined through statistical analysis.

[0112] The version spoofing identifier is a range of one or more version numbers used to identify which cryptographic library version identifiers are statistically highly unlikely to be generated by legitimate client software, and are therefore more likely to be abnormal version numbers used by attackers for spoofing. Its core principle is the assumption that the version numbers of legitimate software are continuous or concentrated in time or version sequence, while version numbers outside this concentrated range are more suspicious.

[0113] Specifically, the system can determine the range of spoofing by calculating the central tendency and dispersion of historical versions. A common implementation is as follows: First, each version number in the historical version identifier list is parsed into a comparable numerical value (e.g., converting "3.0.2" to the numerical value 302). Then, the arithmetic mean μ (mu) and standard deviation σ (sigma) of these values ​​are calculated. Finally, the range of values ​​is defined as follows: The range outside this range is defined as the version spoofing identifier range, where n is a preset multiple (e.g., 2 or 3) used to control the strictness of the range. This means that if the value corresponding to a version number falls within this range, it is considered "common"; if it falls outside this range, it is considered "abnormal / spoofed". For example, suppose the mean of the values ​​after parsing the historical version list is 300 and the standard deviation is 5. Taking n=2, the calculated common range is... ,Right now Therefore, the default version spoofing identifier range is a numerical range of less than or equal to 290 and greater than or equal to 310. The corresponding version number may include "GmSSL 2.8.0" (value 280) or "GmSSL 3.3.0" (value 330).

[0114] The password database version identifier of the object to be identified is matched and judged against the preset version disguise identifier range.

[0115] Among them, the matching judgment is to perform a simple set attribute check to determine whether the version identifier to be verified falls within the preset abnormal range.

[0116] Specifically, the system can make judgments using a range comparison algorithm. First, the system parses the cryptographic library version identifier of the object to be verified (e.g., "OpenSSL 1.1.0a") into a numerical value according to the same rules as the statistical analysis described above. Then, the parsed numerical value is compared with a pre-calculated and stored range of version spoofing identifiers (i.e., one or more numerical ranges) in the system. If the numerical value falls within any of the preset abnormal ranges, the result of the match judgment is "yes" (i.e., determined to be spoofing); otherwise, the result is "no". For example, the cryptographic library version identifier of the object to be verified is "GmSSL 2.8.5", which parses into the numerical value 285. The system's preset range of version spoofing identifiers is a range of numerical values... or numerical value Since 285 is less than or equal to 290, it falls within this range, so the matching result is "yes", and the system determines that this version identifier is suspected to be a fake version.

[0117] Therefore, according to the above implementation method, the system can automatically identify abnormal cryptographic library version declarations that deviate from the normal distribution in a statistical sense based on data-driven analysis of the legitimate ecosystem version distribution. This method can effectively detect the suspicions revealed by attackers during the impersonation handshake due to arbitrary fabrication or use of outdated / advanced version libraries without prior knowledge of the characteristics of all malicious samples, thus enhancing the intelligence and coverage of the version verification process.

[0118] In some embodiments, a comprehensive judgment is made on the object to be identified based on the results of feature matching calculation and version spoofing verification to filter out potential counterfeit communication components, including: Receive the feature matching metric generated by feature matching calculation and the verification status generated by version spoofing verification.

[0119] The feature matching metric is a specific numerical value representing the similarity between the handshake fingerprint features of the object to be verified and a certain record in the spoofing traffic feature database. It is usually expressed as a floating-point number, ranging from 0.0 to 1.0, where 1.0 represents a perfect match. The verification status is a Boolean value (true / false) or an enumerated value (such as "pass" or "fail"), indicating whether the cryptographic database version identifier of the object to be verified is determined to be a spoofed version.

[0120] Specifically, the system can receive these two inputs via function calls or message passing. After completing the calculation, the feature matching calculation module outputs the calculated similarity score (e.g., 0.87) and the matched spoofing feature record ID (e.g., "APT_Tool_X_v1"). After completing the judgment, the version spoofing verification module outputs a status code (e.g., 1 for "pass" and 0 for "fail"). The system binds these outputs with the unique identifier of the object to be identified and passes them to the comprehensive judgment module. For example, for the object to be identified "Session_0x8A2F", the system receives a feature matching metric of 0.91 and a version spoofing verification status of "pass" (status code 1).

[0121] Based on preset decision rules, logical decisions are made on feature matching metrics and verification status.

[0122] The decision rules are a predefined set of conditional logics used to map quantified matching degrees and binary verification states to a final judgment result regarding whether the object is a potential spoofing communication component. Logical decision is the process of executing this set of conditional logics.

[0123] Specifically, the system can be implemented using an if-else statement or decision table that includes threshold comparisons and logical operators. A typical decision rule is: if the feature matching metric is greater than or equal to a preset matching decision threshold (e.g., 0.85), and the verification status is "passed," then the logical decision output is "yes" (determined as a potential threat); otherwise, the output is "no." For example, the system's preset matching decision threshold is 0.85. For the object "Session_0x8A2F," its feature matching metric... The verification status is "passed". According to the "AND" logic, the output result after the system executes the logical decision is "yes".

[0124] When the output of the logical decision meets the preset component identification conditions, an identifier entry for the potential counterfeit communication component is generated.

[0125] The component identification condition is the specific value that the logical decision output must meet, usually a "yes" result. An identifier entry is a structured data object used to uniquely identify and record a network session identified as a potential spoofing communication component. It typically includes the session's core network identifier, the basis for the determination, and associated threat intelligence identifiers.

[0126] Specifically, the system can trigger the generation process of the identifier entry when the logical decision output is "yes". The system will create a new data structure and fill in the following information: (1) The original session identifier (such as TCP 5-tuple) of the object to be identified; (2) The ID of the matched counterfeit feature record (e.g., "APT_Tool_X_v1"); (3) Feature matching metric (e.g., 0.91); (4) Determine the timestamp. This filled data structure is an identifier for a potential counterfeit communication component.

[0127] For example, when the logical decision for "Session_0x8A2F" outputs "Yes", the system generates an identifier entry with the following content: {Session ID:{Source IP:2001:db8::1,Source Port:55555,Destination IP:2001:db8::80,Destination Port:443},Matching Feature ID:"CobaltStrike_4.7",Matching Degree:0.91,Timestamp:2023-10-28T14:30:00Z}.

[0128] Therefore, according to the above implementation method, the system can automatically and comprehensively make decisions based on evidence from two dimensions: feature similarity analysis and version declaration rationality verification, through explicit and configurable decision rules. This avoids the limitations of a single indicator, forming a reliable filter that ensures only those objects that are highly suspicious in both behavioral characteristics and attribute declarations are officially marked as potential spoofing communication components, thereby improving the accuracy and reliability of threat screening results.

[0129] In another embodiment, such as Figure 2 This illustration demonstrates the implementation process of end-to-end identification and interception of covert APT attacks in IPv6 encrypted traffic based on handshake fingerprint characteristics. To more intuitively illustrate the actual operating logic and collaborative mechanism of this solution, this embodiment constructs a specific practical application scenario. The scenario is set as follows: Suppose that in the core office network of a large energy company, the security operations center captures a batch of abnormal outbound traffic. This traffic all points to a suspicious domain name overseas, and traditional detection engines based on plaintext features or static blacklists have failed to identify any obvious threat. This solution integrates with the traffic auditing platform to perform deep probing on this batch of IPv6 encrypted traffic (using the TLS 1.3 protocol). The specific implementation process is detailed in the attached document. Figure 2 The steps shown are explained below: S101: National Cryptographic Handshake Message Acquisition and Initial Data Construction The traffic auditing platform is deployed as a bypass listening node at the network boundary. The system uses Deep Packet Inspection (DPI) technology to accurately capture all ClientHello messages sent to external servers from massive amounts of IPv6 network traffic. In this process, the system extracts the order of the cipher suites carried in each message (e.g., The system extracts features such as the combination of extension fields and the cryptographic library compilation version identifier (e.g., GmSSL1.0.0). These extracted features are then structured and encapsulated to form an initial handshake fingerprint dataset, which serves as the foundation for subsequent analysis.

[0130] S102: Sequence alignment and legality compliance assessment For the initial handshake fingerprint data set generated in the previous step, the system uses a sequence alignment algorithm to perform a bit-by-bit matching calculation between the cipher suite arrangement of each ClientHello message and a preset range of legitimate browser variants (such as common permutations in mainstream browsers like Chrome, Firefox, and Edge). Simultaneously, the system also integrates the combination of extended fields and the cipher library compilation version identifier for a comprehensive comparison. For example, if the cipher suite arrangement of a certain traffic declaration is extremely rare and has never appeared in the standard variant library of legitimate browsers, the system calculates and determines that the cipher suite arrangement match level of this message is low (e.g., a match score of only 0.3, far below the legitimate threshold of 0.8).

[0131] S103: Cross-validation and Potential Component Locking Based on the suite matching level determined in S102, the system filters out messages with a matching level below a preset legal threshold (e.g., 0.8) and marks them as objects to be investigated. Subsequently, for these objects to be investigated, the system matches their handshake fingerprint characteristics against an internally maintained APT implant spoofing traffic signature database (containing specific fingerprint templates of known malware families such as CobaltStrike and Metasploit). Simultaneously, the system performs cross-validation using the cryptographic library's compiled version identifier. For example, if the system finds that the cryptographic suite characteristics of an object to be investigated are highly similar to the "CobaltStrike_v4.8" template in the APT library, and its declared version identifier GmSSL 1.0.0 falls within the system's predefined abnormal spoofing version range (e.g., less than 1.0.2), the system determines that the communication component is highly suspicious, locks it as a potential spoofing communication component, and determines its specific identifier list in the network.

[0132] S104: Fine-sequence similarity and difference analysis and homology clustering After identifying a series of potential spoofing communication components, the system performs clustering analysis on the handshake fingerprints corresponding to each identifier in the list to determine subtle differences in arrangement, in order to identify related traffic hidden under the same attack campaign. By quantitatively analyzing the subtle permutation differences of the cipher suites (for example, some components are generally identical, but there are subtle swaps in the last few weak cipher suites), the system groups identifiers that exhibit highly consistent permutation characteristics into the same source group. For example, after clustering analysis, the system found that the handshake features of 50 potential spoofing communication components clustered in the same high-density region after dimensionality reduction. These components exhibited extremely strong homology characteristics, and the system thus identified this as a subset of covert spoofing traffic generated by a group of APT attacks.

[0133] S105: Targeted Interception and Strategy Distribution Finally, based on the subset of spoofed traffic identified by S104, the system extracts the core characteristics of each traffic within the subset (such as source address, specific cipher suite combinations, and extended field structures) and session identifiers. The system transforms these extracted characteristics into precise blocking rules and writes them into the dynamic blocking rule table of the traffic auditing platform. Once new traffic matching these characteristics reappears in the network, the platform will directly drop or reset the connection, thereby achieving precise and targeted interception of APT spoofing communication within IPv6 encrypted traffic and effectively curbing the risk of data leakage.

[0134] According to a second aspect, the present invention provides an IPv6 encrypted traffic advanced persistent threat attack identification system, including a processor and a memory. The memory stores a computer program, which, when executed by the processor, implements any of the IPv6 encrypted traffic advanced persistent threat attack identification methods in the embodiments of the present invention.

[0135] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0136] According to embodiments of the present invention, the above-described method of the present invention can be applied to an electronic device and a readable storage medium.

[0137] Figure 3 A schematic block diagram of an electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0138] like Figure 3 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0139] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0140] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a method for identifying advanced persistent threat attacks (APS) on IPv6 encrypted traffic. For example, in some embodiments, a method for identifying APS on IPv6 encrypted traffic can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of a method for identifying APS on IPv6 encrypted traffic described above can be performed. Alternatively, in other embodiments, computing unit 601 may be configured by any other suitable means (e.g., by means of firmware) to perform an advanced persistent threat attack identification method for IPv6 encrypted traffic.

[0141] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0146] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0147] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.< / suspectedobject>

Claims

1. A method for identifying advanced persistent threat attacks on IPv6 encrypted traffic, characterized in that, include: Extract the cipher suite arrangement sequence, extended field type set, and cipher library version identifier from the transport layer handshake message of the IPv6 encrypted link; Based on the cipher suite arrangement sequence, the extended field type set, and the cipher library version identifier, a sequence comparison and tolerance determination are performed with a preset legitimate client variant fingerprint library to determine the suite arrangement matching level; The tolerance determination refers to the operation of judging whether the position offset is within the reasonable deviation range allowed by the system, specifically including: Obtain the reference sequence corresponding to the permutation sequence of the cipher suite; The sequence alignment calculation is performed between the cipher suite permutation sequence and the baseline sequence to obtain the initial matching degree; Based on the initial matching degree, and combined with the extended field type set and the cryptographic library version identifier, the initial matching degree is corrected and calculated to obtain the comprehensive matching degree; The overall matching degree is compared with a preset tolerance threshold to determine the suitability level of the kit arrangement; specifically including: The overall matching degree is matched with preset high matching threshold and low matching threshold for matching judgment; Tolerance determination is performed on the positional offset of sequence elements between the cipher suite permutation sequence and the reference sequence; Based on the results of the matching decision and the tolerance determination, the kit arrangement matching level is determined; Based on the matching level of the kit arrangement, messages with matching levels lower than a preset legal threshold are filtered out to generate a set of objects to be identified; The handshake fingerprint features in the set of objects to be identified are matched and verified with a known database of spoofing traffic features. Cluster analysis is then performed based on the permutation feature vector of the set of objects to be identified to identify a subset of hidden spoofing traffic from the same source.

2. The method according to claim 1, characterized in that, The step of arranging the matching levels according to the suite, filtering out messages to be identified whose matching levels are lower than a preset legal threshold, and generating a set of objects to be identified includes: The suite arrangement matching level of each message is read from the handshake fingerprint data set; The kit arrangement matching level is compared with the preset valid threshold; Messages whose suite arrangement matching level is lower than the preset legal threshold are marked as objects to be identified, and the session identifier and source address information of the objects to be identified are extracted to generate the set of objects to be identified.

3. The method according to claim 2, characterized in that, The step of matching and verifying the handshake fingerprint features in the set of objects to be identified with a known database of spoofing traffic features, and performing cluster analysis based on the permutation feature vectors of the set of objects to be identified to identify a subset of hidden spoofing traffic from the same source, includes: The handshake fingerprint features in the set of objects to be identified are matched and verified with the spoofing traffic feature library to filter out potential spoofing communication components. Cluster analysis is performed on the arrangement feature vectors of the potential counterfeit communication components to identify homogeneous groups that exhibit consistent arrangement features; Traffic belonging to the same source group is identified as the subset of the hidden spoofing traffic.

4. The method according to claim 2, characterized in that, The step of marking messages whose suite arrangement matching level is lower than the preset legal threshold as objects to be identified, and extracting the session identifier and source address information of the objects to be identified to generate the set of objects to be identified, includes: The session identifier and source address information are parsed from the handshake fingerprint record corresponding to the message to be identified; Based on the parsed session identifier and source address information, an identifier entry for the object to be identified is constructed; The constructed identifier entries are added to the list to be identified, and the list to be identified is output as the set of objects to be identified.

5. The method according to claim 3, characterized in that, The step of matching and verifying the handshake fingerprint features in the set of objects to be identified with the spoofing traffic feature database to filter out potential spoofing communication components includes: Obtain the spoofing feature records from the spoofing traffic feature database and perform feature matching calculations with the handshake fingerprint features of the object to be identified; Perform version spoofing verification on the cryptographic library version identifier of the object to be identified, and determine whether the cryptographic library version identifier falls within the spoofing range defined by the counterfeit feature record; Based on the results of the feature matching calculation and the results of the version spoofing verification, a comprehensive judgment is made on the object to be identified in order to filter out the potential counterfeit communication components.

6. The method according to claim 5, characterized in that, The step of performing version spoofing verification on the cryptographic library version identifier of the object to be verified includes: Retrieve a list of historical cryptographic library version identifiers for known legitimate client variants; Based on the historical cryptographic database version identifier list, a preset range of version spoofing identifiers is determined through statistical analysis; The password database version identifier of the object to be identified is matched and judged with the preset version spoofing identifier range.

7. The method according to claim 5, characterized in that, The result of the feature matching calculation and the result of the version spoofing verification are used to make a comprehensive judgment on the object to be identified, so as to filter out the potential counterfeit communication components, including: Receive the feature matching metric generated by the feature matching calculation and the verification status generated by the version spoofing verification; Based on preset decision rules, logical decisions are made on the feature matching metric and the verification status; When the output of the logical decision meets the preset component identification conditions, an identification entry for the potential counterfeit communication component is generated.

8. A system for identifying advanced persistent threat attacks on IPv6 encrypted traffic, characterized in that, It includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • System environment identification method and device based on encrypted malicious traffic attack

    CN116318827A

  • Network attack homology analysis method fusing network flow characteristics and threat intelligence

    CN117955745A

  • Suspicious traffic analysis method and system

    CN120768666A

  • Server data interaction network security monitoring processing method and device

    CN121356766A