A large language model driven multi-modal attack tracing method, system, storage medium and program product
Patent Information
- Application Number
- CN202611046229.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]然而,采用上述基于多模态特征融合的攻击溯源方式,在复杂加密网络攻防对抗场景下,当多源异构数据中掺杂攻击者伪造的干扰信息时,直接拼接融会使大语言模型被误导(产生数据偏离或幻觉),此时,仅依赖统计学特征关联和时序匹配进行攻击链路还原,难以准确识别多源异构数据中掺杂的伪造干扰信息,使得伪造干扰信息作为有效证据参与攻击链路推导,从而造成溯源链路虚假、攻击源头误判以及溯源结果不可解释等溯源失效情形,进而导致相关技术在复杂加密网络攻防对抗场景下进行攻击溯源的可靠性较差
[0022]第四方面,本申请实施例提供一种包含指令的计算机程序产品,当上述计算机程序产品在多模态攻击溯源系统上运行时,使得上述多模态攻击溯源系统执行如第一方面以及第一方面中任一可能的实现方式描述的方法。
Smart Images

Figure CN122845236A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and in particular to a method, system, storage medium and program product for tracing multimodal attacks driven by a large language model. Background Technology
[0002] As network attack and defense confrontations continue to escalate, advanced network attacks are exhibiting characteristics such as multi-level springboards, concealment, and long durations. Traditional single-dimensional, rule-based attribution techniques are no longer adequate for complex real-world needs. In recent years, multimodal data fusion and large language modeling technologies have gradually become the mainstream development direction for intelligent network attack attribution, aiming to achieve automated reconstruction of attack chains by aggregating multi-source network data.
[0003] In related technologies, attack attribution tracing methods based on multimodal feature fusion are commonly used. Specifically, firstly, heterogeneous data from multiple sources, such as traffic and logs, are collected from the target network environment; then, features are extracted from various types of data using multiple models, and the extracted features are concatenated and fused through attention mechanisms and other methods; finally, the concatenated fused features are input into a deep learning classification model or a pre-trained language model, and classification and identification are performed based on the statistical feature associations and temporal matching of the data, thereby outputting the final attack path or attribution conclusion.
[0004] However, when using the aforementioned attack tracing method based on multimodal feature fusion in complex encrypted network attack and defense scenarios, direct splicing and fusion can mislead the large language model (causing data deviation or illusion) when multi-source heterogeneous data is mixed with forged interference information by attackers. In this case, relying solely on statistical feature correlation and temporal matching to reconstruct the attack chain makes it difficult to accurately identify the forged interference information mixed in multi-source heterogeneous data. This results in the forged interference information being used as valid evidence in attack chain deduction, leading to tracing failures such as false tracing chains, misjudgment of the attack source, and inexplicable tracing results. Consequently, the reliability of related technologies for attack tracing in complex encrypted network attack and defense scenarios is poor. Summary of the Invention
[0005] This application provides a method, system, storage medium, and program product for tracing multimodal attacks driven by a large language model, which can improve the reliability of attack tracing in complex encrypted network attack and defense scenarios.
[0006] Firstly, this application provides a large language model-driven multimodal attack tracing method, applied to a multimodal attack tracing system. The method includes: acquiring multimodal heterogeneous data of the target network environment, performing credibility screening on the multimodal heterogeneous data to obtain target tracing data, and performing parallel feature extraction on the target tracing data according to its data type to obtain multimodal feature vectors; determining multiple tracing entities represented by the multimodal feature vectors, performing cross-modal alignment on the multiple tracing entities to determine the cross-modal association relationships between the multiple tracing entities, and based on the multiple... The project constructs a temporal multimodal attack evidence graph by tracing the source entity, cross-modal relationships, and multimodal feature vectors. This graph is then transformed into query vectors. Target knowledge entries matching the query vectors are retrieved from the target attack and defense knowledge base. A scenario-based reasoning constraint template is generated based on the target knowledge entries and a multi-level reasoning constraint boundary system. A thought chain reasoning prompt word sequence is constructed based on the scenario-based reasoning constraint template and the temporal multimodal attack evidence graph. Finally, a target large language model is used to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt word sequence, generating causal attack chain data.
[0007] By employing the aforementioned technical solution, credibility screening is performed on multimodal heterogeneous data. Interference data forged by attackers is filtered from the data source, ensuring that all data participating in subsequent inference possesses quantifiable credibility guarantees. A temporal multimodal attack evidence graph is constructed, structurally linking fragmented source tracing evidence across modalities and time periods. This provides a complete evidence carrier for causal inference. Subsequently, authoritative knowledge entries matching the target's attack and defense knowledge base are retrieved to generate scenario-based inference constraint templates. This constrains the target's large language model to perform step-by-step causal inference within verifiable knowledge boundaries, eliminating the illusion of large-model inference. Each step of the source tracing inference is based on credible verified evidence and causal basis verified by authoritative knowledge, making the attack chain reconstruction results traceable, verifiable, and resistant to interference. This solves the technical problem of poor reliability in attack source tracing under complex encrypted network attack and defense scenarios, achieving the technical effect of improving the reliability of attack source tracing in complex encrypted network attack and defense scenarios.
[0008] Optionally, before converting the temporal multimodal attack evidence graph into query vectors, the process includes: constructing a target attack and defense knowledge base based on attack and defense tactical framework data, vulnerability database data, attack organization characteristic data, and historical source tracing case data; constructing basic rule constraint boundaries based on vulnerability database data and preset network attack and defense logic rules; constructing attack behavior characteristic constraint boundaries based on attack and defense tactical framework data; constructing attack chain logic constraint boundaries based on attack organization characteristic data; constructing case paradigm constraint boundaries based on historical source tracing case data; and constructing a multi-level inference constraint boundary system by combining the basic rule constraint boundaries, attack behavior characteristic constraint boundaries, attack chain logic constraint boundaries, and case paradigm constraint boundaries.
[0009] By adopting the above technical solutions, a four-level inference constraint boundary is constructed based on attack and defense tactical technical framework data, vulnerability database data, attack organization characteristic data, and historical source tracing case data, forming a multi-level constraint system from basic rules to case paradigms. This ensures that the target large language model is subject to comprehensive constraints from authoritative attack and defense knowledge when performing causal inference, avoiding the generation of inference conclusions that violate the basic logic of network attack and defense or are detached from real combat scenarios, thereby improving the authority and verifiability of causal inference results.
[0010] Optionally, the temporal multimodal attack evidence graph is transformed into a query vector. Target knowledge entries matching the query vector are retrieved from the target attack and defense knowledge base. A scenario-based inference constraint template is generated based on the target knowledge entries and the multi-level inference constraint boundary system. This includes: vector encoding the temporal multimodal attack evidence graph to generate a query vector matching the dimensions of the target attack and defense knowledge base; using the query vector as the retrieval benchmark, a similarity-ranked retrieval is performed in the target attack and defense knowledge base to obtain a set of candidate knowledge entries; the similarity between the query vector and the knowledge vector of each candidate knowledge entry in the candidate knowledge entry set is calculated, and candidate knowledge entries with similarity below a preset similarity threshold are removed from the candidate knowledge entry set to obtain a set of target knowledge entries; based on the knowledge type of each target knowledge entry in the target knowledge entry set, each target knowledge entry is loaded into the multi-level inference constraint boundary system to generate a scenario-based inference constraint template.
[0011] By adopting the above technical solution, vector encoding is performed on the temporal multimodal attack evidence graph and similarity ranking retrieval is executed. This allows for the accurate selection of authoritative knowledge entries that highly match the current attack scenario from the target attack and defense knowledge base. Furthermore, a similarity threshold filtering mechanism is used to remove low-relevance interference knowledge. Based on the knowledge type of the target knowledge entries, they are loaded into the corresponding level of the multi-level inference constraint boundary system to generate a dedicated inference constraint template for the current attack scenario. This achieves dynamic adaptation between retrieval enhancement and inference constraints, ensuring that the inference constraints accurately match the characteristics of the current attack scenario.
[0012] Optionally, a thought chain reasoning prompt sequence is constructed based on the scenario-based reasoning constraint template and the temporal multimodal attack evidence graph. Then, a target large language model is used to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt sequence to generate causal attack chain data. This includes: generating a step-by-step reasoning task sequence based on the scenario-based reasoning constraint template, which includes evidence sorting, temporal rearrangement, causal deduction, interference removal, link reconstruction, and source localization tasks; encoding the temporal multimodal attack evidence graph, the step-by-step reasoning task sequence, and the scenario-based reasoning constraint template into a thought chain reasoning prompt sequence; and using the target large language model to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt sequence. Perform step-by-step causal reasoning operations according to the thought chain reasoning prompt sequence to obtain step-by-step causal reasoning results and reasoning process logs; perform constraint rule verification on the step-by-step causal reasoning results to obtain constraint rule verification results, including evidence validity verification, temporal logic compliance verification, behavioral causal matching degree verification, false link judgment rationality verification, attack progression logic verification, and source feature compliance verification; if the step-by-step causal reasoning results are determined to pass the constraint rule verification based on the constraint rule verification results, generate causal attack chain data based on the step-by-step causal reasoning results, reasoning process logs, scenario-based reasoning constraint templates, and target knowledge item set.
[0013] By adopting the above technical solution, the step-by-step reasoning task sequence is encoded into a thought chain reasoning prompt word sequence according to a fixed logical order of evidence sorting, temporal rearrangement, causal inference, interference removal, link restoration, and source location. The target large language model is forced to perform reasoning operations step by step according to the preset causal inference logic. After each reasoning step is completed, the corresponding constraint rules are verified. When the reasoning result of any step violates the constraint rules, backtracking correction is performed. This ensures that the final generated causal attack chain data is verified by authoritative knowledge at each reasoning stage, realizing full controllability, traceability, and verifiability of the step-by-step causal reasoning process.
[0014] Optionally, credibility screening is performed on multimodal heterogeneous data to obtain target source tracing data, including: standardizing and preprocessing the multimodal heterogeneous data to obtain a standard dataset; analyzing each standard data in the standard dataset using a credibility scoring model to obtain the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor for each standard data; determining the current security confrontation scenario of the target network environment, and determining the target adjustment factor from the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor based on the current security confrontation scenario, and adjusting the weight coefficients of the target adjustment factor to obtain the target weight configuration; performing credibility analysis on the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor based on the target weight configuration to obtain the credibility score for each standard data; and selecting target source tracing data from the standard dataset based on the credibility score.
[0015] By adopting the above technical solution, a credibility scoring model is used to quantitatively evaluate each standard data from four dimensions: data source credibility, time series rationality, data integrity, and historical false alarm rate. The weight coefficients of each dimension are dynamically adjusted according to the current security confrontation scenario, so that the credibility screening mechanism can adapt to the differences in data quality under different confrontation intensities, accurately filter low credibility data that has been tampered with or forged by attackers, and ensure the purity and reliability of the evidence on which subsequent feature extraction and causal inference are based from the data source.
[0016] Optionally, based on the current security confrontation scenario, target adjustment factors are determined from data source trust factor, time series reasonableness factor, data integrity factor, and historical false alarm factor. The target adjustment factors are then weighted to obtain a target weight configuration. This includes: when the current security confrontation scenario is determined to be a high-adversarial tampering scenario, using the data integrity factor and historical false alarm factor as target adjustment factors; increasing the first weight coefficient corresponding to the data integrity factor in the baseline weight configuration of the trustworthiness scoring model to the second weight coefficient, and decreasing the third weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration to the fourth weight coefficient, to obtain the target weight configuration; or, when the current security confrontation scenario is determined to be a high-frequency false alarm field... In one scenario, the historical false alarm factor and the time-series reasonable factor are used as target adjustment factors; the fifth weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is increased to the sixth weight coefficient, and the seventh weight coefficient corresponding to the time-series reasonable factor in the baseline weight configuration is decreased to the eighth weight coefficient, to obtain the target weight configuration; or, when the current security confrontation scenario is determined to be a boundary security tracing scenario, the data source trust factor and the historical false alarm factor are used as target adjustment factors; the ninth weight coefficient corresponding to the data source trust factor in the baseline weight configuration is increased to the tenth weight coefficient, and the eleventh weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is decreased to the twelfth weight coefficient, to obtain the target weight configuration.
[0017] By adopting the above technical solution, differentiated weight adjustment strategies are set for high-counter tampering scenarios, high-frequency false alarm scenarios, and boundary security tracing scenarios. In high-counter tampering scenarios, the weight of the data integrity factor is increased to enhance the ability to identify and filter data tampered with fields. In high-frequency false alarm scenarios, the weight of the historical false alarm factor is increased to strengthen the punishment of high-frequency false alarm device data. In boundary security tracing scenarios, the weight of the data source trust factor is increased to prioritize the retention of high-security-level device data. This enables the trustworthiness screening mechanism to be accurately adapted to different security confrontation scenarios and avoids the problem of degraded screening effect in specific scenarios caused by using fixed weights.
[0018] Optionally, the credibility scoring model includes a data source evaluation module, a time series verification module, a completeness statistics module, and a false alarm attenuation module. The credibility scoring model analyzes each standard data point in the standard dataset to obtain the data source credibility factor, time series reasonableness factor, data completeness factor, and historical false alarm factor for each standard data point. This includes: using the data source evaluation module to obtain the security protection level score of the data acquisition device corresponding to each standard data point, and using the ratio of the security protection level score to the preset highest level score as the data source credibility factor for the corresponding standard data point; using the time series verification module to calculate the absolute difference between the trigger timestamp of each standard data point and the historical normalized time series mean of the data acquisition device, and determining the time series reasonableness factor for the corresponding standard data point based on the absolute difference and the historical time series standard deviation of the data acquisition device; using the completeness statistics module to count the number of valid core fields for each standard data point, and using the ratio of the number of valid core fields to the preset total number of core fields as the data completeness factor for the corresponding standard data point; using the false alarm attenuation module to obtain the total amount of invalid data and the total amount of reported data from the data acquisition device within a preset historical time window, and determining the historical false alarm factor for the corresponding standard data point based on the preset penalty coefficient, the total amount of invalid data, and the total amount of reported data.
[0019] By adopting the above technical solutions, the data source evaluation module performs quantitative evaluation based on the security protection level score of the data acquisition equipment, the time sequence verification module performs time sequence rationality verification based on the deviation between the trigger timestamp and the historical normal time sequence baseline, the integrity statistics module performs data integrity statistics based on the proportion of valid core fields, and the false alarm attenuation module performs dynamic attenuation penalty based on the proportion of invalid data within a preset historical time window. The four modules respectively quantify and score the standard data from four independent dimensions: equipment credibility, time sequence regularity, field validity, and historical data quality, so as to realize fine-grained, multi-dimensional, and computable quantitative evaluation of the credibility of multimodal heterogeneous data.
[0020] Secondly, embodiments of this application provide a multimodal attack tracing system, which includes: one or more processors and a memory; the memory is coupled to one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and one or more processors call the computer instructions to cause the multimodal attack tracing system to perform the methods described in the first aspect and any possible implementation thereof.
[0021] Thirdly, embodiments of this application provide a computer-readable storage medium including program instructions that, when executed on a multimodal attack tracing system, cause the multimodal attack tracing system to perform the method described in the first aspect and any possible implementation thereof.
[0022] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a multimodal attack tracing system, cause the multimodal attack tracing system to execute the method described in the first aspect and any possible implementation thereof.
[0023] By employing the aforementioned technical solution, credibility screening is performed on multimodal heterogeneous data. Interference data forged by attackers is filtered from the data source, ensuring that all data participating in subsequent inference possesses quantifiable credibility guarantees. A temporal multimodal attack evidence graph is constructed, structurally linking fragmented source tracing evidence across modalities and time periods. This provides a complete evidence carrier for causal inference. Subsequently, authoritative knowledge entries matching the target's attack and defense knowledge base are retrieved to generate scenario-based inference constraint templates. This constrains the target's large language model to perform step-by-step causal inference within verifiable knowledge boundaries, eliminating the illusion of large-model inference. Each step of the source tracing inference is based on credible verified evidence and causal basis verified by authoritative knowledge, making the attack chain reconstruction results traceable, verifiable, and resistant to interference. This solves the technical problem of poor reliability in attack source tracing under complex encrypted network attack and defense scenarios, achieving the technical effect of improving the reliability of attack source tracing in complex encrypted network attack and defense scenarios. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating a multimodal attack tracing method driven by a large language model, as described in this application.
[0025] Figure 2 This is a schematic diagram of the physical device structure of a multimodal attack tracing system in the embodiments of this application. Detailed Implementation
[0026] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0027] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0028] This application provides a method for tracing the source of multimodal attacks driven by a large language model. (See [reference]) Figure 1 , Figure 1 This is a flowchart illustrating a multimodal attack tracing method driven by a large language model, as described in this application, which includes the following steps:
[0029] Step S101: Obtain multimodal heterogeneous data of the target network environment, perform credibility screening on the multimodal heterogeneous data to obtain target source data, and perform parallel feature extraction on the target source data according to the data type of the target source data to obtain multimodal feature vectors;
[0030] Step S102: Determine multiple source entities represented by multimodal feature vectors, perform cross-modal alignment on multiple source entities to determine the cross-modal association between multiple source entities, and construct a temporal multimodal attack evidence map based on multiple source entities, cross-modal association, and multimodal feature vectors;
[0031] Step S103: Convert the temporal multimodal attack evidence graph into a query vector, retrieve the target knowledge entries that match the query vector from the target attack and defense knowledge base, and generate a scenario-based reasoning constraint template based on the target knowledge entries and the multi-level reasoning constraint boundary system.
[0032] Step S104: Construct a thought chain reasoning prompt word sequence based on the scenario-based reasoning constraint template and the temporal multimodal attack evidence graph, and use the target large language model to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt word sequence to generate causal attack chain data.
[0033] The target network environment refers to the network infrastructure environment to be analyzed for attack attribution, including but not limited to enterprise internal networks, cloud data center networks, and critical information infrastructure networks, which deploy various network and security devices. Multimodal heterogeneous data refers to raw attribution data collected from the target network environment, containing different data modalities and structures. Specifically, it includes four data modalities: network traffic data, system log data, network topology data, and security alarm data. Each data modality has different data formats, collection methods, and feature dimensions. Credibility screening refers to the data quality control operation of evaluating the credibility of each piece of data in the multimodal heterogeneous data through a quantitative scoring mechanism and removing data with credibility scores below a preset screening threshold. Target attribution data refers to the set of highly credible data retained after credibility screening, whose credibility scores reach the preset screening threshold, used as input data for subsequent feature extraction and causal inference. Parallel feature extraction refers to the operation of simultaneously encoding features of various types of data using feature extraction models adapted to the characteristics of each data type, based on the data type of the target attribution data. Multimodal feature vectors refer to a set of fixed-dimensional numerical vectors obtained after parallel feature extraction from various target tracing data. They are used to characterize the attack behavior features contained in various tracing data.
[0034] In this context, "source entity" refers to the entity objects related to attack tracing identified from multimodal feature vectors, including but not limited to device nodes, IP addresses, attack behaviors, vulnerability exploits, and security alerts. "Cross-modal alignment" refers to entity-level matching and association operations on feature vectors representing the same source entity or attack event in different data modalities to determine the semantic correspondence between different modal data. "Cross-modal association" refers to the association relationships between multiple source entities identified after cross-modal alignment, manifested in different data modalities, including communication access relationships, behavioral temporal relationships, vulnerability exploitation relationships, and risk association relationships. "Temporal multimodal attack evidence graph" refers to a structured attack evidence network with source entities as nodes, cross-modal association relationships as edges, and carrying temporal attributes and credibility weight attributes. It is used to connect fragmented attack evidence across time periods, devices, and network segments into a complete evidence carrier. "Query vector" refers to transforming the temporal multimodal attack evidence graph into a numerical vector consistent with the dimensions of the target attack and defense knowledge base through a vector encoding model, used for similarity retrieval within the target attack and defense knowledge base. The target attack and defense knowledge base refers to a standardized vector knowledge base that integrates authoritative attack and defense tactical and technical frameworks, vulnerability databases, attack organization characteristics, and historical attribution cases, and is constructed through vector encoding. A target knowledge entry refers to a set of authoritative knowledge entries retrieved from the target attack and defense knowledge base whose similarity to the query vector exceeds a preset similarity threshold.
[0035] The multi-level reasoning constraint boundary system refers to a hierarchical reasoning constraint framework composed of four levels: basic rule constraint boundary, attack behavior feature constraint boundary, attack link logic constraint boundary, and case paradigm constraint boundary. This framework is used to limit the reasoning logic boundaries of the target large language model. The scenario-based reasoning constraint template refers to a set of exclusive reasoning constraint rules generated for the current attack scenario after loading target knowledge items according to knowledge type into the corresponding level of the multi-level reasoning constraint boundary system. The thought chain reasoning prompt word sequence refers to the input prompt word sequence generated by encoding the temporal multimodal attack evidence graph, the step-by-step reasoning task sequence, and the scenario-based reasoning constraint template according to a preset format. This sequence drives the target large language model to perform step-by-step causal reasoning in a fixed logical order. The target large language model refers to a pre-trained large language model with natural language understanding and logical reasoning capabilities, used to perform step-by-step causal reasoning operations based on the thought chain reasoning prompt word sequence. The step-by-step causal reasoning operation refers to the process by which the target large language model, according to the preset reasoning steps in the thought chain reasoning prompt word sequence, progressively executes reasoning tasks such as evidence sorting, temporal rearrangement, causal deduction, interference removal, link restoration, and source location. Causal attack chain data refers to complete attack chain data generated through step-by-step causal reasoning operations, which includes the causal dependencies between each stage of the attack, including the attack source, attack path, causal relationships at each stage, reasoning basis, and reasoning process logs.
[0036] In the above embodiment, the tracing scenario of an advanced persistent threat attack on an enterprise's internal network is used as an example. The target network environment is an internal network of an enterprise that has deployed a perimeter firewall, core switch, intrusion detection system, multiple business servers, and terminal devices. The multimodal attack tracing system collects four types of multimodal heterogeneous data from the target network environment: The first type is network traffic data, including Transmission Control Protocol (TCP) traffic and Transport Layer Security (TLS) encrypted traffic collected from the perimeter firewall and core switches. TCP traffic includes a five-tuple of source address, destination address, source port, destination port, and protocol type, as well as the message payload content. TLS encrypted traffic includes statistical hidden features such as the packet length sequence of encrypted sessions, packet time intervals, and handshake message characteristics. The second type is system log data, including operating system logs and network service logs collected from business servers, recording information such as account login events, file read / write operations, process creation operations, network external connections, and permission change operations. The third type is network topology data, including the interconnection relationships of all network device nodes, IP address segment mappings, virtual LAN partition information, and routing rules. The fourth type is security alarm data, including brute-force attack alarms, abnormal external connection alarms, and vulnerability exploitation alarms generated by the intrusion detection system. Each alarm data includes alarm level, alarm time, associated device identifier, and alarm description information.
[0037] In the above embodiments, the process of extracting statistical latent features from encrypted traffic of transport layer security protocols is described in detail. A non-decryption, non-intrusive collection mechanism is adopted for encrypted traffic; the key is not cracked, and the message content is not parsed. Multidimensional statistical latent features are extracted solely through statistical algorithms. The specific extraction algorithms and mathematical formulas for each statistical latent feature are as follows: The first statistical latent feature is the packet length sequence distribution characteristic of the encrypted traffic. Taking a single encrypted session as the statistical unit, the byte length of all data packets within that encrypted session is collected to construct a packet length time sequence. ,in This represents the total number of data packets in this encrypted session. The following core features are extracted through probability distribution statistics: mean packet length. The calculation formula is: ; Package variance The calculation formula is: ; Package length distribution skewness The calculation formula is: ;in, The length in bytes of the i-th data packet; This reflects the overall traffic volume of the encrypted session; Reflects the fluctuation range of package length; Reflecting the asymmetry of packet length distribution, the packet length of encrypted attack tunnels fluctuates greatly and is skewed, while the packet length of normal encrypted services is evenly distributed and fluctuates smoothly. The skewness of packet length distribution can effectively distinguish between the two.
[0038] In the above embodiment, the second statistical latent feature is the statistical feature of data packet time intervals. The time difference between the transmission and reception of adjacent data packets within the encrypted session is extracted to construct a time interval sequence. The temporal rhythm of flow is characterized by second-order statistics: mean of time intervals. The calculation formula is: Time interval variance The calculation formula is: ;in, The time interval between sending and receiving the i-th data packet and the (i+1)-th data packet is in milliseconds. Reflects the frequency of traffic transmission; Reflecting the stability of traffic transmission rhythm, encrypted jump-board attacks by advanced persistent threat organizations generally exhibit characteristics of disordered time interval fluctuations and sudden, dense bursts. The third statistical latent feature is the traffic information entropy feature. Information entropy is calculated based on the probability distribution of packet length, quantifying the disorder and randomness of encrypted traffic: Traffic Information Entropy. The calculation formula is: ;in, The number of categories for discrete values of the packet length; Let be the probability of the k-th type of packet length appearing in the packet length time sequence; As for traffic entropy, the entropy value of normal encrypted business is stable, while the entropy value of attack traffic such as encryption penetration and data leakage is significantly abnormal.
[0039] In the above embodiment, the fourth statistical latent feature is the uplink and downlink byte ratio feature. It involves statistically analyzing the total uplink and downlink bytes in a single encrypted session to quantify the traffic transmission direction preference: uplink ratio. The calculation formula is: Downward proportion The calculation formula is: ;in, This represents the total number of upstream bytes in this encrypted session; This represents the total downlink bytes for the encrypted session; uplink traffic is significantly higher for data theft and lateral movement attacks, while downlink traffic dominates for regular business access. The fifth statistical feature is the traffic burstiness characteristic. This quantifies the degree of short-term traffic bursts and identifies attacker behaviors such as bulk data transmission or high-frequency probing: Traffic Burstiness. The calculation formula is: ;in, This is the maximum single packet byte length within this encrypted session; The burstiness coefficient represents the burstiness of traffic, and a higher coefficient indicates a more pronounced burst characteristic. This is a core feature for identifying batch data transmission from encrypted tunnels. The sixth statistical latent feature is the session heartbeat cycle and handshake feature sequence. The heartbeat cycle is the average interval between idle keep-alive messages in a statistical encrypted session. The heartbeat cycle of a malicious encrypted tunnel is fixed and extremely simple, unlike the dynamic heartbeat of normal business. The handshake feature sequence is constructed by extracting the number of messages, message size, handshake time, and number of retries during the transport layer security protocol handshake phase, creating a fixed-dimensional handshake feature vector that uniquely identifies the attributes of the encrypted connection.
[0040] In the above embodiments, each piece of multimodal heterogeneous data is quantitatively scored from four dimensions: data source credibility, temporal reasonableness, data integrity, and historical false alarm rate, to calculate the credibility score of each piece of data. For example, a piece of network traffic data collected from a border firewall has a data source credibility factor of 0.9 (high firewall security level), a temporal reasonableness factor of 0.735 (the standardized deviation of the trigger time from the historical business temporal baseline of the device is 0.36 times the standard deviation), a data integrity factor of 1.0 (the five-tuple, timestamp, and traffic statistics fields are all complete), and a historical false alarm factor of 0.973 (the firewall has 3% invalid data in the past 7 days). The overall credibility score calculated according to the baseline weight configuration is approximately 0.895, which is higher than the preset screening threshold of 0.5. Therefore, this piece of data is retained as target source tracing data. A system log data report from a certain terminal device had a data source credibility factor of 0.5 (low security level for ordinary terminals), a timing reasonableness factor of 0.3 (trigger time significantly deviates from the device's historical normal timing), a data integrity factor of 0.4 (multiple core fields are missing), and a historical false alarm factor of 0.45 (high proportion of invalid data in the past 7 days for this terminal). The calculated comprehensive credibility score was 0.41, lower than the preset screening threshold of 0.5. This data was judged as suspected tampering and was removed. After credibility screening, the retained high-credibility data constituted the target traceability data.
[0041] In the above embodiments, for target source tracing data of network traffic type, a convolutional neural network is used to extract traffic spatial features and temporal fluctuation features, outputting a 128-dimensional traffic feature vector; for target source tracing data of system log type, a pre-trained language model fine-tuned by security log corpus is used to identify attack entities and extract behavioral relationships, outputting a 256-dimensional log semantic feature vector; for target source tracing data of network topology type, a graph neural network is used to learn node connectivity relationships and link propagation path features, outputting a 64-dimensional topology feature vector; for target source tracing data of security alarm type, alarm level, alarm type, alarm frequency, and associated risk degree are quantized and encoded, outputting a 32-dimensional alarm risk feature vector. These various feature vectors together constitute a multimodal feature vector.
[0042] In the above embodiments, communication connection entities (including source device node IP address 192.168.1.100, destination server node IP address 10.0.0.50, and external IP address 203.0.113.88) are identified from traffic feature vectors, behavioral entities (including abnormal login behavior, privilege escalation behavior, and file transfer behavior) are identified from log semantic feature vectors, network path entities (including cross-network segment access paths and routing jump nodes) are identified from topology feature vectors, and risk event entities (including brute-force attack alarm events and vulnerability exploitation alarm events) are identified from alarm risk feature vectors. Cross-modal alignment is performed on the source entities identified in the different modalities mentioned above. For example, the communication connection entity with source address 192.168.1.100 in the traffic data is matched at the entity level with the abnormal login behavior entity of device 192.168.1.100 recorded in the log data to determine that they point to the same source entity. Similarly, the vulnerability exploitation alarm event entity associated with device 10.0.0.50 in the alarm data is matched and associated with the communication connection entity with destination address 10.0.0.50 in the traffic data. Cross-modal alignment determines the cross-modal association relationships between multiple source entities, including the communication access relationship between 192.168.1.100 and 10.0.0.50, the behavioral sequence relationship between abnormal login behavior and privilege escalation behavior, and the vulnerability exploitation relationship between vulnerability exploitation alarm events and privilege escalation behavior. The multimodal attack tracing system constructs a temporal multimodal attack evidence graph based on multiple tracing entities, cross-modal relationships, and multimodal feature vectors. This temporal multimodal attack evidence graph uses tracing entities as nodes and cross-modal relationships as edges. Each node carries a corresponding multimodal feature vector and a credibility weight attribute, and each edge carries a temporal attribute to mark the order in which the associated events occur.
[0043] In the above embodiments, the temporal multimodal attack evidence graph is encoded into a 768-dimensional query vector using a vector encoding model. Using this query vector as the retrieval benchmark, a similarity-ranked retrieval is performed in a pre-built target attack and defense knowledge base to obtain the top 10 candidate knowledge entries based on similarity. For example, the retrieved candidate knowledge entries include: entries matching the "initial access - exploitation of public-facing applications" technique in the attack and defense tactical framework (cosine similarity 0.89), entries matching a remote code execution vulnerability in a publicly available vulnerability database (cosine similarity 0.85), entries matching the typical attack chain characteristics of a certain advanced persistent threat attack organization (cosine similarity 0.82), and entries matching a historical attribution case (cosine similarity 0.78). The cosine similarity between the query vector and the knowledge vector of each candidate knowledge entry is calculated. Candidate knowledge entries with a similarity lower than a preset similarity threshold of 0.6 are removed, and entries with a similarity higher than 0.6 are retained as the target knowledge entry set. Subsequently, based on the knowledge type of each target knowledge item (vulnerability rule type, attack behavior characteristic type, attack chain logic type, case paradigm type), it is loaded into the corresponding level of the basic rule constraint boundary, attack behavior characteristic constraint boundary, attack chain logic constraint boundary, and case paradigm constraint boundary in the multi-level inference constraint boundary system, generating a scenario-based inference constraint template for the current attack scenario.
[0044] In the above embodiments, the structured representation of the temporal multimodal attack evidence graph, including the step-by-step reasoning task sequence of evidence sorting, temporal rearrangement, causal inference, interference elimination, link restoration and source location tasks, and the scenario-based reasoning constraint template are encoded according to a preset format to generate a thought chain reasoning prompt word sequence. After receiving the thought chain reasoning prompt sequence, the target large language model performs step-by-step causal reasoning operations according to a fixed logical order: First, it performs an evidence sorting task, sorting all highly credible source entities and their cross-modal relationships in the temporal multimodal attack evidence graph to confirm the scope of valid evidence for reasoning; second, it performs a temporal rearrangement task, rearranging all attack events in chronological order based on the temporal attributes carried by each source entity and cross-modal relationship; third, it performs a causal deduction task, deriving the causal dependencies between rearranged attack events based on the attack behavior feature constraints and attack chain logic constraints loaded in the scenario-based reasoning constraint template, and determining whether preceding events are necessary preconditions for subsequent events; fourth, it performs an interference removal task, identifying and removing temporally adjacent but causally dependent events as interference events; fifth, it performs a chain reconstruction task, connecting the causally verified attack events into a complete multi-level attack propagation chain according to causal dependencies; sixth, it performs a source localization task, locating the starting node of the causal chain in the reconstructed attack propagation chain to determine the initial intrusion point and the true attack source. After each of the above reasoning steps is completed, the multimodal attack tracing system uses the corresponding constraint rules in the scenario-based reasoning constraint template to perform compliance verification on the reasoning result of that step. When the reasoning result violates the constraint rules, a backtracking correction is triggered until the reasoning result passes the constraint rule verification. Finally, causal attack chain data is generated. This causal attack chain data contains a complete causal attack chain with the attack source being the external IP address 203.0.113.88, the initial intrusion point being the business server 10.0.0.50, and the attack path being "external attackers breach the boundary through remote code execution vulnerability → gain server control → exploit privilege escalation vulnerability to gain administrator privileges → lateral movement to the internal network terminal 192.168.1.100 → execute data transfer operation", as well as the reasoning basis and reasoning process logs for each reasoning step.
[0045] Through the above steps, multimodal heterogeneous data undergoes credibility screening, filtering out attacker-forged interference data from the data source to ensure that all data participating in subsequent inference has quantifiable credibility guarantees. A temporal multimodal attack evidence graph is constructed, structurally linking fragmented source tracing evidence across modalities and time periods, providing a complete evidence carrier for causal reasoning. Then, matching authoritative knowledge entries are retrieved from the target attack and defense knowledge base to generate scenario-based reasoning constraint templates. This constrains the target large language model to perform step-by-step causal reasoning within verifiable knowledge boundaries, eliminating the illusion of large model reasoning. Each step of the source tracing reasoning conclusion is based on credible verified evidence and causal basis verified by authoritative knowledge, making the attack chain reconstruction results traceable, verifiable, and interference-resistant. This solves the technical problem of poor reliability in attack source tracing in complex encrypted network attack and defense scenarios, achieving the technical effect of improving the reliability of attack source tracing in complex encrypted network attack and defense scenarios.
[0046] The entity executing the above steps can be a system, such as a multimodal attack tracing system, or a device, or a controller or processor in the system or device, or a separate controller or processor, or other processing devices or processing units with similar processing functions, but is not limited to these.
[0047] In an optional embodiment, before converting the temporal multimodal attack evidence graph into a query vector, the process includes: constructing a target attack and defense knowledge base based on attack and defense tactical framework data, vulnerability database data, attack organization characteristic data, and historical source tracing case data; constructing basic rule constraint boundaries based on vulnerability database data and preset network attack and defense logic rules; constructing attack behavior characteristic constraint boundaries based on attack and defense tactical framework data; constructing attack chain logic constraint boundaries based on attack organization characteristic data; constructing case paradigm constraint boundaries based on historical source tracing case data; and constructing a multi-level inference constraint boundary system by combining the basic rule constraint boundaries, attack behavior characteristic constraint boundaries, attack chain logic constraint boundaries, and case paradigm constraint boundaries.
[0048] The attack and defense tactical framework data refers to standardized attack and defense knowledge data covering various attack tactics and techniques, compiled based on a pre-defined network attack and defense framework. Each attack technique data includes structured fields such as technique number, technique name, tactical stage, attack behavior description, preconditions, post-attack traces, and mitigation measures. Vulnerability database data refers to vulnerability information data compiled based on publicly available vulnerability databases and / or pre-defined internal vulnerability databases. Each vulnerability data includes structured fields such as vulnerability number, vulnerability type, scope of impact, exploitation preconditions, attack vector, and severity level. Attack organization characteristic data refers to typical attack behavior characteristic data of known advanced persistent threat (APS) attack organizations, including organization identifiers, common attack tactical chains, common vulnerability exploitation methods, common malicious tools, and typical target industries. Historical source tracing case data refers to verified historical real-world network attack source tracing case data, including attack scenario descriptions, attack chain reconstruction results, attack source location conclusions, and source tracing reasoning process records. Vector encoding refers to converting the above textualized attack and defense knowledge data into fixed-dimensional numerical vector representations through a pre-trained embedding model to support subsequent high-speed retrieval operations based on vector similarity. The basic rule constraint boundary refers to the first-level constraint layer constructed from vulnerability database data and basic network attack and defense logic. It limits the generation of inference conclusions that violate vulnerability exploitation principles, temporal logic, permission logic, and network connectivity logic during the target's large language model inference process. The attack behavior characteristic constraint boundary refers to the second-level constraint layer constructed based on attack and defense tactical technology framework data. It requires that all attack behaviors inferred by the target's large language model must match corresponding attack technology characteristics; inference results without matching characteristics are deemed invalid. The attack chain logic constraint boundary refers to the third-level constraint layer constructed based on typical attack causal chain templates in attack organization characteristic data. It limits the progressive relationship and causal dependency of each stage of the inference chain, prohibiting inference chains with reversed temporal order or causal inversion. The case paradigm constraint boundary refers to the fourth-level constraint layer constructed based on historical source tracing case data. It provides reference constraints for inference in similar attack scenarios, limiting the reasonable range of source tracing conclusions and source characteristics.
[0049] In the above embodiments, the process of constructing the target attack and defense knowledge base and the multi-level reasoning constraint boundary system is explained, continuing the tracing scenario of attacks on the enterprise's internal network. The following four categories of attack and defense knowledge data are acquired: The first category is attack and defense tactical framework data, which extracts all attack techniques from the MITRE ATT&CK framework, including initial access, execution, persistence, privilege escalation, defense evasion, credential access, discovery, lateral movement, collection, command and control, data infiltration, and impact. Examples include technique T1190 exploiting public-facing applications, technique T1068 exploiting vulnerabilities for privilege escalation, and technique T1021 lateral movement of remote services. Each technique entry includes a complete behavioral description and pre- and post-conditions. The second category is vulnerability database data, which extracts vulnerability entries from publicly available vulnerability databases and / or pre-set internal vulnerability databases, including vulnerability IDs, vulnerability types, attack vectors, impact components, and exploitation conditions. The third category is attack organization characteristic data, including the habitual attack tactical sequences and tool characteristics of known advanced persistent threat (APS) attack organizations. The fourth category is historical attribution case data, containing verified real attribution cases and their complete reasoning processes. The multimodal attack tracing system uses a pre-trained embedding model to vectorize the four types of attack and defense knowledge data, transforming each piece of knowledge data into a 768-dimensional knowledge vector. It also constructs a high-speed vector index based on an approximate nearest neighbor search algorithm to form a target attack and defense knowledge base.
[0050] In the above embodiments, basic rule constraint boundaries are constructed based on vulnerability database data, clarifying the preconditions for exploiting each vulnerability (e.g., a remote code execution vulnerability requires the target device to run a specific version of a network service and that the service port must be reachable), attack behavior characteristics (e.g., exploiting the vulnerability will inevitably generate an abnormal request message in a specific format), and scope of impact (e.g., successful exploitation can obtain server user privileges). Temporal constraint rules (the timestamp of the attack behavior must be later than the timestamp of the vulnerability exploitation behavior), privilege constraint rules (the precondition for privilege escalation behavior must be prior acquisition of low-privilege access), and network connectivity constraint rules (cross-network segment access must have a corresponding routable path) are also constructed. The multimodal attack tracing system constructs attack behavior characteristic constraint boundaries based on attack and defense tactical technology framework data, classifying 198 attack techniques according to stages such as detection, breach, privilege escalation, lateral movement, and data transmission, defining the standard behavioral characteristics and dependencies for each stage. The multimodal attack attribution system constructs logical constraint boundaries for attack chains based on attack organization characteristic data. It defines the standard progressive order of each attack stage as initial access → execution → privilege escalation → lateral movement → data leakage, prohibiting inference chains that skip levels (e.g., executing lateral movement directly without initial access) or have reversed time sequences (e.g., data leakage occurring before initial access). The system also constructs case paradigm constraint boundaries based on historical attribution case data, extracting characteristic patterns from attribution conclusions of similar attack scenarios to provide a reasonable range of references for inferences about new attacks. These four levels of constraints together constitute a multi-level inference constraint boundary system.
[0051] In an optional embodiment, the temporal multimodal attack evidence graph is converted into a query vector. Target knowledge entries matching the query vector are retrieved from the target attack and defense knowledge base. A scenario-based inference constraint template is generated based on the target knowledge entries and the multi-level inference constraint boundary system. This includes: vector encoding the temporal multimodal attack evidence graph to generate a query vector matching the dimensions of the target attack and defense knowledge base; using the query vector as the retrieval benchmark, performing a similarity-ranked retrieval in the target attack and defense knowledge base to obtain a set of candidate knowledge entries; calculating the similarity between the query vector and the knowledge vector of each candidate knowledge entry in the candidate knowledge entry set, and removing candidate knowledge entries with similarity below a preset similarity threshold from the candidate knowledge entry set to obtain a set of target knowledge entries; and loading each target knowledge entry into the multi-level inference constraint boundary system according to the knowledge type of each target knowledge entry in the target knowledge entry set to generate a scenario-based inference constraint template.
[0052] The vector encoding model refers to the same pre-trained embedding model used when constructing the target attack and defense knowledge base. This ensures that the query vector encoded by the temporal multimodal attack evidence graph and the knowledge vectors in the target attack and defense knowledge base are in the same vector space and have comparable semantic similarity. Similarity ranking retrieval refers to the operation of calculating the cosine similarity between the query vector and all knowledge vectors in the target attack and defense knowledge base, ranking them from highest to lowest similarity score, and returning the top K candidate knowledge entries. The candidate knowledge entry set refers to the set of the top K knowledge entries returned by the similarity ranking retrieval, where K is a preset parameter for the number of retrieved items. The preset similarity threshold is the lower limit of the cosine similarity score used to filter low-relevance candidate knowledge entries. Candidate knowledge entries with a cosine similarity score below this threshold are deemed insufficiently relevant to the current attack scenario and are removed. Knowledge type refers to the category label of the target knowledge item in the target attack and defense knowledge base, including four types of knowledge: vulnerability rule type, attack behavior characteristic type, attack chain logic type, and case paradigm type. Target knowledge items of different knowledge types are loaded into different constraint levels in the multi-level reasoning constraint boundary system.
[0053] In the above embodiments, continuing the source tracing scenario, the process of retrieving target knowledge entries from the target attack and defense knowledge base and generating scenario-based reasoning constraint templates is described. The feature vectors, cross-modal relationship attribute information, and temporal attribute information of all source entities in the temporal multimodal attack evidence graph are encoded using the same pre-trained embedding model as when constructing the target attack and defense knowledge base, generating a 768-dimensional query vector. This query vector comprehensively represents the traffic behavior characteristics, log semantic characteristics, topological path characteristics, and alarm risk characteristics of the current attack event. The cosine similarity between the query vector and each knowledge vector in the target attack and defense knowledge base is calculated. The formula for calculating the cosine similarity is: Where Q is the query vector. Let be the knowledge vector of the i-th knowledge entry in the target attack and defense knowledge base. To query the inner product of the vector and the knowledge vector, and These are the Euclidean norms of the query vector and the knowledge vector, respectively. The top 10 candidate knowledge entries are sorted by cosine similarity score from highest to lowest and returned as a candidate knowledge entry set. This candidate knowledge entry set includes: knowledge entries in the attack and defense tactical technology framework such as technology number T1190 exploiting a public-facing application (cosine similarity 0.89); knowledge entries in the attack and defense tactical technology framework such as technology number T1068 exploiting a vulnerability for privilege escalation (cosine similarity 0.86); knowledge entries in the vulnerability database of a remote code execution vulnerability (cosine similarity 0.85); knowledge entries in the attack organization characteristic data of a typical attack chain of an advanced persistent threat organization (cosine similarity 0.82); and knowledge entries in historical source tracing cases of a similar scenario (cosine similarity 0.78), etc.
[0054] In the above embodiments, the cosine similarity between the query vector and the knowledge vector of each candidate knowledge entry in the candidate knowledge entry set is calculated, and candidate knowledge entries with a cosine similarity lower than a preset similarity threshold of 0.6 are eliminated. In this embodiment, the cosine similarity of all 10 candidate knowledge entries in the candidate knowledge entry set is higher than 0.6, so all of them are retained as target knowledge entries. Target knowledge entries of the vulnerability rule type (the exploitation preconditions, attack behavior characteristics, and impact scope of a remote code execution vulnerability) are loaded to the basic rule constraint boundary; target knowledge entries of the attack behavior characteristic type (the standard behavior characteristics, dependency conditions, and post-traces of technical numbers T1190 and T1068) are loaded to the attack behavior characteristic constraint boundary; target knowledge entries of the attack chain logic type (the progressive order and causal dependency relationship of each stage of a typical attack chain of an advanced persistent threat organization) are loaded to the attack chain logic constraint boundary; and target knowledge entries of the case paradigm type (the source conclusion characteristics and source characteristic patterns of a similar scenario source tracing case) are loaded to the case paradigm constraint boundary. Once loaded, the four constraint levels together form a scenario-based inference constraint template for the current attack scenario.
[0055] In an optional embodiment, a thought chain reasoning prompt sequence is constructed based on a scenario-based reasoning constraint template and a temporal multimodal attack evidence graph. Then, a target big language model is used to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt sequence to generate causal attack chain data. This includes: generating a step-by-step reasoning task sequence based on the scenario-based reasoning constraint template, whereby the step-by-step reasoning task sequence includes evidence sorting, temporal rearrangement, causal deduction, interference removal, link reconstruction, and source location tasks; encoding the temporal multimodal attack evidence graph, the step-by-step reasoning task sequence, and the scenario-based reasoning constraint template into a thought chain reasoning prompt sequence; and using the target big language model... The reasoning model performs step-by-step causal reasoning operations according to the thought chain reasoning prompt sequence, obtaining step-by-step causal reasoning results and reasoning process logs. The step-by-step causal reasoning results are then subjected to constraint rule verification, resulting in constraint rule verification results. The constraint rule verification includes verification of evidence validity, temporal logic compliance, behavioral causal matching degree, reasonableness of false link determination, attack progression logic, and source feature compliance. If the step-by-step causal reasoning results pass the constraint rule verification based on the constraint rule verification results, causal attack chain data is generated based on the step-by-step causal reasoning results, reasoning process logs, scenario-based reasoning constraint templates, and target knowledge item set.
[0056] The step-by-step reasoning task sequence refers to an ordered sequence of sub-tasks with fixed logical progression, which decompose step-by-step causal reasoning operations. Specifically, it includes six sub-tasks arranged in execution order: evidence sorting, temporal rearrangement, causal deduction, interference removal, link reconstruction, and source localization. The thought chain reasoning prompt sequence is an input text sequence formed by combining and encoding the structured description of the temporal multimodal attack evidence graph, the task instructions of the step-by-step reasoning task sequence, and the constraint rules of the scenario-based reasoning constraint template according to a preset format. This sequence drives the target large language model to perform reasoning operations according to the specified logical order and constraint rules. Constraint rule verification refers to the automated verification of the compliance of the reasoning result of each step after the target large language model completes its reasoning step, using the constraint rules corresponding to that step in the scenario-based reasoning constraint template. Specifically, it includes six verification rules: evidence validity verification, temporal logic compliance verification, behavioral causal matching degree verification, false link judgment rationality verification, attack progression logic verification, and source feature compliance verification. Backtracking correction refers to the process where, when the reasoning result of a certain reasoning step fails the constraint rule verification, the target large language model reverts to that step and re-executes the reasoning until a reasoning result that conforms to the constraint rule is generated, before continuing to execute the error correction operation for subsequent reasoning steps.
[0057] In the above embodiments, continuing the source tracing scenario, the process of constructing the thought chain reasoning prompt word sequence and executing step-by-step causal reasoning operations is described. The temporal multimodal attack evidence graph is transformed into a structured text description, including all source entities (external IP address 203.0.113.88, business server 10.0.0.50, internal network terminal 192.168.1.100, etc.) and their cross-modal relationships and temporal attributes; the step-by-step reasoning task sequence is encoded into six task instructions to be executed sequentially; and the scenario-based reasoning constraint template is encoded into constraint rule text that must be followed in each reasoning step. The above three parts are concatenated and encoded according to the format of "evidence description → task instructions → constraint rules" to generate the thought chain reasoning prompt word sequence. The thought chain reasoning prompt sequence is input into the target large language model. The target large language model executes the step-by-step causal reasoning operation according to the logical order of the step-by-step reasoning task sequence. The specific execution process is as follows: First, the target large language model performs the evidence sorting task, sorting out all source entities and their cross-modal relationships in the temporal multimodal attack evidence graph. The scope of valid evidence is confirmed to include: the communication access relationship between the external IP address 203.0.113.88 and the business server 10.0.0.50, the abnormal login behavior and privilege escalation behavior recorded on the business server 10.0.0.50, the lateral movement communication relationship between the business server 10.0.0.50 and the internal network terminal 192.168.1.100, the data transmission behavior recorded on the internal network terminal 192.168.1.100, and the vulnerability exploitation alarm generated by the intrusion detection system. The multimodal attack source tracing system uses the evidence validity verification rules to verify the results of this step, confirming that all evidence included in the reasoning comes from the target source data after credibility screening, and the verification passes.
[0058] In the above embodiment, the second step involves the target large language model performing a time-series rearrangement task, arranging the events in time according to their timestamp attributes: At time T1, external IP address 203.0.113.88 initiates an abnormal access request to business server 10.0.0.50 → At time T2, business server 10.0.0.50 triggers a vulnerability exploitation alarm → At time T3, business server 10.0.0.50 records abnormal login behavior → At time T4, business server 10.0.0.50 records privilege escalation behavior → At time T5, business server 10.0.0.50 initiates lateral movement communication to internal network terminal 192.168.1.100 → At time T6, internal network terminal 192.168.1.100 records data outgoing behavior. The multimodal attack tracing system uses time-series logic compliance verification rules to verify that the time sequence T1 < T2 < T3 < T4 < T5 < T6 increases, and the time interval of each event conforms to the reasonable execution cycle of the attack behavior; therefore, the verification passes. The third step involves the target large language model performing a causal inference task. Based on the attack behavior characteristic constraints and attack chain logic constraints in the scenario-based reasoning constraint template, it infers the causal dependencies between adjacent events pair by pair: an external abnormal access request (T1) is a causal prerequisite for triggering a vulnerability exploitation alert (T2) (exploiting a remote code execution vulnerability requires initiating an attack request first); successful vulnerability exploitation (T2) is a causal prerequisite for abnormal login (T3) (initial access is obtained after successful vulnerability exploitation); abnormal login (T3) is a causal prerequisite for privilege escalation (T4) (privilege escalation requires initial access with low privileges); privilege escalation to obtain administrator privileges (T4) is a causal prerequisite for lateral movement (T5) (lateral movement across devices requires high-privilege credentials); lateral movement to obtain terminal control (T5) is a causal prerequisite for data transmission (T6) (data transmission requires read and network transmission permissions on the target terminal). The multimodal attack tracing system uses behavioral causal matching degree verification rules to verify that each pair of causal relationships matches the pre- and post-condition constraints of the corresponding technologies in the attack and defense tactical technology framework, and the verification passes.
[0059] In the above embodiment, in the fourth step, the target large language model performs an interference removal task, checking whether there are any temporally adjacent but causally dependent coincidental events in the temporal multimodal attack evidence graph. In this embodiment, there is a routine intranet business communication record at time T3.5 (routine data synchronization between a business server and a database server). This event is time-series prior to privilege escalation (T4), but it has no causal dependency with any event in the attack chain. The target large language model identifies it as an interference event and removes it. The multimodal attack tracing system uses the false link determination rationality verification rules to verify that the removed event does not meet the behavioral characteristics of any attack technique in the attack behavior characteristic constraint boundary, and the verification passes. The fifth step involves the target large language model performing a link reconstruction task. This task connects the causally verified attack events into a complete multi-level attack propagation link based on causal dependencies: External IP address 203.0.113.88 exploits a remote code execution vulnerability to breach the business server 10.0.0.50 → gains initial access to the business server 10.0.0.50 → exploits a local privilege escalation vulnerability to gain administrator privileges → uses administrator credentials to lateral movement to the internal network terminal 192.168.1.100 → executes data leakage from the internal network terminal 192.168.1.100. The multimodal attack tracing system uses attack progression logic verification rules to confirm that the progression order of each stage of the link conforms to the standard progression logic of "initial access → privilege escalation → lateral movement → data leakage" in the attack link logic constraint boundary, and the verification passes. Step 6: The target large language model performs source localization. In the reconstructed attack propagation chain, the starting node of the causal link is located as the external IP address 203.0.113.88. The initial intrusion point is determined to be the remote code execution vulnerability in the business server 10.0.0.50, and the actual attack source is the external IP address 203.0.113.88. The multimodal attack tracing system uses source feature compliance verification rules to verify that the located attack source is the upstream node of the attack propagation chain and that there are no earlier causal pre-events at this node; the verification passes. The inference results and processes of all six inference steps are summarized to generate causal attack chain data. This causal attack chain data includes a complete attack propagation chain, a description of the causal dependencies between each stage, the attack source localization conclusion, the reasoning basis for each inference step, and the inference process log.
[0060] In an optional embodiment, the process of performing credibility screening on multimodal heterogeneous data to obtain target source tracing data includes: performing standardization preprocessing on the multimodal heterogeneous data to obtain a standard data set; analyzing each standard data included in the standard data set using a credibility scoring model to obtain the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor for each standard data; determining the current security confrontation scenario of the target network environment, and determining the target adjustment factor from the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor based on the current security confrontation scenario, and adjusting the weight coefficients of the target adjustment factor to obtain the target weight configuration; performing credibility analysis on the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor based on the target weight configuration to obtain the credibility score for each standard data; and filtering the target source tracing data from the standard data set based on the credibility score.
[0061] Standardization preprocessing refers to data governance operations that uniformly address issues such as format chaos, redundancy, partial missing data, and anomalous noise in multimodal heterogeneous data. Specifically, it includes removing redundant data within a preset time window, repairing short-term missing data segments, removing extreme anomalous data exceeding a preset statistical deviation range, and standardizing timestamp and device identification encoding formats to align multimodal heterogeneous data at the format, temporal, and entity identification levels. Standard data refers to standardized data obtained after standardization preprocessing, characterized by unified format, temporal alignment, and entity identification alignment. The credibility scoring model is a quantitative evaluation model that calculates the credibility score of each standard data item through weighted summation based on four evaluation dimensions: data source credibility factor, temporal reasonableness factor, data integrity factor, and historical false alarm factor. Current security confrontation scenarios refer to the categories of security threats currently faced by the target network environment, including three scenario categories: high-confrontation tampering scenarios, high-frequency false alarm scenarios, and boundary security tracing scenarios. The preset screening threshold refers to the lower limit of the credibility score used to determine whether standard data is qualified to enter the subsequent feature extraction process. Standard data with a credibility score higher than or equal to the threshold is retained as target traceability data, while standard data with a credibility score lower than the threshold is removed.
[0062] In the above embodiments, continuing the source tracing scenario, the detailed execution process of credibility screening is explained. Deduplication is performed on all collected multimodal heterogeneous data. A sliding window algorithm is used to detect and remove duplicate and redundant data with the same source and content within a 1-second time window. Temporal interpolation repair is performed on short-term missing segments in log and traffic data, and missing data is reasonably supplemented based on the statistical regularity of data from adjacent time periods. Based on the 3-standard-deviation statistical criterion, extreme abnormal data deviating from the mean of similar data by more than 3 standard deviations is identified as noise data and removed. Finally, the timestamps of all network data are uniformly converted to Coordinated Universal Time (UTC) standard format, IP addresses are uniformly converted to a four-segment decimal standard format, and device identifiers are uniformly converted to a unique encoding format. Standard data is obtained after standardization preprocessing. Based on the current threat situation of the target network environment, the current security confrontation scenario is determined to be a high-confrontation tampering scenario (based on the detection of log tampering traces and traffic spoofing signs in the target network). In high-adversarial tampering scenarios, the four-dimensional weight coefficients of the credibility scoring model are adjusted as follows: data source credibility factor weight is 0.35, time series reasonableness factor weight is 0.25, data integrity factor weight is 0.3, and historical false alarm factor weight is 0.1 (the weight of the data integrity factor is increased to enhance the ability to identify and filter data tampered with fields).
[0063] In the above embodiments, the complete mathematical formula system of the credibility scoring model is explained in detail. The formula for calculating the overall credibility score is as follows: The constraints are: Furthermore, all weight coefficients are non-negative real numbers. As a data source credibility factor, As a time-series reasonable factor, For data completeness factor, Historical misreporting factor; , , , These are the weight coefficients for the data source credibility factor, time series reasonableness factor, data integrity factor, and historical false alarm factor, respectively. The baseline weight configuration (general scenario) is as follows: Adaptive dynamic weight adjustment rules: In high-adversarial tampering scenarios, increase the weight of the data integrity factor. Simultaneous reduction In high-frequency false alarm scenarios, increase the weight of historical false alarm factors. Simultaneous reduction In boundary security tracing scenarios, increase the weight of data source factors. Simultaneous reduction The normalization constraint is always satisfied after the weight adjustment.
[0064] In the above embodiments, the data source reliability factor The calculation formula for (equipment level quantitative scoring) is: ;in, The security protection level score of the current data acquisition equipment is determined according to the preset equipment security level score mapping table (firewalls and security gateways have a score of 10, intrusion prevention systems and intrusion detection systems have a score of 9, core switches have a score of 8, business servers have a score of 7, and ordinary terminals have a score of 5). The preset maximum score is 10; the output value range is... This is to prevent valid data from being mistakenly rejected due to low data scores on terminal devices.
[0065] In the above embodiments, the timing rationality factor The calculation formula for (time series deviation quantization verification) is as follows: ;in, This is the trigger timestamp for the current standard data; This represents the historical, normalized time-series average value for the same equipment, the same service, and the same time period. This represents the historical time series standard deviation of the device; when the trigger timestamp perfectly matches the historical normalized time series baseline. The greater the deviation, the lower the score; extremely disordered data... →0. Data integrity factor The formula for calculating (field validity quantification) is: ;in, This represents the number of valid and complete core fields in the current standard data. The total number of core fields preset for this type of data; When all fields are complete and valid When all are missing Historical misreporting factor The formula for calculating (dynamic decay penalty) is: ;in, A fixed penalty coefficient; This refers to the total amount of false alarms, tampering, and invalid data collected by the data acquisition device within a preset historical time window (the past 7 days). This refers to the total number of data reports submitted by the data acquisition device within a preset historical time window (the past 7 days); when the device has no historical false alarms... The higher the false alarm rate The lower the score, the better. The screening logic is as follows: calculate the final credibility score (Score) of a single standard data point. When the Score is ≥ 0.5, the standard data point passes the screening, is retained as target traceability data, and is admitted to the subsequent modeling process; when the Score is < 0.5, the standard data point fails the screening and is directly removed.
[0066] In the above embodiments, factor scores in four dimensions are calculated for each piece of standard data: taking a piece of transport layer security protocol encrypted traffic data collected by a border firewall as an example, the data source trust factor is calculated based on the firewall device security level score of 10 (out of 10), resulting in a data source trust factor score of 0.9; the time series reasonableness factor is calculated based on the deviation between the trigger timestamp of this data and the historical service time series average of the same device during the same period. The standardized deviation of the trigger time from the historical normalized time series baseline is 0.36 times the standard deviation, which is substituted into the time series reasonableness factor calculation formula. The time-series reasonableness factor score is 0.735; the data integrity factor is calculated based on the proportion of the number of valid and complete core fields of this data (all 5 core fields including the 5-tuple, timestamp, packet length sequence, time interval sequence, and traffic entropy value) to 5% of the preset total number of core fields, resulting in a data integrity factor score of 1.0; the historical false alarm factor is calculated based on the proportion of the total number of false alarms and invalid data in the firewall in the past 7 days to the total number of reported data. The proportion of invalid data in the past 7 days is 3%, resulting in a historical false alarm factor score of 0.973. The comprehensive credibility score of this data is calculated according to the baseline weight configuration as follows: 0.35×0.9+0.25×0.735+0.25×1.0+0.15×0.973≈0.895.
[0067] In the above embodiment, the overall credibility score is compared with a preset screening threshold of 0.5. The credibility score of this data, 0.938, is higher than the preset screening threshold of 0.5, and therefore it is retained as target traceability data. For another piece of system log data from a terminal device, its data integrity factor score is only 0.3 (multiple core fields have abnormal blanks and format errors, suspected of being tampered with), and the calculated overall credibility score is 0.42, which is lower than the preset screening threshold of 0.5. This data is judged as suspected tampered data and is removed. Through the above credibility screening process, low-credibility data forged or tampered by attackers is filtered out from the data source, and the retained target traceability data all have quantifiable and verifiable credibility guarantees.
[0068] In an optional embodiment, a target adjustment factor is determined from the data source trust factor, time series reasonableness factor, data integrity factor, and historical false alarm factor based on the current security confrontation scenario. The target adjustment factor is then adjusted by adjusting its weight coefficients to obtain a target weight configuration. This includes: when the current security confrontation scenario is determined to be a high-confrontation tampering scenario, using the data integrity factor and historical false alarm factor as target adjustment factors; increasing the first weight coefficient corresponding to the data integrity factor in the baseline weight configuration of the trustworthiness scoring model to the second weight coefficient, and decreasing the third weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration to the fourth weight coefficient, to obtain the target weight configuration; or, when the current security confrontation scenario is determined to be a high-confrontation tampering scenario, the target adjustment factor is determined by adjusting the weight coefficients of the data source trust factor, time series reasonableness factor, data integrity factor, and historical false alarm factor. In scenarios with frequent false alarms, historical false alarm factors and time-series reasonable factors are used as target adjustment factors; the fifth weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is increased to the sixth weight coefficient, and the seventh weight coefficient corresponding to the time-series reasonable factor in the baseline weight configuration is decreased to the eighth weight coefficient to obtain the target weight configuration; or, when the current security confrontation scenario is determined to be a boundary security tracing scenario, the data source trust factor and historical false alarm factors are used as target adjustment factors; the ninth weight coefficient corresponding to the data source trust factor in the baseline weight configuration is increased to the tenth weight coefficient, and the eleventh weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is decreased to the twelfth weight coefficient to obtain the target weight configuration.
[0069] Among these, high-adversarial tampering scenarios refer to security scenarios in which attackers actively tamper with system logs, forge network traffic, or inject false alarms in the target network environment, resulting in a significantly higher risk of data integrity loss compared to other scenarios. High-frequency false alarm scenarios refer to security scenarios in the target network environment where security devices generate a large number of false alarms due to improper rule configuration or environmental noise interference. In these scenarios, it is necessary to increase the penalty for historical false alarm rates to improve filtering accuracy. Boundary security attribution scenarios refer to attribution analysis scenarios that focus on network boundary devices (firewalls, security gateways, etc.). In these scenarios, the data credibility priority of high-level boundary security devices is higher than that of internal terminal devices. Increasing the weight of the data integrity factor in high-adversarial tampering scenarios involves increasing the weight coefficient of the data integrity factor in the credibility scoring model relative to the baseline weight configuration. This enhances the sensitivity of the credibility scoring model in assessing the integrity of data fields, thereby more effectively identifying and filtering data where core fields have been tampered with or are missing. Increasing the weight of historical false alarm factors refers to raising the weight coefficient of historical false alarm factors in the credibility scoring model relative to the baseline weight configuration in high-frequency false alarm scenarios, thereby strengthening the penalty for data reported by devices with high recent false alarm rates. Increasing the weight of data source credibility factors refers to raising the weight coefficient of data source credibility factors in the credibility scoring model relative to the baseline weight configuration in boundary security tracing scenarios, thereby prioritizing the retention of data from high-security-level boundary devices.
[0070] In the above embodiments, the weight adjustment strategies under three different security confrontation scenarios are illustrated as examples. In the first scenario, the current security confrontation scenario is a high-confrontation tampering scenario. The multimodal attack tracing system detects signs of tampering in the target network environment, such as the clearing of key fields in system logs and the presence of forged normal business packets in network traffic, and determines that the current security confrontation scenario is a high-confrontation tampering scenario. In this scenario, the multimodal attack tracing system increases the weight of the data integrity factor from the baseline value of 0.25 to 0.30, while decreasing the weight of the historical false alarm factor from the baseline value of 0.15 to 0.10. The weight of the data source trust factor remains unchanged at 0.35, and the weight of the time series reasonable factor remains unchanged at 0.25, ensuring that the total weight remains 1.0. By increasing the weight of the data integrity factor, the trust scoring model assigns lower trust scores to data with missing or abnormal core fields, thereby more effectively filtering and eliminating tampered data.
[0071] In the above embodiments, in the second scenario, the current security confrontation scenario is a high-frequency false alarm scenario. The multimodal attack tracing system detects a significant increase in the recent false alarm rate of intrusion detection devices in the target network environment, with a large number of invalid alarms interfering with normal tracing analysis. In this scenario, the multimodal attack tracing system increases the weight of the historical false alarm factor from the baseline value of 0.15 to 0.20, while decreasing the weight of the time-series reasonable factor from the baseline value of 0.25 to 0.20. The weight of the data source trust factor remains unchanged at 0.35, and the weight of the data integrity factor remains unchanged at 0.25, ensuring that the total weight remains 1.0. By increasing the weight of the historical false alarm factor, data reported by devices with high recent false alarm rates are more severely penalized in the trustworthiness score, and their overall trustworthiness score is more likely to fall below the preset screening threshold and be eliminated.
[0072] In the above embodiments, in the third scenario, the current security confrontation scenario is a perimeter security attribution scenario. The multimodal attack attribution system determines that the current attribution analysis focuses on network perimeter breach behavior, requiring priority to ensure the credibility weight of data from perimeter security devices. In this scenario, the multimodal attack attribution system increases the weight of the data source credibility factor from the baseline value of 0.35 to 0.40, while decreasing the weight of the historical false alarm factor from the baseline value of 0.15 to 0.10. The weight of the time-series reasonable factor remains unchanged at 0.25, and the weight of the data integrity factor remains unchanged at 0.25, ensuring that the total weight remains 1.0. By increasing the weight of the data source credibility factor, data from high-security-level perimeter devices such as firewalls and security gateways obtains a higher credibility score, while the credibility score of data from internal low-security-level terminal devices is relatively reduced. Thus, in the perimeter security attribution scenario, high-credibility data from perimeter devices is prioritized.
[0073] In an optional embodiment, the credibility scoring model includes a data source evaluation module, a time series verification module, a completeness statistics module, and a false alarm attenuation module. The credibility scoring model analyzes each standard data point in the standard dataset to obtain a data source credibility factor, a time series reasonableness factor, a data completeness factor, and a historical false alarm factor for each standard data point. This includes: using the data source evaluation module to obtain the security protection level score of the data acquisition device corresponding to each standard data point, and using the ratio of the security protection level score to a preset maximum level score as the data source credibility factor for the corresponding standard data point; using the time series verification module to calculate the absolute difference between the trigger timestamp of each standard data point and the historical normalized time series mean of the data acquisition device, and determining the time series reasonableness factor for the corresponding standard data point based on the absolute difference and the historical time series standard deviation of the data acquisition device; using the completeness statistics module to count the number of valid core fields for each standard data point, and using the ratio of the number of valid core fields to the preset total number of core fields as the data completeness factor for the corresponding standard data point; and using the false alarm attenuation module to obtain the total amount of invalid data and the total amount of reported data from the data acquisition device within a preset historical time window, and determining the historical false alarm factor for the corresponding standard data point based on a preset penalty coefficient, the total amount of invalid data, and the total amount of reported data.
[0074] The data source assessment module, within the credibility scoring model, is responsible for evaluating the security protection level of data acquisition equipment. Based on a pre-defined equipment security level score mapping table, this module maps the security protection level of the data acquisition equipment to a data source credibility factor score. The equipment security protection level score is a pre-set quantitative score based on the type and security protection capabilities of the data acquisition equipment; higher security protection levels result in higher scores. The time-series verification module, also within the credibility scoring model, assesses the reasonableness of data triggering times. This module quantifies the reasonableness of the data in the time-series dimension by calculating the standardized deviation between the current data trigger timestamp and the historical normalized time-series baseline of the equipment. The trigger timestamp refers to the specific time when the data was generated or collected, recorded in the standard data. The historical normalized time-series baseline refers to the statistical average of historical data trigger times for the same equipment within the same business period, reflecting the time-series pattern during normal business operations.
[0075] The completeness statistics module within the credibility scoring model is responsible for assessing the completeness rate of core data fields. This module quantifies the validity of data in terms of fields by statistically analyzing the proportion of valid and complete core fields in the standard data relative to the total number of preset core fields for that type of data. Valid and complete core fields refer to necessary fields in the standard data that are not empty, have a standardized format, and whose values are within a reasonable range. The total number of preset core fields refers to the number of core fields that must be included in that type of data, as predefined according to the data type. The false alarm attenuation module is responsible for dynamically penalizing devices based on their recent data quality within the credibility scoring model. This module applies dynamic attenuation penalties to devices with deteriorating data quality by statistically analyzing the proportion of invalid data (false alarms, tampered data, and data with abnormal formats) generated by the device within a preset historical time window relative to all reported data. The preset historical time window refers to the time range used by the false alarm attenuation module to statistically analyze the historical data quality of the device, used to assess the credibility of the device's recent data.
[0076] In the above embodiment, continuing the source tracing scenario, we will use a security alert from an intrusion detection system as an example to illustrate the specific scoring process of the four modules in the credibility scoring model. This security alert is a vulnerability exploitation alert generated by the intrusion detection system at time T2, indicating that the business server 10.0.0.50 is suspected of being subjected to a remote code execution attack. The data acquisition device for this data is the intrusion detection system. According to the preset device security level score mapping table, the device security protection level score of the intrusion detection system is 9 (out of 10). The data source evaluation module substitutes the data source credibility factor calculation formula... The trigger timestamp for this data entry is T2. The time-series verification module obtains the statistical mean (historical normalized time-series baseline) and historical time-series standard deviation of the historical alarm trigger times of the intrusion detection system within the same business period. The absolute deviation between the trigger timestamp T2 and the historical normalized time-series baseline is calculated, and then divided by the historical time-series standard deviation to obtain a standardized deviation value of 0.51. This standardized deviation value of 0.51 is then substituted into the time-series reasonable factor calculation formula. Since the deviation value of 0.51 is less than twice the standard deviation, the time series is reasonably reasonable, and the reasonableness factor score is [value missing]. .
[0077] In the above embodiment, the total number of preset core fields for security alarm data is 6, including alarm number, alarm level, alarm time, associated device identifier, alarm type, and alarm description. The completeness statistics module checks the above 6 core fields of the data, confirming that all 6 fields are not empty, have the correct format, and their values are within a reasonable range, and the number of valid complete core fields is 6. The completeness statistics module substitutes the data completeness factor calculation formula into the data completeness statistics module. The false alarm attenuation module counts a total of 500 alarms reported by the intrusion detection system within a preset historical time window (the past 7 days), of which 40 were subsequently verified as false alarms or invalid. The false alarm attenuation module will apply a fixed penalty coefficient. Set to 0.9, calculate the percentage of invalid data. Substitute into the historical false alarm factor calculation formula The credibility scoring model is configured with weights corresponding to the current security adversarial scenario (assuming a high-adversarial tampering scenario, with weights of 0.35, 0.25, 0.3, and 0.1 respectively). The overall credibility score for this security alert data is calculated as: 0.35 × 0.9 + 0.25 × +0.3×1.0 + 0.1×0.928 = 0.873. This score is higher than the preset screening threshold of 0.5, so this security alarm data is retained as target tracing data and participates in the subsequent parallel feature extraction and causal reasoning process.
[0078] It should be noted that the examples of all the specific values mentioned above are merely exemplary embodiments, and the specific values are not limited to the examples mentioned above.
[0079] Through the embodiments of this application, credibility screening is performed on multimodal heterogeneous data, filtering out interference data forged by attackers from the data source, ensuring that all data participating in subsequent reasoning has quantifiable credibility assurance, constructing a temporal multimodal attack evidence map, and structurally linking fragmented source tracing evidence across modalities and time periods, providing a complete evidence carrier for causal reasoning, and then retrieving matching authoritative knowledge entries from the target attack and defense knowledge base to generate scenario-based reasoning constraint templates, which can constrain the target large language model to perform step-by-step causal reasoning within verifiable knowledge boundaries, eliminate the illusion of large model reasoning, and ensure that each step of the source tracing reasoning conclusion is based on credibility-verified evidence and causal basis verified by authoritative knowledge, making the attack chain reconstruction results traceable, verifiable and resistant to interference.
[0080] The multimodal attack tracing system in this application embodiment is described below from a hardware processing perspective. (See attached document.) Figure 2 , Figure 2 This is a schematic diagram of the physical device structure of a multimodal attack tracing system in the embodiments of this application.
[0081] It should be noted that, Figure 2 The structure of the multimodal attack tracing system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0082] like Figure 2As shown, the multimodal attack tracing system includes a Central Processing Unit (CPU) 201, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 202 or programs loaded from storage section 208 into Random Access Memory (RAM) 203, such as executing the methods described in the above embodiments. The RAM 203 also stores various programs and data required for platform operation. The CPU 201, ROM 202, and RAM 203 are interconnected via bus 204. An I / O interface 205 is also connected to bus 204. The following components are connected to the I / O interface 205: an input section 206 including audio input devices, push-button switches, etc.; an output section 207 including a Liquid Crystal Display (LCD) and audio output devices, indicator lights, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including network interface cards such as LAN (Local Area Network) cards, modems, etc. The communication section 209 performs communication processing via a network such as the Internet. The drive 210 is also connected to the I / O interface 205 as needed. Removable media 211, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive 210 as needed so that computer programs read from them can be installed into the storage section 208 as needed.
[0083] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 209, and / or installed from removable medium 211. When the computer program is executed by central processing unit (CPU) 201, it performs the various functions defined in this application.
[0084] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution platform, apparatus, or device.
[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation that may be implemented in platforms, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than that shown in the drawings.
[0086] Specifically, the multimodal attack tracing system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the large language model-driven multimodal attack tracing method provided in the above embodiment.
[0087] In another aspect, this application also provides a computer-readable storage medium, which may be included in the multimodal attack tracing system described in the above embodiments; or it may exist independently and not assembled into the multimodal attack tracing system. The storage medium carries one or more computer programs, which, when executed by a processor of the multimodal attack tracing system, cause the multimodal attack tracing system to implement the large language model-driven multimodal attack tracing method provided in the above embodiments.
[0088] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for tracing the source of multimodal attacks driven by a large language model, characterized in that, include: Acquire multimodal heterogeneous data of the target network environment, perform credibility screening on the multimodal heterogeneous data to obtain target source data, and perform parallel feature extraction on the target source data according to the data type of the target source data to obtain multimodal feature vectors; The process involves identifying multiple source entities represented by the multimodal feature vectors, performing cross-modal alignment on these entities to determine cross-modal relationships, and constructing a temporal multimodal attack evidence graph based on the multiple source entities, the cross-modal relationships, and the multimodal feature vectors. The temporal multimodal attack evidence graph is transformed into a query vector. Target knowledge entries that match the query vector are retrieved from the target attack and defense knowledge base. A scenario-based reasoning constraint template is then generated based on the target knowledge entries and the multi-level reasoning constraint boundary system. Based on the scenario-based reasoning constraint template and the temporal multimodal attack evidence graph, a thought chain reasoning prompt word sequence is constructed, and the target large language model is used to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt word sequence to generate causal attack chain data.
2. The method according to claim 1, characterized in that, Before converting the temporal multimodal attack evidence graph into a query vector, the process includes: The target attack and defense knowledge base is constructed based on attack and defense tactical and technical framework data, vulnerability database data, attack organization characteristic data, and historical source tracing case data. Based on the vulnerability database data and preset network attack and defense logic rules, construct basic rule constraint boundaries; Construct attack behavior feature constraint boundaries based on the aforementioned offensive and defensive tactical technology framework data; Construct logical constraint boundaries for the attack chain based on the attack organization's characteristic data; Construct case paradigm constraint boundaries based on the historical source case data; The multi-level reasoning constraint boundary system is constructed by the basic rule constraint boundary, the attack behavior feature constraint boundary, the attack link logic constraint boundary, and the case paradigm constraint boundary.
3. The method according to claim 2, characterized in that, The step of converting the temporal multimodal attack evidence graph into a query vector, retrieving target knowledge entries matching the query vector from the target attack and defense knowledge base, and generating a scenario-based inference constraint template based on the target knowledge entries and the multi-level inference constraint boundary system includes: The temporal multimodal attack evidence graph is vector-encoded to generate a query vector that matches the dimension of the target attack and defense knowledge base; Using the query vector as the retrieval benchmark, a similarity-ranked retrieval is performed in the target attack and defense knowledge base to obtain a set of candidate knowledge entries; Calculate the similarity between the query vector and the knowledge vector of each candidate knowledge entry in the candidate knowledge entry set, and remove the candidate knowledge entries with similarity lower than a preset similarity threshold from the candidate knowledge entry set to obtain the target knowledge entry set; Based on the knowledge type of each target knowledge item in the target knowledge item set, each target knowledge item is loaded into the multi-level reasoning constraint boundary system to generate a scenario-based reasoning constraint template.
4. The method according to claim 3, characterized in that, The step involves constructing a thought chain reasoning prompt sequence based on the scenario-based reasoning constraint template and the temporal multimodal attack evidence graph, and then using the target large language model to perform step-by-step causal reasoning operations according to the thought chain reasoning prompt sequence to generate causal attack chain data, including: A step-by-step reasoning task sequence is generated based on the scenario-based reasoning constraint template. The step-by-step reasoning task sequence includes evidence sorting task, temporal rearrangement task, causal inference task, interference removal task, link restoration task, and source location task. The temporal multimodal attack evidence graph, the step-by-step reasoning task sequence, and the scenario-based reasoning constraint template are encoded into the thought chain reasoning prompt word sequence; The step-by-step causal reasoning operation is performed using the target large language model according to the thought chain reasoning prompt word sequence to obtain the step-by-step causal reasoning result and reasoning process log; The step-by-step causal reasoning results are subjected to constraint rule verification to obtain constraint rule verification results. The constraint rule verification includes evidence validity verification, temporal logic compliance verification, behavioral causal matching degree verification, false link determination rationality verification, attack progression logic verification, and source feature compliance verification. If the step-by-step causal reasoning result passes the constraint rule verification based on the constraint rule verification result, the causal attack chain data is generated based on the step-by-step causal reasoning result, the reasoning process log, the scenario-based reasoning constraint template, and the target knowledge item set.
5. The method according to claim 2, characterized in that, The process of performing credibility screening on the multimodal heterogeneous data to obtain target source tracing data includes: The multimodal heterogeneous data is subjected to standardization preprocessing to obtain a standard data set; The credibility scoring model is used to analyze each standard data included in the standard dataset to obtain the data source credibility factor, time series rationality factor, data integrity factor and historical false alarm factor for each standard data. The current security confrontation scenario of the target network environment is determined, and a target adjustment factor is determined from the data source trust factor, the time series reasonable factor, the data integrity factor and the historical false alarm factor based on the current security confrontation scenario. The weight coefficients of the target adjustment factor are adjusted to obtain the target weight configuration. Based on the target weight configuration, a credibility analysis is performed on the data source credibility factor, the time series rationality factor, the data integrity factor, and the historical false alarm factor to obtain a credibility score for each standard data. The target tracing data is selected from the standard dataset based on the credibility score.
6. The method according to claim 5, characterized in that, The step of determining a target adjustment factor from the data source trust factor, the time series reasonableness factor, the data integrity factor, and the historical false alarm factor based on the current security confrontation scenario, and adjusting the weight coefficients of the target adjustment factor to obtain the target weight configuration includes: When the current security challenge scenario is determined to be a high-risk tampering scenario, the data integrity factor and the historical false alarm factor are used as the target adjustment factors; the first weight coefficient corresponding to the data integrity factor in the baseline weight configuration of the credibility scoring model is increased to the second weight coefficient, and the third weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is decreased to the fourth weight coefficient, to obtain the target weight configuration; or, When the current security confrontation scenario is determined to be a high-frequency false alarm scenario, the historical false alarm factor and the time-series reasonable factor are used as the target adjustment factor; the fifth weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is increased to the sixth weight coefficient, and the seventh weight coefficient corresponding to the time-series reasonable factor in the baseline weight configuration is decreased to the eighth weight coefficient to obtain the target weight configuration; or, When the current security confrontation scenario is determined to be a boundary security tracing scenario, the data source trust factor and the historical false alarm factor are used as the target adjustment factors; the ninth weight coefficient corresponding to the data source trust factor in the baseline weight configuration is increased to the tenth weight coefficient, and the eleventh weight coefficient corresponding to the historical false alarm factor in the baseline weight configuration is decreased to the twelfth weight coefficient, to obtain the target weight configuration.
7. The method according to claim 5, characterized in that, The credibility scoring model includes a data source evaluation module, a time series verification module, a completeness statistics module, and a false alarm decay module. The process of analyzing each standard data point in the standard dataset using a credibility scoring model yields a data source credibility factor, a time-series reasonableness factor, a data integrity factor, and a historical false alarm factor for each standard data point, including: The security protection level score of the data acquisition device corresponding to each standard data is obtained by the data source evaluation module, and the ratio of the security protection level score to the preset highest level score is used as the data source credibility factor of the corresponding standard data. The timing verification module is used to calculate the absolute difference between the trigger timestamp of each standard data and the historical normal timing mean of the data acquisition device, and the timing rationality factor of the corresponding standard data is determined based on the absolute difference and the historical timing standard deviation of the data acquisition device. The integrity statistics module is used to count the number of valid core fields for each standard data, and the ratio of the number of valid core fields to the total number of preset core fields is used as the data integrity factor of the corresponding standard data. The false alarm attenuation module is used to obtain the total amount of invalid data and the total number of reported data of the data acquisition device within a preset historical time window, and the historical false alarm factor of the corresponding standard data is determined according to the preset penalty coefficient, the total amount of invalid data and the total number of reported data.
8. A multimodal attack tracing system, characterized in that, The multimodal attack tracing system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors invoke the computer instructions to cause the multimodal attack tracing system to perform the method as described in any one of claims 1-7.
9. A computer-readable storage medium comprising program instructions, characterized in that, When the program instructions are run on the multimodal attack tracing system, the multimodal attack tracing system performs the method as described in any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method as described in any one of claims 1-7 is implemented.