A security metric evaluation method and system based on a large language model
Patent Information
- Application Number
- CN202610640573.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]为解决现有技术中安全度量评估结果难以适应网络环境的动态变化问题,本发明提出了一种基于大语言模型的安全度量评估方法及系统
[0021]与现有技术相比,本发明的有益效果至少包括:
Smart Images

Figure CN122845162A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network security technology, and in particular relates to a security measurement and evaluation method and system based on a large language model. Background Technology
[0002] With the continuous development of cyberattack methods and the increasing complexity of network systems, traditional network security protection technologies are facing more and more challenges. Current security assessment technologies mainly rely on rule-based detection systems, static vulnerability scanning, and penetration testing.
[0003] For example, patent application CN120342697A extracts features from multi-source security monitoring data and inputs them into a security posture model to determine multi-dimensional security assessment indicators (such as asset exposure and attack complexity), improving the accuracy of security assessments. Another example is patent application CN120337225A, which employs multimodal learning and deep learning technologies to integrate vulnerability descriptions, code structure, and runtime data to predict the exploitability and risk level of vulnerabilities, improving the accuracy and reliability of assessing potential and actual security threats to the operating system. Finally, patent application CN119691754A extracts vulnerability data, calculates vulnerability values, obtains computer system risk values through curve fitting, and dynamically adjusts the operation and maintenance cycle based on this, thereby reducing overall operation and maintenance costs and minimizing business interruptions.
[0004] However, existing technologies largely rely on preset static rules, correction coefficients, and thresholds, making the evaluation results highly sensitive to parameters and unable to truly adapt to the dynamic changes in the network environment. At the same time, these systems, which are trained based on historical data and known models, struggle to quickly and accurately identify and evaluate zero-day vulnerabilities or new attack patterns that have not appeared in the training data, limiting the comprehensiveness and real-time nature of their evaluations and preventing them from achieving deep reasoning and real-time quantification of complex, multi-stage attack paths. Summary of the Invention
[0005] To address the problem that existing security measurement and evaluation results are difficult to adapt to the dynamic changes in the network environment, this invention proposes a security measurement and evaluation method and system based on a large language model.
[0006] The present invention adopts the following technical solution. The first aspect of the present invention provides a security measurement and evaluation method based on a large language model, comprising the following steps: Multi-source heterogeneous security data is collected from network security systems, and the security data is preprocessed to generate security event data. The security event data is then input into a large language model for automated analysis, and the security event data is mapped to a network security knowledge base and model framework to generate tagged structured event data. Based on structured event data, an aggregated subject is constructed with subject ID and session ID as composite primary keys; the structured event data is grouped by the aggregated subject to form event streams; each event stream is divided into independent attack sessions to construct an attack session sequence; the environmental context features, result feedback features, the technical sequence of occurrence, and execution parameter features of the attack session sequence are extracted to fuse and generate training samples for training large language models. The large language model is trained and optimized using training samples; the trained large language model is called to generate multiple candidate attack paths and the feasibility confidence of each candidate attack path is calculated; based on the feasibility confidence of each candidate attack path, quantitative risk indicators including path success rate, potential business loss and comprehensive risk score are calculated for each candidate attack path. Based on the quantitative risk indicators and the candidate attack paths, the existing defense strategies are evaluated, and a list of defense optimization suggestions ranked by priority is generated based on the evaluation results.
[0007] Furthermore, the generation of security event data includes: The system collects data from multiple security systems, including Security Information and Event Management (SIEM), Endpoint Detection and Response (EDR), Network Detection and Response (NDR), vulnerability scanners, and asset management, covering attack behavior logs, vulnerability and asset information. The collected data undergoes preprocessing, including data cleaning, format normalization, and semantic alignment, to generate semantically unified security event data.
[0008] Furthermore, the generation of tagged structured event data includes: Construct a prompt word template that includes character settings, input fields, output fields, and template requirements; The security event data is input into the large language model according to the prompt word template format; The large language model maps each security incident data to multiple tactical and technical TTPs in the MITRE ATT&CK framework and outputs the mapping result of each security incident data and its corresponding confidence score. The large language model analyzes and outputs tagged structured event data according to the format of the prompt word template; the structured event data includes the original event ID, normalized time, subject ID, ATT&CK tactical and technical number, confidence score and EDR.
[0009] Furthermore, the structured event data is grouped based on the composite primary key to form an event stream, including: Based on the structured event data, the main assets associated with the event are matched and determined from the preset asset list; Obtain the unique identifier pre-stored in the asset list for the main asset, and use it as the main asset. ; Extract the login session identifier from the structured event data to identify the session. ; use =Main Body Username Session Constructing the aggregate entity ; by For identification purposes, the structured event data is grouped, and the event data within each group is arranged in chronological order to form a continuous event stream.
[0010] Furthermore, the step of dividing each event stream into independent attack sessions and constructing an attack session sequence includes: Traverse the event stream and calculate adjacent events. , Time difference ; Preset time interval threshold ,like Then The event marked as the start of a new session; If the event If it includes attack success and failure feedback events, then it will The event marked as the start of a new session; Based on the above two rules, the dividing point identified in the continuous event stream is the split point; Events that are located between two split points and are time-continuous are aggregated in chronological order to form an attack session sequence corresponding to an independent attack session.
[0011] Further, the generation of training sample pairs for training the large language model includes: Extract the sequence of techniques that have occurred in the attack session sequence; Extract the execution parameter features corresponding to each technology in the already occurred technology sequence; For the aggregated entities corresponding to the attack session sequence, environmental context features are extracted from asset profiles, vulnerability lists, and defense information; Based on the logs and structured events in the security system data, extract the results associated with the in-session events, including blocking status, detection delay, and response actions, as result feedback features; The occurrence of the technical sequence, the execution parameter features, the environmental context features, and the result feedback features are integrated to form training samples for training large language models. The occurrence of the technical sequence includes extracting multiple technical numbers that are identified and mapped to a pre-set attack framework in chronological order from the attack session sequence to form tactical technical sequence features; The extraction of execution parameter features includes, for each technology in the sequence of occurrences, extracting one or more of the following: command line parameters, script content, payload features, protocol type, port, URL, file path, request header, target object identifier, account information, and permission features, to form a set of technical execution parameters; The tactical technology sequence is characterized by a technology number sequence within the MITRE ATT&CK framework.
[0012] Furthermore, the calculation of the feasibility confidence of each candidate attack path includes: The large language model is trained and optimized based on the training samples. Based on a well-trained large language model, multiple candidate attack paths are generated through probabilistic decoding; For each generated candidate attack path, an environmental consistency determination is performed. The determination includes matching the set of preconditions for each technical TTP in the attack path with the set of vulnerability evidence in the current environment. Calculate the joint conditional probability for each candidate attack path; the joint conditional probability is the product of the conditional probabilities of all techniques in the candidate attack path. The calculated joint conditional probability is the feasibility confidence of the corresponding candidate attack path; Feasibility confidence With threshold When comparing, ≥ If the condition is met, the candidate attack path will be retained; otherwise, it will be eliminated. Each generated candidate attack path corresponds to a joint conditional probability.
[0013] Furthermore, the matching of the set of preconditions for each technique TTP in the attack path with the set of vulnerability evidence in the current environment includes: Extracting the first candidate attack path from the pre-defined TTP precondition knowledge base A set of preconditions for each TTP ; A vulnerability evidence set is constructed based on asset profiling, vulnerability lists, and defense deployment status using environmental context features. ; Perform for each TTP and The matching operation yields the matching indicator. ; If and only if exist When there are conditions that satisfy the conditions and there is no logical conflict with the defense deployment status or firewall access control rules. ,otherwise ; When the candidate attack path satisfies that all TTPs in the corresponding path are available If the candidate attack path meets the environmental consistency requirements, it is retained; otherwise, it is discarded.
[0014] Furthermore, the quantitative risk indicators for calculating each candidate attack path include path success rate, potential business loss, and comprehensive risk score, including: For each technical step in the candidate attack path, calculate its technical feasibility probability, defense blocking probability, and environmental reachability probability. The success probability of the technical step is obtained by jointly calculating the probability of technical feasibility, the probability of defense blocking, and the probability of environmental reachability; the success rate of the candidate attack path is obtained by jointly calculating the success probabilities of all technical steps in the path.
[0015] Furthermore, the technical feasibility probability is based on the matching result between the service fingerprint of the target asset and the affected version range of the vulnerability associated with the technology, as well as the vulnerability's public vulnerability score, exploitability score, and historical success rate. The defense blocking probability is the model blocking probability predicted based on the current defense configuration of the target asset, and the historical statistical blocking rate obtained from the historical alarm database under similar technologies and defense configurations. The environmental reachability probability, based on the analysis of network topology, routing policies, and access control rules, determines whether the communication path from the host in the previous step to the target asset required in the current technical step is reachable.
[0016] Furthermore, the calculation includes quantitative risk indicators such as path success rate, potential business losses, and comprehensive risk score, and also includes: The potential business loss is obtained by using asset value quantification mapping rules, based on asset role classification, data sensitivity classification, and service continuity requirement classification to obtain the maximum value of the standardized loss score of the asset.
[0017] Furthermore, the calculation includes quantitative risk indicators such as path success rate, potential business losses, and comprehensive risk score, and also includes: The comprehensive risk score is calculated by weighting the path success rate and potential business losses and multiplying it by a duration penalty function; The duration penalty function is negatively correlated with the attack duration. The attack duration is calculated by traversing each technical step in the candidate attack path, summing the average execution time and the dependency waiting time, and then calculating the attack duration.
[0018] Furthermore, the assessment of existing defense strategies includes: The evaluation results of existing defense strategies depend on the overall blocking confidence level; For each candidate attack path, calculate the blocking probability for each attack step; Calculate the overall blocking confidence of the path based on the probability that all steps in the acquisition path are not blocked; The effectiveness of existing defense strategies against the current attack path is determined based on the overall blocking confidence level.
[0019] Furthermore, the step of determining whether the existing defense strategy is effective against the current attack path based on the comprehensive blocking confidence level includes: Comprehensive blocking confidence With preset blocking threshold Compare; like If it is valid, it is considered valid; otherwise, it is considered invalid or partially valid.
[0020] Furthermore, the generated and prioritized list of defense optimization suggestions includes: Set a risk screening threshold, and select attack paths with a comprehensive risk score higher than the risk screening threshold as high-risk attack paths; For the technologies in the high-risk attack paths where the defense measures are determined to be ineffective or partially effective, corresponding defense reinforcement suggestions are generated; The defense reinforcement recommendations are ranked based on a multi-dimensional priority ranking model. The evaluation dimensions of the ranking model include at least risk urgency, repair cost, and multi-path coverage gain. The risk urgency is based on the product of the success rate of the corresponding attack path and the potential business loss; the remediation cost is quantified based on the complexity level of the strategy adjustment and whether downtime is required; the multi-path coverage gain is the number of high-risk attack paths that the defense reinforcement suggestion can cover and the resulting risk reduction. A second aspect of the invention provides a security measurement and evaluation system based on a large language model, running the aforementioned security measurement and evaluation method, including: Data Acquisition and Labeling Module: Collects multi-source heterogeneous security data from network security systems, preprocesses the security data to generate security event data, inputs the security event data into a large language model for automated analysis, and maps the security event data to a network security knowledge base and model framework to generate structured event data with labels; The data processing feature module constructs an aggregated subject with subject ID and session ID as composite primary keys based on structured event data; groups the structured event data into event streams using the aggregated subject; and divides each event stream into independent attack sessions to construct an attack session sequence; extracts the environmental context features, result feedback features, the technical sequence of events, and execution parameter features of the attack session sequence to fuse and generate training samples for training large language models. The training optimization and generation module uses training samples to train and optimize the large language model; it calls the trained large language model to generate multiple candidate attack paths and calculates the feasibility confidence of each candidate attack path; based on the feasibility confidence of each candidate attack path, it calculates quantitative risk indicators for each candidate attack path, including path success rate, potential business loss, and comprehensive risk score. The evaluation and optimization module is implemented to evaluate existing defense strategies based on quantitative risk indicators and the candidate attack paths, and to generate a list of defense optimization suggestions sorted by priority based on the evaluation results.
[0021] Compared with the prior art, the beneficial effects of the present invention include at least the following: (1) It enhances the dynamic adaptability of security assessment. By utilizing the generation and reasoning capabilities of large language models, it can infer new attack paths beyond the predefined rule base based on the input real-time security data, thus overcoming the problem of poor adaptability of existing technologies to unknown attack patterns.
[0022] (2) The security assessment process has been highly automated. By automatically generating candidate attack paths and automatically generating and sorting defense optimization suggestions, the penetration testing and defense strategy analysis work that traditionally relied on security experts to perform manually has been automated, which has significantly shortened the time cycle from threat perception to defense response.
[0023] (3) It provides multi-dimensional quantitative security metrics. By calculating quantitative risk assessment metrics, defense coverage and blocking confidence, the original qualitative security situation is transformed into a series of calculable and comparable quantitative metrics, providing data-driven decision-making basis for the optimization of security strategies.
[0024] (4) It realizes the transformation from post-event response to pre-event defense. By simulating and generating attack paths and predicting the defense effect before an attack occurs, it can prioritize repairing the most threatening defense weaknesses, thereby proactively reducing the risk of the system being successfully invaded. Attached Figure Description
[0025] Figure 1 This is a flowchart of a security metric evaluation method based on a large language model.
[0026] Figure 2This is a flowchart of a security measurement and evaluation system based on a large language model. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this invention are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.
[0028] As an embodiment of the present invention, a specific implementation method of a security measurement and evaluation method based on a large language model is disclosed, referring to... Figure 1 The method specifically includes the following steps: Step 1: Multi-source security data collection and preprocessing; the data comes from multiple security systems, including SIEM, EDR, NDR, vulnerability scanners, asset management systems, etc., covering attack behavior logs, vulnerability and asset information, etc.
[0029] 1.1: Attack behavior data includes: Simulated attack behavior captured from honeypot system attack logs; Various security incident information is collected and aggregated from the logs of the SIEM (Security Information and Incident Management) platform; The EDR (Endpoint Detection and Response) system log records process behavior, command activity, etc. on the endpoint; Network traffic sessions and abnormal network behavior captured by NDR (Network Detection and Response) probes; The vulnerability scanner outputs alerts about exploit attempts, as well as abnormal login or configuration change records associated with the asset management system.
[0030] 1.2: Vulnerability and asset information includes: Host attributes, including operating system type, service version, open ports, and a list of running software, form an asset fingerprint; Policies for security devices such as firewalls and WAFs: including firewall access control rules, WAF rule sets, etc. EDR protection policies: Includes enabled protection policies; Cloud security group configuration: Security group configuration in a cloud environment.
[0031] 1.3: In a further implementation, existing open-source large language models (such as Llama-3, Qwen, Mistral, etc.) are invoked, combined with prompt engineering techniques, to perform data preprocessing on the aforementioned raw data. The preprocessing steps include: Data cleaning: Generate behavioral fingerprints or hash values based on event content to remove duplicate alerts; complete missing fields through contextual reasoning; and filter invalid or benign events using known whitelist rules.
[0032] Format normalization: Convert timestamps from different systems to the ISO 8601 standard format; map hostnames and IP addresses to standardized asset identifiers (e.g., unify "WIN-ABC123" or "192.168.1.10" into asset ID "A-1001"); encode protocol and port combinations (e.g., TCP / 445) into semantic tags (e.g., "SMB"). Semantic alignment: For the different descriptions of the same attack behavior by different security vendors (e.g., vendor A describes it as "suspicious PowerShell" while vendor B labels it as "malicious executable script"), the semantic understanding capabilities of the large language model are used to map them to a unified general security semantic space, eliminating ambiguity in vendor terminology.
[0033] Step 2: Automated analysis and ATT&CK annotation; The data output from step 1 (i.e., various types of security events) is input into an open-source large language model, which uses natural language understanding (NLU) and knowledge reasoning capabilities for automated analysis. The security events are then mapped to tactics and techniques (TTPs) in the MITRE ATT&CK framework, thereby generating structured event data with ATT&CK tags.
[0034] 2.1: Constructing Prompt Templates; Designing structured prompt templates to ensure consistency between input and output. Prompt templates include: Role definition: Define the role of the large language model in task processing; Input fields: Define the event text, associated asset ID, asset type, deployment status of the defense strategy, and event occurrence time, etc. Output fields: The large language model is required to output the event's ATT&CK tactics, technical numbers (TTPs), technical names, and a brief analysis of the attack intent, etc. Template requirements: Include input and output examples to help large language models understand how to process data and produce correct output.
[0035] 2.2: Process the security event data output from Step 1 according to the prompt word template format and input it into the large language model. The large language model performs deep natural language understanding on the input text through its Transformer architecture, and analyzes the keywords, verbs, and parameters in the event text by combining its internal knowledge base trained on the security corpus (i.e., model parameters). Based on the semantic analysis results of the event text, determine the stage of the event in the attack chain, the purpose and intent of the attack.
[0036] Furthermore, the large language model retrieves the most matching ATT&CK TTPs from the knowledge base based on the analyzed intent. An event may be mapped to multiple TTPs (techniques) or tacticalities, and the model outputs each mapping result and its corresponding confidence score.
[0037] As an example, given the input "A host executed cmd.exe / c powershell.exe -Enc XXXX", the model identifies it as a command and control (C2) and execution intent, mapped to T1059.001 (PowerShell), and gives a confidence score of 0.95.
[0038] 2.3: Output the annotation results generated by the large language model in a unified structured format.
[0039] The model strictly adheres to the JSON or XML format requirements defined in the cue word project when outputting analysis results. Output fields include: original event ID, normalized time, subject ID, ATT&CK tactical / technical number, EDR, model-generated confidence score, and a brief summary of attack intent.
[0040] Among them, the original event ID is the unique identifier of the original event, and the subject ID is the unique identifier of the subject asset associated with the original event in the asset list. Each original event is associated with and outputs only one subject ID, which is an existing asset identifier rather than a new identifier generated by the large language model.
[0041] As an optional implementation, a threshold is set, and TTPs with confidence levels below the threshold are considered low-quality labels; if the model outputs multiple conflicting TTPs (such as being labeled as "file deletion" and "sensitive file reading" at the same time), the labels are considered to be conflicting and require further processing.
[0042] Furthermore, when the model outputs low-confidence or conflict-ridden annotations, the system triggers a retrieval enhancement generation process. Using the original event text and low-confidence TTPs as queries, it retrieves higher-confidence TTP descriptions from external knowledge bases (such as the MITRE official knowledge base or security expert experience bases). If the results after retrieval correction are still below the threshold, the event will be marked as "pending manual review" or directly rejected, ensuring that only high-quality, high-confidence annotations proceed to subsequent steps.
[0043] Step 3: Construction of attack session sequences and generation of training data; Based on the labeled events output in step 2, a “context-attack path” training sample pair and a dense vector representation are generated following the process of unique subject → temporal aggregation → session segmentation → sequence construction → feature extraction → sample generation. These are used for subsequent attack path prediction, similarity retrieval, or model training.
[0044] 3.1: Based on the subject ID and session ID output in step 2.3, an aggregate subject ID is generated for each labeled event using a composite primary key mechanism; The subject ID is a unique identifier for the subject asset associated with the tagged event in the asset list; the session ID is a system-level login session identifier or communication session identifier extracted from EDR logs or network traffic; the aggregate subject ID is calculated from a composite primary key consisting of at least the subject ID and the session ID. As an example, the aggregate subject ID ( Generate as follows: ; Determining a unique entity ID ensures that events generated by different users or different login sessions on the same host can be accurately distinguished, preventing confusion between attack paths.
[0045] 3.2: According to All structured, labeled events are grouped, and the data within each group is arranged into a continuous event stream according to time series. In one optional implementation, for each group... implement Generate a subject-level event stream data structure, for example: ; In a further implementation, each continuous event stream is segmented into independent attack sessions. Specifically, a sequence segmentation algorithm based on time windows and state transitions is employed.
[0046] In a further specific implementation, the system traverses the ordered event stream within each subject group and calculates adjacent events. and time interval The preset time interval threshold is ,like Then The event is marked as the termination event of the current session. An event marked as the start of a new session indicates that the attacker may have completed the current phase of their action, been blocked, or switched attack tools / context; if the event... If the attack includes a success / failure feedback event (e.g., target process crashes, service completely stops, or EDR explicitly reports successful blocking), then The event is marked as the termination event of the current session. This is marked as the start of a new session. The preset time threshold can be set empirically based on domain knowledge or the average dwell time of historical attacks.
[0047] 3.3: Construct a complete attack session sequence, aggregating each independent attack session after segmentation according to time-sequential events located between the segmentation points. The core is to form a time-ordered chain of ATT&CK TTPs from the segmented session sequence. Complete the potential attack path.
[0048] 3.4: Extract tactical / technical layer features; Extracting consecutive ATT&CK TTPs (such as T1059.001) from the attack session sequence, each TTP number corresponds to a technical TTP in the MITRE ATT&CK framework, and each technical TTP is associated with its tactical identifier and tactical name to form the tactical / technical sequence characteristics corresponding to the attack session sequence.
[0049] Within the same attack session, a sequence of historical technique TTP numbers, identified and mapped to the MITREATT&CK framework, arranged chronologically by event occurrence, serves as the sequence of occurred TTPs. This sequence characterizes the attack steps already implemented by the attacker in the current session. This sequence of occurred TTPs can then be used as historical input for subsequent attack path prediction tasks.
[0050] Extract the execution parameters of each technical TTP in the sequence of occurred TTPs. The execution parameters refer to the parameter information used to characterize the specific execution method of the corresponding technical TTP, including but not limited to command line parameters, script content characteristics, payload characteristics, protocol type, communication port, URL or file path, request header fields, target object identifier, account information, permission characteristics, and other context fields related to the execution process of the technology.
[0051] Furthermore, the execution parameters of the TTPs are feature-encoded, converting textual, categorical, or structured parameter information into computable feature vectors using TF-IDF, Word2Vec, BERT, or other vectorized encoding methods. This forms a parameter vector sequence aligned positionally with the sequence of TTPs that have occurred. Each parameter vector in the parameter vector sequence corresponds one-to-one with a specific technical TTP at that position, collectively characterizing historical attack behaviors and their execution details.
[0052] 3.5: Context / Result Feedback Feature Extraction; For each attack session sequence, extract context / result feedback features from the aggregated subject ID. The aggregated subject ID is composed of a subject ID and a session ID and is used to uniquely identify the subject-session object of the attack sequence.
[0053] Environmental context features: Extract environmental context features related to the subject assets and their sessions corresponding to the aggregated subject ID from the asset / vulnerability / defense information collected in step 1.2, including: target asset type, network topology location, and defense policy deployment status.
[0054] Result feedback features: Based on the security system logs in step 1.1, the defense strategy status in step 1.2, and the structured event output in step 2, attack result fields associated with each event or session segmentation point within the session are extracted, and these attack result fields are encoded into computable features, specifically including: blocking status features, detection latency features, and response action features; wherein: Blocking status characteristics: Extract the "blocked / unblocked / unknown" blocking status from the handling result fields of security systems such as EDR / WAF / firewall, and encode it as a discrete numerical feature. Where 1 corresponds to blocked, 0 corresponds to unblocked, and -1 corresponds to unknown; Detection latency characteristics: Detection latency is calculated based on the attack occurrence timestamp and alarm issuance timestamp for the same event. By standardizing the time unit and performing truncation and normalization, numerical characteristics are obtained. ,in ,and The preset maximum delay threshold; Response Action Characteristics: Extract response action text from SOAR work orders, handling records, or safety personnel response logs, and map the response action to one or more category codes from a preset response action category set. One-hot or multi-hot encoding is then used to obtain the action vector. Alternatively, the response action text can be represented as a dense vector using BERT encoding; a result feedback vector can be generated based on the blocking state features, detection delay features, and response action features. In step 3.6, these features are concatenated by field along with the environmental context features and TTPs sequence features, or used as structured field inputs for training samples.
[0055] 3.6: Generate "context-attack path" training data; By integrating environmental context features, the sequence of TTPs that have occurred, the sequence of TTP execution parameter vectors aligned with the sequence of TTPs that have occurred, and the result feedback features, an instruction input is formed to guide the model to understand the current scene, and the subsequent TTPs or the next prediction step of the session sequence are taken as the expected output of the model.
[0056] The goal of this step is to convert the structured feature data into JSON format samples that conform to the training specifications for large language models.
[0057] For example: 1. Define the sample structure template use<Instruction,Input,Output> The triplet structure: Instruction: Used to specify the task type, taking the fixed text "Context-based prediction of the next attack TTP".
[0058] Input: A structured JSON object containing at least the following fields: session_id: Session identifier; asset_id: Asset identifier; env_context: A collection of environment context fields, including asset attributes, defense status, and vulnerability background; history_ttps: A sequence of historical TTP numbers, arranged in chronological order. history_params_vec: The execution parameter vector for each TTP, a sequence of parameter vectors aligned with history_ttps by position; feedback: A collection of result feedback fields, including blocking status, detection delay, and response action encoding results; Output: Supervision label field, which is next_ttp (next real TTPs number) when the task is single-step prediction; and future_path (real subsequent TTPs number sequence) when the task is path completion. When the training set is generated, one of the two is selected and the corresponding task scope is fixed in the Instruction.
[0059] Based on the attack session sequence and its corresponding TTPs sequence obtained in step 3.3, a training sample is generated at the starting point of each window according to the preset window length and step size. Among them, history_ttps is the sequence of historical TTP numbers arranged in chronological order within the window, history_params_vec is the sequence of parameter vector subsequences aligned with it in position, and Output is the next real TTP number in the single-step prediction task, so as to realize the next attack prediction based on context and historical behavior details.
[0060] 2. Encoding and Fusion Rules for Result Feedback Features block_state: The blocking status is encoded as a discrete value, with 1 for blocked, 0 for unblocked, and -1 for unknown; detect_delay_norm: Detects delays based on a preset maximum delay threshold. Perform truncation and normalization, and take... ; response_action_code: The response action is mapped to a preset action category code and written as a one-hot or multi-hot vector, or the response action text is written as a dense vector obtained by a preset encoding model; block_state, detect_delay_norm and response_action_code are stored as independent fields in Input.feedback, and different types of fields are concatenated in the form of unencoded raw text.
[0061] 3. Sample generation and alignment rules Based on the attack session sequence obtained in step 3.3 and their corresponding TTPs sequences Set the window length to Step size is For each window starting point Generate a training sample: history_ttps ; history_params_vec takes a subsequence of parameter vectors aligned with the position of history_ttps; env_context is a constant environment context field associated with this session; The feedback field is taken as the result feedback field that is aligned with the event level or session level. In a single-step prediction task, the output is set to next_ttp= .
[0062] 4. Example assumption: Step 3.3 constructs the attack session: session_id=S-9988, asset_id=A-1001, TTPs sequence: T1190→T1059.001→T1003. When =1, When =1, sample A and sample B are generated and written to the training set as JSON objects, as shown in the example below (the example field values are only for illustrating the field format): Sample A: history_ttps=[“T1190”], Output.next_ttp=“T1059.001”; Sample B: history_ttps=["T1059.001"], Output.next_ttp="T1003".
[0063] Step 4: Use the training data obtained after processing in Step 3 to train and optimize the large language model; 4.1: Selection of Large Language Model This invention uses open-source large language models based on Transformer, such as Llama-3 and Mistral, as its foundation. These models can handle high semantic density security event sequences, attack paths, TTPs sequences, and environmental contexts, and can capture complex temporal dependencies and environmental influences, thus meeting the needs of multi-step reasoning (attack path prediction) and numerical regression (risk scoring).
[0064] 4.2: Fine-tuning of domain data The base model is fine-tuned using general security domain corpora from the first phase (such as CVE reports, threat intelligence, and security blogs) in an unsupervised or weakly supervised manner. This allows the model to learn security domain terminology, conceptual associations, and basic grammatical structures. The word embeddings and attention mechanisms within the model are adjusted to adapt it to the security context. A relatively low learning rate is used to prevent the model from forgetting general knowledge.
[0065] 4.3: Fine-tuning of training samples The model is fine-tuned with supervised instructions using the generated training samples to minimize the cross-entropy loss between the output and the target sequence, so that the model accurately follows the safety instructions.
[0066] 4.4: Reinforcement Learning and Human Feedback (RLHF) Mechanisms By incorporating the rationality scores of the generated paths from red team experts, the output quality of the model in realistic attack and defense contexts is optimized. The red team experts score the rationality of the candidate attack paths generated by the model. This scoring occurs after instruction fine-tuning (SFT) and serves as input data for the reinforcement learning (RLHF) stage.
[0067] Step 5: Probabilistic decoding and feasibility confidence; A probabilistic decoding strategy (such as Top-k sampling or kernel sampling) is adopted to generate multiple candidate attack paths with different probabilities based on the constraints of the global environment context vector. The feasibility confidence of each candidate attack path is calculated. The feasibility confidence corresponds to the conditional probability distribution of each candidate attack path output by the model. The global environment is the real-time environment context during decoding, which includes at least asset profiles, vulnerability lists, and defense status, and may further include network topology and access control information related to the assets. 5.1: Data basis for generating candidate attack paths using a probabilistic decoding strategy The data foundation consists of the model parameters (i.e., model knowledge) after training and the real-time input of asset profiles, vulnerability lists, and defense status (i.e., Prompt Context). Based on these inputs and its training knowledge, the model generates the probability distribution of the next TTPs.
[0068] 5.2: Feasibility confidence of each candidate attack path The feasibility confidence level is directly related to the conditional probability output by the probabilistic decoding strategy. In the autoregressive generation process, the joint conditional probability of the path... Defined as the product of the conditional probabilities of all TTPs in the path: ; in, It is the position number of the TTP currently being calculated in the sequence. This represents the total number of TTPs included.
[0069] Probabilistic decoding strategies (such as Top-k sampling and kernel sampling) generate joint conditional probabilities for multiple attack paths based on the probability distribution of the model output. Candidate attack paths. Joint conditional probability of candidate attack paths. It is used as the confidence level of the feasibility of this path in the current environment.
[0070] Step 6: Calculate the quantitative risk index based on the generated set of candidate attack paths.
[0071] 6.1: Determination of environmental consistency and screening of feasibility confidence. The system is based on a large language model and employs a probabilistic decoding strategy, including Top-k sampling, kernel sampling, or beam search algorithms, to generate multiple candidate attack paths. These candidate attack paths are sequences of ATT & CK TTPs arranged in chronological order.
[0072] The "preferred path" in the candidate attack paths refers to a path that meets any of the following conditions: (1) During the bundle search process, the top N positions are ranked according to feasibility confidence; (2) Its feasibility confidence level is not lower than the preset path feasibility threshold. The preset path feasibility threshold A feasibility confidence threshold used to determine whether a candidate attack path should be included in the reserve set; preferably, 0.6 is acceptable.
[0073] The top N ranking rule and the threshold comparison rule are not two different threshold judgments, but rather two ways for candidate paths to enter the preferred path set: the former is a ranking-based retention mechanism, and the latter is a threshold-based retention mechanism. The quantitative risk index of the candidate attack path can be comprehensively calculated by integrating prior information from the knowledge base, global environmental context information, and the feasibility confidence of the model output.
[0074] 6.1.1: Ensure the attack path environment based on the processed data. The asset profile, vulnerability list, and defense deployment status are used as the initial context input to the large language model to generate candidate attack paths. For each candidate attack path output by the model, an environment consistency determination is first performed; for candidate attack paths that meet the environment consistency requirements, their joint conditional probability is calculated based on step 5.2, and the joint conditional probability is used as the feasibility confidence of the path.
[0075] Furthermore, the feasibility confidence of the candidate attack path is compared with a preset path feasibility threshold. Compare; when ≥ When, retain the candidate attack path; when < If the candidate attack path is not found, it is eliminated. Therefore, the environmental consistency determination and the feasibility confidence threshold comparison together constitute the final retention conditions for candidate attack paths.
[0076] The environmental consistency determination includes: extracting the first candidate attack path from the pre-set TTP precondition knowledge base. A set of preconditions for each TTP ,in At a minimum, this includes required service and port conditions, affected software version range or associated CVE number conditions, required permissions or credentials conditions, required network reachability conditions, and observable technical characteristics conditions; a set of vulnerability evidence is constructed based on the asset profile, vulnerability list, and defense deployment status. ,in At a minimum, this includes open ports and service versions, unpatched CVE IDs and their CVSS base score and availability metrics, rule set IDs or policy IDs of defensive measures and their action types, as well as network segmentation and access control rules; execute for each TTP. and The matching operation is used to obtain the matching indicator. , where if and only if exist When there are conditions that satisfy the conditions and there is no logical conflict with the defense deployment status or access control rules. =1, otherwise =0; when the candidate attack path satisfies that it has a probability of success for all TTPs within the path. When the value is 1, the candidate attack path is determined to meet the environmental consistency requirements and is retained; otherwise, it is discarded.
[0077] 6.1.2: Obtain the attack path sequence and feasibility confidence level through the decoding module; In the process of generating candidate attack paths, random sampling, kernel sampling, Top-k sampling, or beam search decoding strategies are employed to obtain the joint conditional probability of the path given by the model for each candidate attack path. and the path joint conditional probability The confidence level of the corresponding candidate attack path is used for subsequent sorting, screening and risk assessment.
[0078] For candidate attack paths generated using beam search, if their feasibility confidence ranks among the top N in the current decoding results, then the candidate attack path can be inserted into the candidate attack path set; for candidate attack paths generated using a sampling strategy, if their feasibility confidence is greater than or equal to a preset path feasibility threshold... If so, the candidate attack path can be inserted into the candidate attack path set. In other words, a candidate attack path satisfies either "being in the top N positions of the bundle search ranking" or "having a feasibility confidence level not lower than the path feasibility threshold". "Any one of these two conditions can be inserted into the candidate attack path set."
[0079] Furthermore, by traversing the candidate attack path set, the ATT&CK TTPs sequence corresponding to each candidate attack path, arranged in chronological order, and the corresponding initial feasibility confidence level are output, providing input for subsequent environmental consistency determination, path success rate calculation, potential business loss estimation, and comprehensive risk scoring.
[0080] Preferably, the preset path feasibility threshold The first threshold is a preset constant used for quality control of the sampled path generation; the second threshold, the top N-bit retention rule for beam search, is used for size control of the beam search generated path. Both are used for the initial screening of candidate attack paths and do not imply the existence of two different feasibility confidence thresholds.
[0081] 6.2: Quantitative Calculation of Path Success Rate The path success rate is derived from the success rate of similar TTP combinations in historical attack data, the degree of defense deficiencies in the target system, and environmental reachability, through joint inference by the model. The attack duration estimate is based on the cumulative modeling of the average execution time and dependencies of TTPs at each stage. The potential business loss is mapped to a standardized loss score according to the importance of the affected assets (e.g., database servers have a higher weight than ordinary terminals), data sensitivity, and service continuity requirements.
[0082] 6.2.1: Calculation Formula and Factor Solution Path success rate It is the joint result of the success probabilities of all TTPs in the path: ; The feasibility probability of TTP is generated based on the matching results between asset service fingerprints and the range of versions affected by vulnerabilities. The system generates a service fingerprint (SF) based on the list of running software, service versions, and open ports obtained in step 1.2. The service fingerprint includes at least the protocol, port, service name, vendor, product, version, and CPE identifier. This is used for candidate attack steps. Search its associated vulnerability set And the affected version range (CVE) for each vulnerability, and perform version matching determination. , where there exists if and only if satisfy When the port corresponding to the service is open, it is determined that the target asset exists with... Corresponding exploitable weaknesses; when matching fails or the service is not running, The constraint is 0 or close to 0; when a match is successful, the system reads the EPSS score, CVSS exploitability component, PoC public status and in-the-wild exploitation frequency of the vulnerability from the historical attack knowledge base, and statistically analyzes similar attacks from the historical attack event set. Success rate under the same or similar asset fingerprint conditions The above factors are normalized and then input into the calibration function. To generate, ,and The output is limited to the interval [0,1].
[0083] : Probability of defense blocking. For the attack steps to be evaluated. The defense strategy configuration for the target asset is parsed into a set of defense measures. Each of these defensive measures It should at least include a defense type identifier (firewall / ACL, WAF, EDR / NDR, etc.), rule set ID or policy ID, effective status, action type (block / alarm only / allow), matching conditions (protocol, port, URL characteristics, command line characteristics, process tree characteristics, file hash / path characteristics, etc.), and applicable scope (source / destination network segment, asset tag, user / process subject). Analysis into a set of technical features The set of technical features At least including the Required protocols and ports, payload and parameter types (e.g., PowerShell code execution, script suffixes, deserialization characteristics, SQL injection payload patterns), trigger points (network layer / host layer / application layer), and preconditions and observable indicators (e.g., process creation, command-line arguments, network connection destination port, HTTP request path and header). Based on a pre-built TTP-defense rule mapping knowledge base. And semantic matching mechanisms, for and Perform a hit determination to obtain the specific defensive measures. Hit indicator and hit strength Hit strength is used to characterize the Key technical features and the aforementioned defense measures The degree of consistency between the matching conditions; when the defense measures When the action type is "alarm only", mark its blocking contribution as unblockable and apply a blocking reduction factor to the hit strength. To obtain the effective hit strength of the block .
[0084] The blocking probability is decomposed into a weighted calibration result of "model-predicted blocking probability" and "historical statistical blocking rate", wherein the model-predicted blocking probability The large language model trained in step 4, when inputting < , The output is given under the condition of a vulnerability list >, and the output is the conditional probability of the "attack blocked" event; the historical statistical blocking rate Based on the historical alarm database, similar requirements are met. And the defense configuration and A set of identical or similar events The statistics show that: ; in, For smoothing, similarity determination is based at least on consistency of rule set ID / policy ID, rule version, scope of application, and hit strength range. The two types of probabilities are fused according to weighted coefficients to obtain the defense blocking probability: ; in, , This is a calibration function used to limit the fusion result to the [0, 1] interval and to achieve temperature scaling or logistic regression calibration; when When the sample size is below the preset minimum threshold, Adjust to a lower value and increase accordingly. When the set of defense measures There is a clear "blocking" action and the following conditions are met: and When the rule is not lower than the hit strength threshold, the above will be applied. by As a constraint factor, it is adjusted upwards; when Zhongyu Inconsistent triggering levels (e.g., only network layer ACLs exist, while the aforementioned...) When it is a purely host-local behavior, the above Penalty coefficient for hierarchical inconsistency A reduction is applied to reflect the impact of the defense strategy configuration on the... The actual blocking capability.
[0085] Environmental reachability probability. Based on environmental context such as network segmentation and firewall ACLs, assess whether an attacker can reach the target asset required for the current TTP from the host in the previous step.
[0086] The source host to the target asset in the A set of reachability determination rules for the required communication elements, wherein the set of rules includes at least the following: Communication requirements standard: to include the aforementioned Parsing into a communication demand vector The communication requirement vector includes at least: a set of source endpoints (IP / subnet / security group of the host in the previous step), a set of destination endpoints (IP / subnet / security group of the target asset), a protocol type (TCP / UDP / ICMP / HTTP(S) etc.), a destination port or port range, a direction (inbound / outbound), a connection initiator (client / server role), and whether it depends on an intermediate stepping stone or proxy.
[0087] Network Segmentation and Routing Reachability Criteria: Constructing a Directed Graph Based on Network Topology and Segmentation Context The nodes must include at least a subnet, network segment, availability zone, VLAN, and VRF / routing domain, and the edges must include at least a routing entry, a Layer 3 forwarding relationship, and a NAT / gateway connectivity relationship. The "routing reachability" condition is met if and only if there exists a routing path from the source endpoint to the destination endpoint, and the routing domains of each segment on the path are consistent or there is an explicit cross-domain routing / gateway forwarding table entry. The "routing unreachable" condition is met if there is a black hole route, a missing default route, VRF isolation, or an explicit segmentation policy that prohibits cross-segment communication.
[0088] Necessary intermediate node criteria: when the When it is specified that access must be via intermediate nodes such as jump hosts, bastion hosts, proxies, API gateways, and service mesh entry points, the reachability determination is extended to a series determination of multiple links.
[0089] 6.2.2: Joint Inference of Models Based on the path success rate definition in step 6.2.1, the system analyzes each attack step in the path. Calculate the feasibility probability separately , probability of defense blocking and the probability of environmental accessibility The component probabilities are then subjected to prior fusion and calibration processing, and the calibrated component probabilities are substituted into the joint calculation formula in step 6.2.1 to obtain the path success rate. .in, Prior fusion and calibration of feasible components: targeting The associated vulnerability set is matched with the service fingerprint results to generate vulnerability priors. The vulnerability prior term is obtained by normalizing and mapping the CVSS exploitability component, EPSS score, PoC public state, and in-the-wild exploitation frequency; the model outputs a feasibility prediction term based on the input features. Performing a weighted fusion on the two and restricting them to the [0,1] interval using a calibration function, we obtain: ; in, As a feasible activation function, the result is mapped to the interval [0,1]. The weights of the feasibility prediction items, The weight of the vulnerability prior.
[0090] Statistical calibration of the blocking component: The blocking prediction term is output from the model trained in step 4. The statistical blocking rate under similar TTPs and defense configurations is obtained from the historical alarm database. The two are fused using adaptive weighting based on sample size and then calibrated to obtain: ; in, As the activation function, it maps the result to the interval [0,1]. for The weights of statistical items are adjusted; when the sample size of events used for statistics is lower than a preset threshold, the weights of statistical items are reduced and the weights of model items are increased.
[0091] Reachability component evaluation by rules: based on the communication requirement vector described in step 6.2.1. Network topology diagram The access control policy set is used to perform route reachability and policy hit determination to obtain the reachability determination result. ;Will Mapped to reachability probability ,in Take 1 at time. Take 0 at time, The value is set to (0,1) according to the preset missing value reduction factor.
[0092] 6.3: Attack Market Estimation ( Modeling By using a time-series accumulation model, a time-dimensional quantitative indicator is added to the path.
[0093] 6.3.1: Duration Modeling and Dependencies Attack duration ( ) is the average execution time of TTPs at each stage in the path ( ) and the dependency waiting time between each stage ( The sum of ( ). Attack duration The formula is: ; in, This is the stage number. , This represents the total number of stages contained in the entire attack path. (Average execution time): This comes from historical attack data statistics or threat intelligence knowledge bases. For example, vulnerability exploitation (T1190) may only take a few seconds, while credential dumping (T1003) may take tens of seconds.
[0094] (Dependency / Waiting Time): Models the time required for an attacker to think, scan, or wait for feedback between completing one TTP and executing the next TTP. This time can be determined by historical data distribution or a fixed threshold (e.g., 1 minute).
[0095] 6.4: Potential Business Losses ( Standardized calculation Transform technological risks into quantifiable losses at the business level.
[0096] 6.4.1: Standardized Score Mapping The affected assets are mapped to a standardized loss score (e.g., in the range of 0-10) based on their business value and data sensitivity.
[0097] Asset Importance: Assign high weight to critical business systems and low weight to ordinary office terminals.
[0098] Data sensitivity: Assets involving core intellectual property or citizens' privacy data are given additional weight.
[0099] Service continuity requirements: Assets that cannot be downtime must be assigned the highest loss weight.
[0100] In one specific implementation, the system pre-sets an "asset value quantification mapping rule" to map each affected asset... Mapped to standardized loss score The mapping is determined by three types of fields: asset role, data sensitivity, and service continuity requirements.
[0101] Table 1. Asset Role Classification and Definition Standards
[0102] Table 2. Data Sensitivity Classification and Definition Standards
[0103] Table 3. Service Continuity Requirements Classification and Definition Standards
[0104] Table 4. Formula for Standardized Loss Score
[0105] assets Standardized loss score definition: ; in, For assets Data sensitivity weighted increment, For assets The weighted increment corresponding to service continuity, For assets The corresponding asset / role level.
[0106] 6.4.2: Determining the Loss Value Define the set of assets affected after a successful attack path as follows: It includes at least: assets that are targeted or accessed / controlled in the attack path, and assets that can be further accessed under the conditions of path final state permissions and network reachability.
[0107] The final potential business loss value is the largest standardized loss score among the affected asset sets: .
[0108] 6.5: Comprehensive Risk Score ( ) Calculation and classification Comprehensive Risk Score It is the success rate of the path. and potential business losses The weighted summation, and the introduction of duration The product of functions. Comprehensive risk score.
[0109] , These are weighting coefficients, adjusted according to business needs (e.g.) (The weight is usually higher).
[0110] Duration penalty function: , Attack duration The shorter the attack duration, the more efficient and difficult the attack is to respond to, and the higher the risk.
[0111] Based on comprehensive risk score Values are used to classify risk levels: if If so, the risk level is high; if If the risk level is 0, the risk level is medium; otherwise, it is low. This risk level definition provides quantifiable risk input for subsequent defense optimization suggestions and decisions.
[0112] Step 7: Based on the quantitative risk indicators of attack paths and real-time threat intelligence, dynamically quantify the effectiveness of the currently deployed defense strategies, and generate a list of defense optimization suggestions based on a multi-dimensional priority ranking model.
[0113] Output-based quantitative risk indicators (R-value, , It identifies attack paths and dynamically evaluates the effectiveness of existing defense strategies, outputting priority-based optimization suggestions with quantifiable risk reduction effects.
[0114] 7.1: Quantification of Defense Effectiveness For each preferred attack path, calculate the defense coverage rate and blocking confidence. The defense coverage rate is the proportion of TTPs in the attack path that are covered by existing defense measures (such as WAF rules, EDR behavior detection policies, and network ACLs). The blocking confidence is a prediction of the probability of the defense measures taking effect in the current specific environmental context based on a large language model. This prediction can be weighted or calibrated by combining the success rate statistics of defense measures in historical attack events. 7.1.1: Defense Coverage ( )calculate The calculation formula is ; The coverage determination module is used to check each technical / tactical element in the attack path item by item based on the pre-set TTPs defense rule mapping knowledge base: a. The knowledge base stores the mapping relationship between TTP identifiers and defense rules in a structured manner, such as establishing a unique mapping entry between T1059.001 (PowerShell) and the EDR policy "Suspicious Command Line Execution Detection"; b. During the calculation phase, the system traverses the TTP sequence contained in the attack path and retrieves the set of defense measures deployed on the target asset, including but not limited to ACL, WAF, and EDR policies; c. If a rule in the defense measure set has a clear mapping relationship with any TTP, then the TTP is marked as "covered"; otherwise, it is marked as "uncovered," and the path risk score is updated accordingly. The mapping knowledge base supports dynamic updates. When a new defense rule or TTP version is added, the system expands the mapping entries through incremental writing without recompiling the core evaluation logic.
[0115] 7.1.2: Blocking confidence ( )calculate Blocking confidence This is a quantitative metric for a single attack path. (The metric pertains to the path.) The system first processes each step Calculate its step-level blocking probability under the current environment and current defense configuration. Then, the path-level blocking confidence is obtained by aggregation. .
[0116] 1. Calculation of step-level blocking probability Weighted or calibrated based on large language model prediction results and historical statistics: ; in, , These are the weighting coefficients. .
[0117] Model-predicted blocking probability: The output of the blocking probability prediction model obtained through step 4, whose input includes at least < , ,env_context>,where The set of defensive measures effective against the target asset, collected and parsed from step 1.2 (including rule set / policy ID, action type "block / alarm only / allow", matching conditions, and applicable scope, etc.), is defined by `env_context`, which represents the environmental context such as asset profile / vulnerability list / network segmentation. The model outputs the conditional probability of the event "attack in this step is blocked," i.e.: ; (Historical statistical blocking rate): Filtering from the historical security event database and... Similar in type and with the same defense configuration A set of identical or similar events The statistics show that: ; in, For a single event in the historical security event database, For the event State attributes, As a smoothing factor, to prevent the denominator from being 0, This represents the total number of events.
[0118] It should be noted that similarity determination is based at least on the consistency of rule set ID / policy ID, rule version consistency, scope of application consistency, and similarity of hit conditions.
[0119] This is a calibration function used to limit the fusion results to the [0,1] interval and perform temperature scaling or logistic regression calibration; when Reduce when below the minimum sample threshold ,improve .
[0120] 2. Aggregation of Path-Level Blocking Confidence Aggregating step-level blocking probabilities into path-level blocking confidence, and adopting the assumption that "blocking at any step leads to path failure," then: ; in, For the current stage of calculation, This represents the total number of stages contained in the attack path.
[0121] 7.1.3: Validity Determination Overall blocking confidence of comparing attack paths With preset blocking threshold .
[0122] Threshold setting: Preset blocking threshold It is typically set to >= 0.5. Below this threshold, it means the model considers the defense measures to have a high probability of failure.
[0123] Judgment result: If < If the defense strategy is deemed "ineffective" or "partially effective" in this scenario, then the TTPs corresponding to this path have a defense gap.
[0124] 7.2: Automatic generation of defense optimization suggestions Based on the evaluation results in section 7.1, actionable defense optimization suggestions are automatically generated for attack paths with risk scores exceeding a set threshold.
[0125] 7.2.1: Risk-driving mechanisms and thresholds Risk score ( Corresponding comprehensive risk score ( ).
[0126] The threshold can be dynamically adjusted based on the organization's risk tolerance. A risk screening threshold is set. (For example ).only > Only then does it enter the suggestion generation process.
[0127] 7.2.2: Suggestions for Large Model Generation The TTPs sequence of high-risk paths and the corresponding defense gaps ( < The TTPs and real-time threat intelligence are used as Prompt inputs to the large language model.
[0128] The inference engine is configured to invoke a trained security knowledge base to perform targeted inference on gap defense measures: when the blocking confidence of T1059.001 is detected to be lower than the threshold, the model outputs at least one defense reinforcement suggestion to fill the gap, including: enabling deep detection rules for encoded PowerShell commands in the EDR policy, or restricting PowerShell execution permissions for non-administrator accounts; the suggestion is automatically generated based on historical successful blocking records and policy-attack mapping relationship, and is dynamically updated as the confidence changes.
[0129] 7.3: Multidimensional Priority Ranking Model A multi-dimensional priority ranking model is introduced to rank the generated defense optimization suggestions by comprehensively considering risk urgency, remediation cost, and multi-path coverage gain. Wherein, the urgency of the risk = path success rate ( ) × Potential loss ( The repair cost includes the complexity of strategy adjustment and whether downtime is required. The multi-path coverage gain is the number of high-risk attack paths that the defense reinforcement suggestion can cover and the amount of risk reduction it brings.
[0130] 7.3.1: Ranking Models and Weights Model formula: using a weighted scoring function :
[0131] Weight setting basis: weight ( This reflects the risk-driven principle. Typically... ;For example: =0.5, =0.5, =0.1, =0.1.
[0132] 7.3.2: Quantification of Ranking Factors Risk urgency : ; Reflected in the ranking model: Give the highest weight to recommendations that ensure a high success rate and a high loss rate.
[0133] Repair costs : ; Complexity classification ( ): A 1-5 level quantification standard is adopted. For example: Level 1 (automation / policy fine-tuning); Level 3 (requires application restart / simple code modification); Level 5 (requires hardware deployment / architecture refactoring).
[0134] Does the machine need to be shut down? ): Requires shutdown =1; No need to stop =0; Effect on cost: (Stop weight) should be much higher than (Complexity weight), for example =0.8, =0.2.
[0135] Multipath Coverage Gain : ; in, This is a set of high-risk paths. The path risk score, This indicates that the recommendation is being applied. The amount of reduction in post-path risk (e.g., making Decrease or make An increase thus leading to decline), Lower the contribution threshold for risk.
[0136] Reflected in the ranking model: for Higher-scoring recommendations will receive bonus points, and will be given priority in recommending remedial measures that can significantly reduce risks across multiple high-risk paths.
[0137] 7.4: Output a list of optimization suggestions Output a list of defense optimization suggestions sorted in descending order of priority, with the expected risk reduction indicated.
[0138] Table 5, the recommended list format is as follows:
[0139] As an embodiment of the present invention, such as Figure 2 As shown, a security measurement and evaluation system based on a large model is disclosed. The system employs a specific implementation of the aforementioned security measurement and evaluation method and includes: Data Acquisition and Labeling Module: Collects multi-source heterogeneous security data from network security systems, preprocesses the security data to generate security event data, inputs the security event data into a large language model for automated analysis, and maps the security event data to a network security knowledge base and model framework to generate structured event data with labels; The data processing feature module constructs an aggregated subject with subject ID and session ID as composite primary keys based on structured event data; groups the structured event data into event streams using the aggregated subject; and divides each event stream into independent attack sessions to construct an attack session sequence; extracts the environmental context features, result feedback features, the technical sequence of events, and execution parameter features of the attack session sequence to fuse and generate training samples for training large language models. The training optimization and generation module uses training samples to train and optimize the large language model; it calls the trained large language model to generate multiple candidate attack paths and calculates the feasibility confidence of each candidate attack path; based on the feasibility confidence of each candidate attack path, it calculates quantitative risk indicators for each candidate attack path, including path success rate, potential business loss, and comprehensive risk score. The evaluation and optimization module is implemented to evaluate existing defense strategies based on quantitative risk indicators and the candidate attack paths, and to generate a list of defense optimization suggestions sorted by priority based on the evaluation results.
[0140] As an embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it employs the specific implementation described in the above-described security measurement and evaluation method.
[0141] As an embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, employs the specific implementation described in the above-described security measurement and evaluation method.
[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A security measurement and evaluation method based on a large language model, characterized in that, Includes the following steps: Collect multi-source heterogeneous security data from network security systems, and preprocess the security data to generate security event data; Security incident data is input into a large language model for automated analysis, and the security incident data is mapped to a cybersecurity knowledge base and model framework to generate tagged structured incident data. Based on structured event data, an aggregated subject is constructed with subject ID and session ID as composite primary keys; The structured event data is grouped by the aggregation subject to form an event stream; each event stream is then divided into independent attack sessions to construct an attack session sequence; the environmental context features, result feedback features, the technical sequence that has occurred, and the execution parameter features of the attack session sequence are extracted and fused to generate training samples for training large language models; The large language model is trained and optimized using training samples; the trained large language model is called to generate multiple candidate attack paths and the feasibility confidence of each candidate attack path is calculated; based on the feasibility confidence of each candidate attack path, quantitative risk indicators including path success rate, potential business loss and comprehensive risk score are calculated for each candidate attack path. Based on the quantitative risk indicators and the candidate attack paths, the existing defense strategies are evaluated, and a list of defense optimization suggestions ranked by priority is generated based on the evaluation results.
2. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The generated security event data includes: The system collects data from multiple security systems, including Security Information and Event Management (SIEM), Endpoint Detection and Response (EDR), Network Detection and Response (NDR), vulnerability scanners, and asset management, covering attack behavior logs, vulnerability and asset information. The collected data undergoes preprocessing, including data cleaning, format normalization, and semantic alignment, to generate semantically unified security event data.
3. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The generation of tagged structured event data includes: Construct a prompt word template that includes character settings, input fields, output fields, and template requirements; The security event data is input into the large language model according to the prompt word template format; The large language model maps each security incident data to multiple tactical and technical TTPs in the MITRE ATT&CK framework and outputs the mapping result of each security incident data and its corresponding confidence score. The large language model analyzes and outputs tagged structured event data according to the format of the prompt word template; the structured event data includes the original event ID, normalized time, subject ID, ATT&CK tactical and technical number, confidence score and EDR.
4. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The structured event data is grouped based on a composite primary key to form an event stream, including: Based on the structured event data, the main assets associated with the event are matched and determined from the preset asset list; Obtain the unique identifier pre-stored in the asset list for the main asset, and use it as the main asset. ; Extract the login session identifier from the structured event data to identify the session. ; use =Main Body Username_session Constructing the aggregate entity ; by For identification purposes, the structured event data is grouped, and the event data within each group is arranged in chronological order to form a continuous event stream.
5. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The step of dividing each event stream into independent attack sessions and constructing an attack session sequence includes: Traverse the event stream and calculate adjacent events. , Time difference ; Preset time interval threshold ,like Then The event marked as the start of a new session; If the event If it includes attack success and failure feedback events, then it will The event marked as the start of a new session; Based on the above two rules, the dividing point identified in the continuous event stream is the split point; Events that are located between two split points and are time-continuous are aggregated in chronological order to form an attack session sequence corresponding to an independent attack session.
6. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The generation of training sample pairs for training large language models includes: Extract the sequence of techniques that have occurred in the attack session sequence; Extract the execution parameter features corresponding to each technology in the already occurred technology sequence; For the aggregated entities corresponding to the attack session sequence, environmental context features are extracted from asset profiles, vulnerability lists, and defense information; Based on the logs and structured events in the security system data, extract the results associated with the in-session events, including blocking status, detection delay, and response actions, as result feedback features; The occurrence of the technical sequence, the execution parameter features, the environmental context features, and the result feedback features are integrated to form training samples for training large language models. The occurrence of the technical sequence includes extracting multiple technical numbers that are identified and mapped to a pre-set attack framework in chronological order from the attack session sequence to form tactical technical sequence features; The extraction of execution parameter features includes, for each technology in the sequence of occurrences, extracting one or more of the following: command line parameters, script content, payload features, protocol type, port, URL, file path, request header, target object identifier, account information, and permission features, to form a set of technical execution parameters; The tactical technology sequence is characterized by a technology number sequence within the MITRE ATT&CK framework.
7. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The calculation of the feasibility confidence of each candidate attack path includes: The large language model is trained and optimized based on the training samples. Based on a well-trained large language model, multiple candidate attack paths are generated through probabilistic decoding; For each generated candidate attack path, an environmental consistency determination is performed. The determination includes matching the set of preconditions for each technical TTP in the attack path with the set of vulnerability evidence in the current environment. Calculate the joint conditional probability for each candidate attack path; the joint conditional probability is the product of the conditional probabilities of all techniques in the candidate attack path. The calculated joint conditional probability is the feasibility confidence of the corresponding candidate attack path; Feasibility confidence With threshold When comparing, ≥ If the condition is met, the candidate attack path will be retained; otherwise, it will be eliminated. Each generated candidate attack path corresponds to a joint conditional probability.
8. The security measurement and evaluation method based on a large language model according to claim 7, characterized in that, The process of matching the set of preconditions for each technique TTP in the attack path with the set of vulnerability evidence in the current environment includes: Extracting the first candidate attack path from the pre-defined TTP precondition knowledge base A set of preconditions for each TTP ; A set of vulnerability evidence is constructed based on asset profiling, vulnerability lists, and defense deployment status using environmental context features. ; Perform for each TTP and The matching operation yields the matching indicator. ; If and only if exist When there are conditions that satisfy the conditions and there is no logical conflict with the defense deployment status or firewall access control rules. ,otherwise ; When the candidate attack path satisfies that all TTPs in the corresponding path are available If the candidate attack path meets the environmental consistency requirements, it is retained; otherwise, it is discarded.
9. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The quantitative risk indicators used to calculate each candidate attack path include path success rate, potential business loss, and comprehensive risk score, including: For each technical step in the candidate attack path, calculate its technical feasibility probability, defense blocking probability, and environmental reachability probability. The success probability of the technical step is obtained by jointly calculating the probability of technical feasibility, the probability of defense blocking, and the probability of environmental reachability; the success rate of the candidate attack path is obtained by jointly calculating the success probabilities of all technical steps in the path.
10. The security measurement and evaluation method based on a large language model according to claim 9, characterized in that, The probability of technical feasibility is based on the matching result between the service fingerprint of the target asset and the affected version range of the vulnerability associated with the technology, as well as the vulnerability's public vulnerability score, exploitability score and historical success rate. The defense blocking probability is the model blocking probability predicted based on the current defense configuration of the target asset, and the historical statistical blocking rate obtained from the historical alarm database under similar technologies and defense configurations. The environmental reachability probability, based on the analysis of network topology, routing policies, and access control rules, determines whether the communication path from the host in the previous step to the target asset required in the current technical step is reachable.
11. The security measurement and evaluation method based on a large language model according to claim 9, characterized in that, The calculation includes quantitative risk indicators such as path success rate, potential business losses, and comprehensive risk score, and also includes: The potential business loss is obtained by using asset value quantification mapping rules, based on asset role classification, data sensitivity classification, and service continuity requirement classification to obtain the maximum value of the standardized loss score of the asset.
12. The security measurement and evaluation method based on a large language model according to claim 9, characterized in that, The calculation includes quantitative risk indicators such as path success rate, potential business losses, and comprehensive risk score, and also includes: The comprehensive risk score is calculated by weighting the path success rate and potential business losses and multiplying it by a duration penalty function; The duration penalty function is negatively correlated with the attack duration. The attack duration is calculated by traversing each technical step in the candidate attack path, summing the average execution time and the dependency waiting time, and then calculating the attack duration.
13. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The assessment of existing defense strategies includes: The evaluation results of existing defense strategies depend on the overall blocking confidence level; For each candidate attack path, calculate the blocking probability for each attack step; Calculate the overall blocking confidence of the path based on the probability that all steps in the acquisition path are not blocked; The effectiveness of existing defense strategies against the current attack path is determined based on the overall blocking confidence level.
14. The security measurement and evaluation method based on a large language model according to claim 13, characterized in that, The determination of whether the existing defense strategy is effective against the current attack path based on the comprehensive blocking confidence level includes: Comprehensive blocking confidence With preset blocking threshold Compare; like If it is valid, it is considered valid; otherwise, it is considered invalid or partially valid.
15. The security measurement and evaluation method based on a large language model according to claim 1, characterized in that, The generated and priority-sorted list of defense optimization suggestions includes: Set a risk screening threshold, and select attack paths with a comprehensive risk score higher than the risk screening threshold as high-risk attack paths; For the technologies in the high-risk attack paths where the defense measures are determined to be ineffective or partially effective, corresponding defense reinforcement suggestions are generated; The defense reinforcement recommendations are ranked based on a multi-dimensional priority ranking model. The evaluation dimensions of the ranking model include at least risk urgency, repair cost, and multi-path coverage gain. The risk urgency is based on the product of the success rate of the corresponding attack path and the potential business loss; the repair cost is quantified based on the complexity level of the strategy adjustment and whether downtime is required; the multi-path coverage gain is the number of high-risk attack paths that the defense reinforcement recommendation can cover and the amount of risk reduction it brings.
16. A security measurement and evaluation system based on a large language model, executing the security measurement and evaluation method as described in any one of claims 1-15, characterized in that, The system includes: Data Acquisition and Labeling Module: Collects multi-source heterogeneous security data from network security systems, preprocesses the security data to generate security event data, inputs the security event data into a large language model for automated analysis, and maps the security event data to a network security knowledge base and model framework to generate structured event data with labels; The data processing feature module constructs an aggregated subject with subject ID and session ID as composite primary keys based on structured event data; groups the structured event data into event streams using the aggregated subject; and divides each event stream into independent attack sessions to construct an attack session sequence; extracts the environmental context features, result feedback features, the technical sequence of events, and execution parameter features of the attack session sequence to fuse and generate training samples for training large language models. The training optimization and generation module uses training samples to train and optimize the large language model; it calls the trained large language model to generate multiple candidate attack paths and calculates the feasibility confidence of each candidate attack path; based on the feasibility confidence of each candidate attack path, it calculates quantitative risk indicators for each candidate attack path, including path success rate, potential business loss, and comprehensive risk score. The evaluation and optimization module is implemented to evaluate existing defense strategies based on quantitative risk indicators and the candidate attack paths, and to generate a list of defense optimization suggestions sorted by priority based on the evaluation results.
17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the security measurement and evaluation method according to any one of claims 1-15.
18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the security measurement and evaluation method according to any one of claims 1-15.
Citation Information
Patent Citations
Information security assessment method and system based on cloud platform
CN119691754A
Multi-modal learning and deep learning technology-based operating system security evaluation system and method
CN120337225A
Security state evaluation method and device of network environment, equipment and storage medium
CN120342697A