LLM-based 6G network automatic security processing method and system

Through the LLM-based 6G network automatic security disposal method, combined with the intrusion detection model and blockchain evidence storage, the problems of high missed reporting rate of advanced threats and long response cycle in existing technologies are solved, and efficient and reliable network security management is achieved.

CN120711397APending Publication Date: 2025-09-26TERMINUSBEIJING TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510901706.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing network security protection technologies cannot effectively parse encrypted traffic or correlate advanced threats in multi-source logs, resulting in a high rate of missed reports; the handling strategies lack understanding of dynamic network context and have a long response cycle; and they lack blockchain evidence storage capabilities, making it difficult to meet the 6G network's requirements for real-time, automated, and trusted evidence storage.

Method used

An LLM-based 6G network automatic security disposal method is adopted. By obtaining network data streams, extracting structured log data, and using pre-trained intrusion detection models and LLM security models to generate network security disposal strategies, the strategies are uploaded to the blockchain system for evidence storage.

Benefits of technology

It realizes intelligent identification and automatic disposal of network attacks, ensures the traceability and data integrity of the disposal process, improves the intelligence and credibility of network security management, and ensures timely response and high security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120711397A_ABST
    Figure CN120711397A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an LLM-based 6G network automatic security processing method and system. The method is applied to the field of network security intelligent protection, and comprises the following steps: acquiring a network data stream, generating a structured log, extracting features, inputting the features into an intrusion detection model to identify attack behaviors, performing format conversion on a result to generate an LLM model, inputting the LLM model, and reasoning to obtain a network security disposal strategy. And extracting key fields, performing structured packaging, uploading the key fields to a block chain to complete evidence storage and integrity verification, and finally executing corresponding security control operation according to a strategy. According to the scheme, intelligent identification and automatic processing of network attacks are realized, after key fields are extracted, the structured strategy data are generated and uploaded to the block chain system for evidence storage and verification, traceability and data integrity of the processing process are ensured, network defense control which is automatic, high in security and timely in response can be realized, and the network attack processing efficiency is improved. And the intelligent and credible level of network security management is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network security intelligent protection, and in particular to a 6G network automatic security handling method and system based on LLM. Background Art

[0002] With the rapid development of 6G network technology, its ultra-high speed, ultra-low latency, and massive connectivity provide revolutionary support for scenarios such as the Industrial Internet, intelligent transportation, and telemedicine. However, it also faces unprecedented network security challenges. The diversification of device types, exponential growth in data traffic, and rapid expansion of the attack surface in 6G networks make it difficult for traditional network security protection methods to cope with complex and changing unknown threats. For example, the intelligence level of zero-day vulnerability attacks, advanced persistent threats (APTs), and distributed denial of service (DDoS) attacks continues to increase. Attackers use AI technology to generate adversarial samples or encrypt malicious traffic, resulting in a significant increase in the false alarm and miss rates of traditional detection systems based on rule matching or shallow machine learning. In addition, the 6G network's requirements for real-time, automated processing, and cross-domain collaboration urgently require a technical framework that can adapt to threat evolution, intelligently generate processing strategies, and ensure reliable operations.

[0003] Existing network security protection technologies primarily rely on two types of solutions: Traditional signature-based intrusion detection systems (IDSs) match known attack patterns through predefined rules but are completely incapable of identifying unknown threats. Anomaly detection models based on traditional machine learning (e.g., support vector machines (SVMs) and random forests) can capture some unknown behaviors, but are limited by the complexity of feature engineering, insufficient model generalization, and static decision thresholds. These models result in high false positive rates and delayed responses in the dynamic environment of 6G networks. During the response phase, mainstream solutions rely on manual log analysis and configuration of firewall rules or isolation devices, resulting in a fragmented and inefficient process. Some solutions attempt to integrate expert systems to generate response instructions, but rule base updates lag behind the iteration of attack methods, causing the effectiveness of these strategies to rapidly decline. Furthermore, existing technologies generally lack reliable evidence and audit capabilities for response operations, making it difficult to trace responsibility or restore the system status in the event of an erroneous operation or malicious tampering.

[0004] The detection mechanisms of existing network security protection technologies are limited by feature signatures or shallow machine learning, and are unable to parse encrypted traffic or correlate advanced threats in multi-source logs, resulting in a high rate of missed reports of zero-day vulnerabilities and APT attacks. Disposal strategies rely on static rule bases or simple decision trees, lacking semantic understanding of dynamic network contexts, making it difficult to generate precise cross-domain collaborative strategies, and manual intervention results in response cycles of up to tens of minutes. At the same time, traditional solutions lack blockchain evidence storage capabilities, and disposal operations are not traceable, making it difficult to meet the core requirements of 6G networks for real-time, automation, and trusted evidence. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present disclosure provides a 6G network automatic security disposal method and system based on LLM. This disclosure addresses the following issues: the detection mechanism of existing network security protection technologies is limited by feature signatures or shallow machine learning, and is unable to parse encrypted traffic or correlate advanced threats in multi-source logs, resulting in a high rate of missed reports of zero-day vulnerabilities and APT attacks; the disposal strategy relies on static rule bases or simple decision trees, lacks semantic understanding of dynamic network context, and is difficult to generate precise cross-domain collaborative strategies. Manual intervention leads to response cycles of up to tens of minutes. At the same time, traditional solutions lack blockchain evidence storage capabilities, and disposal operations are not traceable, making it difficult to meet the core requirements of 6G networks for real-time, automated, and trusted evidence storage.

[0006] According to a first aspect of the present disclosure, there is provided a 6G network automatic security handling method based on LLM, comprising: acquiring network data flow data, and determining structured log data in JSON format based on the network data flow data;

[0007] Performing feature extraction on structured log data to obtain input feature data, inputting the input feature data into a pre-trained intrusion detection model for classification processing to obtain network attack detection results;

[0008] Performing instruction fine-tuning format conversion processing on the network attack detection result to obtain input data of a pre-trained LLM security model, inputting the input data into the pre-trained LLM security model for reasoning, and obtaining a network security disposal strategy;

[0009] Extracting key fields of the network security disposal policy, and packaging the network security disposal policy and the key fields into structured data;

[0010] The structured data is uploaded to the blockchain system for evidence storage, the integrity of the structured data content is verified, and network security control operations are performed according to the network security disposal strategy.

[0011] According to a second aspect of the present disclosure, there is provided a 6G network automatic security handling system based on LLM, which is used to perform the method according to the first aspect, including: a log generation module, which is used to obtain network data flow data and determine structured log data in JSON format based on the network data flow data;

[0012] The intrusion detection module is used to perform feature extraction processing on the structured log data to obtain input feature data, input the input feature data into the pre-trained intrusion detection model for classification processing, and obtain network attack detection results;

[0013] A policy reasoning module is used to convert the network attack detection results into an instruction fine-tuning format to obtain input data for a pre-trained LLM security model, input the input data into the pre-trained LLM security model for reasoning, and obtain a network security disposal strategy;

[0014] A data packaging module, configured to extract key fields of the network security disposal policy and package the network security disposal policy and the key fields into structured data;

[0015] The on-chain evidence storage module is used to upload the structured data to the blockchain system for evidence storage, verify the integrity of the structured data content, and perform network security control operations according to the network security disposal strategy.

[0016] According to a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: a memory and a processor. The memory stores a computer program. When the processor executes the program, the method described above is implemented.

[0017] In the LLM-based 6G network automatic security disposal method and system provided above, the disclosed embodiment realizes intelligent identification and automatic disposal of network attacks by extracting structured logs from network data streams and combining them with pre-trained intrusion detection models and LLM security big models. After extracting key fields, structured policy data is generated and uploaded to the blockchain system for evidence storage and verification. This not only ensures the traceability and data integrity of the disposal process, but also realizes automated, highly secure, and timely responsive network defense control, significantly improving the intelligence and credibility of network security management. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of a 6G network automatic security handling method based on LLM according to an embodiment of the present disclosure is shown;

[0019] Figure 2 A schematic diagram of a 6G network automatic security handling method based on LLM according to an embodiment of the present disclosure is shown;

[0020] Figure 3 A schematic block diagram of a 6G network automatic security handling system based on LLM according to an embodiment of the present disclosure is shown;

[0021] Figure 4 A block diagram of an exemplary electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0022] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0023] Those skilled in the art will understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and do not represent any specific technical meaning, nor do they represent the necessary logical order between them. It should also be understood that in the embodiments of the present disclosure, "multiple" may refer to two or more, and "at least one" may refer to one, two or more. It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly defined or given a contrary revelation in the context. In addition, the term "and / or" in the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are in an "or" relationship. It should also be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and the same or similar aspects thereof can be referenced to each other. For the sake of brevity, they will not be described one by one.

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0025] Figure 1 A flowchart of a 6G network automatic security handling method based on LLM provided in an embodiment of the present disclosure is provided. The method of the embodiment of the present disclosure aims to achieve accurate detection of large and small objects in images.

[0026] S101, obtaining network data stream data, and determining structured log data in JSON format according to the network data stream data.

[0027] Network data flow data can refer to the original data packet stream transmitted through the network, including source IP / destination IP, source port / destination port, protocol (such as TCP, UDP), payload, timestamp, packet length and other timing information.

[0028] The JSON format can be a lightweight data exchange format and is often used to represent structured data.

[0029] Structured log data can refer to data log entries with uniform field names and structures, which is different from ordinary text logs and is convenient for machine processing and analysis.

[0030] To obtain network data flow data, you can use Pcap or nfstream packet capture tools to generate real-time network logs with a set sampling period T (in seconds). The log data format refers to the CICIDS2017 dataset structure. The sampling period T can be dynamically adjusted according to the frequency of network attacks: if the attacks are frequent, the period is set short, and if it is stable, it can be set long. The collected raw log data will undergo the following cleaning and structuring steps: Formatting: unify the log format (such as converting to a table or JSON), and remove useless columns (such as empty fields, features not related to intrusion). Missing value and outlier processing: delete infinite values, duplicate data, and fill or remove missing fields. Encoding and structuring: such as label encoding or one-hot encoding of categorical features to adapt them to the input requirements of machine learning models. Standardization: numerical features (such as traffic size, duration) can be normalized / standardized. The resulting structured data is organized in JSON format, as shown in the following example:

[0031] {

[0032] "SourceIP": "192.168.0.1",

[0033] "DestinationIP": "192.168.0.102",

[0034] "SourcePort": 36894,

[0035] "DestinationPort": 4000,

[0036] "Protocol": "TCP",

[0037] FlowDuration: 3456,

[0038] "TotalFwdPackets": 20,

[0039] "TotalBwdPackets": 18,

[0040] "Label": "BruteForce"

[0041] }

[0042] Based on the above technical solution, optionally, determining structured log data in JSON format according to the network data stream data includes:

[0043] Performing real-time packet capture and adaptive sampling processing on the network data stream data to obtain formatted network log data;

[0044] The formatted network log data is cleaned and structured to obtain structured log data in JSON format.

[0045] In this scenario, network log data refers to the log records generated by capturing and processing network data streams, which record network communication behavior, connection status, communication content characteristics, and other information. This data typically includes timestamps, source IP addresses, destination IP addresses, source ports, destination ports, protocol types, packet sizes, flow directions, connection status, and flags, and is a key data foundation for network behavior analysis, attack detection, and security audits.

[0046] This process begins with real-time network data stream collection using high-performance packet capture tools such as Pcap or nfstream. This collection process employs an adaptive sampling mechanism, dynamically adjusting the sampling period T based on the current network load and anomaly risk level. For example, during periods of frequent network attacks or high risk, T is set shorter to improve detection sensitivity, while during periods of network stability, T is appropriately extended to reduce system resource consumption. The collected data initially consists of raw network traffic data. This data is then formatted, converting each network connection or flow into a record with a standard structure, forming preliminary network log data. This log data is then cleaned, including removing empty or redundant fields, deleting duplicate records, correcting or eliminating invalid values ​​(such as infinity or NaN), standardizing field formats (such as IP and time formats), and filtering key fields relevant to security analysis, such as the communicating IP and port numbers, protocol, packet size, and flow duration. Finally, the processed data is organized into structured log data in a unified JSON format, providing a high-quality data foundation for subsequent feature extraction, attack identification, and strategy generation.

[0047] In this solution, by performing real-time packet capture and adaptive sampling processing on network data streams, and cleaning and structuring the formatted network log data, the accuracy, integrity and availability of network log data can be significantly improved, giving it a unified format and high-quality features, which is conducive to the subsequent intrusion detection and intelligent analysis systems to efficiently identify security threats, reduce false alarm rates, and ensure the real-time nature of data processing and the optimization of resource utilization.

[0048] Based on the above technical solution, the formatted network log data can be optionally cleaned and structured to obtain structured log data in JSON format, including:

[0049] Perform redundant information removal on the formatted network log data to obtain preliminary cleansed log data;

[0050] Perform abnormal numerical processing on the preliminary cleaned log data to obtain numerically standardized log data;

[0051] Perform classification feature encoding on the numerical specification log data to obtain encoded log data;

[0052] Convert the encoded log data into JSON format to obtain structured log data in JSON format.

[0053] In this solution, the initial cleansing of log data can be performed after removing redundant information from formatted network log data. This process typically involves removing empty columns, irrelevant fields, redundant blank rows, and duplicate records, while retaining fields meaningful for security analysis, such as source / destination IP addresses, ports, protocols, traffic volume, and timestamps. This makes the log data more concise and focused, and reduces the computational overhead of subsequent processing.

[0054] Numerical normalization of log data involves further processing abnormal values ​​(such as replacing or removing infinities, NaNs, and outliers) based on preliminary log data cleaning to ensure that the data conforms to the format requirements of model input and improve data quality and consistency. This step stabilizes and reliables the numerical features in the logs, helping to improve the accuracy of model predictions.

[0055] Encoding log data can be to encode the categorical (non-numeric) feature fields (such as protocol type, flag, service type, etc.) in the numerical specification log data (such as one-hot encoding, label encoding, etc.), convert the text data into a numerical form acceptable to the model, and generate a data structure suitable for machine learning model training and reasoning.

[0056] When removing redundant information from formatted network log data, we first identify and delete irrelevant fields in the log (such as log record numbers, empty value columns, useless identifiers, etc.), remove redundant blank lines and duplicate records, and retain fields closely related to network behavior and attack detection, such as timestamps, source / destination IP addresses, ports, protocol types, traffic characteristics, etc., to obtain compact preliminary cleaned log data. Subsequently, we normalize abnormal values ​​in the data, including replacing infinity (inf), non-numbers (NaN), empty values, or illegal characters with reasonable default values ​​(such as 0 or the mean), and deleting obviously erroneous outliers, to generate numerically normalized log data. Next, we convert non-numeric category fields in the data, such as "protocol type," "service type," and "connection status," into a numerical format that can be processed by machine learning models using label encoding or one-hot encoding to obtain encoded log data. Finally, we convert these encoded data into a structure consisting of key-value pairs according to a unified field standard and serialize them into JSON format, thus forming JSON-formatted structured log data with a clear structure and complete fields, which is convenient for subsequent model input and on-chain evidence storage.

[0057] In this solution, by removing redundant information, normalizing abnormal values ​​and encoding classification features, the neatness, consistency and availability of network log data are significantly improved, making the log structure more conducive to the accurate identification of network attacks by machine learning models. At the same time, it is converted into a unified JSON format for easy storage, transmission, trusted evidence and efficient retrieval in the blockchain.

[0058] S102: performing feature extraction processing on the structured log data to obtain input feature data, inputting the input feature data into a pre-trained intrusion detection model for classification processing to obtain a network attack detection result.

[0059] The input feature data can be the key fields extracted from structured network logs for model classification, including traffic-related features: such as the total number of traffic bytes (TotalFwd / BwdBytes), flow rate (FlowDuration), network protocol features: such as protocol type (TCP / UDP / ICMP), IP and port information: source IP, destination IP, source port, destination port, packet number and flag features: such as FwdPackets, BwdPackets, FlagCount, time features: such as timestamp, interval time, statistical features: such as average packet length, maximum and minimum packet length, etc.

[0060] The pre-trained intrusion detection model can be an XGBoost classification model, which is trained using datasets such as CICIDS2017 and CICIDS2018. The model type is a gradient boosting decision tree, and the task is multi-classification: determine which type of attack the current network traffic belongs to or whether it is normal traffic.

[0061] The network attack detection result can be the result of model classification output, formatted as JSON, specifically including source IP (SourceIP), target IP (DestinationIP), system type (such as Windows, Linux, etc.), attack type (such as DDoS, SQLi, BruteForce, etc.), attack probability or confidence (such as 98%), label (normal / attack), and timestamp (Timestamp).

[0062] Feature extraction can be performed on this structured log data. Network traffic behavior features such as flow duration, forward / backward packet count, average packet length, and traffic direction statistics are extracted. Protocol and connection features such as protocol type, port number, connection direction, and number of flag bits are extracted. Host behavior features such as the access frequency of the source / destination IP addresses and the density of data exchange per unit time are also extracted. These extracted input features are converted into numerical vectors (such as NumPy arrays or Pandas DataFrames) and fed into a pre-trained XGBoost intrusion detection model. The XGBoost model is a multi-classifier trained on standard attack datasets such as CICIDS2017 and CICIDS2018. The model has learned the characteristic patterns of common attack behaviors. The model classifies the input features and ultimately outputs network attack detection results, including source / destination IP addresses, attack type (such as DDoS, Brute Force, XSS, and SQL injection), attack probability (model confidence), label (normal / abnormal), system type (such as Windows), and event timestamp. The output is again encapsulated as JSON, for example:

[0063] {

[0064] "source_ip": "192.168.0.1",

[0065] "destination_ip": "192.168.0.102",

[0066] "attack_type": "BruteForce",

[0067] "system": "Windows",

[0068] "confidence": 0.987,

[0069] ″label″:″Attack″,

[0070] "timestamp": "2025-05-31T14:33:45Z"

[0071] }.

[0072] The pre-trained intrusion detection model is trained using traditional machine learning methods, mainly based on public network attack datasets such as CICIDS2017 and CICIDS2018. First, the original network traffic data is preprocessed, including cleaning redundant fields, processing missing values, standardizing numerical fields, label encoding, etc. Then, key features such as traffic size, transmission protocol, source / destination IP and port number, session duration, etc. are extracted from the cleaned data to construct training samples. The samples are input into classification models such as XGBoost, RandomForest, and LightGBM for supervised learning training. The goal is to accurately identify different types of network attacks, such as DDoS, SQL injection, XSS, port scanning, etc., and output information such as attack type, probability, timestamp, IP address, etc. The model tunes hyperparameters on the validation set to improve classification performance, and finally forms a stable intrusion detection classifier.

[0073] Based on the above technical solution, optionally, feature extraction processing is performed on the structured log data to obtain input feature data including:

[0074] Perform flow size feature extraction on structured log data to obtain flow feature data;

[0075] Perform protocol type feature extraction on structured log data to obtain protocol feature data;

[0076] Extract source / destination IP address features from structured log data to obtain address feature data;

[0077] Perform port number feature extraction on structured log data to obtain port feature data;

[0078] The flow characteristic data, protocol characteristic data, address characteristic data and port characteristic data are combined and processed to obtain input characteristic data.

[0079] In this solution, traffic feature data may refer to network traffic-related indicators extracted from structured logs, typically including the number of bytes per connection (such as source-to-destination traffic, destination-to-source traffic), number of data packets, duration, etc., which are used to reflect the scale and behavior pattern of network connections.

[0080] Protocol feature data can refer to the communication protocol type (such as TCP, UDP, ICMP, etc.) information extracted from logs. By converting protocol identifiers into processable numerical values ​​or one-hot encoding, it helps the model identify the differences in communication behaviors of different protocols.

[0081] Address feature data can refer to features related to the source IP address and the destination IP address. It can be the original IP, or it can be further converted into the location, network segment or IP type (intranet, extranet) to identify the communication subject.

[0082] Port feature data can refer to the relevant information of the source port number and the destination port number, which is used to determine the service type used for communication (such as HTTP, FTP, SSH, etc.). It can also be classified or standardized to enhance the discrimination ability of the model.

[0083] When extracting features from structured log data, first extract network traffic-related fields from each log, such as FlowBytes / s, FlowPackets / s, TotalFwd / BwdBytes, etc., and normalize or standardize them to obtain traffic feature data reflecting the size of the connection traffic; secondly, extract the Protocol field, and perform numerical mapping or one-hot encoding on different protocols (such as TCP, UDP, ICMP, etc.) to obtain protocol feature data; then, extract SourceIP and DestinationIP, and optionally convert the IP address into the corresponding regional label, network segment category, or internal and external network identifier to form address feature data; then extract SourcePort and DestinationPort, classify them as common service ports (such as 80, 443, 21, 22, etc.) or map them to service types to obtain port feature data; finally, the above four types of features are spliced, format aligned, and dimensionally unified to form complete input feature data.

[0084] In this solution, by performing multi-dimensional feature extraction and unified coding processing on structured log data, the behavioral characteristics of network traffic can be comprehensively characterized, providing accurate and standardized input data for intrusion detection models, which helps to improve the model's ability to identify abnormal traffic and the accuracy of classification, thereby achieving more efficient and intelligent network security protection.

[0085] S103, performing instruction fine-tuning format conversion processing on the network attack detection result to obtain input data of the LLM security model, inputting the input data into the pre-trained LLM security model for inference, and obtaining a network security disposal strategy.

[0086] The pre-trained LLM security large model can refer to a model that is fine-tuned with LoRA based on an existing large language model. The specific description is as follows: Basic model: The DeepSeekR1 model is used, which is a multi-language general large model with strong understanding and reasoning capabilities. Fine-tuning method: Using LoRA (Low-Rank Adaptation) fine-tuning technology, it focuses on enhancing the command response capability, which is convenient for deployment on edge devices or local servers. Training data set: The training corpus comes from multiple authoritative network security data sets: CICIDS2017, CICIDS2018, CICIoT2022, CICIoMT, CICIoT-DIAD2024, CIC-BCCC-NRC2024. The data set is processed by the intrusion detection subsystem to convert the model classification output into the instruction fine-tuning format:

[0087] {

[0088] "instruction": "PleaseGive the security protection strategy."

[0089] "input": "FlowIdr: [XXXXXXX]\SourceIP: [192.168.0.1]\DestinationIP: [192.168.0.102]\SourcePort: [3689 4]\DestinationPort:

[4000] \Protocol:[TCP]\AttackType:[BruteForce]\Timetamp[2025-05-31T14:33:45Z]″,

[0090] ″response″:[

[0091] "DetectedaBruteForceInternetTrafficattackat[Windows][2025-05-31T14:33:45Z]",

[0092] "Block IP192.168.0.1",

[0093] "Check SSH configuration",

[0094] "Enable two-factor authentication" ]

[0096] }

[0097] Deployment platform: After training is completed, the model is deployed on the Tesloop Deepseek all-in-one machine to run locally to support localized reasoning and privacy protection.

[0098] The input data can be a structured string after format conversion and instruction fine-tuning of the JSON results (such as attack type, IP, port, protocol, system, etc.) output by the intrusion detection subsystem, which is used to input into the LLM security model.

[0099] Raw JSON example (output by a pre-trained intrusion detection model):

[0100] {

[0101] "source_ip": "192.168.0.1",

[0102] "destination_ip": "192.168.0.102",

[0103] ″source_port″: 36894,

[0104] "destination_port": 4000,

[0105] "protocol": "TCP",

[0106] "attack_type": "BruteForce",

[0107] "system": "Windows",

[0108] "timestamp": "2025-05-31T14:33:45Z"

[0109] }

[0110] The input format for conversion to LLM model is as follows:

[0111] {

[0112] "input": "FlowId: [XXXXXXX]\\SourceIP: [192.168.0.1]\\DestinationIP: [192.168.0.102]\\SourcePort:

[36894] \\DestinationPort:

[4000] \\Protocol: [TCP]\\AttackTyp e: [BruteForce]\nTimeStamp[″2025-05-31T14:33:45Z″]Abstract: DetectedaBruteForceInternetTrafficattackat[Windows]″

[0113] }

[0114] The network security disposal strategy can be a specific security response plan for a detected network attack event generated by the LLM security model. This strategy has the following characteristics: structured output (response field), the format is as follows:

[0115] {

[0116] ″response″:[

[0117] "Block IP192.168.0.1",

[0118] "Check SSH configuration",

[0119] "Enable two-factor authentication" ]

[0121] }

[0122] The policy content covers (automatically adjusted according to the type of attack): Isolating the source of the attack: such as blocking the IP, disconnecting, restricting the port, etc. System reinforcement measures: such as strengthening login authentication, upgrading system patches, and adjusting configuration files. Log auditing and tracking: such as saving attack logs, uploading to the blockchain, and generating traceability evidence. Automatic response control: controlling the firewall, IPS / IDS, antivirus engine, etc. by calling the Agent subsystem. Policy usage: provided to the Agent subsystem for executing automatic protection actions; uploaded to the blockchain subsystem to prevent tampering and ensure traceability; realizing a complete closed loop from "intrusion detection → large model reasoning → secure execution response".

[0123] In order to meet the input requirements of the reasoning interface of large language models (such as the LLM security large model based on DeapseekR1), the network attack detection results need to be converted into a structured natural language instruction format, namely the "instruction fine-tuning format". Format conversion processing logic:

[0124] defformat_attack_result_to_instruction(data: dict)->dict:

[0125] ″″″

[0126] Format the network attack detection results into the instruction fine-tuning input format required by the LLM large model.

[0127] parameter:

[0128] data: the raw output of the intrusion detection model (JSON)

[0129] return:

[0130] Input dictionary for instruction fine-tuning format

[0131] ″″″

[0132] instruction_input = (

[0133] f″FlowId: [{data.get('flow_id', 'XXXXXXX')}]\\″

[0134] f′SourceIP: [{data['source_ip']}]\\″

[0135] f′DestinationIP: [{data['destination_ip']}]\\″

[0136] f′SourcePort: [{data['source_port']}]\\″

[0137] f″DestinationPort: [{data['destination_port']}]\\″

[0138] f″Protocol: [{data['protocol']}]\\″

[0139] f″AttackType: [{data['attack_type']}]\n″

[0140] f″TimeStamp: [{data['timestamp']}\n″

[0141] f″Abstract: Detected a {data['attack_type']} Internet Traffic attack at [{data['system']}]″ )

[0143] return {″input″: instruction_input}

[0144] Input example:

[0145] attack_json = {

[0146] ″flow_id″: ″123456789″,

[0147] ″source_ip″: ″192.168.0.1″,

[0148] "destination_ip": "192.168.0.102",

[0149] ″source_port″: 36894,

[0150] "destination_port": 4000,

[0151] "protocol": "TCP",

[0152] "attack_type": "BruteForce",

[0153] "system": "Windows",

[0154] "timestamp": "2025-05-31T14:33:45Z"

[0155] }

[0156] formatted_input=format_attack_result_to_instruction(attack_json)

[0157] Sample output:

[0158] {

[0159] "input": "FlowId: [123456789]\\SourceIP: [192.168.0.1]\\DestinationIP: [192.168.0.102]\\SourcePort:

[36894] \\DestinationPort:

[4000] \\Pro tocol: [TCP]\\AttackType: [BruteForce]\nTimeStamp[″2025-05-31T14:33:45Z″]Abstract: DetectedaBruteForceInternetTrafficattackat[Windows]″

[0160] }

[0161] The converted data is then fed as prompts into a LoRA-tuned LLM security model deployed locally (e.g., on a Terminus DeepSeek appliance) for inference. The model generates corresponding network security policy responses. Here's an example inference call (pseudocode / API logic), assuming the deployed inference interface supports inference via HTTP API or local Python:

[0162] importrequests

[0163] defcall_llm_security_model(formatted_input: dict)->list:

[0164] ″″″

[0165] Call the deployed LLM large model for reasoning and output the security disposal strategy.

[0166] parameter:

[0167] formatte d _input: Input data in instruction fine-tuning format

[0168] return:

[0169] Security policy list output by the LLM model

[0170] ″″″

[0171] response=requests.post("http: / / localhost:8000 / infer", json=formatted_input)

[0172] returnresponse.json().get("response",[])

[0173] Expected return:

[0174] {

[0175] ″response″:[

[0176] "Block IP192.168.0.1",

[0177] "Check SSH configuration",

[0178] "Enable two-factor authentication" ]

[0180] }

[0181] The pre-trained LLM security model is based on the Deepseek-R1 language model and is fine-tuned on LoRA on multiple security datasets processed by intrusion detection models through instruction tuning. The training datasets include CICIDS2017, CICIoT2023, CIC-DIAD2024, CIC-BCCC-NRC2024, etc., and constructs training samples with a unified format. The "instruction" field defines the task objective, the "input" field contains attack traffic summary information (such as source / destination IP, port, protocol, attack type, system type, etc.), and the "response" field is the corresponding network security disposal strategy (such as blocking IP, modifying firewall configuration, enabling authentication mechanism, etc.). The model learns the logical relationship between different attack types and defense responses on tens of millions of samples, and masters network situational awareness and policy generation capabilities. After training, it is deployed in the local inference device for real-time generation of personalized security policies.

[0182] On the basis of the above technical solution, optionally, the network attack detection result is subjected to instruction fine-tuning format conversion processing to obtain input data of the pre-trained LLM security model, including:

[0183] Extract and process the attack type, source IP, target IP, system type, and timestamp fields in the network attack detection result to obtain key detection data;

[0184] Formatting the key detection data according to a preset instruction template to obtain formatted detection data;

[0185] Adding an instruction prompt prefix to the formatted detection data to obtain formatted detection data with the instruction prefix;

[0186] The formatted detection data with instruction prefixes is encapsulated in JSON structure to obtain the input data of the pre-trained LLM security model.

[0187] In this solution, the attack type can refer to the specific attack behavior type identified in the network intrusion detection results, such as DDoS, PortScan, BruteForce, SQL injection, XSS, etc., which is used to illustrate the specific threat type of the current abnormal network traffic.

[0188] The source IP address can refer to the IP address of the host that initiates a network attack. It is the source identifier of network traffic and is usually used to track the source of the attacker.

[0189] The target IP address can refer to the IP address of the host under attack. It is the target identifier of abnormal traffic and is used to determine the victim system or service of the attack.

[0190] The system type can refer to the target host's operating system or device type, such as Windows, Linux, or IoT devices, which helps generate targeted defense strategies.

[0191] The timestamp can record the specific time when the network attack occurred, which is convenient for tracing, statistical analysis and evidence preservation.

[0192] Key detection data can refer to the core field information (i.e., attack type, source IP, target IP, system type, and timestamp) extracted from the network attack detection results. This information is the necessary basis for the subsequent generation of disposal strategies and constitutes the core content of the LLM security model input.

[0193] The preset instruction template can be a format framework designed to convert key detection data into natural language prompts. An example is as follows:

[0194] "FlowId: [XXX]\nSourceIP: [192.168.1.1]\nDestinationIP: [192.168.1.100]\nSourcePort:

[45678] \nDestinationPort:

[80] \nProtocol: [TCP]\nAttackType: [SQLInjection]\nAbstract: DetectedSQLInjectionattackat[Windows]"

[0195] Formatted detection data can be text content after key detection data is filled in according to a preset instruction template. Its form is network attack event information described in natural language, with a clear structure, and is suitable as input material for large models.

[0196] Formatting test data with a command prefix can be done by adding a large model prompt or a task command prefix before the formatted test data, for example:

[0197] "Please give the security protection strategy:" + format the detection data to guide the large model to understand the task objectives.

[0198] From the structured output of the network attack detection model, key fields highly relevant to security response are precisely extracted, including attack type (such as DDoS, SQL injection), source IP (attack originating address), target IP (victim host address), system type (target system operating system or device type, such as Linux, Windows, IoT device), and timestamp (the specific time the attack occurred). This information together constitutes key detection data. Next, based on a pre-set instruction template, the key detection data is mapped into a natural language structure, for example: "Detected at [timestamp], attack type [attack type], originating from IP address [source IP], target [target IP], running on a host with system type [system type]." This resulting sentence is formatted detection data, which is highly human-readable and contextually expressive. Subsequently, a semantically explicit instruction prefix is ​​added to this formatted text, such as: "Please generate a targeted security response strategy based on the following network attack information:" to create formatted detection data with an instruction prefix. This step significantly improves the comprehension capabilities of large models in reasoning tasks. Finally, the prefixed data is used as the "input" field, together with the task description (such as "instruction: output response strategy"), and is structured and encapsulated in a standard JSON format to generate input data suitable for the pre-trained LLM secure large model inference interface.

[0199] In this solution, the original detection results can be automatically converted into a standardized input format that can be understood by the large model, achieving seamless connection from attack identification to intelligent response strategy generation, improving the automation level and response efficiency of network security disposal, while ensuring that the data semantics are clear, the structure is standardized, and it is easy to trace and verify later.

[0200] S104: Extract key fields of the network security handling policy, and package the network security handling policy and the key fields into structured data.

[0201] Key fields can refer to elements that have operational significance and execution controllability in network security disposal strategies, and fields that can directly participate in subsequent system automation execution and control.

[0202] Structured data can be the conversion of natural language policy responses (such as "block IPxxx") into dictionaries or JSON data with clear fields and fixed formats to support system processing, matching, execution, or uploading to the blockchain.

[0203] First, the response generated by the LLM security model is parsed to extract action instructions with clear security implications, such as "Block IP 192.168.0.1," "Check SSH configuration," and "Enable two-factor authentication." To ensure enforceability and traceability of these policies, the system extracts and associates several key fields, including but not limited to: source IP address (source_ip), destination IP address (destination_ip), source port (source_port), destination port (destination_port), communication protocol (protocol), attack type (attack_type), attacked system type (system_type), attack time (timestamp), attack identification label (attack_label), model ID (model_id), device MAC address (mac_address), the original prompt input, and the inferred response content. Extracting these fields is typically based on a unified JSON structure template, with field separation achieved by parsing the LLM input and output formats. The system then encapsulates the policy and the extracted key fields into a structured data object.

[0204] S105: Upload the structured data to the blockchain system for evidence storage, verify the integrity of the structured data content, and perform network security control operations according to the network security disposal policy.

[0205] The structured data is uploaded to the blockchain system for evidence storage. First, the structured data containing the network security disposal strategy (such as source IP, attack type, detection time, model response strategy, MAC address, model ID, etc.) is written into the blockchain as non-tamperable transaction information to ensure that every step of the network security disposal process is traceable and verifiable. The blockchain system serves as a decentralized trusted data storage and sharing platform with anti-tampering, traceability, and strong transparency. It can ensure that the disposal strategy output by the LLM large model is not modified or forged in subsequent operations. When uploading structured data, the system will also use a hash verification mechanism to verify the integrity of the data content to ensure that the data has not been tampered with before and after uploading.

[0206] After completing the blockchain evidence storage, the system extracts the policy portion of the evidence data and inputs it into the Agent subsystem as an execution instruction. The Agent subsystem triggers the corresponding network security control operation based on the policy content. Network security control operation refers to the automatic call of defensive tools or configuration items integrated in the system according to the disposal strategy, such as implementing access bans for specific IPs, modifying firewall rules, strengthening identity authentication mechanisms (such as enabling two-factor authentication), closing high-risk ports, restricting certain types of protocol communications, or calling antivirus software to scan and clean up affected nodes. The Agent subsystem will also retrieve the policy record with the corresponding timestamp and ID from the blockchain system again, and perform a consistency check with the currently received policy output to ensure that the policy of the execution operation is completely consistent with the evidence information on the chain, and finally complete the closed-loop control of the entire process from attack detection, model reasoning, data evidence to security disposal.

[0207] In an embodiment of the present application, network data flow data is obtained, and structured log data in JSON format is determined based on the network data flow data; feature extraction processing is performed on the structured log data to obtain input feature data, and the input feature data is input into a pre-trained intrusion detection model for classification processing to obtain a network attack detection result; instruction fine-tuning format conversion processing is performed on the network attack detection result to obtain input data of a pre-trained LLM security model, and the input data is input into the pre-trained LLM security model for inference to obtain a network security disposal strategy; key fields of the network security disposal strategy are extracted, and the network security disposal strategy and the key fields are packaged into structured data; the structured data is uploaded to a blockchain system for notarization, the integrity of the structured data content is verified, and network security control operations are performed according to the network security disposal strategy. Through the above-mentioned LLM-based 6G network automatic security disposal method, by extracting structured logs from network data streams and combining them with pre-trained intrusion detection models and LLM security big models, intelligent identification and automatic disposal of network attacks are achieved. After extracting key fields, structured policy data is generated and uploaded to the blockchain system for storage and verification. This not only ensures the traceability and data integrity of the disposal process, but also realizes automated, highly secure, and timely responsive network defense control, significantly improving the intelligence and credibility of network security management.

[0208] On the basis of the above technical solution, optionally, performing a network security control operation according to the network security disposal policy includes:

[0209] Analyze the network security disposal strategy and extract specific operation instructions;

[0210] Perform matching processing based on the preset rule base and specific operation instructions to determine the security tools that need to be called;

[0211] Call the security tools as needed, call the firewall tool to perform IP blocking processing, and generate IP blocking execution results;

[0212] Based on the security tools that need to be called, call the system configuration tool to perform two-factor authentication activation processing and generate the authentication configuration results;

[0213] The security tool called as needed calls the port management tool to perform port closing processing and generate port management results;

[0214] The IP blocking execution results, authentication configuration results and port management results are summarized and processed to generate a network security control execution report.

[0215] In this solution, specific operational instructions can refer to standardized security action commands parsed from the network security policy and directly executable by the system, such as "block IP address 192.168.0.1," "enable two-factor authentication," or "close port 4000." These are meaningful operational entities within the LLM security model output.

[0216] A preset rule base is a set of predefined rules that describe different types of operations, their corresponding processing logic, the security tools to be invoked, and their configuration parameters. For example, a rule might specify that "blocking an IP address" should invoke the firewall module, while "enabling two-factor authentication" should invoke the system configuration module.

[0217] The security tools that need to be called can refer to software modules or system functions that should be actually called in security control, determined according to specific operating instructions and rule bases, such as firewall tools, system configuration tools, port management tools, etc.

[0218] A firewall tool can be a software or system module used to implement network access control policies, supporting operations such as IP blocking, protocol filtering, inbound / outbound rule configuration, etc., and is usually deployed at the network boundary or host end.

[0219] System configuration tools can refer to automated modules used to adjust host or operating system security policies, supporting functions such as enabling two-factor authentication (2FA), modifying login policies, and configuring authentication modules.

[0220] A port management tool is a software module used to manage the port status in a system. It can enable, disable, or restrict communication on specific ports to prevent attackers from invading the system by exploiting open ports.

[0221] The IP blocking execution result may refer to the processing status information returned by the firewall tool after completing the "blocking the specified IP address" operation, such as "IP192.168.0.1 blocked successfully" or "Blocking failed: insufficient permissions".

[0222] The authentication configuration result may refer to the return information of the system configuration tool after executing the instruction to enable two-factor authentication, such as "2FA enabled successfully: SMS verification supported" or "Failed: configuration file missing".

[0223] The port management result may refer to the response data after the port management tool performs an operation such as closing the port, such as "Port 4000 is closed" or "Close failed: The port is occupied by a system service."

[0224] A network security control execution report may refer to a unified report document or data structure generated by summarizing the processing results of all security tools (such as IP blocking results, authentication configuration results, and port management results). It is used to record the execution process, success status, abnormal information, etc. of this security control operation, and supports auditing, tracing, and display.

[0225] After receiving the network security disposal policy generated by the large model, the system first conducts an in-depth analysis of the policy content and extracts the specific operational instructions contained therein. These operational instructions are usually presented in the form of clear security actions, such as "block the source IP address 192.168.1.100," "enable two-factor authentication to enhance login security," and "close port 445 of the target host to prevent file sharing vulnerabilities." After the extraction is complete, the system matches each operational instruction with the built-in preset rule library, which defines a one-to-one mapping relationship between common operational instructions and security tools. For example, blocking an IP requires calling a firewall tool, configuring account authentication requires calling a system settings tool, and managing ports requires calling a port control tool. Through this rule matching process, the system accurately determines the type of security tool that each instruction requires. The system then calls the corresponding security tool modules to carry out the actual execution: If the instruction is to block an IP address, the system will call the firewall tool to block the specified IP address and generate a specific IP blocking execution result, recording whether it was successful, the time, the target address, etc.; if the instruction is to enable two-factor authentication, the system will call the system configuration tool to complete the configuration process and generate the authentication configuration result, including the enabled status, application object, authentication method, etc.; if the instruction is to close a port, the system will use the port management tool to close the specified port and output the port management result, including the port number, closed status, response time, etc. After all the execution results are collected and organized, the system will automatically summarize them, standardize the format, and semantically encapsulate them, ultimately generating a complete network security control execution report. This report records in detail all the operation steps, tool calls, processing results, and corresponding timestamps during the security response process, providing reliable and traceable data support for subsequent security audits, risk tracking, and policy assessments, while also significantly improving the automation level and execution efficiency of network security responses.

[0226] This solution implements a closed-loop process from security policy analysis to automatic execution. It can accurately match security tools and automatically complete operations such as IP blocking, two-factor authentication activation, and port closure, effectively improving response speed and execution efficiency, reducing human intervention, and enhancing the real-time and intelligent level of network protection.

[0227] Figure 2 A flowchart of a 6G network automatic security handling method based on LLM provided in an embodiment of the present disclosure. The method may include the following steps:

[0228] S201 , obtaining network data stream data, and determining structured log data in JSON format according to the network data stream data.

[0229] S202 , performing feature extraction processing on the structured log data to obtain input feature data, inputting the input feature data into a pre-trained intrusion detection model for classification processing, and obtaining a network attack detection result.

[0230] S203, performing instruction fine-tuning format conversion processing on the network attack detection result to obtain input data of the pre-trained LLM security model, inputting the input data into the pre-trained LLM security model for inference, and obtaining a network security disposal strategy.

[0231] S204: Extract key fields of the network security handling policy, and package the network security handling policy and the key fields into structured data.

[0232] S205: Upload the structured data to the blockchain system for evidence storage, verify the integrity of the structured data content, and perform network security control operations according to the network security disposal policy.

[0233] S206: Perform statistical feature difference analysis on the structured log data to obtain feature importance scores.

[0234] Feature importance scoring can be used to evaluate the contribution of each feature to attack detection by analyzing the relationship between each feature in structured log data and network attack outcomes. Common methods include information gain, Gini index, SHAP value, and mutual information. The higher the score, the more critical the feature is to the model's discriminative ability.

[0235] To perform statistical feature difference analysis on structured log data, it is first necessary to quantitatively evaluate the relationship between each feature and the target variable (such as the attack type label). The specific method is: for classification features, statistical methods such as chi-square test and mutual information (Mutual Information) can be used to evaluate their correlation with the attack label; for numerical features, the degree of their influence on the classification results can be determined through analysis of variance (ANOVA), information gain (Information Gain), Pearson correlation coefficient, etc. Next, the above indicators are calculated for each feature and normalized to form a feature importance score with a unified dimension. The higher the value, the stronger the discriminative ability of the feature in distinguishing different attack behaviors. The final generated feature importance score can serve as an important basis for subsequent dynamic feature selection and model input optimization.

[0236] S207: Perform dynamic feature selection processing according to the feature importance scores to obtain an optimal feature subset.

[0237] The optimal feature subset can be a set of features that most significantly improves model performance, selected based on feature importance scores through dynamic feature selection algorithms (such as forward selection, L1 regularization, and recursive feature elimination). This subset can reduce redundant information, improve training efficiency, and enhance model generalization.

[0238] According to the feature importance score, dynamic feature selection processing is performed. First, all features are sorted from high to low according to their importance score. Then, the number of features to be retained is dynamically determined in combination with model requirements, system resource limitations or preset feature quantity thresholds. Then, strategies such as threshold methods (such as only retaining features with scores higher than a set value), forward selection, recursive feature elimination (RFE) or greedy algorithms are used to gradually screen out the feature set that contributes most to the improvement of model performance. The entire process can be combined with cross-validation to evaluate the performance of the model under different feature combinations in each round of selection to ensure that the selected feature set has the best generalization ability. The optimal feature subset finally obtained not only retains the key features that are most discriminative for intrusion detection, but also effectively eliminates redundant or noise features, thereby improving the training efficiency and detection accuracy of the model.

[0239] S208: Perform attack behavior time series analysis on the structured log data to obtain attack behavior pattern features.

[0240] Attack behavior pattern features can be extracted by analyzing the time series, request path, access frequency, attack duration, attack stage evolution and other temporal features of network attack behaviors in log data, and used to help the model identify complex or persistent attacks.

[0241] To perform attack behavior time series analysis on the structured log data, it is first necessary to sort the log data according to the timestamp field and construct a time series view to restore the order of occurrence of various events in the network. Subsequently, the log events are segmented and analyzed through a sliding time window to extract specific behavior sequences that occur frequently in a short period of time, such as the same source IP initiating connection requests to multiple target ports within a few seconds, or different protocols appearing alternately. Combined with predefined attack behavior templates (such as scanning, blasting, lateral movement, etc.) or through unsupervised algorithms (such as sequence clustering, change point detection) to identify abnormal behavior sequence patterns, further explore potential attack chains and behavior evolution trajectories. The attack behavior pattern features finally extracted can reflect key information such as the temporal characteristics, strategy changes, continuity and stages of attack activities, and provide high-value semantic feature support for subsequent accurate detection and response.

[0242] S209: Acquire historical attack tag data, perform co-occurrence frequency analysis on the historical attack tag data, and obtain tag association relationships.

[0243] Historical attack label data refers to label information collected from previously annotated cybersecurity incidents, such as "DDoS attack," "SQL injection," and "Trojan horse implantation." This data contains a historical record of attack types and serves as the foundation for label learning and semantic modeling.

[0244] To obtain historical attack label data, we first extract historical attack samples from archived network security event databases or intrusion detection system logs, including the attack type label corresponding to each sample, such as "DDoS attack," "port scan," and "malware download." After cleaning and standardizing these label data, we construct a label co-occurrence matrix, which counts the frequency of label pairs that appear simultaneously in the same attack event. By performing frequency analysis and normalization on this co-occurrence matrix, we can quantify the degree of association between different labels, identify frequently occurring label combinations (for example, "lateral movement" and "privilege escalation" frequently co-occur), and form a correlation relationship map between labels. This label correlation relationship can be used to enrich label semantics, infer unknown attack chains, and improve the generalization ability of multi-label classification models.

[0245] S210: Perform label semantic enhancement processing based on the label association relationship and attack behavior pattern characteristics to obtain a refined attack type label.

[0246] Tag association can be achieved by statistically analyzing the co-occurrence frequencies of historical attack tags to uncover dependencies, combinations, or evolutionary paths. For example, "scanning behavior" often appears together with "brute force attack," indicating a potential complex attack chain.

[0247] Refined attack type labels can be more specific and context-sensitive labels obtained by semantically enhancing the original attack labels by combining attack behavior pattern characteristics and label association relationships. For example, "Web attack" can be refined into "path traversal + file reading type Web attack."

[0248] Label semantic enhancement is performed based on label associations and attack behavior pattern features. First, the label association graph is fused with the temporal pattern features of the attack behavior. The semantic context between labels is modeled through a graph neural network (GNN) or attention mechanism to capture the potential dependency structure of multiple labels co-occurring in attack events. Combined with the characteristic sequence of attack behaviors (such as changes in connection frequency, protocol conversion sequence, IP jump path, etc.), the original labels are contextually expanded and semantically refined. For example, an attack label originally labeled "remote login anomaly" can be semantically enhanced to a more specific "SSH brute force + privilege escalation attack" when it is detected to be accompanied by behavioral patterns such as "weak password attempts" and "privilege escalation." Finally, by aggregating label co-occurrence probabilities and behavioral pattern weights, a refined attack type label with richer semantics is generated, providing a more targeted judgment basis for subsequent intelligent responses.

[0249] S211 , merging the optimal feature subset with the refined attack type label to obtain enhanced feature data.

[0250] Enhanced feature data can be the final training input data formed by fusing the optimal feature subset (highly discriminative structured features) with refined attack type labels (semantically enhanced supervisory signals). This type of data possesses both strong discriminative power and semantic richness, helping to improve the recognition accuracy and generalization capabilities of intrusion detection models.

[0251] The optimal feature subset is merged with the refined attack type labels. First, the structured log features in the optimal feature subset (such as traffic size, protocol type, IP / port information, etc.) are standardized or vectorized to ensure feature dimension consistency. At the same time, the refined attack type labels are numerically encoded (such as One-Hot encoding or label embedding vector representation) to make them suitable for modeling. Subsequently, feature splicing or multimodal fusion is used to align the dimensions of the numerical attack label features and fuse them at the vector level with the optimal log features to construct enhanced feature data that contains contextual semantics and data behavior features. This enhanced feature data not only retains highly important structural features, but also incorporates more semantically discriminative attack label information, which helps improve the recognition accuracy and generalization ability of the intrusion detection model for complex attack types.

[0252] In this embodiment, by performing feature difference analysis, temporal behavior mining and label semantic enhancement on structured log data, not only the key attack features are accurately extracted, but also the expressive ability of attack type labels is improved, and ultimately more discriminative enhanced feature data is formed, which effectively improves the detection accuracy and intelligent response capabilities of the intrusion detection model in complex and changeable attack scenarios.

[0253] Based on the above technical solution, optionally, after obtaining the enhanced feature data, the method further includes:

[0254] The enhanced feature data is input into a preset intrusion detection classification model to obtain a detection result including attack type, attack probability, source IP, target IP, communication protocol, system type and timestamp.

[0255] In this solution, the preset intrusion detection classification model can refer to a pre-trained classification model for identifying network attack behaviors. Common ones include: traditional machine learning models: such as decision trees, random forests, support vector machines (SVM), K-nearest neighbors (KNN), etc.; deep learning models: such as multi-layer perceptron (MLP), convolutional neural networks (CNN), recurrent neural networks (RNN), Transformer, etc.; integrated learning models: such as XGBoost, LightGBM, etc.; usually, labeled network traffic data (such as KDD99, NSL-KDD, UNSW-NB15) is used for supervised training to classify the behavior of network traffic, determine whether it is an attack, and identify the specific attack type.

[0256] The attack type indicates the type of network attack detected. Common types include: DoS (Denial of Service), Probe (Probing Scan), U2R (User to Root), R2L (Remote to Local), Botnet, DDoS, and malicious code injection.

[0257] The attack probability can represent the confidence or possibility that the behavior is judged as an attack. It is usually a floating point number between 0 and 1. For example: attack probability = 0.93 means that the model has 93% confidence that the behavior is an attack.

[0258] Source IP (SourceIP) can represent the address of the device that initiates network communication, that is, the IP address of the source host of the attack behavior in the network.

[0259] The target IP (Destination IP) can represent the address of the device receiving network communication or being attacked, that is, the target host IP of the attack behavior.

[0260] The communication protocol can refer to the protocol type used in this network behavior, such as TCP, UDP, ICMP, HTTP, DNS, FTP, etc.

[0261] The system type can indicate the operating system type used by the attacked or communicating host, such as Windows, Linux, Unix, Android, etc.

[0262] The timestamp can be used to record the time when the communication or attack behavior occurred, usually accurate to seconds or milliseconds, for subsequent tracing analysis and time series modeling.

[0263] The detection result can be the comprehensive judgment information output by the intrusion detection classification model after analyzing the enhanced feature data, which is used to indicate whether a certain network behavior is an attack behavior and related details. It usually includes the following: Attack type: the model determines which form of attack the network behavior belongs to (such as DoS, scanning, malicious login, etc.); Attack probability: the model's confidence in this judgment, that is, the likelihood of an attack; Source IP: the address of the device that initiated the network behavior; Target IP: the address of the device affected or attacked by the behavior; Communication protocol: the type of protocol used by this network behavior (such as TCP, UDP, etc.); System type: the type of operating system involved in the communication or attack target (such as Linux, Windows, etc.); Timestamp: records the time when the behavior occurred, which is used for timeline tracing and log correlation analysis.

[0264] Enhanced feature data can be fed into a pre-defined intrusion detection and classification model. First, ensure that the model has been fully trained based on historical attack samples and feature data, enabling it to identify and classify network intrusions. Enhanced feature data consists of an optimal feature subset and refined attack type labels. It encompasses key network behavior parameters and their semantically enhanced context, enabling a high degree of differentiation between normal and abnormal traffic. The input process typically involves standardizing the enhanced feature data to ensure it aligns with the input format used during model training. The processed data is then fed into the model in batches or in real-time. Within the model, this feature data is analyzed and inferred layer by layer using a neural network or other machine learning architecture, outputting multi-dimensional predictions. Ultimately, the model determines whether the input data represents an attack based on its characteristic patterns. If so, it further identifies the specific attack type and calculates a confidence score (i.e., attack probability) for this judgment. It also extracts the associated source and destination IP addresses, the network protocol used for communication, the type of affected system, and the timestamp of the behavior. All of this information constitutes the final detection result, providing a structured basis for subsequent response decisions.

[0265] The training process for the pre-defined intrusion detection classification model primarily involves data preparation, feature engineering, model selection, training optimization, and evaluation and validation. First, a large amount of historical network log data, encompassing both normal and various types of network attacks, is collected. This data is then cleaned, structured, and labeled to form a sample set labeled with attack types. Next, effective features are extracted through feature extraction and selection techniques to construct an input vector suitable for the classification model. Next, a suitable machine learning model (such as a random forest, support vector machine, or deep neural network) is selected and fed into the model for supervised learning. Model parameters are continuously adjusted to improve classification performance by minimizing the loss function. Cross-validation, regularization, and hyperparameter tuning are incorporated into the training process to prevent overfitting and enhance generalization. After training, the model's performance is evaluated using an independent test set to ensure high detection accuracy, recall, and stability in real-world environments. Ultimately, the model is deployed as the core analysis engine of the intrusion detection system.

[0266] In this solution, by training the preset intrusion detection classification model, the recognition accuracy and response speed of network attack behaviors can be effectively improved, and automatic identification and traceability analysis of various attack types can be achieved, thereby enhancing the system's intelligent protection capabilities, reducing the cost of manual intervention, and improving the overall network security defense efficiency and effectiveness.

[0267] Figure 3 A schematic block diagram of a 6G network automatic security handling system based on LLM provided in an embodiment of the present disclosure is characterized in that the system includes:

[0268] The log generation module 301 is used to obtain network data stream data and determine structured log data in JSON format according to the network data stream data;

[0269] The intrusion detection module 302 is used to perform feature extraction processing on the structured log data to obtain input feature data, input the input feature data into the pre-trained intrusion detection model for classification processing, and obtain network attack detection results;

[0270] The policy reasoning module 303 is used to convert the network attack detection results into an instruction fine-tuning format to obtain input data for a pre-trained LLM security model, input the input data into the pre-trained LLM security model for reasoning, and obtain a network security disposal policy;

[0271] A data packaging module 304 is configured to extract key fields of the network security handling policy and package the network security handling policy and the key fields into structured data;

[0272] The on-chain evidence storage module 305 is used to upload the structured data to the blockchain system for evidence storage, verify the integrity of the structured data content, and perform network security control operations according to the network security disposal strategy.

[0273] Figure 4 A schematic block diagram of an electronic device 400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0274] The electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a ROM 402 or a computer program loaded from a storage unit 408 into a RAM 403. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An I / O interface 405 is also connected to the bus 404.

[0275] Multiple components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0276] The computing unit 401 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the 6G network automatic security handling method based on LLM. For example, in some embodiments, the 6G network automatic security handling method based on LLM can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the 6G network automatic security handling method based on LLM described above can be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to execute the LLM-based 6G network automatic security handling method in any other appropriate manner (for example, by means of firmware).

[0277] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0278] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0279] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0280] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0281] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A 6G network automatic security handling method based on LLM, characterized in that: The method comprises: Obtaining network data stream data, and determining structured log data in JSON format based on the network data stream data; Performing feature extraction on structured log data to obtain input feature data, inputting the input feature data into a pre-trained intrusion detection model for classification processing to obtain network attack detection results; Performing instruction fine-tuning format conversion processing on the network attack detection result to obtain input data of a pre-trained LLM security model, inputting the input data into the pre-trained LLM security model for reasoning, and obtaining a network security disposal strategy; Extracting key fields of the network security disposal policy, and packaging the network security disposal policy and the key fields into structured data; The structured data is uploaded to the blockchain system for evidence storage, the integrity of the structured data content is verified, and network security control operations are performed according to the network security disposal strategy.

2. The method according to claim 1, characterized in that in, Determining structured log data in JSON format according to the network data stream data includes: Performing real-time packet capture and adaptive sampling processing on the network data stream data to obtain formatted network log data; The formatted network log data is cleaned and structured to obtain structured log data in JSON format.

3. The method according to claim 2, characterized in that in, Clean and structure the formatted network log data to obtain structured log data in JSON format, including: Perform redundant information removal on the formatted network log data to obtain preliminary cleansed log data; Perform abnormal numerical processing on the preliminary cleaned log data to obtain numerically standardized log data; Perform classification feature encoding on the numerical specification log data to obtain encoded log data; Convert the encoded log data into JSON format to obtain structured log data in JSON format.

4. The method according to claim 1, wherein in, Perform feature extraction on structured log data to obtain input feature data including: Perform flow size feature extraction on structured log data to obtain flow feature data; Perform protocol type feature extraction on structured log data to obtain protocol feature data; Extract source / destination IP address features from structured log data to obtain address feature data; Perform port number feature extraction on structured log data to obtain port feature data; The flow characteristic data, protocol characteristic data, address characteristic data and port characteristic data are combined and processed to obtain input characteristic data.

5. The method according to claim 1, wherein in, The network attack detection result is converted into an instruction fine-tuning format to obtain input data for a pre-trained LLM security model, including: Extract and process the attack type, source IP, target IP, system type, and timestamp fields in the network attack detection result to obtain key detection data; Formatting the key detection data according to a preset instruction template to obtain formatted detection data; Adding an instruction prompt prefix to the formatted detection data to obtain formatted detection data with the instruction prefix; The formatted detection data with instruction prefixes is encapsulated in JSON structure to obtain the input data of the pre-trained LLM security model.

6. The method according to claim 1, wherein in, Executing network security control operations according to the network security handling policy includes: Analyze the network security disposal strategy and extract specific operation instructions; Perform matching processing based on the preset rule base and specific operation instructions to determine the security tools that need to be called; Call the security tools as needed, call the firewall tool to perform IP blocking processing, and generate IP blocking execution results; Based on the security tools that need to be called, call the system configuration tool to perform two-factor authentication activation processing and generate the authentication configuration results; The security tool called as needed calls the port management tool to perform port closing processing and generate port management results; The IP blocking execution results, authentication configuration results and port management results are summarized and processed to generate a network security control execution report.

7. The method according to claim 1, characterized in that in, After determining the structured log data in JSON format according to the network data stream data, the method further includes: Performing statistical feature difference analysis on the structured log data to obtain feature importance scores; Performing dynamic feature selection processing according to the feature importance scores to obtain an optimal feature subset; Performing attack behavior time series analysis on the structured log data to obtain attack behavior pattern characteristics; Obtain historical attack tag data, perform co-occurrence frequency analysis on the historical attack tag data, and obtain tag association relationships; Performing label semantic enhancement processing based on the label association relationship and attack behavior pattern characteristics to obtain a refined attack type label; The optimal feature subset is combined with the refined attack type label to obtain enhanced feature data.

8. The method according to claim 7, characterized in that in, After obtaining the enhanced feature data, the method further includes: The enhanced feature data is input into a preset intrusion detection classification model to obtain a detection result including attack type, attack probability, source IP, target IP, communication protocol, system type and timestamp.

9. A 6G network automatic security handling system based on LLM, used to execute the method according to any one of claims 1 to 8, characterized in that: The system comprises: A log generation module is used to obtain network data flow data and determine structured log data in JSON format based on the network data flow data; The intrusion detection module is used to perform feature extraction processing on the structured log data to obtain input feature data, input the input feature data into the pre-trained intrusion detection model for classification processing, and obtain network attack detection results; A policy reasoning module is used to convert the network attack detection results into an instruction fine-tuning format to obtain input data for a pre-trained LLM security model, input the input data into the pre-trained LLM security model for reasoning, and obtain a network security disposal strategy; A data packaging module, configured to extract key fields of the network security disposal policy and package the network security disposal policy and the key fields into structured data; The on-chain evidence storage module is used to upload the structured data to the blockchain system for evidence storage, verify the integrity of the structured data content, and perform network security control operations according to the network security disposal strategy.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Key evidence construction, research and judgment method in automatic safety operation

    CN121441548A

  • Computer network big data security protection method

    CN121907576A