A network security field large model training method based on a generative large model
By preprocessing, improving and fine-tuning power network security data, and combining intelligent plug-in mechanisms and collaborative attack and defense, the problems of data acquisition and labeling difficulties and poor adaptability of generative large models in power network security have been solved, achieving more efficient network security protection and intelligent linkage.
Patent Information
- Application Number
- CN202411742847.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2044-11-29
AI Technical Summary
In the field of cybersecurity, especially in the area of power network security, difficulties in data acquisition and annotation, insufficient model generalization ability, and poor adaptability to application scenarios make it difficult to effectively apply generative large models.
By acquiring and preprocessing network security data from the power system, customizing and pre-training a general large model, and fine-tuning it using LORA instruction fine-tuning technology and prompt optimization technology, a large model foundation for power network security analysis is constructed. An intelligent plug-in mechanism is designed to interface with network security devices to achieve collaborative handling of attack and defense.
It improves the intelligence and automation of network security protection, enhances the generalization ability and application scenario adaptability of the model, can more accurately discover potential threats and attack patterns, realize intelligent linkage and collaborative defense between devices, and improve the ability to respond to complex threats.
Smart Images

Figure CN119917851B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power network security technology, and in particular to a method for training large network security models based on generative large models. Background Technology
[0002] With the rapid development of information technology and the internet, digital transformation has become an inevitable trend in the power industry. Against this backdrop, cybersecurity issues are increasingly prominent, becoming a key factor restricting the stable operation of digital power grids. Especially during the construction of smart grids, the deep integration of critical infrastructure such as industrial control systems and distribution automation systems with the internet has led to a continuous expansion of cyberattack threats, with attack methods becoming increasingly intelligent and covert. Traditional cybersecurity protection methods based on rule matching and feature recognition are no longer sufficient to cope with the increasingly complex and ever-changing cyber threats. In this context, artificial intelligence technology, especially generative large models, provides a new solution for cybersecurity protection. Generative large models, with their powerful knowledge representation and reasoning capabilities, can better understand the patterns and characteristics of cyberattacks and provide more intelligent security protection solutions. Therefore, developing a large-scale model training method for cybersecurity based on generative large models to improve the intelligence level of cybersecurity has become an urgent technical challenge.
[0003] While some applications of generative large models have been attempted in the field of cybersecurity, especially in power network security, numerous challenges remain. The primary issue is the difficulty in data acquisition and annotation. Due to the scarcity and time-sensitive nature of cyberattack data, and the need for accurate annotation by professionals, obtaining high-quality training data is costly. Secondly, existing models lack generalization ability and struggle to cope with unknown types of cyberattacks, particularly targeted attacks against power systems. Furthermore, the significant differences in network environments and business scenarios among different power companies result in poor adaptability of the models to various application scenarios, hindering rapid deployment and migration. These problems severely restrict the effective application of generative large models in the cybersecurity field. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the problem to be solved by this invention is to provide a method for training large models in the cybersecurity field based on generative large models, aiming to improve the current cybersecurity field, especially in power network security, which seriously restricts the effective application of generative large models in the cybersecurity field, such as difficulties in data acquisition and annotation, insufficient model generalization ability, and poor adaptability to application scenarios.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, embodiments of the present invention provide a method for training a large-scale model in the field of cybersecurity based on a generative large-scale model. The method includes acquiring and preprocessing cybersecurity data from a power system to obtain a training dataset; using a general large-scale model as a base model, and improving and pre-training the base model according to the characteristics of power network security to obtain a pre-trained model; fine-tuning the pre-trained model using the training dataset to obtain a fine-tuned model; performing intelligent correlation analysis on cybersecurity data based on the fine-tuned model to obtain analysis results; and designing an intelligent plug-in mechanism based on the analysis results, and interfacing the intelligent plug-in mechanism with cybersecurity devices to perform collaborative attack and defense countermeasures.
[0008] As a preferred embodiment of the generative large model-based large model training method for cybersecurity described in this invention, the preprocessing includes cleaning, classification, labeling, privacy correction of data, and compliance checks.
[0009] As a preferred embodiment of the generative large model-based network security large model training method described in this invention, the improvement and pre-training of the basic model according to the characteristics of power network security includes the following steps: selecting a general large model as the basic model; customizing and improving the basic model according to the characteristics of power network security to obtain a customized improved model; pre-training the customized improved model using a large-scale dataset to obtain a pre-trained model; and using LORA instruction fine-tuning technology to adapt to specific tasks in the power network security field by adding trainable parameters while keeping the parameters of the pre-trained model unchanged.
[0010] As a preferred embodiment of the generative large model-based large model training method for cybersecurity described in this invention, the method for obtaining the fine-tuned model includes the following steps: fine-tuning the pre-trained model using the training dataset, adjusting the model parameters to obtain the fine-tuned model; designing prompts to guide the fine-tuned model in reasoning and judgment based on the characteristics and requirements of cybersecurity tasks using prompt optimization techniques; training the fine-tuned model using the designed prompts, and adjusting the structure and content of the prompts based on the training results.
[0011] As a preferred embodiment of the large-scale model training method for cybersecurity based on generative large-scale models described in this invention, the intelligent correlation analysis includes the following steps: constructing a large-scale power network security analysis base based on the fine-tuned model; acquiring general knowledge, network security knowledge, and power network security attack and defense confrontation data, and fusing them with the large-scale power network security analysis base to obtain a fused large-scale power network security analysis base; and using the fused large-scale power network security analysis base to perform intelligent correlation analysis on network security data to obtain analysis results.
[0012] As a preferred embodiment of the generative large model-based large model training method for cybersecurity described in this invention, the analysis results include potential threats and attack patterns.
[0013] As a preferred embodiment of the generative large-scale model training method for cybersecurity described in this invention, the collaborative attack and defense response includes the following steps: designing an intelligent plug-in mechanism based on the analysis results to achieve seamless integration with cybersecurity devices and establishing an intelligent linkage mechanism between cybersecurity devices to achieve collaborative defense; designing an attack and defense analysis framework based on a generative large-scale model to achieve intelligent linkage and collaborative defense between devices; introducing a collaborative defense mechanism to link multiple cybersecurity devices or systems, automatically generating corresponding defense strategies when a potential attack is detected, and distributing the defense strategies to relevant devices for execution; and constructing an attack and defense knowledge base to accumulate and share experience and knowledge during the attack and defense process.
[0014] Secondly, to further address security issues in power network security, this invention provides a large-scale network security model training system based on a generative large-scale model. The system includes: a data acquisition module for acquiring and preprocessing network security data from the power system to obtain a training dataset; a model training module for using a general large-scale model as a base model and customizing and pre-training the base model according to the characteristics of power network security to obtain a pre-trained model; a model fine-tuning module for fine-tuning the pre-trained model using the training dataset to obtain a fine-tuned model; an intelligent analysis module for performing intelligent correlation analysis on network security data based on the fine-tuned model to uncover potential threats and attack patterns; and an attack-defense confrontation module for designing an intelligent plug-in mechanism based on the analysis results, connecting the intelligent plug-in mechanism with network security devices, and conducting collaborative attack-defense confrontation.
[0015] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements any step of the large model training method for the cybersecurity field based on generative large models as described in the first aspect of the present invention.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the large model training method for the cybersecurity field based on generative large models as described in the first aspect of the present invention.
[0017] The beneficial effects of this invention are as follows: This invention proposes a generative large-scale model training method for the cybersecurity field. By introducing a generative large-scale model, it achieves intelligent analysis and decision support for cybersecurity data, which not only improves the efficiency and accuracy of cybersecurity protection but also makes cybersecurity management more intelligent and automated. Traditional manual analysis methods are often time-consuming and labor-intensive and cannot comprehensively cover all potential threats, while the generative large-scale model can quickly detect and respond to cybersecurity risks through rapid processing and analysis of massive amounts of data. Through data acquisition and preprocessing, model base selection and pre-training, and fine-tuning and optimization steps, the professionalism and practicality of the model in the cybersecurity field are significantly improved, especially for the power industry. Customized improvements were made to the model to better adapt to the specific characteristics of the power network environment, enhancing its generalization ability and application scenario adaptability. Intelligent correlation analysis and collaborative attack-defense response steps were designed to further enhance the synergy and integrity of network security protection. By constructing a large-scale power network security analysis model foundation and integrating general knowledge, network security knowledge, and power network security attack-defense data, automatic identification and correlation of scattered network security data were achieved, enabling more accurate discovery of potential threats and attack patterns. Simultaneously, an attack-defense analysis framework based on a generative large-scale model was designed, realizing intelligent linkage and collaborative defense among devices, improving the ability to respond to complex threats. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0019] Figure 1 This is an overall flowchart of the large model training method for the cybersecurity field based on generative large models in Example 1.
[0020] Figure 2 This is a schematic diagram of the computer device in Example 3. Detailed Implementation
[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0022] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0023] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0024] Example 1
[0025] Reference Figure 1 This is the first embodiment of the present invention, which provides a method for training large models in the field of cybersecurity based on generative large models.
[0026] Existing methods for training large-scale models in the cybersecurity field suffer from the following main problems: First, the difficulty in data acquisition and annotation is a major issue. Due to the scarcity and time-sensitivity of cyberattack data, and the need for accurate annotation by professionals, the cost of acquiring high-quality training data is high. Second, existing models lack generalization ability and struggle to cope with unknown types of cyberattacks, especially targeted attacks against power systems. Furthermore, the network environments and business scenarios of different power companies vary significantly, resulting in poor adaptability of the models to different application scenarios and making rapid deployment and migration difficult. These problems severely restrict the effective application of generative large-scale models in the cybersecurity field.
[0027] This application provides a method that can effectively solve the problems mentioned above. The following will describe in detail how to implement this method for training large models in the cybersecurity field based on generative large models, using several embodiments.
[0028] Figure 1 The overall flowchart of a generative large model-based method for training large models in the cybersecurity field is shown, including:
[0029] S1: Obtain network security data of the power system and preprocess it to obtain the training dataset.
[0030] Preferably, network security data includes log data, traffic data, and event log data.
[0031] Specifically, the formula for a dataset of cybersecurity data is as follows:
[0032] D = {d1, d2, ..., dn}
[0033] Where D is the dataset of network security data; dn is the nth piece of network security data.
[0034] Preferably, preprocessing includes cleaning, sorting, labeling, privacy correction of the data, and compliance checks to ensure the legality and security of the data.
[0035] Specifically, data cleaning refers to removing invalid or erroneous data to ensure its accuracy and integrity. The specific formula for the cleaned dataset is as follows:
[0036] D' = filter(D)
[0037] Where D′ is the cleaned dataset.
[0038] Specifically, classification refers to categorizing data by type or severity, including attack type and impact level. The specific formula is as follows:
[0039] Dc = classify(D')
[0040] Where Dc is the classified dataset.
[0041] Specifically, annotation refers to adding labels to cybersecurity data to facilitate subsequent model training and analysis. The specific formula is as follows:
[0042] Dl = label(Dc)
[0043] Where Dl is the labeled dataset.
[0044] It should be noted that, in order to enhance the diversity of data collection and preprocessing, in the data collection stage, in addition to collecting data from within the power system, data can also be obtained from external cybersecurity databases or publicly available security incident reporting channels to increase the diversity and comprehensiveness of the data; or data can be generated through simulated attack experiments, which can simulate various attack scenarios and methods, providing more practical samples for model training; in the data cleaning stage, more advanced data cleaning technologies can be introduced, such as machine learning-based anomaly detection algorithms, to more accurately identify and remove invalid or erroneous data; in the data annotation stage, semi-automated or automated annotation tools can be used in combination with manual review to improve annotation efficiency and accuracy.
[0045] Furthermore, privacy verification and compliance checks on data include ensuring that data collection complies with relevant laws and regulations and protects user privacy, ensuring the legality and security of data, and avoiding legal disputes or privacy leaks during model training and use. This includes the following steps: identifying potential privacy risks and vulnerabilities in data processing practices, including a comprehensive review of the collection, storage, use, transmission, and disclosure of data.
[0046] Conduct a compliance assessment of data processing practices in accordance with applicable privacy regulations and standards (such as GDPR and CCPA), including checking whether data processing complies with regulatory requirements and whether appropriate security measures have been taken to protect personal data.
[0047] Based on the assessment results, improvement measures and recommendations are proposed to enhance the compliance and security of data processing and protection practices, including strengthening data encryption, restricting data access permissions, improving data backup and recovery strategies, etc.
[0048] Preferably, this invention ensures high-quality input data by utilizing cleaning, classification, and annotation techniques, thereby improving the effectiveness of the training dataset and solving the problem of complex and ever-changing cybersecurity data. At the same time, compliance checks and privacy correction techniques ensure data legality, prevent data leakage during model use, and improve the reliability and legal compliance of the model.
[0049] S2: Using the general large model as the base model, and improving and pre-training the base model according to the characteristics of power network security, a pre-trained model is obtained.
[0050] Preferably, improving and pre-training the basic model according to the characteristics of power network security includes the following steps: selecting a general large model as the basic model, wherein the general large model includes BERT and GPT.
[0051] The basic model was customized and improved based on the characteristics of power network security, including adjusting the model structure and adding a vocabulary for specific domains, resulting in a customized and improved model. The specific formula is as follows:
[0052] M' = customize(M)
[0053] Where M′ is the customized improved model; M is the general large model.
[0054] The customized and improved model is pre-trained using a large-scale dataset to obtain a pre-trained model that possesses basic language understanding and analysis capabilities. The specific formula is as follows:
[0055] Mp = pretrain(M')
[0056] Where Mp is the pre-trained model.
[0057] By employing the LORA instruction fine-tuning technique, the parameters of the pre-trained model are kept unchanged, and trainable parameters are added to adapt to specific tasks in the field of power network security.
[0058] Ideally, by employing the LORA instruction fine-tuning technique, it is possible to adapt to specific tasks in the field of power network security by adding a small number of trainable parameters while keeping the original model parameters unchanged. The LORA instruction fine-tuning technique can reduce the time and resource consumption of model training, while maintaining the stability and performance of the original model.
[0059] Specifically, the LORA instruction fine-tuning technique includes the following steps: selecting a task-related pre-trained model as a starting point, which is trained on a large dataset and has strong representation capabilities and generalization performance.
[0060] During fine-tuning, most of the weights of the pre-trained model are frozen, and only a small portion of the trainable layers (such as the rank-factor matrix) are updated, reducing the number of training parameters required and lowering computational resources and time costs.
[0061] Trainable layers are injected into each Transformer block to learn relevant knowledge for a specific task. By training these layers, the pre-trained model can be fine-tuned to better adapt to new datasets and tasks.
[0062] The model is trained and its performance is evaluated using a new dataset. Based on the evaluation results, the training parameters and the structure of the trainable layers are adjusted to further improve the model's performance.
[0063] It should be noted that, in order to diversify the choice of model bases, in addition to general large models such as BERT and GPT, it is also possible to consider choosing specialized large models more suitable for the cybersecurity field as bases, such as specialized models for security text analysis. Different large model bases can be selected for different security needs to achieve more refined security analysis and prediction. In terms of pre-training, self-supervised learning or contrastive learning pre-training strategies can be introduced to enable the model to learn useful feature representations on unlabeled data, thereby improving the model's generalization ability. At the same time, customized pre-training tasks can be designed in combination with specific tasks in the cybersecurity field, including security text classification and security event prediction, to enhance the model's adaptability to the cybersecurity field.
[0064] S3: Fine-tune the pre-trained model using the training dataset to obtain the fine-tuned model.
[0065] Preferably, obtaining the fine-tuned model includes the following steps: fine-tuning the pre-trained model using the training dataset, adjusting the model parameters, and obtaining the fine-tuned model to improve the model's professionalism and practicality in the field of cybersecurity, such as improving the ability to identify specific attack types or improving the accuracy of analysis.
[0066] Based on the characteristics and requirements of cybersecurity tasks, prompt optimization technology is adopted to design appropriate prompts to guide the fine-tuning model in reasoning and judgment.
[0067] The model is trained using designed prompts, and the structure and content of the prompts are adjusted based on the training results to improve the model's ability to understand and analyze cybersecurity data.
[0068] Specifically, the formula for fine-tuning the model is as follows:
[0069] Mf = finetune(Mp, Dl)
[0070] Where Mf is the fine-tuned model; Mp is the pre-trained model; and Dl is the labeled dataset.
[0071] Furthermore, before fine-tuning the pre-trained model using the training dataset, differentiated fine-tuning strategies can be designed based on different cybersecurity application scenarios.
[0072] Specifically, the differentiated fine-tuning strategies include malware detection task strategies and network attack prediction task strategies.
[0073] Specifically, the malware detection task strategy refers to increasing the training data containing malicious code samples and adjusting the model's weights and parameters to make it more sensitive to training data containing malicious code samples. This allows for a focus on optimizing the model's ability to learn malicious code features, enabling it to more accurately identify the characteristics of malicious code, such as specific code patterns and behavioral patterns.
[0074] Specifically, the network attack prediction task strategy refers to the use of historical attack data to construct time series models or pattern recognition models for network attack prediction tasks, enabling the models to predict future attack behaviors, thereby enhancing the models' ability to identify attack patterns and trends, and thus identifying network attack patterns and trends.
[0075] Furthermore, the prompt optimization technique can be used to improve the model's ability to understand and analyze cybersecurity data. This technique involves designing appropriate prompts to guide the model in better understanding and analyzing cybersecurity data, thereby enhancing the model's practicality and reliability. The steps include: designing appropriate prompts to guide the model in reasoning and judgment based on the characteristics and requirements of the cybersecurity task. The prompts include problem descriptions, contextual information, and examples, which help the model better understand the task objectives and data characteristics.
[0076] Using a Prompt to train the model and adjusting its structure and content based on the training results, the model's ability to identify and analyze cybersecurity data can be improved through continuous optimization of the Prompt.
[0077] The optimized model is applied to real-world cybersecurity tasks and its performance is evaluated. Based on the evaluation results, the Prompt and model parameters are further adjusted to achieve better performance.
[0078] It should be noted that, to improve model training efficiency and performance, transfer learning methods can be introduced to utilize knowledge learned in other domains to assist in fine-tuning models in the cybersecurity field, thereby accelerating learning on new tasks. Secondly, in terms of optimization techniques, in addition to prompt optimization, hyperparameter tuning, model pruning, and knowledge distillation can be considered to further improve model performance and efficiency. Hyperparameter tuning involves adjusting the model's hyperparameters, such as learning rate and batch size, to find the optimal model configuration. Model pruning reduces model complexity and computational cost by removing redundant parameters or neurons while maintaining model performance. Knowledge distillation compresses the knowledge of a large model into a smaller model while maintaining model performance; this can be achieved by training a smaller model to mimic the output of a larger model.
[0079] Ideally, by introducing prompt optimization techniques, the model can be guided to more accurately identify and analyze data in task-specific contexts, improving its accuracy in identifying attack patterns, optimizing the model's adaptability, reducing the training cycle, and enhancing security in practical applications through targeted fine-tuning.
[0080] S4: Based on the fine-tuning model, intelligent correlation analysis is performed on network security data to obtain analysis results.
[0081] Preferably, intelligent correlation analysis includes the following steps: constructing a large-scale power network security analysis base based on a fine-tuning model.
[0082] Obtain general knowledge, network security knowledge, and power network security attack and defense confrontation data, and fuse them with the power network security analysis large model base to obtain the fused power network security analysis large model base.
[0083] Use the fused power network security analysis large model base to perform intelligent correlation analysis on network security data, obtain analysis results, and achieve automatic identification and correlation of dispersed network security data.
[0084] Specifically, the analysis results include potential threats and attack patterns, and the specific formula is as follows:
[0085] R = analyze(Mf, D_large)
[0086] Where R is the analysis result; Mf is the fine-tuning model; D_large is the massive network security data.
[0087] It should be noted that for the deepening and improvement of intelligent correlation analysis, a more complex correlation analysis model can be constructed, such as a model based on graph neural networks, to better capture the correlation relationships between network security data; in addition, real-time correlation analysis technology can also be introduced to achieve real-time monitoring and rapid response to network security data.
[0088] Furthermore, by constructing a power network security analysis large model base, fusing general knowledge, network security knowledge, and power network security attack and defense confrontation data, and achieving automatic identification and correlation of dispersed network security data, the following steps are included: Collect power network security-related data, including network traffic data, log data, and vulnerability information, and integrate and clean these data to eliminate redundant and error information.
[0089] Fuse general knowledge, network security knowledge, and power network security attack and defense confrontation data into the large model base, which can be achieved through knowledge graph or ontology technology, and helps the model better understand the concepts and relationships in the network security field.
[0090] Use the integrated data to train the model, and adjust the model structure and parameters according to the training results. By continuously optimizing the model, improve its ability to identify and analyze power network security data.
[0091] Apply the trained model to actual power network security tasks, verify its performance, and further adjust the model and optimize the algorithm according to the verification results to achieve better performance.
[0092] It should be noted that there are no specific formulas or data format requirements in this process. The key is how to effectively integrate and utilize various data sources and knowledge resources to build a large model foundation, and improve the model's performance through training and optimization. Integrating multiple knowledge sources can enhance the model's ability and accuracy in analyzing complex cybersecurity issues, providing security teams with more comprehensive and in-depth threat intelligence.
[0093] Ideally, by utilizing the integrated power network security analysis big data platform for intelligent correlation analysis of multi-source data, potential threats and attack patterns can be identified in real time, complex attack patterns and potential security threats can be captured, and different security data can be automatically correlated. This fusion of deep learning and knowledge graph can effectively improve the ability to predict attack behavior and provide stronger data support for decision-making.
[0094] S5: Based on the analysis results, design an intelligent plug-in mechanism, connect the intelligent plug-in mechanism with network security devices, and carry out collaborative attack and defense confrontation.
[0095] Preferably, the collaborative attack and defense response includes the following steps: designing an intelligent plug-in mechanism based on the analysis results to achieve seamless integration with network security devices, and establishing an intelligent linkage mechanism between network security devices to achieve collaborative defense. The specific formula for the collaborative defense strategy is as follows:
[0096] S=defense_strategy(Mf,Devices)
[0097] Where S represents the collaborative defense strategy; Devices represents existing network security devices.
[0098] Design an attack and defense adversarial analysis framework based on generative large models to achieve intelligent linkage and collaborative defense between devices, and improve the ability to respond to complex threats. This framework uses sequence generation technology to construct attack scenarios and adversarial training technology to optimize attack identification.
[0099] By introducing a collaborative defense mechanism, multiple network security devices or systems can be linked together. When a potential attack is detected, a corresponding defense strategy is automatically generated and distributed to relevant devices for execution, thereby achieving more efficient defense and response.
[0100] Build an attack and defense knowledge base to accumulate and share experience and knowledge from the attack and defense process in order to improve the overall security level.
[0101] Specifically, the formula for coordinated offensive and defensive countermeasures is as follows:
[0102] T=attack_defense(Mf,Threats)
[0103] Where T represents the intelligent attack and defense response results; Threats represents the detected network threats.
[0104] Furthermore, a generative large model-based attack and defense analysis framework is designed to achieve intelligent linkage and collaborative defense between devices. This includes the following steps: clarifying the specific scenarios and objectives of attack and defense, including attack types and defense strategies, which helps to determine the design requirements and functional requirements of the framework.
[0105] Generative large models can be trained using offensive and defensive adversarial data to generate various possible attack and defense strategies. This can be achieved through sequence generation or adversarial training techniques.
[0106] By integrating generative large models into the communication and collaboration mechanisms between devices, intelligent linkage and collaborative defense between devices can be achieved. When a potential attack is detected, the framework can automatically generate corresponding defense strategies and distribute them to relevant devices for execution.
[0107] Based on the actual results and feedback data of offensive and defensive confrontations, the generative large model and collaborative defense mechanism are continuously optimized through online learning or reinforcement learning techniques to improve the adaptability and robustness of the framework.
[0108] It should be noted that the design and implementation of the attack and defense confrontation analysis framework based on generative large models need to comprehensively consider the complexity of attack and defense confrontation, the communication and collaboration mechanisms between devices, and the performance and scalability factors of generative large models. Through continuous optimization and improvement, a more efficient, intelligent, and reliable collaborative attack and defense confrontation response solution can be achieved. The attack and defense confrontation analysis framework based on generative large models can improve the overall defense level of network security, achieve rapid response and effective handling of complex threats, and promote the collaborative work between different security devices, thereby improving overall security effectiveness.
[0109] It should be noted that, regarding the overall system architecture and implementation of this invention, in terms of system architecture design, a microservice-based system architecture is designed to decouple and independently deploy the data acquisition, model training, intelligent correlation analysis, and attack and defense collaborative handling modules, thereby improving the system's scalability and flexibility; a suitable technology stack for big data processing and machine learning is selected, such as Hadoop, Spark, TensorFlow, or PyTorch, to achieve efficient data processing and model training; cloud-native technologies, such as cloud computing, cloud storage, and cloud monitoring, are introduced to improve the system's availability and reliability; and containerization technologies, such as Docker and Kubernetes, are introduced, which are crucial for achieving rapid system deployment and automated management.
[0110] Furthermore, Docker is used to create, deploy, and manage containerized applications, packaging applications and their dependencies into a single container for rapid deployment and consistent operation; Kubernetes, as an open-source container orchestration system, automates the deployment, scaling, and management of containers, managing the lifecycle of containerized applications, including deployment, monitoring, scaling, and fault recovery; Hadoop provides a reliable and cost-effective solution for massive data storage, achieving parallel data processing and fault tolerance through its distributed file system HDFS and MapReduce programming model; Spark, with its speed, ease of use, and in-memory computing advantages, has become... Ideal for handling complex data streams and real-time analysis, cloud computing provides highly abstract data processing interfaces, making data processing logic more concise and efficient. TensorFlow and PyTorch, as deep learning frameworks, support efficient model training and inference, forming the foundation for implementing machine learning algorithms. Cloud computing provides on-demand computing resources and services, including virtual machines, containers, and databases, supporting elastic scaling and pay-as-you-go pricing. Cloud storage provides highly available and scalable data storage services, supporting data backup, recovery, and cross-region replication. Cloud monitoring provides real-time monitoring and alerting of system performance and operational status, helping to promptly identify and address potential problems.
[0111] In summary, this invention proposes a generative large-scale model training method for the cybersecurity field. By introducing a generative large-scale model, it achieves intelligent analysis and decision support for cybersecurity data, improving not only the efficiency and accuracy of cybersecurity protection but also making cybersecurity management more intelligent and automated. Traditional manual analysis methods are often time-consuming, labor-intensive, and unable to comprehensively cover all potential threats, while the generative large-scale model can quickly identify and respond to cybersecurity risks through rapid processing and analysis of massive amounts of data. Through data acquisition and preprocessing, model base selection and pre-training, and fine-tuning and optimization steps, the professionalism and practicality of the model in the cybersecurity field are significantly improved, especially for power network security. Customized improvements were made to the model to better adapt to the specific characteristics of the power network environment, improving its generalization ability and application scenario adaptability. Intelligent correlation analysis and collaborative attack-defense response steps were designed to further enhance the synergy and integrity of network security protection. By constructing a large-scale power network security analysis model foundation and integrating general knowledge, network security knowledge, and power network security attack-defense data, automatic identification and correlation of scattered network security data were achieved, enabling more accurate discovery of potential threats and attack patterns. Simultaneously, an attack-defense analysis framework based on a generative large-scale model was designed, realizing intelligent linkage and collaborative defense between devices, improving the ability to respond to complex threats.
[0112] Example 2, an embodiment of the present invention, provides a large-scale model training system for the cybersecurity field based on a generative large-scale model, comprising: a data acquisition module for acquiring and preprocessing cybersecurity data of a power system to obtain a training dataset; a model training module for using a general large-scale model as a base model and customizing and pre-training the base model according to the characteristics of power network security to obtain a pre-trained model; a model fine-tuning module for fine-tuning the pre-trained model using the training dataset to obtain a fine-tuned model; an intelligent analysis module for performing intelligent correlation analysis on cybersecurity data based on the fine-tuned model to uncover potential threats and attack patterns; and an attack and defense confrontation module for designing an intelligent plug-in mechanism based on the analysis results, connecting the intelligent plug-in mechanism with cybersecurity devices, and performing collaborative attack and defense confrontation.
[0113] Example 3 is an embodiment of the present invention, which differs from the previous embodiment in that:
[0114] like Figure 2 As shown, if the function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0115] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0116] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0117] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0118] Example 4 is an embodiment of the present invention, which provides a method for training large models in the field of cybersecurity based on generative large models. In order to verify the beneficial effects of the present invention, a simulation experiment is conducted for scientific demonstration.
[0119] This example demonstrates how a microservice-based system architecture is designed to decouple and independently deploy data acquisition, model training, intelligent correlation analysis, and collaborative attack and defense modules, thereby improving the system's scalability and flexibility. In selecting a suitable technology stack for big data processing and machine learning, the speed of Hadoop and Spark technologies in task processing is compared to choose the stack most suitable for this invention. Specific data is shown in Table 1.
[0120] Table 1 Comparison of Data Processing Performance
[0121] Dataset size Processing tasks Spark (seconds) Hadoop (seconds) 1GB Sort 10 30 10GB polymerization 30 120 100GB filter 120 480
[0122] By introducing cloud-native technologies, such as cloud computing, cloud storage, or cloud monitoring, the availability and reliability of the system can be improved. Tables 2 and 3 show the performance data for cloud storage and cloud monitoring, respectively.
[0123] Table 2 Performance data of cloud storage
[0124] Storage services Read / write speed (MB / s) Delay (ms) Availability (%) AWS S3 200 10 99.99
[0125] Table 3 Performance Data of Cloud Monitoring
[0126] Monitoring indicators Threshold setting Number of alarms CPU utilization >80% 2 times Memory usage >90% 1 time Network bandwidth >1Gbps 3 times
[0127] In the data collection and preprocessing steps, in order to ensure the legality and security of the data, privacy verification and compliance checks are also performed on the network security data. Table 4 shows the results of the privacy verification and compliance checks.
[0128] Table 4. Results of Privacy Correction and Compliance Check
[0129] Data types Privacy Risk Level Compliance inspection results User Personal Information high pass Transaction records middle pass System Log Low pass
[0130] When fine-tuning the pre-trained model using the training dataset to obtain the fine-tuned model, the prompt optimization technique is used to design appropriate prompts to guide the model to better understand and analyze cybersecurity data, thereby improving the model's ability to understand and analyze cybersecurity data and thus improving the model's practicality and reliability. Table 5 shows the experimental results of the Prompt optimization.
[0131] Table 5. Results of the Prompt optimization experiment
[0132] Prompt Design Accuracy improvement (%) Basic Prompt 0 Optimize Prompt1 +5 Optimize Prompt2 +8
[0133] Finally, by using a fine-tuning model to perform intelligent correlation analysis on network security data, the analysis results were obtained. Based on the analysis results, an intelligent plug-in mechanism was designed and connected with network security devices to carry out collaborative attack and defense. Table 6 shows the experimental results of collaborative attack and defense.
[0134] Table 6. Results of the Collaborative Offensive and Defensive Response Experiment
[0135] Attack type Defense strategy Success rate (%) SQL injection Input Validation 100 DDoS attack Flow cleaning 98 Malware Behavioral Analysis 95
[0136] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for training a large model in the field of network security based on a generative large model, characterized in that: The method comprises the following steps: Obtain network security data of a power system and preprocess the data to obtain a training data set; Use a general large model as a base model, and improve and pre-train the base model according to the characteristics of the power network security to obtain a pre-trained model; Fine-tune the pre-trained model using the training data set to obtain a fine-tuned model; Intelligently associate and analyze the network security data based on the fine-tuned model to obtain an analysis result; Design an intelligent plug-in mechanism based on the analysis result, and connect the intelligent plug-in mechanism with network security devices for attack-defense confrontation collaborative disposal; The attack-defense confrontation collaborative disposal comprises the following steps: designing an intelligent plug-in mechanism based on the analysis result, realizing seamless connection with the network security devices, establishing an intelligent linkage mechanism between the network security devices, and realizing collaborative defense. The specific formula of the collaborative defense strategy is as follows: ; wherein, for a cooperative defense strategy; for an existing network security device; Introduce a collaborative defense mechanism to link multiple network security devices or systems, automatically generate corresponding defense strategies when potential attacks are detected, and distribute the defense strategies to related devices for execution; The specific formula of the attack-defense confrontation collaborative disposal is as follows: ; wherein, is an intelligent attack-defense confrontation handling result; is a detected network threat; Train a generative large model using attack-defense confrontation data to enable it to generate attack and defense strategies, and realize this through sequence generation or adversarial training technology.
2. The large model training method for the network security field based on the generative large model according to claim 1, characterized in that: The preprocessing comprises cleaning, classification, labeling, privacy revision of data, and compliance check.
3. The large model training method for the network security field based on the generative large model according to claim 2, characterized in that: Improving and pre-training the base model according to the characteristics of the power network security comprises the following steps: Select a general large model as a base model; Customize and improve the base model according to the characteristics of the power network security to obtain an improved model; Pre-train the improved model using a large-scale data set to obtain a pre-trained model; Use LORA instruction fine-tuning technology to adapt to the tasks in the power network security field by increasing trainable parameters while keeping the parameters of the pre-trained model unchanged.
4. The large model training method for the network security field based on the generative large model according to claim 3, characterized in that: The fine-tuned model comprises the following steps: Fine-tune the pre-trained model using the training data set, adjust the model parameters, and obtain a fine-tuned model; According to the characteristics and requirements of network security tasks, use prompt optimization technology to design a prompt to guide the fine-tuned model to reason and judge; Train the fine-tuned model using the designed prompt, and adjust the structure and content of the prompt according to the training results.
5. The large model training method for the network security field based on the generative large model according to claim 4, characterized in that: The intelligent association analysis comprises the following steps: Build a power network security analysis large model base based on the fine-tuned model; Obtain general knowledge, network security knowledge, and attack-defense confrontation data of the power network security, and fuse them with the power network security analysis large model base to obtain a fused power network security analysis large model base; Use the fused power network security analysis large model base to intelligently associate and analyze the network security data to obtain an analysis result.
6. The large model training method for the network security field based on the generative large model according to claim 5, characterized in that: The analysis result comprises potential threats and attack patterns.
7. The large model training method for the network security field based on the generative large model according to claim 6, characterized in that: The attack-defense confrontation collaborative disposal comprises the following steps: Design an intelligent plug-in mechanism based on the analysis result, realize seamless connection with the network security devices, establish an intelligent linkage mechanism between the network security devices, and realize collaborative defense. An attack-defense confrontation analysis framework based on a generative large model is designed to realize intelligent linkage and collaborative defense between devices. A collaborative defense mechanism is introduced to link multiple network security devices or systems, automatically generate corresponding defense strategies when potential attacks are detected, and distribute the defense strategies to relevant devices for execution. An attack-defense knowledge base is constructed to accumulate and share experience and knowledge in the attack-defense process.
8. A large model training system for the network security field based on a generative large model, based on the large model training method for the network security field based on a generative large model according to any one of claims 1-7, characterized in that: The method comprises the following steps: a data acquisition module is used to acquire network security data of a power system and perform preprocessing to obtain a training data set; a model training module is used to use a general large model as a base model, and customize and improve the base model and pre-train it according to the characteristics of the power network security to obtain a pre-trained model; a model fine-tuning module is used to fine-tune the pre-trained model using the training data set to obtain a fine-tuned model; an intelligent analysis module is used to perform intelligent correlation analysis on the network security data based on the fine-tuned model to mine potential threats and attack patterns; an attack-defense confrontation module is used to design an intelligent plug-in mechanism based on the analysis results, and interface the intelligent plug-in mechanism with network security devices for attack-defense confrontation and collaborative disposal. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is characterized in that: The processor executes the computer program to realize the steps of the network security field large model training method based on the generative large model according to any one of claims 1-7.
10. A computer readable storage medium having stored thereon a computer program, characterized in that: The computer program is executed by the processor to realize the steps of the network security field large model training method based on the generative large model according to any one of claims 1-7.
Citation Information
Patent Citations
Customer service scene-oriented generation matching type large model construction method, medium and equipment
CN117709969A
Intelligent Web application protection method based on AI semantics
CN119030732A