LLM-based electric power industrial control network simulation verification scene generation method

Through the generation method of the power industrial control network simulation verification scenario based on large language models, the problem of cumbersome construction of power industrial control network simulation scenarios in the existing technology is solved, and the effect of rapidly building complex attack scenarios and improving response capabilities is achieved.

CN119989912APending Publication Date: 2025-05-13TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140523.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing method of building a simulation scenario for power industrial control networks mainly relies on manual debugging, the process is cumbersome and time-consuming, it is difficult to respond quickly to complex attack methods, and lacks intelligent support, so it is impossible to effectively predict and simulate potential attack paths and security vulnerabilities.

Method used

The power industrial control network simulation verification scenario generation method is adopted based on large language model (LLM), and the training data set is constructed by collecting and preprocessing the input data and label data of the power industrial control network topology, and the large language model is used to generate the model and output the directed graph matrix and alarm information of the network topology, evaluating the ability to generate the model, and finally building a virtual simulation scenario.

Benefits of technology

It has realized the rapid construction of complex and real attack scenarios, improved the response capabilities of the power industrial control network in the face of actual attacks, enhanced the virtualized power industrial control network counterattack capabilities of security experts, reduced the dependence of human debugging, and improved the adaptability and coverage of simulation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989912A_ABST
    Figure CN119989912A_ABST
Patent Text Reader

Abstract

The invention discloses a power industrial control network simulation verification scene generation method based on LLM, and the method comprises the steps: collecting paired input data and label data based on a power industrial control network topology to generate a training data set, and converting the training data set into a model input vector through preprocessing and vector coding; constructing a generation model based on a large language model by using the vectorized training data set, and obtaining an output result of a directed graph matrix containing network topology and alarm information through the trained generation model; evaluating the alarm information generation capability and the matrix generation capability according to the output result to obtain an evaluation result; and converting a square matrix and a vertex array output by the generative model into a directed graph form according to an evaluation result, and constructing a virtual simulation scene in a graph node mapping mode. According to the invention, the problems of single and fixed network scene required by simulation verification of the electric power industrial control network, incapability of rapid adjustment and the like can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power network security, and in particular to a method for generating a power industrial control network simulation verification scenario based on LLM. Background Art

[0002] The stable operation of the power system is the foundation of national economic development, and power network security has been raised to the national level. In recent years, confrontation risks at home and abroad have continued to intensify. Once a power safety accident occurs, it will inevitably have a direct impact on social production and residents' lives. Since the power production system has extremely high requirements for continuity and stability, it is impossible to directly carry out safety-related tests and research on the production system. Therefore, it is necessary to establish a flexible and highly simulated simulation verification environment and confrontation verification environment. By constructing a power network security simulation verification environment, the operating status of the power system can be simulated, potential security threats can be predicted and evaluated, and a scientific basis can be provided for the safety protection of the power system. At present, both at home and abroad, security simulation verification environments including network security, big data security, and power network security have been completed or begun to be built, playing an important role in technical research, scientific evaluation, competition and evaluation exercises, and talent training.

[0003] The construction of power industrial control network simulation scenarios is a key link in the development of power network security simulation platforms, and has far-reaching significance for the security protection and technological innovation of the power industry. By building realistic simulation scenarios, researchers can conduct in-depth research on new attack methods and corresponding defense technologies, thereby improving the effectiveness of network security strategies. This virtualized environment provides security experts with a test field for testing and optimizing security strategies. It can not only simulate attacks and defenses in various scenarios, but also help identify the strengths and weaknesses of existing defense measures, and further optimize the design and implementation of security defense systems. LLM (Large Language Model) Large language model, LLM can understand and generate various complex forms of human language, including grammar, semantics, contextual relationships, etc. by training on large-scale text data sets. The rise of LLM is due to the improvement of computing power, the accumulation of big data, and the advancement of deep learning algorithms. These models have achieved significant performance improvements in multiple tasks, such as text generation, question-answering systems, machine translation, text classification, summary generation, etc. It has shown great potential in the generation of power network security simulation scenarios.

[0004] Through repeated simulation and testing of simulation scenarios, researchers can accurately evaluate the performance of various security strategies in different threat scenarios, thereby improving the ability of power industrial control networks to respond to actual attacks. In addition, this simulation practice also provides an important platform for cultivating power industry security professionals with practical capabilities. By training in a simulation environment, professionals can master emergency response skills in complex network environments and enhance their ability to respond to real network threats. This is especially important for the power system, a critical infrastructure that is related to the lifeline of the national economy, to ensure that it can maintain stable operation under any circumstances.

[0005] However, existing scenario construction methods mainly rely on manual debugging, which is not only cumbersome and time-consuming, but also easily affected by the experience and cognitive limitations of the debuggers. This construction method that relies on traditional experience is often difficult to respond quickly when faced with increasingly complex attack methods, resulting in limited adaptability and coverage of simulation scenarios. Especially in the face of the ever-evolving automated penetration technology, traditional manual debugging methods seem to be unable to meet the needs of quickly generating diverse simulation scenarios.

[0006] In addition, existing methods lack intelligent support when constructing complex scenarios and cannot effectively predict and simulate potential attack paths and security vulnerabilities. This not only limits researchers' exploration of new attack methods, but also makes it difficult for security experts to identify and assess potential risks in advance in a simulation environment.

[0007] Through repeated simulation and testing of simulation scenarios, researchers can accurately evaluate the performance of various security strategies in different threat scenarios, thereby improving the ability of power industrial control networks to respond to actual attacks. In addition, this simulation practice also provides an important platform for cultivating power industry security professionals with practical capabilities. By training in a simulation environment, professionals can master emergency response skills in complex network environments and enhance their ability to respond to real network threats. This is especially important for the power system, a critical infrastructure that is related to the lifeline of the national economy, to ensure that it can maintain stable operation under any circumstances.

[0008] However, existing scenario construction methods mainly rely on manual debugging, which is not only cumbersome and time-consuming, but also easily affected by the experience and cognitive limitations of the debuggers. This construction method that relies on traditional experience is often difficult to respond quickly when faced with increasingly complex attack methods, resulting in limited adaptability and coverage of simulation scenarios. Especially in the face of the ever-evolving automated penetration technology, traditional manual debugging methods seem to be unable to meet the needs of quickly generating diverse simulation scenarios.

[0009] In addition, existing methods lack intelligent support when constructing complex scenarios and cannot effectively predict and simulate potential attack paths and security vulnerabilities. This not only limits researchers' exploration of new attack methods, but also makes it difficult for security experts to identify and assess potential risks in advance in a simulation environment. Summary of the invention

[0010] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0011] The present invention proposes a method for generating a power industrial control network simulation verification scenario based on LLM, so as to solve the problems that the network scenario required for the power industrial control network simulation verification is single, fixed, and cannot be quickly adjusted.

[0012] To achieve the above-mentioned purpose, the present invention proposes, on one hand, a method for generating a simulation verification scenario of an electric power industrial control network based on LLM, comprising:

[0013] Based on the power industrial control network topology, paired input data and label data are collected to generate a training data set, and the training data set is converted into a model input vector through preprocessing and vector encoding;

[0014] The vectorized training data set is used to build a generative model based on a large language model, and the output results of the directed graph matrix containing the network topology and the alarm information are obtained through the trained generative model;

[0015] Evaluate the generation capability of the alarm information and the matrix generation capability according to the output result to obtain an evaluation result;

[0016] According to the evaluation results, the square matrix and vertex array output by the generated model are converted into a directed graph form, and a virtual simulation scene is constructed through graph node mapping.

[0017] The method for generating a power industrial control network simulation verification scenario based on LLM in the embodiment of the present invention may also have the following additional technical features:

[0018] In one embodiment of the present invention, paired input data and label data are collected based on the power industrial control network topology to generate a training data set, and the training data set is converted into a model input vector through preprocessing and vector encoding, including:

[0019] Taking a text sentence describing the topological structure characteristics of the power industrial control network as input data, and obtaining matching label data constructed for the text sentence to obtain a training data set;

[0020] The word segmenter is used to segment the text sentences in the training data set into individual words or characters, convert the words or characters into vector representations, and perform position encoding to obtain the model input vector.

[0021] In one embodiment of the present invention, the text sentence describing the topological structure characteristics of the power industrial control network is used as input data, and matching tag data constructed for the text sentence is obtained, including:

[0022] Clean, classify and re-label the network topology, number the power equipment names with duplicate names, and delete the redundant equipment information contained in the network topology;

[0023] The filtered label data are sorted in the order from the entrance to the exit of the power industrial control network to obtain the directed graph matrix structure of the power industrial control network topology, and defined as label data;

[0024] Add a description statement to each directed graph matrix to obtain the input data corresponding to the label data.

[0025] In one embodiment of the present invention, the vectorized training data set is used to construct a generative model based on a large language model, and the trained generative model is used to output a directed graph matrix of the network topology and alarm information, including:

[0026] Constructing a network architecture of a generative model based on a large language model, and defining an input vector and an output vector of the generative model; wherein the generative model includes a pre-trained large model and a decoder network;

[0027] Define the model parameters of the pre-trained large model, and use the back propagation of the gradient to update the parameters, and input the defined input vector into the pre-trained large model, and copy the vector output by the pre-trained large model into two copies. The first copy outputs the alarm information or the vertex array of the directed graph matrix through probability distribution analysis;

[0028] The copied output vector is input into the decoder network to output the square matrix part of the directed graph matrix representing the topology of the power industrial control network;

[0029] The input vectors of all models in the training data set are sequentially input into the pre-trained large model and the decoder network, and the generated model is trained to update the model parameters.

[0030] In one embodiment of the present invention, the model parameters of the pre-trained large model are defined, and the parameters are updated by using the back propagation of the gradient, and the defined input vector is input into the pre-trained large model, and the vector output by the pre-trained large model is copied into two copies, and the first copy outputs the alarm information or the vertex array of the directed graph matrix through probability distribution analysis, including:

[0031] The model parameters before and after the penultimate decoder layer in the pre-trained large model are frozen and do not participate in parameter updates during training. The model parameters after the penultimate decoder layer are non-frozen and updated through gradient back propagation during training.

[0032] The input vector of the model is first fed into the frozen part of the pre-trained large model for dimension conversion, and then fed into the non-frozen part, and the output vector of the large model is copied into two copies;

[0033] Through the linear layer and activation layer of the non-frozen part, the first large model output vector is converted into a probability distribution and mapped into text through the vocabulary. If there is an error description statement in the model input vector, an alarm message is output. If there is no error description statement in the model input vector, the vertex array of the directed graph matrix is ​​output.

[0034] In one embodiment of the present invention, the copied output vector is input into the decoder network to output the square matrix part of the directed graph matrix for representing the topology of the power industrial control network, including:

[0035] The copied output vector of the other large model is input into each layer of the decoder network, and the optimized group query attention mechanism is used to capture the relationship between the model input vectors.

[0036] The copied output vector of the other large model is input into the fully connected layer for mapping, and after the mapping is completed, a reshaping operation is performed to convert it into a two-dimensional matrix that conforms to the shape of the directed graph matrix to output the square matrix part of the directed graph matrix.

[0037] In one embodiment of the present invention, the step of sequentially inputting the input vectors of all models in the training data set into the pre-trained large model and the decoder network, and training the generative model to update the model parameters includes:

[0038] Input the model input vectors into the pre-trained large model in sequence for training. When there is an error description statement in the input model input vector, generate an alarm prompt message describing the error content, calculate the loss function between the model output statement and the label data, and then update the model parameters of the non-frozen part of the pre-trained large model through back propagation;

[0039] When there is no error description statement in the input model input vector, the vertex array of the directed graph matrix is ​​generated, the loss function between the model output vertex array and the vertex array label data is calculated, and then the model parameters of the non-frozen part of the pre-trained large model are updated, and the loss function between the square matrix of the directed graph matrix generated by the decoder network and the square matrix label is calculated, and the network parameters of the decoder network and the model parameters of the non-frozen part of the pre-trained large model are updated.

[0040] In one embodiment of the present invention, the expression of the loss function between the square matrix of the directed graph matrix generated by the calculation decoder network and the square matrix label is:

[0041]

[0042] Where L is the value of the loss function; N 2 is the size of the square matrix row; y′ i,j is the soft label of the square matrix of row i and column j; p i,j It is the probability that the sample in row i and column j of the decoder output in the pre-trained large model is a positive class.

[0043] In one embodiment of the present invention, the step of evaluating the alarm information generation capability and the matrix generation capability according to the output result to obtain the evaluation result includes:

[0044] Based on the power industrial control network topology, input data is obtained and a test set is constructed. The directed graph matrix output by the generative model is tested using the test set. The precision, recall and F1 score of the generative model are calculated respectively. The matrix generation capability of the generative model is evaluated based on the calculation results.

[0045] The warning information output by the generation model is calculated using a text generation evaluation index to evaluate the warning information generation capability of the generation model.

[0046] In one embodiment of the present invention, the square matrix and vertex array output by the generation model are converted into a directed graph form according to the evaluation result, and a virtual simulation scene is constructed by graph node mapping, including:

[0047] Convert the square matrix and vertex array of the directed graph matrix output by the generation model into a graph form, map the graph nodes to physical devices and virtualized components, define the edges as the data flow and control flow between physical devices and virtualized components, and then configure the communication address and entry and exit gateways;

[0048] Extract the elements of the node array in the directed graph in sequence, use the elements to define the device name, and search for virtual device components in the device resource pool according to the device name. Then set the connection relationship between the virtual device components, create and manage the virtual network, and form a virtual simulation scene.

[0049] The LLM-based power industrial control network simulation verification scenario generation method of the embodiment of the present invention can quickly construct complex and realistic attack scenarios by integrating advanced algorithms and large models and utilizing the generation and fitting capabilities of artificial intelligence technology, thereby providing security experts and practitioners with a wealth of virtualized power industrial control network counterattack scenarios.

[0050] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0052] Figure 1 is a flowchart of a method for generating a power industrial control network simulation verification scenario based on LLM according to an embodiment of the present invention;

[0053] Figure 2 is a network schematic diagram of a method for generating a power industrial control network simulation scenario according to an embodiment of the present invention;

[0054] Figure 3 is an example diagram of a network topology re-marking method according to an embodiment of the present invention;

[0055] Figure 4 is a schematic diagram of a method for constructing a training data set according to an embodiment of the present invention;

[0056] Figure 5 It is a modular flow chart of a method for generating a power industrial control network scenario based on a large model according to an embodiment of the present invention;

[0057] Figure 6 is a schematic diagram of a method for optimizing attention of grouped queries in a decoder structure according to an embodiment of the present invention;

[0058] Figure 7 It is a schematic diagram of converting a directed graph matrix into a simulation scenario according to an embodiment of the present invention. DETAILED DESCRIPTION

[0059] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0060] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0061] The following describes a method for generating a simulation verification scenario for an electric power industrial control network based on LLM according to an embodiment of the present invention with reference to the accompanying drawings.

[0062] Figure 1 FIG. 1 is a flow chart of a method for generating a simulation verification scenario of an electric power industrial control network based on LLM according to an embodiment of the present invention. Figure 1 , Figure 2 and Figure 5 As shown, the method includes:

[0063] S1, based on the power industrial control network topology, collects paired input data and label data to generate a training data set, and converts the training data set into a model input vector through preprocessing and vector encoding.

[0064] It can be understood that the present invention is based on the existing power industrial control network topology, collects paired input data and label data, generates a training data set, and converts it into a model input vector through preprocessing and vector encoding.

[0065] Specifically, the input data is a sentence describing the topological structure characteristics of the power industrial control network. After preprocessing such as error correction and filtering, vector encoding is performed to convert the input sentence into a vector representation that can be used by the model as the input of the large language model. After the network topology is preprocessed by methods such as deduplication and renumbering, the existing network topology is converted into a square matrix and vertex array of a directed graph matrix to construct a training data set. At this time, the vertex array and square matrix corresponding to the directed graph matrix are the positive labels corresponding to the prompt sentence in the training phase of the large language model, which are the content that needs to be generated by the large language model.

[0066] In addition, the training data set contains positive and negative samples. When the input is an incorrect description sentence, the label is a warning description sentence; when the input is a correct description sentence, the label is a vertex array and a corresponding square matrix. During the training phase, when there are logical errors in the input text, the target output to be fitted is a warning sentence, and the vertex array and square matrix are not output; when there are no logical errors in the input text, the target function to be fitted is a vertex array and square matrix, and no warning information is output. After the model training is completed, it will have the ability to output target sentences and square matrices that meet the requirements.

[0067] In one embodiment of the present invention, step S1 may specifically include the following steps:

[0068] S11. Obtain a text sentence describing the topological structure characteristics of the power industrial control network as input data, and obtain matching label data constructed for the text sentence to form a training data set.

[0069] Specifically, the input data is a text description (correct or incorrect) constructed based on the network topology structure, including the equipment type, equipment quantity, connection method, relative position, etc. of the safety protection equipment, network equipment, control equipment, monitoring equipment, communication equipment, etc. in the production control area (Area 1 and Area 2) and the information management area (Area 3 and Area 4).

[0070] The label data is obtained through pre-processing operations such as deduplication, renumbering, and conversion based on the network topology of different scenarios. The label part converts the existing network topology into a directed graph matrix (the directed graph matrix includes two parts: the node array and the square matrix). If an error description statement is added, the label part is a warning statement.

[0071] From the topology diagram of the same power network simulation scenario, multiple input and output positive and negative sample data pairs are constructed through expert annotation. The negative sample labels are warning information sentences, and the positive sample labels are array lists and square matrices of directed graph matrices.

[0072] After preprocessing such as deduplication and renumbering, the devices in the topology diagram are converted into a vertex array of a directed graph matrix and a square matrix representing edge connectivity as label data for training the generative model. The vertex array corresponds to the device name of the network topology, and the "0" and "1" in the square matrix represent the connectivity of the devices represented by the corresponding positions of the corresponding vertices. "1" indicates mutual connection, and "0" indicates no connection.

[0073] Specifically, the existing network topology is converted into a directed graph matrix. First, list all nodes in the network and arrange them into an array in a certain order (such as alphabetical order, ID order, etc.). This array will serve as the index of the rows and columns of the directed graph matrix. Create a square matrix based on the length of the node array. The size of the square matrix is ​​equal to the length of the node array. Traverse all edges in the network, and for each directed edge, set the elements in the square matrix to 1 (or other non-zero values).

[0074] In some embodiments of the present invention, S11 obtains a text sentence describing the topological structure characteristics of the power industrial control network as input data, and obtains matching tag data constructed for the text sentence, which may specifically include the following steps:

[0075] S111. Clean, classify and re-label the network topology, number the names of power equipment with duplicate names, and delete redundant equipment information contained in the network topology.

[0076] For example, topology re-marking preprocessing such as Figure 3 shown.

[0077] S112, sorting and combing the filtered label data in the order from the entrance to the exit of the power industrial control network to form a directed graph matrix structure of the power industrial control network topology, and defining it as label data.

[0078] For example, the processed directed graph matrix is ​​shown in Table 1 below.

[0079] Table 1

[0080]

[0081] S113. Add a description statement to each directed graph matrix to form input data corresponding to the label data.

[0082] Specifically, the label data constructed at this time is:

[0083] vertex = {Internet portal, vpn, firewall-3-1, intrusion prevention-3-1,

[0084] Internet access management-3-1, log audit-3-1, switch-3-1,

[0085] Reverse isolation-2-3, forward isolation-2-3, wind power SIS-2-1, ...}

[0086] The square matrix y is:

[0087]

[0088] The method for constructing positive and negative samples of the dataset is as follows Figure 4 shown.

[0089] S12. Use a word segmenter to segment the text sentences in the input data into individual words or characters, convert the words or characters into vector representations, and perform position encoding to obtain a model input vector.

[0090] S2, uses the vectorized training data set to build a generative model based on the large language model, and obtains the output results of the directed graph matrix containing the network topology and the alarm information through the trained generative model.

[0091] It can be understood that, in S2 of the embodiment of the present invention, using the vectorized training data set to build a generation model based on a large language model, and then outputting a directed graph matrix of the network topology and alarm information through the trained generation model, specifically may include the following steps:

[0092] S21. Based on the large language model, build the network architecture of the generative model and define the model input and model output. The generative model includes a pre-trained large model and a decoder network.

[0093] S22. Define the model parameters of the pre-trained large model, update the parameters using the back propagation of the gradient, input the model input vector into the pre-trained large model, copy the vector output by the pre-trained large model into two copies, one of which is analyzed through the probability distribution of the vector to output an alarm message or a vertex array of a directed graph matrix.

[0094] In one embodiment of the present invention, S22 defines the model parameters of the pre-trained large model, updates the parameters using the back propagation of the gradient, inputs the model input vector into the pre-trained large model, copies the vector output by the pre-trained large model into two copies, one of which is analyzed by the probability distribution of the vector, and outputs the alarm information or the vertex array of the directed graph matrix, including the following steps:

[0095] S221, the model parameters of the penultimate decoder layer and before in the pre-trained large model are frozen and do not participate in the parameter update during the training process; the model parameters after the penultimate decoder layer are treated as non-frozen parts, and the parameters are updated through gradient back propagation during the training process;

[0096] S222, the model input vector first enters the frozen part of the pre-trained large model, performs dimension conversion, and then enters the non-frozen part, and copies the large model output vector into two copies;

[0097] S223. Convert one of the large model output vectors into a probability distribution through the linear layer and activation layer of the non-frozen part, and map it into text through the vocabulary. If there is an error description statement in the model input vector, output a warning message. If there is no error description statement in the model input vector, output the vertex array of the directed graph matrix.

[0098] S23. Input the copied output vector of the other large model into the decoder network, and after being decoded by the decoder network, output the square matrix part of the directed graph matrix used to represent the topology of the power industrial control network.

[0099] Specifically, the decoder network is mainly composed of a Transformer Decoder and a fully connected layer. In the decoder network, Values ​​and Keys are grouped separately to optimize the group query attention. The specific method is to first group the query vectors, each group of query matrices uses a common Key, and then group the Keys, each group of Keys uses a common Value, further reducing the number of parameters and reducing the computing power requirements of the device. The improved group query attention structure is as follows: Figure 6 shown.

[0100] In one embodiment of the present invention, S23 inputs the copied output vector of the other large model into the decoder network, and after being decoded by the decoder network, outputs the square matrix part of the directed graph matrix for representing the topology of the power industrial control network, including the following steps:

[0101] S231, input the large model output vector into each layer of the decoder in the decoder network, and use the optimized group query attention mechanism to capture the relationship between the model input vectors;

[0102] S232, input the large model output vector to the fully connected layer for further mapping, and after the mapping is completed, perform a reshaping operation to convert it into a two-dimensional matrix that conforms to the shape of the directed graph matrix, and output the square matrix part of the directed graph matrix. Specifically, when there is no error description statement in the model input vector, the network output includes two parts, namely, vertex array vertex = {v1, v2, v3...v n}; and square matrix Abbreviated as where y i,j Is 0 or 1.

[0103] Specifically, the grouped query attention mechanism generates a query vector at each layer of the decoder and assigns the query vector to each group, or allows the query vector to query all groups at the same time. Within each group, an attention mechanism (such as dot product attention, multi-head attention, etc.) is used to calculate the similarity or correlation between the query vector and the input vector within the group, and the attention outputs of all groups are merged to form the final input of the current layer of the decoder.

[0104] S24. Input all model input vectors in the training data set into the pre-trained large model and the decoder network in sequence, train the generated model, and update the model parameters.

[0105] In one embodiment of the present invention, S24 sequentially inputs all model input vectors in the training data set into the pre-trained large model and the decoder network to train the generated model, and updating the model parameters includes the following steps:

[0106] S241, inputting the model input vectors into the pre-trained large model in sequence for training. During the training process, when there is an error description statement in the input model input vector, an alarm prompt message describing the error content is generated, and the loss function between the model output statement and the label data is calculated, and then the model parameters of the non-frozen part of the pre-trained large model are updated through back propagation;

[0107] S242. When there is no error description statement in the input model input vector, a vertex array of the directed graph matrix is ​​generated, and the loss function between the model output vertex array and the vertex array label data is calculated. Then, the model parameters of the non-frozen part of the pre-trained large model are updated, and the loss function between the square matrix of the directed graph matrix generated by the decoder network and the square matrix label is calculated. The network parameters of the decoder network and the model parameters of the non-frozen part of the pre-trained large model are updated.

[0108] The label data is converted on the existing public network topology map, using a directed graph matrix consisting of vertex arrays and square matrices to represent the original network topology. The "0" and "1" at different positions in the square matrix are labels representing the connection status. In order to avoid overconfidence in the model during the training and generation process, the probability distribution of the real label is "smoothed", that is,

[0109] where y i, j is the true label of the square matrix of the i-th row and j-th column, which is 0 or 1, and ε is a manually set hyperparameter. When calculating the loss function between the square matrix label, the binary cross entropy function is used to calculate the loss and update the parameters of the non-frozen part and the decoder part. The formula for calculating the loss function between the square matrix of the directed graph matrix generated by the decoder network and the square matrix label is as follows:

[0110]

[0111] Where L is the value of the loss function; N2 is the size of the matrix; y′ i,j is the soft label of the square matrix of row i and column j; p i,j It is the probability that the sample in row i and column j of the decoder output in the pre-trained large model is a positive class.

[0112] S3, evaluating the generation capability of alarm information and the matrix generation capability according to the output result to obtain an evaluation result.

[0113] In one embodiment of the present invention, evaluating the alarm information generation and matrix generation capabilities of the generation model according to the output results includes the following steps:

[0114] S31. Based on the existing power industrial control network topology, obtain input data to build a test set, use the test set to test the directed graph matrix output by the generation model, calculate the precision, recall rate and F1 score of the generation model respectively, and evaluate the matrix generation ability of the generation model based on the calculation results.

[0115] Specifically, the precision, recall, and F1 score of the generated matrix are calculated. The calculation formula is as follows:

[0116]

[0117] Among them, TP is a true positive example, FN is a false negative example, and FP is a false positive example.

[0118] S32. Calculate the warning information output by the generation model using the text generation evaluation index to evaluate the warning information generation capability of the generation model.

[0119] Specifically, the Rougle-L indicator is used to evaluate the quality of the generated alarm information. The calculation formula is as follows:

[0120]

[0121] Where LCS is the longest common subsequence, β is a hyperparameter. Rlcs indicates the proportion of LCS in the generated warning information that matches the reference text. Plcs indicates the proportion of LCS in the generated warning information that matches the reference text, and Flcs is the harmonic mean of Rlcs and Plcs.

[0122] S4, converts the square matrix and vertex array output by the generated model into a directed graph form according to the evaluation results, and constructs a virtual simulation scene through graph node mapping.

[0123] It can be understood that based on the principles of graph theory, the square matrix and vertex array output by the generated model are converted into a directed graph form, and the virtual simulation scene is constructed through graph node mapping, and then secondary packaging is performed.

[0124] like Figure 7 As shown in FIG. 1 , it is a schematic diagram of converting a directed graph matrix into a simulation scenario, which specifically includes the following steps:

[0125] S41, converting the square matrix and vertex array of the directed graph matrix output by the generation model into a graph form, mapping the graph nodes to physical devices and virtualized components, defining the edges as data flows and control flows between physical devices and virtualized components, and then configuring the communication addresses and entry and exit gateways;

[0126] S42, extracting the elements of the node array in the directed graph in sequence, using the elements to define the device name, searching for the virtual device component in the device resource pool according to the device name, and then setting the connection relationship between the virtual device components, creating and managing the virtual network, and forming a virtual simulation scene.

[0127] In addition, the scenario simulation is created based on the cloud computing management platform. The elements in the node array are first extracted in sequence. The elements contain the device name and device number. The corresponding virtual device component is searched in the pre-uploaded device resource pool according to the device name, and the extracted element is used as the device name; then the connection between the devices is determined according to the numbers "1" and "0" in the square matrix. The network between the connected devices is created and managed by nova-network (virtual machine network), and the host and the external network can access each other.

[0128] According to the LLM-based power industrial control network simulation verification scenario generation method of the present invention, the power network topology is used to build its own data set to train the generation model, and the input and output data training data that conforms to a specific format is added and parsed from the existing network topology, including the production control area and the information management area equipment information and connection mode, etc.; the generation model outputs the vertices and square matrix of the directed graph matrix containing the network topology structure information; and the matrix is ​​converted into a simulation scenario through the scenario generation module. The present invention can quickly build complex and real attack scenarios by integrating advanced algorithms and large models, and using the generation and fitting capabilities of artificial intelligence technology, and provide security experts and practitioners with a wealth of virtualized power industrial control network counterattack scenarios; effectively using pre-trained large models to develop new generation algorithms can reduce development time and reduce the requirements for graphics card floating-point computing capabilities during actual use; at the same time, a large number of power industrial control network simulation scenarios that meet compliance requirements can be automatically built, and the constructed scenarios are not constrained by human habits and personal cognitive limitations, which is convenient for researchers to study the security of different network topologies.

[0129] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0130] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In one embodiment of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

Claims

1. A method for generating a simulation verification scenario for an electric power industrial control network based on LLM, characterized in that: include: Based on the power industrial control network topology, paired input data and label data are collected to generate a training data set, and the training data set is converted into a model input vector through preprocessing and vector encoding; The vectorized training data set is used to build a generative model based on a large language model, and the trained generative model is used to obtain the output results of a directed graph matrix containing network topology and alarm information; Evaluate the generation capability of the alarm information and the matrix generation capability according to the output result to obtain an evaluation result; According to the evaluation results, the square matrix and vertex array output by the generated model are converted into a directed graph form, and a virtual simulation scene is constructed through graph node mapping.

2. The method according to claim 1, characterized in that Based on the power industrial control network topology, paired input data and label data are collected to generate a training data set, and the training data set is converted into a model input vector through preprocessing and vector encoding, including: Taking a text sentence describing the topological structure characteristics of the power industrial control network as input data, and obtaining matching label data constructed for the text sentence to obtain a training data set; The word segmenter is used to segment the text sentences in the training data set into individual words or characters, convert the words or characters into vector representations, and perform position encoding to obtain the model input vector.

3. The method according to claim 2, characterized in that The text sentence describing the topological structure characteristics of the power industrial control network is used as input data, and matching label data constructed for the text sentence is obtained, including: Clean, classify and re-label the network topology, number the power equipment names with duplicate names, and delete the redundant equipment information contained in the network topology; The filtered label data are sorted in the order from the entrance to the exit of the power industrial control network to obtain the directed graph matrix structure of the power industrial control network topology, and defined as label data; Add a description statement to each directed graph matrix to obtain the input data corresponding to the label data.

4. The method according to claim 1, characterized in that The vectorized training data set is used to construct a generative model based on a large language model, and the trained generative model is used to output a directed graph matrix of the network topology and alarm information, including: Constructing a network architecture of a generative model based on a large language model, and defining an input vector and an output vector of the generative model; wherein the generative model includes a pre-trained large model and a decoder network; Define the model parameters of the pre-trained large model, and use the back propagation of the gradient to update the parameters, and input the defined input vector into the pre-trained large model, and copy the vector output by the pre-trained large model into two copies. The first copy outputs the alarm information or the vertex array of the directed graph matrix through probability distribution analysis; The copied output vector is input into the decoder network to output the square matrix part of the directed graph matrix representing the topology of the power industrial control network; The input vectors of all models in the training data set are sequentially input into the pre-trained large model and the decoder network, and the generated model is trained to update the model parameters.

5. The method according to claim 4, characterized in that Define the model parameters of the pre-trained large model, and use the back propagation of the gradient to update the parameters, and input the defined input vector into the pre-trained large model, and copy the vector output by the pre-trained large model into two copies. The first copy outputs the alarm information or the vertex array of the directed graph matrix through probability distribution analysis, including: The model parameters before and after the penultimate decoder layer in the pre-trained large model are frozen and do not participate in parameter updates during training. The model parameters after the penultimate decoder layer are non-frozen and updated through gradient back propagation during training. The input vector of the model is first fed into the frozen part of the pre-trained large model for dimension conversion, and then fed into the non-frozen part, and the output vector of the large model is copied into two copies; Through the linear layer and activation layer of the non-frozen part, the first large model output vector is converted into a probability distribution and mapped into text through the vocabulary. If there is an error description statement in the model input vector, an alarm message is output. If there is no error description statement in the model input vector, the vertex array of the directed graph matrix is ​​output.

6. The method according to claim 4, characterized in that The copied output vector is input into the decoder network to output the square matrix part of the directed graph matrix representing the topology of the power industrial control network, including: The copied output vector of the other large model is input into each layer of the decoder network, and the optimized group query attention mechanism is used to capture the relationship between the model input vectors. The copied output vector of the other large model is input into the fully connected layer for mapping, and after the mapping is completed, a reshaping operation is performed to convert it into a two-dimensional matrix that conforms to the shape of the directed graph matrix to output the square matrix part of the directed graph matrix.

7. The method according to claim 4, characterized in that The step of sequentially inputting the input vectors of all models in the training data set into the pre-trained large model and the decoder network, and training the generated model to update the model parameters, includes: Input the model input vectors into the pre-trained large model in sequence for training. When there is an error description statement in the input model input vector, generate an alarm prompt message describing the error content, calculate the loss function between the model output statement and the label data, and then update the model parameters of the non-frozen part of the pre-trained large model through back propagation; When there is no error description statement in the input model input vector, the vertex array of the directed graph matrix is ​​generated, the loss function between the model output vertex array and the vertex array label data is calculated, and then the model parameters of the non-frozen part of the pre-trained large model are updated, and the loss function between the square matrix of the directed graph matrix generated by the decoder network and the square matrix label is calculated, and the network parameters of the decoder network and the model parameters of the non-frozen part of the pre-trained large model are updated.

8. The method according to claim 7, characterized in that The expression of the loss function between the square matrix of the directed graph matrix generated by the computation decoder network and the square matrix label is: Where L is the value of the loss function; N 2 is the size of the square matrix row; y′ i,j is the soft label of the square matrix of row i and column j; p i,j It is the probability that the sample in row i and column j of the decoder output in the pre-trained large model is a positive class.

9. The method according to claim 1, characterized in that: The step of evaluating the generation capability of the alarm information and the matrix generation capability according to the output result to obtain the evaluation result includes: Based on the power industrial control network topology, input data is obtained and a test set is constructed. The directed graph matrix output by the generative model is tested using the test set. The precision, recall and F1 score of the generative model are calculated respectively. The matrix generation capability of the generative model is evaluated based on the calculation results. The warning information output by the generation model is calculated using a text generation evaluation index to evaluate the warning information generation capability of the generation model.

10. The method according to claim 1, characterized in that According to the evaluation results, the square matrix and vertex array output by the generated model are converted into a directed graph form, and a virtual simulation scene is constructed by graph node mapping, including: Convert the square matrix and vertex array of the directed graph matrix output by the generation model into a graph form, map the graph nodes to physical devices and virtualized components, define the edges as the data flow and control flow between physical devices and virtualized components, and then configure the communication address and entry and exit gateways; Extract the elements of the node array in the directed graph in sequence, use the elements to define the device name, and search for virtual device components in the device resource pool according to the device name. Then set the connection relationship between the virtual device components, create and manage the virtual network, and form a virtual simulation scene.