Malicious sample analysis method based on knowledge enhancement neural network intelligent model
By introducing knowledge-enhanced BP neural networks into the large language model system, combining knowledge graphs and two-layer BP neural networks, the security threat of the large language model system is solved, and the efficiency and robustness of malicious sample detection are improved.
Patent Information
- Application Number
- CN202510906655.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
When large language model systems face threats such as data poisoning attacks, counter-sample attacks, model theft attacks and backdoor attacks, there is a risk of failure in security and reliability, and it is difficult for existing technologies to effectively identify malicious samples.
By introducing improved BP neural networks with knowledge enhancement, the original features are integrated into the intelligent model training process in the form of a knowledge graph, and the prior information provided by the knowledge graph is used to enhance the model's understanding ability and text representation ability. Combining semantic features, syntactic features and anomaly detection indicators, a two-layer BP neural network model is built for malicious sample analysis.
This improves the model's analysis ability of malicious samples, reduces the demand for training sample size, improves detection efficiency and robustness, and reduces calculation time.
Smart Images

Figure CN120408621A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence security, especially under the large language model system, and relates to a method for analyzing malicious samples of an intelligent model based on a knowledge-enhanced neural network. Background Art
[0002] In recent years, the remarkable development of large models, such as BERT, SORA, and GPT, has played a key role in a series of applications, showing great potential in downstream tasks from text summarization to code generation, visual question answering, multi-modal machine translation, etc. By performing large-scale pre-training on a large amount of unlabeled data and introducing key technologies such as instruction fine-tuning and human alignment, large models can complete various tasks at a high level according to the characteristics of specific application scenarios. However, the application of large models increases the risks of security attacks and privacy leaks. The "pre-training and fine-tuning" stage of artificial intelligence models may become a new security attack surface and is vulnerable to threats such as data poisoning attacks, adversarial sample attacks, model stealing attacks, backdoor attacks, and instruction attacks, which makes the business security and reliability based on large models face the risk of failure. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for analyzing malicious samples of an intelligent model based on a knowledge-enhanced neural network. Through the knowledge-enhanced method, on the one hand, the improved neural network intelligent model can effectively identify maliciously constructed samples, and on the other hand, by using the prior information provided by the knowledge graph, more meaningful initial features are provided through knowledge embedding, reducing the model's dependence on a large amount of training sample data and making the detection optimization process more efficient. To solve the technical problem of data poisoning attacks existing in the large language model system in the prior art.
[0004] The technical solution adopted by the present invention to solve the above technical problem is to introduce an improved BP neural network with knowledge enhancement, integrate the original features into the intelligent model training process in the form of a knowledge graph, obtain stronger model understanding ability and text representation ability, enhance the robustness for malicious sample analysis, reduce the calculation time, and improve the detection efficiency, including the following steps:
[0005] To solve the above technical problem, the specific technical solution of the present invention is as follows:
[0006] A method for analyzing malicious samples of an intelligent model based on a knowledge-enhanced neural network, the method comprising the following steps:
[0007] Step S1: Collect a large number of data samples, including a set of known malicious samples and a set of normal samples, use the pre-trained language model BERT to extract semantic features from the collected data samples, and select [CLS] the marked corresponding vectors as semantic feature vectors;
[0008] Step S2: Use dependency syntactic analysis to extract syntactic features from the collected data samples, calculate the frequency distribution of dependency relation types, form a row statistical feature vector with the frequency values, and obtain a syntactic feature vector after normalization;
[0009] Step S3: Construct an anomaly detection index vector by calculating the Euclidean distance between the feature vector of the data sample and the feature vectors of each normal sample set;
[0010] Step S4: Introduce an external knowledge graph to enhance the semantic feature vector extracted in Step S1; query entities and relationships related to the sample features from the knowledge graph, generate a knowledge-enhanced feature vector based on the query results, and concatenate the semantic feature vector, syntactic feature vector, anomaly detection index vector, and knowledge-enhanced feature vector to generate a comprehensive feature vector;
[0011] Step S5: Construct a two-layer BP neural network model. The input of the two-layer BP neural network model is the normalized comprehensive feature vector, which includes semantic features, syntactic features, anomaly detection indexes, and knowledge enhancement features. The number of neurons in each layer is adjusted according to the input feature dimension. The activation function uses ReLU, and the output layer uses the softmax function to predict the probability that the sample is malicious. Use the cross-entropy loss function and the Adam optimization algorithm to optimize the network parameters. The pre-trained language model BERT and the two-layer BP neural network model form the final knowledge-enhanced malicious sample detection model;
[0012] Step S6: Collect the text to be detected and input it into the knowledge-enhanced malicious sample detection model to output the malicious detection result.
[0013] Further, Step S1 includes the following steps:
[0014] Step S11: Perform word segmentation preprocessing on the data sample . The data sample is expressed as: , where is the th token of the data sample , and represents the number of tokens in the data sample . Concatenate the start and end of the data sample with the tokens [CLS] and [SEP] respectively to obtain the text . [CLS] represents the start of the whole sentence or text sequence, [SEP] Indicates the end of the current input sequence; the text is input into the pre-trained language model BERT;
[0015] Step S12: The pre-trained language model BERT will output the hidden state vectors corresponding to each token. The output form of the BERT model is expressed as: H= [CLS], h 0 ,⋯, h i ,⋯, h n-1 ,[SEP] ; where is the output sequence of the model, is the hidden state vector corresponding to each token; for the entire text , select the vector corresponding to the token [CLS] tag as the semantic feature vector.
[0016] Further, the step S2 specifically includes the following sub-steps:
[0017] Step S21: First, perform dependency syntactic analysis on the data sample to obtain the dependency relationship between each token and other tokens;
[0018] Step S22: Construct a dependency relationship matrix to represent the dependency relationship between tokens. The element in the dependency relationship matrix represents the dependency relationship type when the th token is the subordinate word and the th token is the governing word;
[0019] Step S23: Extract statistical features according to the dependency relationship matrix. By row statistical method, calculate the global frequency distribution of each dependency relationship type, and form these frequency values into a row statistical feature vector;
[0020] Step S24: Normalize the row statistical feature vector to the range [0, 1] to obtain the normalized row statistical feature vector.
[0021] Further, the step S3 specifically includes the following sub-steps:
[0022] Step S31: The semantic feature vector and the syntactic feature vector form a feature vector. The feature vector of the data sample is expressed as ; calculate the feature vector of the semantic feature vector and the syntactic feature vector for the normal sample set, which is expressed as ; calculate the Euclidean distance between the feature vector of the data sample and the feature vector of each normal sample set;
[0023] Step S32: Calculate the data sample The average distance and standard deviation of the Euclidean distances between the feature vectors of the data sample and the feature vectors of all normal sample sets and standard deviation to construct an anomaly detection index vector.
[0024] Furthermore, the step S4 specifically includes the following sub-steps:
[0025] Step S41: Match the semantic feature vector with the knowledge graph and query the related entities and relationships in the knowledge graph;
[0026] Step S42: Use knowledge graph embedding to map the relationships and entities in the knowledge graph to the vector space, and splice the entity vector and the relationship vector to obtain a knowledge-enhanced feature vector;
[0027] Step S43: Splice the knowledge-enhanced feature vector with the semantic feature vector, syntactic feature vector, and anomaly detection index vector into a comprehensive feature vector, and perform normalization processing on the comprehensive feature vector to obtain the normalized comprehensive feature vector.
[0028] Furthermore, the step S5 specifically includes the following sub-steps:
[0029] Step S51: Construct a two-layer BP neural network model, and use the normalized comprehensive feature vector as the input of the BP neural network;
[0030] Step S52: Use the ReLU function as the activation function of the hidden layer;
[0031] Step S53: Use the softmax function to output the probability of malicious samples;
[0032] Step S54: According to the output result, use the weighted cross-entropy loss function and the Adam optimizer to optimize the model.
[0033] Compared with the prior art, the present invention has the following beneficial technical effects:
[0034] (1) Through the knowledge enhancement and anomaly detection index extraction methods, the present invention provides more meaningful initial features for the model, improving the model's analysis ability for malicious samples.
[0035] (2) By constructing an external knowledge graph, the present invention obtains richer feature information, reduces the model's demand for the amount of training samples, and further improves the detection efficiency of malicious samples.
[0036] (3) By introducing an external knowledge graph to enrich the feature dimensions, the present invention provides more comprehensive information support for malicious sample detection. Introducing the knowledge graph in the pre-training stage of the large model enhances the prior knowledge of the model, enables the model to converge faster, and further improves the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a schematic diagram of the overall framework topology structure of a malicious sample detection method based on a knowledge-enhanced BP neural network of the present invention.
[0039] Figure 2 It is a topological schematic diagram of the pre-trained language model BERT of the present invention.
[0040] Figure 3 It is a topological schematic diagram of the BP neural network model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0042] The present invention proposes a malicious sample analysis method based on a knowledge-enhanced neural network intelligent model, which is applied to a large language model system. By introducing an improved BP neural network with knowledge enhancement, the original features are incorporated into the intelligent model training process in the form of a knowledge graph, obtaining stronger model understanding ability and text representation ability, enhancing the robustness for malicious sample analysis, reducing the calculation time, and improving the detection efficiency. As Figure 1 shown, the method includes the following steps:
[0043] Step S1: Collect a large number of data samples, including a set of known malicious samples and a set of normal samples, use the pre-trained language model BERT to extract semantic features from the collected data samples, and select [CLS] the marked corresponding vectors as semantic feature vectors.
[0044] Specifically, as Figure 2As shown, step S1 includes the following steps:
[0045] Step S11: Data samples Perform word segmentation preprocessing, data sample Expressed as: ,in For data samples No. word, Represents data samples The number of words in the data sample The head and tail are spliced together [CLS] and [SEP] After that, get the text , [CLS] Indicates the beginning of a whole sentence or text sequence, [SEP] Indicates the end of the current input sequence; the text Input into the pre-trained language model BERT.
[0046] Step S12: The pre-trained BERT language model outputs the hidden state vector corresponding to each word. Positional encoding information, sentence segmentation information, and token information are embedded in the position, segment, and tag embeddings, respectively, and the three are summed to form the final input embedding. The input embeddings are fused into word embeddings and fed into a feedforward neural network module with a self-attention mechanism, a fully connected layer, and ReLU activation functions, as well as residual connections and layer normalization. After this processing, the model's hidden layer representation is output through the fully connected layer. The output of the BERT model is represented as: H= [CLS], h 0 ,⋯, h i ,⋯, h n-1 ,[SEP] ; Where H is the output sequence of the model, The hidden state vector corresponding to each word. , you can choose a word [CLS] The vector corresponding to the tag (the tag is usually used to represent the overall semantics of the sentence) is used as the semantic feature vector , the hidden state dimension of the pre-trained language model BERT output is , then the semantic feature vector , represents the set of real numbers.
[0047] Step S2: Use dependency syntactic analysis to extract syntactic features from the collected data samples, calculate the frequency distribution of dependency relationship types, form a row statistical feature vector from the frequency values, and obtain a syntactic feature vector after normalization.
[0048] Specifically, step S2 includes the following steps:
[0049] Step S21: First, perform dependency syntactic analysis on the data sample to obtain the dependency relationships between each token and other tokens, such as subject-predicate relationships, verb-object relationships, etc.
[0050] Step S22: Construct a dependency relationship matrix to represent the dependency relationships between tokens. The elements in the dependency relationship matrix represent the type of dependency relationship between the th token as a subordinate token and the th token as a governing token. For example, the subject-predicate relationship type is 1, the verb-object relationship type is 2, etc. If there is no dependency relationship between the th token and the th token, then is 0. is 0.
[0051] Step S23: Extract statistical features based on the dependency relationship matrix. By row statistical methods, calculate the global frequency distribution of each dependency relationship type, and form a row statistical feature vector with a fixed length, where is the number of dependency relationship types. is the number of dependency relationship types.
[0052] The row statistical formula for the th dependency relationship type is expressed as follows:
[0053]
[0054] where m is the number of dependency relationship types, δ is the indicator function, .
[0055] Furthermore, the global frequency distribution of each dependency relationship type is obtained by counting the number of occurrences of each dependency relationship type in all sentences in the corpus and calculating its proportion.
[0056] The row statistical feature vector is expressed as: F syn ' =[ F syn ' 1 ,…, F syn ' k ,…, F syn ' m ] .
[0057] Step S24: Take the row statistical feature vector Normalize to the range of [0, 1]. The normalization uses min-max normalization, and the normalization method is as follows:
[0058]
[0059] Among them, represents the row statistical feature vector after normalization and serves as the data sample of the syntactic feature vector, with the dimension of .
[0060] Step S3: Construct an anomaly detection index vector by calculating the Euclidean distance between the feature vector of the data sample and the feature vectors of each normal sample set.
[0061] Specifically, Step S3 includes the following steps:
[0062] Step S31: The semantic feature vector and the syntactic feature vector form a feature vector. The feature vector of the data sample is represented as . Calculate the feature vector formed by the semantic feature vector and the syntactic feature vector for the normal sample set, which is represented as . Calculate the Euclidean distance between the feature vector of the data sample and the feature vectors of each normal sample set. The calculation method is as follows:
[0063]
[0064] Among them, represents the Euclidean distance between the data sample and the th normal sample, represents the dimension of the feature vector of the data sample , represents the number of elements in the normal sample set, represents the th component of the feature vector of the data sample , represents the th sample in the normal sample set th component.
[0065] Step S32: Calculate the average distance and the standard deviation of the Euclidean distances between the feature vector of the data sample and the feature vectors of all normal sample sets, and construct an anomaly detection index vector.
[0066] The formula for calculating the average distance is:
[0067]
[0068] The formula for calculating the standard deviation is:
[0069]
[0070] The formula for calculating the anomaly detection metric value is:
[0071] N
[0072] where, represents the th anomaly detection metric value; the dimension of the anomaly detection metric vector is .
[0073] The anomaly detection metric vector is expressed as: F anom =[ F anom 1 ,…, F anom i ,…, F anom N ] .
[0074] Step S4: Introduce an external knowledge graph to enhance the semantic feature vectors extracted in step S1. Query the entities and relationships related to the sample features from the knowledge graph, generate knowledge-enhanced feature vectors based on the query results, and concatenate the semantic feature vectors, syntactic feature vectors, anomaly detection metric vectors, and knowledge-enhanced feature vectors to generate comprehensive feature vectors.
[0075] Specifically, step S4 includes the following steps:
[0076] Step S41: The knowledge graph contains a large number of entities and relationships. By matching the semantic feature vectors with the knowledge graph, hidden attack intentions or malicious patterns can be discovered. Query the entities and relationships (such as attack types, behavior patterns) related to it in the knowledge graph according to the semantic feature vectors.
[0077] Step S42: Use TransE (Translating Embedding for Modeling Multi-relational Data, knowledge graph embedding) to map the relationships and entities in the knowledge graph to the vector space, and concatenate the entity vectors and relationship vectors to obtain the knowledge-enhanced feature vector .
[0078] In TransE, the entity And relation are all mapped into a dimensional vector space. For any triple (head h, relation r, tail t), the following translation relation is satisfied:
[0079]
[0080] The knowledge graph consists of triples:
[0081]
[0082] where represents the knowledge graph, is the head entity, is the relation, is the tail entity, and ε represents the set of all possible entities and relations.
[0083] To minimize the distance difference between positive samples and negative samples, negative sampling is used as the loss function.
[0084] The negative samples are represented as where represents the negative sample head entity, represents the negative sample tail entity.
[0085] The loss function of TransE is expressed as:
[0086]
[0087] where γ is a hyperparameter, represents the general Euclidean distance, and the calculation method is as follows:
[0088]
[0089] where represents the L2 norm;
[0090] [∙] + represents the positive function, which is defined as [x] + =max(0,x) where is the maximum value.
[0091] The gradient descent optimizer is used to adjust the entity vectors and relation vectors to minimize the loss.
[0092] The knowledge-enhanced features are concatenated with the entity vectors and relation vectors after being weighted and averaged using an aggregation function:
[0093]
[0094]
[0095] Among them, represents the knowledge-enhanced feature vector, represents the vector representation of the first entity, represents the vector representation of the first relationship, represents the number of entities, represents the number of relationships, represents the weight of the i-th entity, represents the weight of the j-th relationship, represents the vector representation of the i-th entity, represents the vector representation of the j-th relationship. The dimension of the knowledge-enhanced feature vector is .
[0096] Step S43: Concatenate the knowledge-enhanced feature vector with the semantic feature vector, syntactic feature vector, and anomaly detection index vector to form a comprehensive feature vector, and perform normalization processing on the comprehensive feature vector to obtain the normalized comprehensive feature vector.
[0097] Concatenate the semantic feature vector, syntactic feature vector, anomaly detection index vector with the knowledge-enhanced feature vector to generate a comprehensive feature vector , and the concatenation method is as follows:
[0098]
[0099] After feature fusion, globally normalize the comprehensive feature vector. To balance the influence of each part of the feature on the model input, unified normalization can be performed:
[0100]
[0101] Among them, represents the normalized comprehensive feature vector, and are the mean and standard deviation of the comprehensive feature respectively.
[0102] Step S5: Construct a two-layer BP neural network model. The input of the two-layer BP neural network model is the normalized comprehensive feature vector , which includes semantic features, syntactic features, anomaly detection indicators, and knowledge-enhanced features. The number of neurons in each layer is adjusted according to the input feature dimension. The activation function uses ReLU, and the output layer uses the softmax function to predict the probability that the sample is malicious. The cross-entropy loss function and the Adam optimization algorithm are used to optimize the network parameters. The pre-trained language model BERT and the two-layer BP neural network model constitute the final knowledge-enhanced malicious sample detection model.
[0103] Specifically, as Figure 3 shown, step S5 includes the following steps:
[0104] Step S51: Construct a two-layer BP neural network model, and the normalized comprehensive feature vector is used as the input of the BP neural network, and its dimension is:
[0105]
[0106] where represents the dimension of the comprehensive feature vector, represents the dimension of the semantic feature vector, represents the dimension of the syntactic feature vector, represents the dimension of the anomaly detection index vector, represents the dimension of the knowledge enhancement feature vector.
[0107] Step S52: Use the ReLU function as the activation function of the hidden layer, and the calculation formula of each hidden layer is:
[0108]
[0109] where represents the output of the l-th hidden layer, represents the ReLU activation function, is the weight matrix of the l-th layer, represents the output of the l-th hidden layer, that is, the feature of the previous layer as the input of the current layer, is the bias of the l-th layer.
[0110] Step S53: Use the softmax function to output the probability of malicious samples:
[0111]
[0112] where represents the softmax function, represents the weight matrix of the output layer, represents the output of the last hidden layer, represents the bias vector of the output layer.
[0113] When P(malicious) ≥ 50%, the sample is determined to be malicious.
[0114] Step S54: According to the output result, use the weighted cross-entropy loss function and the Adam optimizer to optimize the model.
[0115] To prevent the problem of class imbalance (unequal ratio of malicious samples to normal samples), weighted cross-entropy loss is adopted:
[0116]
[0117] Among them, represents the total number of samples participating in the calculation of the loss function, and are the weights of positive and negative samples, which can be dynamically adjusted according to the class ratio, represents the true label of the i-th sample, represents the probability that the model predicts the i-th sample as a malicious sample.
[0118] Backpropagation and parameter update, calculate the gradient:
[0119]
[0120] Among them, represents the gradient, and θt is the model parameter.
[0121] Use the Adam optimizer formula to update the parameters:
[0122]
[0123]
[0124]
[0125]
[0126]
[0127] Among them, represents the first moment estimate at time , represents the first moment estimate at the previous time , is the gradient, =0.9 is the first momentum decay factor, represents the second moment estimate at time t, represents the second moment estimate at the previous time , =0.999 is the second momentum decay factor, represents the value after bias correction of the first moment estimate, represents the value after bias correction of the second moment estimate, represents the parameter vector at time t, represents the updated parameter vector, α = 1×10-4 is the learning rate, and ϵ is a smoothing term to prevent the denominator from being zero.
[0128] Step S6: Collect the text to be detected and input it into the knowledge-enhanced malicious sample detection model, and output the malicious detection result.
[0129] After step S5, it also includes model testing and evaluation: Calculate the accuracy rate, recall rate, and F1 score using the validation set to evaluate the model performance.
[0130] The performance metric uses the accuracy rate as:
[0131]
[0132] The recall rate is:
[0133]
[0134] Among them, represents the number of samples that the model correctly predicts as normal text, represents the number of samples that the model correctly predicts as malicious samples, represents the number of samples that the model wrongly predicts as normal text, represents the number of samples that the model wrongly predicts as malicious samples.
[0135] The F1-score is:
[0136]
[0137] The present invention also proposes a knowledge-enhanced BP neural network malicious sample detection system, based on the malicious sample detection method of the above-mentioned knowledge-enhanced BP neural network model, including: a feature extraction module, a knowledge-enhanced feature extraction module, a knowledge-enhanced malicious sample detection model, and a model construction and training module;
[0138] The knowledge-enhanced malicious detection model includes a pre-trained language model BERT and a BP neural network model;
[0139] The feature extraction module is used to extract the semantic features, syntactic features, and anomaly detection metrics of the text;
[0140] The knowledge-enhanced feature extraction module is used to further enhance the text features, obtain the knowledge-enhanced vector through the knowledge graph, and embed it into the text features to obtain the comprehensive feature vector;
[0141] The pre-trained language model BERT is used to extract the semantic feature vector of the text;
[0142] The model construction and training module is used to construct a two-layer BP neural network model. The input of the model is a comprehensive feature vector, which includes semantic features, syntactic features, anomaly detection metrics, and knowledge enhancement features. The number of neurons in each layer is adjusted according to the input feature dimension. The ReLU activation function is used, and the softmax function is used in the output layer to predict the probability that the sample is malicious. The cross-entropy loss function and the Adam optimization algorithm are adopted to optimize the network parameters and obtain the final knowledge-enhanced malicious sample detection model;
[0143] The knowledge-enhanced malicious sample detection model is used to input the text to be detected and output the malicious detection result.
[0144] The present invention makes full use of the advantages of knowledge enhancement. Through the improved BP neural network intelligent model, it can not only effectively analyze maliciously constructed samples, but also reduce the model's demand for the amount of training samples. By introducing an external knowledge graph, the features can be enriched, the model's understanding ability can be enhanced, the detection efficiency can be improved, the calculation time can be reduced, and further the robustness of the intelligent model for malicious sample analysis can be improved.
[0145] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A malicious sample analysis method based on a knowledge-enhanced neural network intelligent model, characterized in that: The method includes the following steps: Step S1: Collect a large number of data samples, including a set of known malicious samples and a set of normal samples, use the pre-trained language model BERT to extract semantic features from the collected data samples, and select the corresponding vectors of the tags as semantic feature vectors; Step S2: Use dependency syntactic analysis to extract syntactic features from the collected data samples, calculate the frequency distribution of dependency relation types, form a row statistical feature vector with the frequency values, and obtain a syntactic feature vector after normalization; Step S3: By calculating the data sample The Euclidean distance between the feature vector of and the feature vector of each normal sample set is used to construct an anomaly detection index vector; Step S4: Introduce an external knowledge graph to enhance the semantic feature vector extracted in Step S1; query entities and relationships related to the sample features from the knowledge graph, generate a knowledge-enhanced feature vector based on the query results, and concatenate the semantic feature vector, syntactic feature vector, anomaly detection index vector, and knowledge-enhanced feature vector to generate a comprehensive feature vector; Step S5: Construct a two-layer BP neural network model. The input of the two-layer BP neural network model is the normalized comprehensive feature vector, which includes semantic features, syntactic features, anomaly detection indexes, and knowledge enhancement features. The number of neurons in each layer is adjusted according to the input feature dimension. The activation function uses ReLU, and the output layer uses the softmax function to predict the probability that the sample is malicious. The cross-entropy loss function and Adam optimization algorithm are used to optimize the network parameters. The pre-trained language model BERT and the two-layer BP neural network model constitute the final knowledge-enhanced malicious sample detection model; Step S6: Collect the text to be detected and input it into the knowledge-enhanced malicious sample detection model to output the malicious detection result.
2. The method for analyzing malicious samples based on a knowledge-enhanced neural network intelligent model according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11: Tokenize and preprocess the data sample The data sample is represented as: where is the th token of the data sample , and represents the number of tokens in the data sample . After concatenating the start and end of the data sample with the tokens and respectively, the text is obtained, indicating the start of the entire sentence or text sequence, indicating the end of the current input sequence; Input the text into the pre-trained language model BERT; Step S12: The pre-trained language model BERT outputs the hidden state vectors corresponding to each token. The output form of the BERT model is expressed as: ; where is the output sequence of the model, is the hidden state vector corresponding to each token; for the entire text , select the vector corresponding to the token mark as the semantic feature vector.
3. The malicious sample analysis method based on the knowledge-enhanced neural network intelligent model according to claim 2, wherein The specific steps of Step S2 include the following sub-steps: Step S21: First, perform dependency parsing on the data sample to obtain the dependency relationships between each token and other tokens; Step S22: Construct a dependency matrix To represent the dependency relationship between word units, the dependency matrix Chinese elements Indicates the When the first word is used as a subordinate word, The dependency type of a word as a dominant word; Step S23: Perform statistical feature extraction according to the dependency relation matrix. Through the row statistical method, calculate the global frequency distribution of each dependency relation type, and form a row statistical feature vector with these frequency values; Step S24: Normalize the line statistical feature vector to the range of [0, 1] to obtain the normalized line statistical feature vector.
4. The malicious sample analysis method based on the knowledge-enhanced neural network intelligent model according to claim 3, wherein The specific steps of Step S3 include the following sub-steps: Step S31: The semantic feature vector and the syntactic feature vector form a feature vector, and the feature vector of the data sample is represented as ; Calculate the feature vector formed by the semantic feature vector and the syntactic feature vector for the normal sample set, which is represented as ; Calculate the Euclidean distance between the feature vector of the data sample and the feature vector of each normal sample set. Step S32: Calculate the data sample and the average distance and standard deviation of the Euclidean distances between the eigenvectors of the data sample and the eigenvectors of all normal sample sets to construct an anomaly detection metric vector.
5. The malicious sample analysis method based on a knowledge-enhanced neural network intelligent model according to claim 4, wherein The specific steps of Step S4 include the following sub-steps: Step S41: Match the semantic feature vector with the knowledge graph and query the relevant entities and relationships in the knowledge graph; Step S42: Use knowledge graph embedding to map the relationships and entities in the knowledge graph to the vector space, and concatenate the entity vector and the relationship vector to obtain a knowledge-enhanced feature vector; Step S43: Concatenate the knowledge-enhanced feature vector with the semantic feature vector, syntactic feature vector, and anomaly detection index vector into a comprehensive feature vector, and perform normalization processing on the comprehensive feature vector to obtain the normalized comprehensive feature vector.
6. The malicious sample analysis method based on the knowledge-enhanced neural network intelligent model according to claim 5, characterized in that The specific steps of Step S5 include the following sub-steps: Step S51: Construct a two-layer BP neural network model, and use the normalized comprehensive feature vector as the input of the BP neural network; Step S52: Use the ReLU function as the activation function of the hidden layer; Step S53: Use the softmax function to output the probability of the malicious sample; Step S54: According to the output result, use the weighted cross-entropy loss function and Adam optimizer to optimize the model.
Citation Information
Patent Citations
Semantic recognition method and system for inscriptions on ancient bronze objects
CN112036189A
Malicious software detection method based on semantic analysis and bidirectional coding representation
CN116432184A
Image processing apparatus, image processing method, and storage medium
US20230137350A1