CPE identification system based on ChatGPT large model and incremental learning technology
By adopting ChatGPT large model and incremental learning technology in the CPE recognition system, combined with preprocessing, semantic coding and label prediction layer design, the problem of identifying rare CPE entities in the prior art is solved, and the recognition accuracy and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510088418.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The difficulty in effectively identifying and responding to rare CPE entities in open source software has led to insufficient generalization and identification accuracy of models in the face of changing security challenges.
A CPE recognition system based on ChatGPT large model and incremental learning technology is adopted, which includes a preprocessing layer, a semantic coding layer and a tag prediction layer. The semantic coding layer performs preliminary processing and sequence analysis through ChatGPT pre-trained large model and gated loop unit network module, while the tag prediction layer uses layer-by-layer improved label attention mechanism module and anti-forget incremental learning module for refined processing and model updates.
It improves the accuracy and training efficiency of vulnerability information identification, enhances the generalization ability of the model, and can dynamically adjust to adapt to new vulnerability data, making the CPE identification model more effective in the face of changing security challenges.
Smart Images

Figure CN120030549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology security technology, and more specifically to a CPE identification system based on a ChatGPT large model and incremental learning technology. Background Art
[0002] The software industry in the digital age is growing rapidly, but related software products are mostly based on open source code bases and components; therefore, many organizations focus on the security risks of open source software. Among them, the National Vulnerability Database (NVD), which is widely recognized by the world and supported by the National Institute of Standards and Technology (NIST), has recorded a large number of new security vulnerabilities, and the number is still growing. Because many organizations rely on NVD to update the security vulnerability database, it is crucial for every organization to quickly analyze and respond to security vulnerabilities in open source software.
[0003] Therefore, improving the model's ability to identify rare CPE entities and enhancing the model's generalization so that it can more effectively respond to ever-changing security challenges are issues that technical personnel in this field urgently need to address. Summary of the invention
[0004] In view of this, the present invention provides a CPE identification system based on the ChatGPT large model and incremental learning technology, which can effectively improve the accuracy of vulnerability information identification and training efficiency.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] A CPE identification system based on the ChatGPT large model and incremental learning technology, comprising: a software security vulnerability information database, a preprocessing layer, a semantic encoding layer, and a label prediction layer connected in sequence;
[0007] The preprocessing layer is used to preprocess the data in the software security vulnerability information database;
[0008] The semantic coding layer includes: a ChatGPT pre-trained large model module and a gated recurrent unit network module; the ChatGPT pre-trained large model module is used to perform preliminary processing on the pre-processed data to obtain text feature data after language understanding; the gated recurrent unit network module is used to perform sequence analysis on the text feature data to obtain dynamic features;
[0009] The label prediction layer includes a label attention mechanism module and an anti-forgetting incremental learning module that are improved layer by layer; the label attention mechanism module that is improved layer by layer is used to refine the dynamic features; the anti-forgetting incremental learning module is used to timely update the CPE identification system.
[0010] Preferably, the ChatGPT pre-trained large model module includes a Transformer model; the Transformer model adopts an encoder-decoder architecture, including an input embedding layer, a multi-head self-attention layer, a feedforward neural network layer, a residual connection layer and layer normalization.
[0011] Preferably, the gated recurrent unit network module includes an update gate, a reset gate, candidate hidden states and a final hidden state; the update gate is used to control the transmission of information flow and capture long-term dependencies; the reset gate is used to combine new inputs with past memories and forget irrelevant historical information; the candidate hidden states are used to provide candidate values for hidden state updates; and the final hidden state is used to combine current and past information to form a current state.
[0012] Preferably, the layer-by-layer improved label attention mechanism module includes a label attention layer, a channel attention layer, a self-attention layer and a global feature aggregation module; the process of fine-tuning the dynamic features includes:
[0013] The label attention layer is used to capture the relationship between labels and features, and different attention weights are assigned to different labels; the channel attention layer enhances the feature expression ability of the model by learning the attention weight of each channel feature; the self-attention layer is used to process the irregularity of the text feature data and obtain the correlation between the elements; the global feature aggregation module obtains an enhanced feature map by aggregating local features and global context information.
[0014] Preferably, the layer-by-layer improved label attention mechanism-based module converts the enhanced feature map into a final recognition result through an output layer to determine whether there is a security vulnerability in the software security vulnerability information database and the severity of the vulnerability.
[0015] Preferably, the anti-forgetting incremental learning module includes a memory replay mechanism, a regularization strategy, feature sparsification and contrastive learning;
[0016] The memory replay mechanism is used to store some old category samples and replay them during incremental training; the regularization strategy adds penalty items to limit excessive updates of the CPE recognition system; the feature sparsification and contrastive learning are used to reduce the noise or irrelevant information of the feature data and to determine the category to which the feature data belongs.
[0017] Preferably, a method of a CPE identification system based on the ChatGPT large model and incremental learning technology is also included, characterized in that it includes:
[0018] Acquire data from a software security vulnerability information database; preprocess the acquired data through a preprocessing layer; input the preprocessed data into a semantic coding layer, wherein the semantic coding layer includes a ChatGPT pre-trained large model module and a gated recurrent unit network module, and perform preliminary processing on the preprocessed data using the ChatGPT pre-trained large model module to obtain text feature data after language understanding, and then perform sequence analysis on the text feature data through the gated recurrent unit network module to obtain dynamic features;
[0019] The obtained dynamic features are then input into a label prediction layer, which includes a label attention mechanism module and an anti-forgetting incremental learning module that are improved layer by layer, and the dynamic features are refined using the label attention mechanism module that is improved layer by layer;
[0020] Finally, the CPE recognition system is updated in time through the anti-forgetting incremental learning module.
[0021] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a CPE identification system based on the ChatGPT large model and incremental learning technology, which can support continuous learning of the model, and retain and utilize existing knowledge while learning new knowledge. There is no need to repeatedly process previous training data, saving a lot of computing resources and storage space. At the same time, the existing knowledge is corrected and enhanced according to the new vulnerability data to match the new vulnerability data. This dynamic adjustment mechanism helps to improve the accuracy and generalization ability of the CPE identification model. In addition, ChatGPT is pre-trained on a large amount of unlabeled data to learn the general features of text language, and then fine-tuned for specific vulnerability information identification tasks to enable it to better adapt to CPE tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0023] Figure 1 A schematic diagram of the structure of the AI-enabled CPE identification system provided by the present invention;
[0024] Figure 2 A schematic diagram of the structure of the ChatGPT pre-trained large model module provided by the present invention;
[0025] Figure 3 A schematic diagram of the structure of the gated recurrent unit network module provided by the present invention;
[0026] Figure 4 This is a workflow diagram of the anti-forgetting incremental learning module provided by the present invention. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] The embodiment of the present invention discloses a CPE identification system based on ChatGPT large model and incremental learning technology.
[0029] like Figure 1 As shown, it includes: a software security vulnerability information database, a preprocessing layer, a semantic encoding layer and a label prediction layer connected in sequence.
[0030] The preprocessing layer is used to preprocess the data in the software security vulnerability information database.
[0031] The semantic coding layer includes: a ChatGPT pre-trained large model module and a gated recurrent unit network module; the ChatGPT pre-trained large model module is used to perform preliminary processing on the pre-processed data to obtain text feature data after language understanding; the gated recurrent unit network module is used to perform sequence analysis on the text feature data to obtain dynamic features.
[0032] The label prediction layer includes a label attention mechanism module and an anti-forgetting incremental learning module that are improved layer by layer; the label attention mechanism module that is improved layer by layer is used to refine the dynamic features; the anti-forgetting incremental learning module is used to timely update the CPE identification system.
[0033] The ChatGPT pre-trained large model module, the gated recurrent unit network module (GRU network module), the layer-by-layer improved label-based attention mechanism module (HLAN module) and the anti-forgetting incremental learning module form a pipeline. Each module further processes and improves the results of CPE recognition based on the previous module. The ChatGPT pre-trained large model module provides basic language understanding capabilities, the GRU network module processes sequence features, the HLAN module focuses on key information through the attention mechanism, and the anti-forgetting incremental learning module ensures the model's continuous learning and ability to adapt to new data. This modular design enables the system to handle complex CPE recognition tasks while evolving and improving over time.
[0034] Specifically, Figure 2 As shown, the ChatGPT pre-trained large model module includes a Transformer model; the Transformer model adopts an encoder-decoder architecture, including an input embedding layer, a multi-head self-attention layer, a feedforward neural network layer, a residual connection layer and layer normalization.
[0035] Specifically, the input embedding layer mainly includes word embedding, position encoding and embedding layer. The input embedding layer is used to convert the text data described by the vulnerability information into a numerical form that can be processed by the model, and retain the necessary position information, providing a basis for subsequent Transformer model processing.
[0036] Furthermore, in the embodiment provided by the present invention, the input embedding layer is the first step in processing sequence data in the deep learning model, and the input discrete symbols (software vulnerability security information characters) are converted into continuous vector representations; this layer uses the pre-trained word embedding Word2Vec method, and is randomly initialized during the training process and learned during the model training process. The corresponding mathematical calculation process is: Embedding(W,x)=Wx; wherein W is the embedding matrix and x is the input index vector.
[0037] Specifically, in the multi-head self-attention layer, multiple attention heads allow the model to focus on different parts of the vulnerability input information at the same time, improving the CPE identification system's ability to understand the context of the vulnerability input information.
[0038] Furthermore, in the embodiment provided by the present invention, the multi-head self-attention layer allows the model to capture dependencies between different positions of the vulnerability information sequence regardless of their position in the sequence; by transforming the query vector, key vector and value vector through different linear layers, and then calculating the attention score and weight, the corresponding mathematical calculation process is:
[0039]
[0040] Among them, Q, K and V are the representations of query vector, key vector and value vector respectively, d k is the dimension of the key.
[0041] Specifically, the feedforward neural network layer further processes and extracts features from the output of the self-attention layer in the Transformer module, and enhances the CPE recognition system's ability to capture complex features by introducing nonlinear transformations.
[0042] Further, in the embodiment provided by the present invention, the feedforward neural network layer is another layer after the self-attention layer, which further processes the output of the self-attention layer; it is usually a combination of two linear transformations with a ReLU or other activation function in the middle, and the corresponding mathematical calculation process is:
[0043] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2 ;
[0044] Among them, W 1 , W 2 and b 1 , b 2 are the parameters of the layer.
[0045] Specifically, the residual connection and normalization layer are used to alleviate the gradient vanishing problem, accelerate model training, improve the stability and flexibility of the CPE recognition system, and prevent network degradation, so that the deep network can be effectively trained and learned.
[0046] Furthermore, in the embodiments provided by the present invention, the residual connection helps information flow through the deep network by adding the input directly to the output of the layer, and the layer normalization normalizes the activation of each sample to accelerate training and improve the stability of the model. The corresponding mathematical calculation process is: LayerNorm(x+Sublayer(x))
[0047] Here, x is the input of the layer, and Sublayer(x) is the output of the feed-forward network layer or the self-attention layer.
[0048] Specifically, the encoder is composed of a plurality of identical stacked layers, each of which contains the self-attention mechanism and feedforward neural network mentioned above. This stacked structure allows the model to capture and learn the complex features of the input vulnerability data at different levels, providing a basis for the decoder to generate an output sequence.
[0049] Furthermore, the encoder typically includes a multi-head self-attention layer, a feedforward neural network layer, residual connections and layer normalization; the encoder layer processes the input sequence and generates a continuous representation for subsequent processing, and the corresponding mathematical calculation process is:
[0050] EncoderLayer(x)=LayerNorm(x+Sublayer(x));
[0051] Among them, Sublayer(x) can be a multi-head self-attention layer or a feed-forward network layer.
[0052] Specifically, the decoder layer gradually generates the output CPE recognition text according to the internal representation generated by the encoder, captures the dependency between the input sequence and the generated sequence through the self-attention sublayer and the encoder-decoder attention sublayer, and processes this information through the feedforward neural network sublayer to finally generate the CPE output sequence.
[0053] Further, in the embodiments provided by the present invention, the decoder generally comprises three main parts: a multi-head self-attention layer for processing the target sequence, a multi-head self-attention layer for processing the encoder output and the output of the previous decoder layer, and a feed-forward neural network layer. Each part is accompanied by residual connections and layer normalization. The corresponding mathematical calculation process of this process is:
[0054] DecoderLayer(x,y)=LayerNorm(x+Attention(x,y)+FFN(x))
[0055] Here, x is the input to the decoder layer and y is the output of the encoder.
[0056] Specifically, Figure 3 As shown, the gated recurrent unit (GRU) network module includes an update gate, a reset gate, a candidate hidden state and a final hidden state; the update gate is used to control the transmission of information flow and capture long-term dependencies; the reset gate is used to combine new inputs with past memories and forget irrelevant historical information; the candidate hidden state is used to provide candidate values for hidden state updates; the final hidden state is used to combine current and past information to form a current state.
[0057] Furthermore, in the embodiment provided by the present invention, the main principle of the GRU network in capturing long-term dependencies in CPE vulnerability identification is through its unique gating mechanism. GRU is specially designed to process sequence data and is suitable for modeling long-term dependent tasks. Unlike LSTM, the structure of GRU is relatively simple, containing only two gates (update gate and reset gate) instead of three gates (input gate, forget gate and output gate). This simplification of the structure enables GRU to improve the computational efficiency while maintaining the effect. Specifically, it is the synergy of the update gate and the reset gate. The following is a detailed explanation of the working of these two gating mechanisms.
[0058] In the update gate, the purpose of the update gate is to determine to what extent the hidden state should be updated by the new information; it outputs a value between 0 and 1 through a σ (sigmoid function), which represents the weight of retaining the hidden state of the previous time step. t The calculation process is:
[0059] z t =σ(W z ·[x t ,h t-1 ]+b z );
[0060] Among them, is x t The input of the current time step, ht-1 is the hidden state of the previous time step, W z and b z are the weight matrix and bias vector of the update gate respectively.
[0061] In the reset gate, the reset gate determines the relationship between the input of the current time step and the hidden state of the previous time step; it also outputs a value between 0 and 1 through a sigmoid function, which represents the weight of retaining the hidden state of the previous time step. The calculation process of the reset gate is:
[0062] r t =σ(W r ·[x t ,h t-1 ]+b r );
[0063] Among them, W r and b r are the weight matrix and bias vector of the reset gate respectively.
[0064] In the candidate hidden state, the candidate hidden state combines the current input and the information of the previous time step (adjusted by the reset gate). The calculation process is as follows:
[0065]
[0066] Among them, ⊙ represents the Hadamard product (element-wise product), W is the weight matrix of the candidate hidden state, b is the bias term of the candidate hidden state, and tanh is the hyperbolic tangent function, whose output value is between -1 and 1.
[0067] Current Hidden State: The current hidden state is the final output, which combines the information of the previous time step and the candidate hidden state; its calculation formula is as follows:
[0068]
[0069] Among them, the update gate z t Controls the hidden state h of the previous time step t-1 and candidate hidden state h t In the current hidden state The proportion in .
[0070] Through these gating mechanisms, GRU can dynamically adjust the flow of information when processing sequence data, so that the model can capture long-term dependencies in vulnerability data. In the CPE vulnerability identification task, this means that GRU can identify time series patterns in software vulnerability reports. For example, a vulnerability may persist in multiple versions of the software. GRU can identify potential security vulnerability CPE information by learning these long-term patterns.
[0071] Specifically, the layer-by-layer improved label attention mechanism module includes a label attention layer, a channel attention layer, a self-attention layer and a global feature aggregation module; the process of fine-tuning the dynamic features includes:
[0072] The label attention layer is used to capture the relationship between labels and features, and different attention weights are assigned to different labels; the channel attention layer enhances the feature expression ability of the model by learning the attention weight of each channel feature; the self-attention layer is used to process the irregularity of the text feature data and obtain the correlation between the elements; the global feature aggregation module obtains an enhanced feature map by aggregating local features and global context information.
[0073] Furthermore, in the embodiment provided by the present invention, the label attention layer aims to capture the relationship between labels and features and assign different attention weights to different labels; the core idea is that different labels have different importance for features, and by learning these weights, the model can better identify features related to specific labels. The calculation process of the label attention layer is:
[0074]
[0075] ContextVector=AttentionWeights·V;
[0076] Among them, Q is the query vector, which is usually the embedded representation of the target label; K is the key vector, which corresponds to the embedded representation of each element in the input data; V is the value vector, which also corresponds to the embedded representation of each element in the input data. By calculating the dot product of the query vector and all key vectors, and then applying the softmax function, the attention weight of each element is obtained, and finally the value vector is weighted and summed with these weights to obtain the context vector.
[0077] The channel attention layer focuses on the importance of different channels (features) and enhances the feature expression ability of the model by learning the weight of each channel. The calculation process of the channel attention layer is:
[0078] ChannelWeights=σ(FC(GlobalAvgPool(X)));
[0079] WeightedFeatures=X·ChannelWeights;
[0080] Among them, X is the feature map, GlobalAvgPool is the global average pooling, FC represents the fully connected layer, and σ is the sigmoid activation function for output channel weights.
[0081] The self-attention layer, also called the internal attention layer, allows the CPE recognition model to dynamically adjust the attention paid to each element when processing sequence data, and allows the model to directly capture dependencies between different positions in the sequence without considering the distance; the self-attention layer dynamically adjusts the attention paid to each element by calculating the correlation score between elements at different positions in the sequence. This mechanism enables the model to capture complex dependencies within the sequence without relying on the absolute position of the elements in the sequence.
[0082]
[0083] Among them, Q, K and V are the query vector, key vector and value vector respectively, which are usually different linear transformations of the input features.
[0084] The global feature aggregation module helps the model capture the global structure and pattern of the entire vulnerability information data by aggregating local features and global context information. This is crucial for identifying complex vulnerabilities that span multiple code segments. The calculation process of the global feature aggregation module is as follows:
[0085] GlobalFeature=Concat(GlobalAvgPool(X),GlobalMaxPool(X));
[0086] Among them, X is the feature map, GlobalAvgPool is the global average pooling, GlobalMaxPool is the global maximum pooling, and Concat represents the connection operation.
[0087] Finally, the HLAN module converts the extracted features into the final recognition results through an output layer, for example, determining whether there is a security vulnerability and the severity of the vulnerability. The output layer usually contains a softmax layer, which is used to convert the output of the model into a probability distribution, thereby assigning a probability value to each possible label (such as vulnerabilities of different types and severity). In this way, the model can give the corresponding signal sequence label and identify potential CPE software security vulnerability information. The calculation formula of the Softmax function is:
[0088]
[0089] Among them, p is the probability distribution after the softmax function, z is the output of the fully connected layer, and C is the total number of categories. Through the above steps, the model can give the corresponding signal sequence label and identify potential security vulnerabilities. This enables the HLAN module to effectively extract security vulnerability information related to CPE from a large amount of software vulnerability data and give the corresponding signal sequence label, thereby improving the accuracy and efficiency of the software security vulnerability identification task.
[0090] Specifically, the layer-by-layer improved label attention mechanism module (HLAN module) converts the enhanced feature map into a final recognition result through an output layer to determine whether there is a security vulnerability in the software security vulnerability information database and the severity of the vulnerability.
[0091] Specifically, Figure 4 As shown, the anti-forgetting incremental learning module includes a memory replay mechanism, a regularization strategy, feature sparsification, and contrastive learning;
[0092] The memory replay mechanism is used to store some old category samples and replay them during incremental training; the regularization strategy adds penalty items to limit excessive updates of the CPE recognition system; the feature sparsification and contrastive learning are used to reduce the noise or irrelevant information of the feature data and to determine the category to which the feature data belongs.
[0093] Furthermore, in the embodiments provided by the present invention, the application of incremental learning in the CPE vulnerability information recognition task is reasonable. It can provide an effective mechanism to handle continuously growing and changing data, while maintaining the memory of old knowledge and improving the accuracy and efficiency of CPE vulnerability information recognition. Incremental learning allows the model to continuously improve its performance when receiving new data, which is a significant advantage for CPE vulnerability information recognition because it can ensure that the model always stays up-to-date to cope with the ever-evolving software vulnerability security threats. At the same time, incremental learning can reduce the risk of forgetting the knowledge of old tasks when learning new tasks, which is particularly important for CPE vulnerability information recognition because it is necessary to retain the memory of old vulnerability information for association and comparison with new vulnerability information.
[0094] In the model initialization stage, collect and organize the initial vulnerability information data D 0 , to form a data set These data will be used to train the initial model. The data set should contain the feature vector x i and the corresponding label y i , where the label represents the category or attribute of the vulnerability; used to train the initial CPE recognition system based on the large ChatGPT model and incremental learning technology; use the initial training data set D 0 and the selected gradient descent optimization algorithm to train the model; during the initialization training process, the model parameters θ 0 will be adjusted according to the gradient of the loss function to minimize the loss, and a validation set is used to evaluate the performance of the model; this helps to prevent overfitting and ensure that the CPE recognition model can also perform well on unseen data; select an appropriate loss function to measure the difference between the model prediction value and the true value. For a linear model, the loss function usually adopts the mean squared error (MSE):
[0095]
[0096] where m is the number of samples, hθ(x) is the model prediction value, and y is the true value.
[0097] In the incremental learning stage, through incremental learning, the model absorbs new knowledge, that is, the ability to recognize new category vulnerability information; after the CPE recognition model is deployed, new software vulnerability information data is monitored, and these data may represent vulnerability types that have not been seen before; for the newly emerged vulnerability information, collect relevant training data to form a new data set
[0098] Use the new data set Dnew to calculate the gradient of the loss function and update the model parameters θ 0 to θ 1 , that is In order to achieve anti-forgetting incremental learning, an L2 regularization term is added to the loss function; the L2 regularization term is the sum of the squares of the model parameters. Its purpose is to limit the size of the model parameters, thereby reducing overfitting and helping the model learn on new tasks while reducing the forgetting of old task knowledge. This method helps to improve the generalization ability and adaptability of the model:
[0099]
[0100] Among them, λ is the regularization parameter that controls the strength of regularization; n is the number of model parameters; θ j is the model parameter; the data loss L data and the regularization loss L reg Combined, we get the total loss function:
[0101]
[0102] In the model parameter update phase, the gradient descent algorithm is used in incremental learning to update the model parameters. First, the gradient of the loss function with respect to the parameter θ is calculated:
[0103]
[0104] The gradient calculation process for the L2 regularization term can be expressed as:
[0105]
[0106] Therefore, the gradient value of the summarized model can be expressed as:
[0107]
[0108] Where X is the feature matrix and y is the target value vector. For the gradient descent update rule, in each iteration, the parameters need to be updated according to the gradient descent rule:
[0109]
[0110] Among them, α is the learning rate; this process is iterative, and as new vulnerability information continues to emerge, the model needs to continuously perform incremental learning to maintain the effectiveness of its recognition ability.
[0111] In the model verification test phase, after updating the CPE identification model, an independent software vulnerability information test set Dtest is used to verify the performance of the CPE identification system model based on the ChatGPT large model and incremental learning technology. Through the above steps, the incremental learning model can continuously receive new data while updating its knowledge base to adapt to new identification tasks, while trying to maintain the recognition ability of old tasks.
[0112] Specifically, it also includes a CPE identification method based on the ChatGPT large model and incremental learning technology, characterized in that data is obtained from a software security vulnerability information database; the obtained data is preprocessed through a preprocessing layer; the preprocessed data is input into a semantic coding layer, the semantic coding layer includes a ChatGPT pre-trained large model module and a gated recurrent unit network module, the ChatGPT pre-trained large model module is used to perform preliminary processing on the preprocessed data to obtain text feature data after language understanding, and then the text feature data is sequenced and analyzed through the gated recurrent unit network module to obtain dynamic features;
[0113] The obtained dynamic features are then input into a label prediction layer, which includes a label attention mechanism module and an anti-forgetting incremental learning module that are improved layer by layer, and the dynamic features are refined using the label attention mechanism module that is improved layer by layer;
[0114] Finally, the CPE recognition system is updated in time through the anti-forgetting incremental learning module.
[0115] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0116] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A CPE identification system based on ChatGPT large model and incremental learning technology, characterized in that: include: A software security vulnerability information database, a preprocessing layer, a semantic encoding layer, and a label prediction layer connected in sequence; The preprocessing layer is used to preprocess the data in the software security vulnerability information database; The semantic coding layer includes: a ChatGPT pre-trained large model module and a gated recurrent unit network module; The ChatGPT pre-trained large model module is used to perform preliminary processing on the pre-processed data to obtain text feature data after language understanding; the gated recurrent unit network module is used to perform sequence analysis on the text feature data to obtain dynamic features; The label prediction layer includes a label attention mechanism module and an anti-forgetting incremental learning module that are improved layer by layer; the label attention mechanism module that is improved layer by layer is used to refine the dynamic features; the anti-forgetting incremental learning module is used to timely update the CPE identification system.
2. According to claim 1, a CPE identification system based on ChatGPT large model and incremental learning technology is characterized in that: The ChatGPT pre-trained large model module includes a Transformer model; the Transformer model adopts an encoder-decoder architecture, including an input embedding layer, a multi-head self-attention layer, a feedforward neural network layer, a residual connection layer and layer normalization.
3. According to claim 1, a CPE identification system based on ChatGPT large model and incremental learning technology is characterized in that: The gated recurrent unit network module includes an update gate, a reset gate, a candidate hidden state and a final hidden state; the update gate is used to control the transmission of information flow and capture long-term dependencies; the reset gate is used to combine new inputs with past memories and forget irrelevant historical information; the candidate hidden state is used to provide candidate values for hidden state updates; the final hidden state is used to combine current and past information to form a current state.
4. According to claim 1, a CPE identification system based on ChatGPT large model and incremental learning technology is characterized in that: The layer-by-layer improved label attention mechanism module includes a label attention layer, a channel attention layer, a self-attention layer and a global feature aggregation module; The process of fine-tuning the dynamic features includes: The label attention layer is used to capture the relationship between labels and features, and different attention weights are assigned to different labels; The channel attention layer enhances the feature expression capability of the model by learning the attention weight of each channel feature; the self-attention layer is used to process the irregularity of the text feature data and obtain the correlation between the elements; the global feature aggregation module obtains an enhanced feature map by aggregating local features and global context information.
5. According to claim 4, a CPE identification system based on ChatGPT large model and incremental learning technology is characterized in that: The layer-by-layer improved label attention mechanism module converts the enhanced feature map into a final recognition result through an output layer to determine whether there is a security vulnerability in the software security vulnerability information database and the severity of the vulnerability.
6. According to claim 1, a CPE identification system based on ChatGPT large model and incremental learning technology is characterized in that: The anti-forgetting incremental learning module includes a memory replay mechanism, a regularization strategy, feature sparsification, and contrastive learning; The memory replay mechanism is used to store some old category samples and replay them during incremental training; the regularization strategy adds a penalty term to limit excessive updates of the CPE recognition system; The feature sparsification and contrastive learning are used to reduce the noise or irrelevant information of the feature data and to determine the category to which the feature data belongs.
7. A method for a CPE identification system based on a ChatGPT large model and incremental learning technology as described in any one of claims 1 to 6, characterized in that: include: Obtain data from the software security vulnerability information database; Preprocess the acquired data through the preprocessing layer; The preprocessed data is input into the semantic coding layer, which includes a ChatGPT pre-trained large model module and a gated recurrent unit network module. The ChatGPT pre-trained large model module is used to perform preliminary processing on the preprocessed data to obtain text feature data after language understanding, and then the gated recurrent unit network module is used to perform sequence analysis on the text feature data to obtain dynamic features; The obtained dynamic features are then input into a label prediction layer, which includes a label attention mechanism module and an anti-forgetting incremental learning module that are improved layer by layer, and the dynamic features are refined using the label attention mechanism module that is improved layer by layer; Finally, the CPE recognition system is updated in time through the anti-forgetting incremental learning module.
Citation Information
Patent Citations
Vulnerability entity extraction method for small sample semantic analysis
CN118070804A
Vulnerability detection method based on code data stream enhanced large model
CN118246029A
Penetration testing system and method based on AI large model and storage medium
CN119065983A
Deep Reinforced Model for Abstractive Summarization
US20180300400A1