Chip manual security attribute analysis and extraction method

The corpus is constructed by crawlers and used pre-training and fine-tuning of the anti-data augmentation strategy and BERT model to solve the problems of inefficiency and misjudgment and missed detection in the chip manual security analysis, realizing accurate extraction of the chip manual security configuration and cross-architecture robustness analysis.

CN120492705APending Publication Date: 2025-08-15Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510553316.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing chip manuals have problems of misjudgment and missed security configurations caused by low efficiency in safety analysis technology, poor generalization of rules and lack of knowledge in general model fields.

Method used

By writing crawlers to build the original corpus, the adversarial augmentation data is generated using random deletion, synonym replacement and random insertion strategies, pre-training and fine-tuning are combined with the BERT model, and hierarchical learning rate and adversarial loss training are used to optimize the noise immunity and task adaptability of the model.

Benefits of technology

It realizes accurate extraction of the chip manual security configuration and cross-architecture robustness analysis, significantly improving the accuracy and efficiency of security analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492705A_ABST
    Figure CN120492705A_ABST
Patent Text Reader

Abstract

The invention relates to a method for analyzing and extracting security attributes of a chip manual, which comprises the following steps of: compiling crawler codes, crawling the chip manual on a chip manual website, acquiring original text information, constructing an original corpus, generating adversarial enhancement data by using one or more of a random deletion strategy RD, a synonym replacement strategy SR and a random insertion strategy RI, and extracting the security attributes of the chip manual according to the adversarial enhancement data. Obtaining an enhanced data set to generate a plurality of models Mi; and pre-training the plurality of models Mi, finally outputting an MBERT1 model, setting hyper-parameters to carry out MBERT1 model fine tuning, predicting an answer position span through argmax decoding, and outputting a final fine tuning model to a specified storage path. According to the method, accurate extraction of security configuration and robustness analysis of a cross-chip architecture are realized through a pre-training algorithm and a fine tuning algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for analyzing and extracting security attributes of a chip manual. Background Art

[0002] Amidst the growing threat landscape of embedded system security, traditional chip manual security configuration analysis technology faces multiple bottlenecks. Manual review and analysis relies on domain experts parsing documents item by item. While this approach can accommodate specific version constraints, it is inefficient, time-consuming, and labor-intensive, making it difficult to rapidly audit large heterogeneous chip clusters. Subjective bias can also lead to systemic misjudgment. Rule-based automated detection extracts parameters through predefined grammars or pattern matching, but its rules generalize poorly, significantly increasing false positive and false negative rates when scaling across chip architectures. It is particularly inadequate for unstructured text (e.g., implicit dependencies, ambiguous timing specifications) and dynamic semantic scenarios, leading to missed detection of critical threats. While general large language models (e.g., ChatGPT)-driven analysis offers advantages in natural language understanding, their lack of domain knowledge can easily lead to semantic illusions (e.g., terminology ambiguity, security mode conflicts), and insufficient contextual modeling of physical constraints such as chip version iterations and hardware backplane topology. This can lead to uncontrollable deviations between the generated strategies and actual hardware behavior. While existing knowledge base systems (such as RAGflow) incorporate retrieval enhancement technology, their static knowledge embedding mechanisms struggle to adapt to dynamically evolving domain knowledge systems, creating the risk of semantic lag. Their closed architectures limit cross-domain analogical reasoning and prevent them from transcending pre-defined knowledge boundaries to address complex threat scenarios. These issues collectively contribute to insufficient security configuration parsing accuracy and increased hardware vulnerability remediation costs. Therefore, there is an urgent need to overcome these technical bottlenecks through a collaborative optimization mechanism combining domain pre-training, adversarial enhancement, and knowledge guidance. Summary of the Invention

[0003] This invention aims to address the problems of security configuration misjudgments and missed detections caused by low efficiency, poor rule generalization, and a lack of domain knowledge in general models in existing chip manual security analysis. A chip manual security attribute analysis and extraction method is proposed. This method uses RD / SR / RI adversarial data augmentation to improve the model's noise resistance. Chip terminology representation is optimized through joint pre-trained masked language modeling and sentence order prediction. The optimal model is selected based on perplexity. A fine-tuning phase employs weighted training with hierarchical learning rates and adversarial loss to enhance task adaptability. This method achieves accurate security configuration extraction and robust analysis across chip architectures.

[0004] To achieve the above objectives, the present invention proposes a chip manual security attribute analysis and extraction method, including a BERT model, characterized by comprising:

[0005] Step 1: Write crawler code, crawl chip manuals on the chip manual website, obtain original text information, and build the original corpus D raw;

[0006] Step 2: Original corpus D raw Generate adversarial enhancement data by using one or more strategies among random deletion strategy RD, synonym replacement strategy SR and random insertion strategy RI, and then obtain an enhanced dataset;

[0007] Using multiple augmented datasets and the original corpus D raw Generate multiple models M based on the BERT model i ;

[0008] Step 3: For multiple models M i Perform pre-training, the pre-training comprising:

[0009] Step 3.1: Jointly optimize the masked language modeling loss L MLM and sentence order prediction loss L SOP , by minimizing L1=E[L MLM +L SOP ]Update model parameters;

[0010] Step 3.2: Use perplexity to evaluate the language modeling performance of multiple pre-trained models, and select the model Mi with the lowest perplexity as the final output M BERT-1 Model;

[0011] Step 4: Build a training dataset Dtra containing question-context-answer triplets in / val , set the hyperparameters for M BERT-1 Model fine-tuning, the M BERT-1 Model fine-tuning includes:

[0012] Step 4.1: Use a hierarchical learning rate strategy and set the top-level parameter learning rate to be greater than the lower-level parameter learning rate;

[0013] Step 4.2: Through M BERT-1 The word segmenter of the model generates the token sequence X and the position code pos, and maps the character offset of the answer a to the start and end position labels after word segmentation. Each round of forward propagation calculates the standard loss L2 and the adversarial loss L adv The weighted sum of L2+λL adv , and update the parameters by back propagation, where λ = 0.2;

[0014] Step 4.3: Train according to the hyperparameters and evaluate the validation set using the ExactMatch and F1Score metrics every 100 training cycles.

[0015] Step 5: Predict the answer position span through argmax decoding and output the final fine-tuned model to the specified storage path.

[0016] Furthermore, the original corpus Draw in step 1 includes constructing a training dataset, a fine-tuning dataset, and an independent test set.

[0017] The training dataset consists of several sentences from the original chip manuals for pre-training the BERT model. The fine-tuning dataset consists of several manually annotated safety information sentences from the chip manuals. An independent test set consists of several annotated sentences from the chip manuals for model evaluation.

[0018] Furthermore, the random deletion strategy RD in step 2 includes: randomly selecting an adjective or adverb from the sentence for deletion with uniform probability while maintaining the integrity of the key semantic components of the sentence;

[0019] The synonym replacement strategy SR in step 2 includes: based on the WoSRNet semantic network, performing synonym replacement on the main verb in the sentence while keeping other parts of speech unchanged;

[0020] The random insertion strategy RI in step 2 includes: using a predefined adverb library to insert non-destructive modifying components at random positions in the sentence.

[0021] Each strategy mitigates the risk of data overfitting and enhances the generalization ability of the model by introducing diversity to the original sentences and increasing the potential vocabulary choices. We train multiple BERT models using different augmented data together with the original data and evaluate their performance.

[0022] Furthermore, the mask language modeling loss L in step 3.1 is MLM As shown in formula (1):

[0023]

[0024] Where i∈masked is a traversal of all masked Token positions, the masking ratio is usually 15%, i represents the position index of the i-th masked Token in the input sequence, and P is the conditional probability that the model predicts that the original Token at the masked position i is wi;

[0025] The sentence order prediction loss L described in step 3.1 SOP As shown in formula (2):

[0026] L SOP =-∑[y i logp(s i )+(1-y i )log(1-p(s i ))](2);

[0027] Among them, yi is the true label, when yi = 1, the sentence pair order is correct, yi = 0: the sentence pair order is wrong, and p(si) is the probability value predicted by the model that the order of the i-th sentence pair is correct.

[0028] L MLM Forces the model to understand the local context, L SOP By judging whether the order of sentence pairs is correct, the ability to model long-range semantic relationships is enhanced.

[0029] Furthermore, the perplexity in step 3.2 is expressed as formula (3):

[0030]

[0031] Among them, P(ω t |ω <t ) is the model’s predicted probability for the T-th word.

[0032] Perplexity is essentially the inverse geometric mean of the probability distribution of the model for the test corpus. In the evaluation of corpus in the field of chip design, the perplexity indicator can effectively reflect the modeling accuracy of the professional terminology sequence.

[0033] Furthermore, the hyperparameters set in step 4 include: batch size of 12, learning rate of 3×10 -5 , the maximum input sequence length is 384, and the number of training rounds is 2;

[0034] Step 4.1 adopts a hierarchical learning rate strategy, and the top-level parameter learning rate is 3×10 -4 , the underlying parameter learning rate is 3×10 -5 .

[0035] Domain knowledge transfer is achieved through a hierarchical parameter update mechanism, where model initialization weights are inherited from a universal language representation space. Gradient backpropagation is then performed on the target dataset to optimize task-relevant features. A dynamic sequence chunking strategy is employed during training, with a maximum sequence length and sliding window step size set to address information truncation issues with long text inputs. A learning rate decay mechanism is also introduced to balance model convergence speed with optimization stability.

[0036] Furthermore, the standard loss L1 in step 4.2 is as shown in formula (4):

[0037] L2=-logPstart(s * )-logPend(e * )(4);

[0038] Among them, s * and e *is the actual start and end position index, Pstart and Pend are the probability distributions of the start and end positions predicted by the model;

[0039] Adversarial loss L adv As shown in formula (5):

[0040] L adv =D KL (Poriginal(X0 / / Pperturbation(X+δ)) (5);

[0041] where Poriginal(X) is the probability distribution predicted by the model given the original input X, Pperturbation(X+δ) is the probability distribution predicted by the model given the perturbed input X+δ, where δ is the perturbation δ=∈· / / g / / 2g, and g is the embedding gradient.

[0042] Furthermore, the ExactMatch in step 4.3 is as shown in formula (6):

[0043]

[0044] Among them, N is the total number of test samples, I is the indicator function, when the predicted answer With any standard answer The value is 1 if it is completely consistent, otherwise it is 0;

[0045] F1Score is shown in formulas (7-8):

[0046]

[0047] Among them, TP is the number of words in common between the predicted answer and the standard answer, FP is the number of predicted redundant words, and FN is the number of marked missing words.

[0048] ExactMatch measures the proportion of model outputs that are completely consistent with the ground truth answer. In question-answering systems, especially fact-based question answering, the correct answer is usually clear and unique. EM directly reflects the model's ability to produce the correct answer.

[0049] The answers generated by the model may contain redundant or missing words, but the core information is correct. F1Score uses bag-of-words matching to tolerate some errors, making it more suitable for practical application needs.

[0050] Through the above technical solution, the beneficial effects of the present invention are:

[0051] The chip manual security attribute analysis and extraction method proposed in this invention constructs an original corpus by writing a crawler, generates adversarial enhancement data using random deletion (RD), synonym replacement (SR), and random insertion (RI) strategies, and generates multiple models based on BERT. Pre-training is performed by jointly optimizing the masked language modeling loss and sentence order prediction loss, and the optimal model, MBERT-1, is selected based on perplexity. A triplet training dataset is then constructed, hyperparameters are set, and domain knowledge transfer is achieved through a hierarchical parameter update mechanism. The model initialization weights are inherited from the universal language representation space, and gradient backpropagation is then performed on the target dataset to optimize task-related features. A dynamic sequence block strategy is adopted during training. By setting the maximum sequence length and sliding window step size, the problem of information truncation during long text input is addressed. A learning rate decay mechanism is introduced to balance the model convergence speed and optimization stability, and a fine-tuned model is ultimately output. This method enhances term representation through domain pre-training, improves noise resistance through adversarial data enhancement, and strengthens task adaptation through layered fine-tuning. It achieves accurate extraction of chip manual security configuration and cross-architecture robustness analysis, effectively solving the problems of low efficiency, false positives and missed detections of traditional methods, and significantly improves the accuracy and efficiency of security analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a schematic diagram of the steps of a chip manual security attribute analysis and extraction method of the present invention;

[0053] Figure 2 This is a ChipGuard-BERT model evaluation diagram for a chip manual security attribute analysis and extraction method according to the present invention.

[0054] Figure 3 This is a comparative evaluation chart of similar models of a chip manual security attribute analysis and extraction method of the present invention. DETAILED DESCRIPTION

[0055] Example 1

[0056] like Figures 1 to 3 The ChipGuard-BERT model shows a method for analyzing and extracting chip manual security attributes, including the BERT model and:

[0057] Step 1: Write crawler code, crawl chip manuals on the chip manual website, obtain original text information, and build the original corpus D raw ;

[0058] Step 2: Original corpus D raw Generate adversarial enhancement data by using one or more strategies among random deletion strategy RD, synonym replacement strategy SR and random insertion strategy RI, and then obtain an enhanced dataset;

[0059] Using multiple augmented datasets and the original corpus D raw Generate multiple models M based on the BERT model i ;

[0060] Step 3: For multiple models M i Perform pre-training, the pre-training comprising:

[0061] Step 3.1: Jointly optimize the masked language modeling loss L MLM and sentence order prediction loss L SOP , by minimizing L1=E[L MLM +L SOP ]Update model parameters;

[0062] Step 3.2: Use perplexity to evaluate the language modeling performance of multiple pre-trained models, and select the model Mi with the lowest perplexity as the final output M BERT-1 Model;

[0063] Step 4: Build a training dataset Dtra containing question-context-answer triplets in / val , set the hyperparameters for M BERT-1 Model fine-tuning, the M BERT-1 Model fine-tuning includes:

[0064] Step 4.1: Use a hierarchical learning rate strategy and set the top-level parameter learning rate to be greater than the lower-level parameter learning rate;

[0065] Step 4.2: Through M BERT-1 The word segmenter of the model generates the token sequence X and the position code pos, and maps the character offset of the answer a to the start and end position labels after word segmentation. Each round of forward propagation calculates the standard loss L2 and the adversarial loss L adv The weighted sum of L2+λL adv , and update the parameters by back propagation, where λ = 0.2;

[0066] Step 4.3: Train according to the hyperparameters and evaluate the validation set using the ExactMatch and F1Score metrics every 100 training cycles.

[0067] Step 5: Predict the answer position span through argmax decoding and output the final fine-tuned model to the specified storage path.

[0068] The original corpus Draw described in step 1 includes constructing a training dataset, a fine-tuning dataset, and an independent test set.

[0069] The random deletion strategy RD in step 2 includes: randomly selecting an adjective or adverb from the sentence with uniform probability for deletion while maintaining the integrity of the key semantic components of the sentence;

[0070] The synonym replacement strategy SR in step 2 includes: based on the WoSRNet semantic network, performing synonym replacement on the main verb in the sentence while keeping other parts of speech unchanged;

[0071] The random insertion strategy RI in step 2 includes: using a predefined adverb library to insert non-destructive modifying components at random positions in the sentence.

[0072] Jointly optimize the masked language modeling loss L as described in step 3.1 MLM As shown in formula (1):

[0073]

[0074] Where i∈masked is a traversal of all masked Token positions, the masking ratio is usually 15%, i represents the position index of the i-th masked Token in the input sequence, and P is the conditional probability that the model predicts that the original Token at the masked position i is wi;

[0075] The sentence order prediction loss L described in step 3.1 SOP As shown in formula (2):

[0076] L SOP =-∑[y i logp(s i )+(1-y i )log(1-p(s i ))] (2);

[0077] Among them, yi is the true label, when yi = 1, the sentence pair order is correct, yi = 0: the sentence pair order is wrong, and p(si) is the probability value predicted by the model that the order of the i-th sentence pair is correct.

[0078] The perplexity in step 3.2 is shown in formula (3):

[0079]

[0080] Among them, P(ω t |ω <t ) is the model’s predicted probability for the T-th word.

[0081] The hyperparameters set in step 4 include: batch size 12, learning rate 3×10 -5 , the maximum input sequence length is 384, and the number of training rounds is 2;

[0082] Step 4.1 adopts a hierarchical learning rate strategy, and the top-level parameter learning rate is 3×10 -4, the underlying parameter learning rate is 3×10 -5 .

[0083] The standard loss L1 described in step 4.2 is shown in formula (4):

[0084] L2=-logPstart(s * )-logPend(e * ) (4);

[0085] Among them, s * and e * is the actual start and end position index, Pstart and Pend are the probability distributions of the start and end positions predicted by the model;

[0086] Adversarial loss L adv As shown in formula (5):

[0087] L adv =D KL (Poriginal(X) / / Pperturbation(X+δ)) (5);

[0088] where Poriginal(X) is the probability distribution predicted by the model given the original input X, Pperturbation(X+δ) is the probability distribution predicted by the model given the perturbed input X+δ, where δ is the perturbation δ=∈· / / g / / 2g, and g is the embedding gradient.

[0089] The ExactMatch described in step 4.3 is shown in formula (6):

[0090]

[0091] Among them, N is the total number of test samples, I is the indicator function, when the predicted answer With any standard answer The value is 1 if it is completely consistent, otherwise it is 0;

[0092] F1Score is shown in formulas (7-8):

[0093]

[0094] Among them, TP is the number of words in common between the predicted answer and the standard answer, FP is the number of predicted redundant words, and FN is the number of marked missing words.

[0095] Example 2

[0096] Experiments were conducted based on a chip manual security attribute analysis and extraction method in Example 1, which proved that the present invention has a good effect.

[0097] Experimental configuration:

[0098] The experiment was conducted on a server running Ubuntu 22.04.5 LTS operating system, which was equipped with Intel(R) Xeon(R) Gold 6338 CPU processor and four NVIDIA L40 graphics cards.

[0099] The training dataset consists of 2,593 sentences from the original chip manual. The fine-tuning dataset consists of 600 manually annotated safety information sentences from the aforementioned documents. The independent test set consists of an additional 200 annotated sentences from the original chip manual.

[0100] First, pre-training as described in step 3 is performed to obtain M BERT-1 Model.

[0101] The original dataset was modified using data augmentation. Three types of augmented datasets were generated using a systematic data augmentation strategy: random deletion (RD), synonym substitution (SR), and random insertion (RI). The modified dataset also included the original dataset. The initial BERT model was pre-trained using the training dataset, random deletion, synonym substitution, and random insertion datasets.

[0102] Table 1

[0103]

[0104] As shown in Table 1, the ablation experiment based on perplexity (PPL) and training efficiency (Time / Epoch) shows that the RI strategy performs best in semantic preservation (PPL=19.2). Therefore, the domain adaptation model (M BERT-1 model) for subsequent experiments.

[0105] Then fine-tune as in step 4, with the main hyperparameters set to: batch size 12, learning rate 3×10 -5 , the maximum input sequence length is 384, the document sliding step is 128, and the training round is 2; the hierarchical learning rate strategy is adopted in step 4.1, and the top parameter learning rate is 3×10 -4 , the underlying parameter learning rate is 3×10 -5 During the training process, large-batch training is implicitly achieved through gradient accumulation. The final output model is named ChipGuard-BERT. Its core architecture maintains the 12-layer Transformer structure of BERT-base, and improves the entity relationship understanding ability in the chip design field through domain adaptation fine-tuning.

[0106] For ease of understanding, the same data processing method as Experiment 1 was used to process the fine-tuning dataset, and then the M BERT-1 Conduct fine-tuning experiments on the model.

[0107] Table 2

[0108]

[0109] As shown in Table 2, the RI strategy achieved the best perplexity, demonstrating that inserting noise words effectively improves the model's robustness to long-term contextual interference. Therefore, we selected the RI-enhanced dataset fine-tuned model (named the ChipGuard-BERT model) for subsequent experiments.

[0110] The ChipGuard-BERT model is evaluated using an independent test set:

[0111] Evaluation protocol design: (1) Baseline model selection: original BERT model, M BERT-1 Model,ChipGuard-BERT model,(2) The evaluation indicators are shown in Table 3:

[0112] Table 3

[0113]

[0114] The experimental results are as follows Figure 2 As shown in Figure 2, the original BERT model lags significantly behind in chip design tasks. The ChipGuard-BERT model demonstrates strong recognition capabilities in the field of chip manual terminology recognition.

[0115] To comprehensively evaluate the practical value of domain-specific models, a heterogeneous model comparison framework was constructed, and three typical architectures were selected for horizontal evaluation: ① ChipGuard-BERT ② General-purpose large language model GPT-4o ③ Retrieval-enhanced generation system RAGflow (locally built). The experiment adopted a cross-controlled variable design:

[0116] Test standard construction: (1) Data source: chip manuals from 5 mainstream manufacturers, (2) Question set: 20 questions were manually constructed, (3) Evaluation protocol: 3 chip design engineers independently judged the accuracy of the answers and took the average value;

[0117] like Figure 3 As shown, ChipGuard-BERT significantly outperforms in specialized terminology matching tasks (such as register bit definitions and timing constraints) involving vendor-specific terminology and structured data extraction. While GPT-4o demonstrates semantic coherence in open-ended questions, it is vulnerable to strict technical specification alignment. RAGflow achieves high accuracy when questions explicitly refer to manual sections, but is limited by the coverage of its local knowledge base and suffers from dynamic knowledge loss.

[0118] The embodiments described above are only preferred embodiments of the present invention and do not limit the scope of implementation of the present invention. Therefore, any equivalent changes or modifications made according to the structure, characteristics and principles described in the patent scope of the present invention should be included in the scope of the patent application of the present invention.

Claims

1. A chip manual security attribute analysis and extraction method, including a BERT model, characterized in that: include: Step 1: Write crawler code, crawl chip manuals on the chip manual website, obtain original text information, and build the original corpus D raw ; Step 2: Original corpus D raw Generate adversarial enhancement data by using one or more strategies among random deletion strategy RD, synonym replacement strategy SR and random insertion strategy RI, and then obtain an enhanced dataset; Using multiple augmented datasets and the original corpus D raw Generate multiple models M based on the BERT model i ; Step 3: For multiple models M i Perform pre-training, the pre-training comprising: Step 3.1: Jointly optimize the masked language modeling loss L MLM and sentence order prediction loss L SOP , by minimizing L1=E[L MLM +L SOP ]Update model parameters; Step 3.2: Use perplexity to evaluate the language modeling performance of multiple pre-trained models, and select the model Mi with the lowest perplexity as the final output M BERT-1 Model; Step 4: Build a training dataset Dtra containing question-context-answer triplets in / val , set the hyperparameters for M BERT-1 Model fine-tuning, the M BERT-1 Model fine-tuning includes: Step 4.1: Use a hierarchical learning rate strategy and set the top-level parameter learning rate to be greater than the lower-level parameter learning rate; Step 4.2: Through M BERT-1 The word segmenter of the model generates the token sequence X and the position code pos, and maps the character offset of the answer a to the start and end position labels after word segmentation. Each round of forward propagation calculates the standard loss L2 and the adversarial loss L adv The weighted sum of L2+λL adv , and update the parameters by back propagation, where λ = 0.2; Step 4.3: Train according to the hyperparameters and evaluate the validation set using the ExactMatch and F1Score metrics every 100 training cycles. Step 5: Predict the answer position span through argmax decoding and output the final fine-tuned model to the specified storage path.

2. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: The original corpus Draw described in step 1 includes constructing a training dataset, a fine-tuning dataset, and an independent test set.

3. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: The random deletion strategy RD in step 2 includes: randomly selecting an adjective or adverb from the sentence with uniform probability for deletion while maintaining the integrity of the key semantic components of the sentence; The synonym replacement strategy SR in step 2 includes: based on the WoSRNet semantic network, performing synonym replacement on the main verb in the sentence while keeping other parts of speech unchanged; The random insertion strategy RI in step 2 includes: using a predefined adverb library to insert non-destructive modifying components at random positions in the sentence.

4. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: Jointly optimize the masked language modeling loss L as described in step 3.1 MLM As shown in formula (1): Where i∈masked is a traversal of all masked Token positions, the masking ratio is usually 15%, i represents the position index of the i-th masked Token in the input sequence, and P is the conditional probability that the model predicts that the original Token at the masked position i is wi; The sentence order prediction loss L described in step 3.1 SOP As shown in formula (2): L SOP ==∑[y i logp(s i )+(1-y i )log(1-p(s i ))] (2); Among them, yi is the true label, when yi = 1, the sentence pair order is correct, yi = 0: the sentence pair order is wrong, and p(si) is the probability value predicted by the model that the order of the i-th sentence pair is correct.

5. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: The perplexity in step 3.2 is shown in formula (3): Among them, P(ω t |ω <t ) is the model’s predicted probability for the T-th word.

6. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: The hyperparameters set in step 4 include: batch size 12, learning rate 3×10 -5 , the maximum input sequence length is 384, and the number of training rounds is 2; Step 4.1 adopts a hierarchical learning rate strategy, and the top-level parameter learning rate is 3×10 -4 , the underlying parameter learning rate is 3×10 -5 .

7. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: The standard loss L1 described in step 4.2 is shown in formula (4): L2=-logPstart(s * )-logPend(e * ) (4); Among them, s * and e * is the actual start and end position index, Pstart and Pend are the probability distributions of the start and end positions predicted by the model; Adversarial loss L adv As shown in formula (5): L adv =D KL (Poriginal(X) / / Pperturbation(X+δ)) (5); where Poriginal(X) is the probability distribution predicted by the model given the original input X, Pperturbation(X+δ) is the probability distribution predicted by the model given the perturbed input X+δ, where δ is the perturbation δ=∈· / / g / / 2g, and g is the embedding gradient.

8. A chip manual security attribute analysis and extraction method according to claim 1, characterized in that: The ExactMatch described in step 4.3 is shown in formula (6): Among them, N is the total number of test samples, I is the indicator function, when the predicted answer With any standard answer The value is 1 if it is completely consistent, otherwise it is 0; F1Score is shown in formulas (7-8): Among them, TP is the number of words in common between the predicted answer and the standard answer, FP is the number of predicted redundant words, and FN is the number of marked missing words.