Verilog row-level code completion method based on non-autoregression model

By using a non-autoregressive model to generate tokens in parallel in Verilog code completion, and combining a mixed syntax-guided sampling strategy, the problem of inefficiency of the autoregressive model is solved, and efficient and accurate Verilog code completion is achieved.

CN120179228APending Publication Date: 2025-06-20DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510240647.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the inefficiency problem caused by token generation of autoregressive models, especially in Verilog code completion, resulting in frequent low-level repetitive coding in digital circuit design and low encoding efficiency.

Method used

A non-autoregressive model is adopted to generate all tokens in parallel, and row-level code completion is achieved, and a mixed syntax-guided sampling strategy is introduced during the training process to improve the model's ability to capture code syntax structure and semantic features.

Benefits of technology

It significantly improves the efficiency and accuracy of Verilog code completion, reduces low-level duplicate encoding work, improves coding efficiency, and ensures the correctness and reliability of generated codes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179228A_ABST
    Figure CN120179228A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of code completion, and relates to a Verilog row-level code completion method, in particular to a Verilog row-level code completion method based on a non-autoregression model. The method is different from a traditional code completion method based on an autoregression model, the non-autoregression model is used, delay is smaller, and the reasoning speed is remarkably increased. Besides, a sampling strategy guided by mixed grammar is adopted, the method is similar to Teacher Forcing, the sampling size can be dynamically adjusted according to the learning effect of the model, and the method can enable the model to better learn grammar and semantic information in the model so as to improve the code completion accuracy. In conclusion, the invention provides the Verilog code completion method which has relatively high precision and can remarkably improve the reasoning speed at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of code completion, and relates to a Verilog line-level code completion method, specifically a Verilog line-level code completion method based on a non-autoregressive model. Background Art

[0002] With the rapid development of information technology, intelligent code generation technology has become a key means to improve software development efficiency. At present, the line-level code completion task for high-level programming languages has been widely concerned, and the technologies and products are relatively mature. However, there is less research on code completion for hardware description languages (such as Verilog HDL, which is a widely used hardware description language), and the available tools are also relatively scarce, and there is no adaptation to the characteristics of the Verilog language. For example, the current VSCode plugin contains a code completion tool for the Verilog language, but it can only implement some simple completion functions based on character matching or rules, and it is at the identifier level.

[0003] Whether it is a high-level programming language or a hardware description language, code completion is often the most practical and important branch in intelligent code generation. Compared with the code generation branch, the recommended fragments for code completion are shorter (such as at the identifier level or line level). Since the recommended code fragments for code generation are relatively long, it is easier to introduce errors. Once an error occurs, developers need to re-read the generated code fragments, locate the errors, and repair the program, resulting in reduced efficiency. However, the recommended fragments for code completion can be selectively accepted by developers, and the main program logic is still controlled manually, so the correctness is more guaranteed, and the improvement of program development efficiency is more stable and reliable. It has now become a widely recognized auxiliary development means.

[0004] Code completion has more practical value, but different code recommendation lengths and the implementation methods of the tool itself will affect the recommendation effect. For example, the code completion tool (VSCode plugin) for Verilog has two problems: on the one hand, due to the relatively simple completion method, the recommended code is not intelligent enough; on the other hand, the identifier level is too short, the completion efficiency is low, and the inspiration for users is not enough, while the recommendation of function-level or longer code fragments is prone to introducing errors. In contrast, line-level code completion is more reasonable.

[0005] Currently, the line-level code completion technology for high-level programming languages has been relatively mature, mostly using the autoregressive model (Autoregressive Model, abbreviated as AR model). It assumes that the value of a variable can be predicted by its previous sequence and is a sequence-to-sequence model (Sequence-to-Sequence Model, abbreviated as Seq2Seq model). However, the decoding mechanism of the autoregressive model has inherent limitations. It must rely on all previously generated Tokens to predict the next Token, making the decoding process have to be executed in a strict order and unable to achieve parallel computing, resulting in limited completion speed.

[0006] In recent years, the non-autoregressive model (Non-Autoregressive Model, abbreviated as NAR model) has become a research hotspot due to its parallel generation ability. Different from the autoregressive model, the non-autoregressive model can predict the entire target sequence simultaneously, thus significantly improving the generation speed. Therefore, the purpose of the present invention is to implement a line-level code completion method for Verilog code based on the non-autoregressive model, so as to reduce the low-level repetitive coding work in digital circuit design and help relevant coders improve coding efficiency. Summary of the Invention

[0007] The present invention aims to provide a Verilog line-level code completion method based on the non-autoregressive model, taking both efficiency and accuracy into account to solve the problem of low efficiency caused by the autoregressive model generating Tokens one by one in the prior art. To greatly improve the code completion efficiency, the non-autoregressive model is adopted in the present invention, and all Tokens can be generated in parallel with only one decoding. To further improve the quality of the generated code, a sampling strategy is introduced during training. This is a mechanism similar to Teacher Forcing in the recurrent neural network, which mixes the real values during the training process. At the same time, rules for guiding the mixed grammar are formulated for the sampled Tokens, especially focusing on those Tokens carrying rich grammar information, enabling the model to better capture the syntax structure and semantic features of the code, thereby improving the training efficiency and performance.

[0008] The technical solution of the present invention is as follows: A Verilog line-level code completion method based on the non-autoregressive model includes the following steps:

[0009] Step (1) is to construct a high-quality dataset suitable for Verilog code completion. First, collect Verilog code files from open-source hardware projects and industrial-level design cases to obtain the original dataset D raw . Filter the original dataset D rawFor the overly long items and duplicate items, clean and remove the non-Verilog code data among them, and finally use the MinHash algorithm to remove the code files with overly similar content, obtaining the processed dataset D composed of several high-quality code file data process 。

[0010] Step (2) traverses the processed dataset D process All the code files in it. For each file, use regular expressions to split its code into a Token sequence, and use the open-source parsing tool Pyverilog dedicated to Verilog to filter comments, extract the keywords (such as module, always), operators (such as <=, &) and identifiers (such as signal names) contained in the code, and then obtain the syntax information of each Token. This step will support the subsequent sampling strategy guided by mixed syntax. Finally, obtain the corresponding Token sequence of each code file in the processed dataset D process and the corresponding syntax category of each Token

[0011] Step (3) traverses the Token sequence and identifies the line break characters in it, which makes the line number information in the original code file retained in the Token sequence, and then generates the input Token-target Token pairs of each code file based on the sliding window method. Specifically, set the sliding window size to N w lines, that is, for each code file, if its total number of code lines N c is less than N w +1 lines, then directly extract the first N c -1 lines and the last line as the input Token-target Token pairs. Otherwise, select the first N w lines of the code as the input Token, and the subsequent line (the N w +1 line) as the target Token. Then slide the window down to generate the next input Token (the 2nd to N w +1 lines) and the corresponding target Token (the N w +2 line), and so on until the window cannot slide anymore. Finally, obtain the structured dataset D structure 。

[0012] Step (4) Design a non-autoregressive encoder-decoder model architecture. The encoder adopts the self-attention mechanism, and its structure is similar to the encoder in Transformer. However, to solve the problem of token-by-token generation delay in autoregressive models, a non-autoregressive decoder is designed, and tokens can be generated in parallel during its decoding process. The non-autoregressive method can greatly improve the generation efficiency because the autoregressive model generates tokens one by one and each step depends on the previously generated tokens, while the non-autoregressive model generates all tokens in parallel and does not depend on historical outputs. In addition, to improve the training quality of the model, a hybrid grammar-guided sampling strategy is introduced. For this purpose, a shared decoder is designed for the model to perform secondary decoding during training, that is, during the training process, the same decoder is used to perform an additional decoding. In this additional decoding, the input of the second decoding comes from the mixture of the input of the first decoding and part of the ground truth, depending on the hybrid grammar-guided sampling strategy rather than completely random sampling. This sampling step can help train the model weights, similar to Teacher Forcing, which is very helpful when the model converges slowly or the generation effect is not good in the initial stage of training, and it only appears in the training stage and will not appear in the inference stage, so it will not leak the ground truth during model inference.

[0013] Step (5) Divide the structured dataset D obtained in step (3) structure into a training set D train , a validation set D val and a test set D test , and use the model architecture designed in step (4) for training. Stop training after reaching the set number of training rounds to obtain the trained model M. The number of training rounds is set according to experience, but it must be greater than the number of rounds when the model just converges. The convergence of the model is specifically manifested as the losses on the training set and the validation set tending to be stable, with a very small fluctuation range within 3 to 5 consecutive training rounds.

[0014] Step (6) Use the trained model M in step (5) to perform inference and generate code on the test set D test .

[0015] Step (7) For the code results generated by inference in step (6), evaluate the quality of the code generated by the model from four dimensions: BLEU-4, edit similarity (ES), exact match accuracy (EM), and latency.

[0016] Furthermore, step (1) specifically includes the following steps:

[0017] 1-1) Collect Verilog code files from open-source hardware projects and industrial-level design cases to obtain the original dataset D containing several Verilog code files raw .

[0018] 1-2) A large number of duplicates are contained in the initially collected code files, and the overly long code is not conducive to processing and inputting into the training model. Therefore, the original dataset D raw After filtering out the overly long code files with more than 20,000 code characters, pairwise comparison is performed to remove the code files with exactly the same content.

[0019] 1-3) Clean it to remove the non-Verilog code data therein. According to observations, usually, Verilog code will contain one or more of the following basic keywords: always, assign, always_ff, always_comb, always_latch. Therefore, remove the code that does not contain any of the above basic keywords. Almost all non-Verilog code can be removed in this step.

[0020] 1-4) Use the MinHash algorithm to remove similar codes, and set the Jaccard similarity threshold to 0.8. This step is mainly to remove the code files that were not filtered out in step 1-1) due to minor differences but whose contents are too close to each other, ensuring the diversity of the training data, thereby improving the model training effect. Finally, the processed dataset D consisting of several high-quality code file data is obtained. process .

[0021] Furthermore, step (2) specifically includes the following steps:

[0022] 2-1) Design a regular expression, regard spaces and special characters as split points, filter spaces and retain special characters. Traverse the processed dataset D process All the code files in it. For each file, use the above-designed regular expression to split the code into a Token sequence. It should be noted that since the newline character is very important in the sliding window method of step (3), it is also retained here. Finally, for each code file in the processed dataset D process , a corresponding code Token sequence set composed of several Token sequences can be obtained.

[0023] 2-2) Use the Python open-source library Pyverilog dedicated to parsing Verilog code to automatically filter comments and parse each Token contained in the corresponding Token sequence obtained for each code file in step 2-1), locate the identifiers and keywords therein, obtain the syntax information of each Token and establish an index corresponding to it with the corresponding Token. Finally, the corresponding Token sequence of each code file in the processed dataset D process And the corresponding syntax category of each Token are obtained.

[0024] Furthermore, specifically describe the design of the non-autoregressive encoder-decoder architecture in step (4). The specific structure is as follows:

[0025] Encoder: Add an additional token, and use the output of this token on the encoder to predict the target length. In addition, there is a set of attention heads on the output of the encoder.

[0026] Soft-copy process: According to the predicted target token length from the output of the encoder, map the output S = {s1, s2,..., s m} so that the length becomes the target token to be used as the input H = {h1, h2,.., h m} for the decoder. The method is shown in the following formulas (1) and (2):

[0027]

[0028] W ji =W j W i (i + s i ) (2)

[0029] where W i 、W j 、W ji are weights. W ji depends on the distance relationship between the source position i and the target position j. These weights will be updated during the model training process, but only the parameters are updated during the second decoding process.

[0030] Decoder with shared parameters: There are two decoding processes during the model training process, but only one decoder with shared parameters is used. There is another set of self-attention heads on the generated code. The first decoding process generates predicted tokens in the following way, as shown in the following formula (3):

[0031]

[0032] where X is the input sequence, is the predicted sequence, f decode (·) is the decoding process, f encode (·) is the encoding process, soft-copy(·) is the soft-copy process, θ1 is the encoder parameter, and θ2 is the decoder parameter.

[0033] In the inference stage, the predicted tokens generated at this time are the predicted target sequence; in the training stage, they are only the initial predicted tokens, and a second-round decoding process is required. The accuracy of the initial predicted tokens is used to evaluate the difficulty of fitting the current target, which is used to determine the number of samples during the sampling process. The input tokens for the second-round decoding come from the mixture of the input tokens of the first-round decoding and the true value sampling. The same decoder is used in the second-round decoding to enable the model to learn the remaining tokens that have not been selected.

[0034] Hybrid grammar-guided sampling strategy: The sampling process is divided into two steps. First, determine the number of samples N, and then sample N tokens from the true value according to the hybrid grammar-guided sampling strategy.

[0035] In the first step, calculate the number of samples N as shown in the following formula (4):

[0036]

[0037] where λ is the sampling rate, which is a hyperparameter, represents the Hamming distance between the target sequence Y and the predicted sequence is the token predicted at the current time step t, and y is the corresponding true value. t is the corresponding true value.

[0038] In the second step, introduce a random probability p into the sampling process. With a probability of 1 - p, the sampling is completely random; for the remaining probability of p, the sampling needs to consider the corresponding grammar types, that is, the ratio of keywords, identifiers, and operators is 2:1:1, and a total of N true values are sampled according to this ratio.

[0039] Furthermore, the training process of step (5) is introduced in detail:

[0040] 5-1) In step (3), the data that conforms to the model input has been obtained and divided into a training set D train , a validation set D val and a test set D test according to a ratio, and the training process begins.

[0041] 5-2) First is the encoding process. Due to the characteristics of the non-autoregressive model itself, it is difficult to control the length of the output token sequence. For this reason, the length of the predicted target token is embedded in the first position of the encoder input, and during the soft copy process, the length of the target token is corrected in a mapped manner through this predicted value. That is, the output S = {s1, s2,..., s m} of the encoder is mapped to the input H = {h1, h2,.., h m}, the mapping process is shown in the following formulas (5) and (6):

[0042]

[0043] W ji = W j W i (i + s i ) (6)

[0044] where W i , W j , W ji are weights. W ji depends on the distance relationship between the source position i and the target position j. All weights are trained during the model training process.

[0045] 5 - 3) Use the result obtained by soft-copying in step 5 - 2) as the input for the first decoding process, and the output obtained is the preliminary predicted Token.

[0046] 5 - 4) In step 5 - 3), the preliminary predicted Token has been obtained, and its accuracy indicates the difficulty of fitting the current target. For example, the lower the accuracy of the preliminary prediction, the greater the prediction difficulty, and it is necessary to further increase the sampling number of the added true values. As the training progresses, the sampling strategy gradually reduces the number of sampled true values, and finally enables the model to learn how to predict the entire code segment in one pass without seeing any true values. The sampling number N is dynamically adjusted according to the Hamming distance between the preliminary predicted Token and the true sequence, and the dynamic adjustment method is shown in the following formula (7):

[0047]

[0048] where λ is the sampling rate, which is a hyperparameter, represents the Hamming distance between the target sequence Y and the predicted sequence , is the Token predicted at the current time step t, and y t is the corresponding true value.

[0049] 5 - 5) After determining the sampling number N, the most direct sampling strategy is to randomly select N Tokens from the true values, but the hybrid syntax-guided sampling strategy with a ratio of keywords, identifiers, and operators of 2:1:1 often has better results. In addition, a random probability p is introduced. It is stipulated that with a probability of 1 - p, the sampling is completely random; for the remaining probability of p, the ratio of keywords, identifiers, and operators in the N sampled true values needs to satisfy 2:1:1. Such a strategy helps the model learn syntax and semantic information, and the trained model is better.

[0050] After completing the sampling process, the input sequence for the second decoding process is obtained by mixing. The model's parameters are only updated during the second decoding process. The ultimate goal is to maximize the following loss function, which is shown in Equation (8) below:

[0051]

[0052] where \(X\) is the input sequence, \(Y\) is the target sequence, is the predicted sequence in the first decoding process, \(T\) is the grammar type corresponding to the target sequence, \(\theta\) is the model parameter, \(y\) t is the token generated at the current time step \(t\), is the subset of tokens not selected in the sampling strategy, is the posterior probability of generating the target token under the condition that the input sequence is \(X\) and the model parameter is \(\theta\), combined with the sampling strategy.

[0053] Step (6) is the inference process of the model. Since step (5) has trained the model to generate the final code sequence in one go, the sampling processes in steps 5 - 4) and 5 - 5) are no longer needed or allowed in step (6). Therefore, the inference process is to repeat steps 5 - 2) and 5 - 3) again, and the token output by step 5 - 3) is the final inference result.

[0054] Furthermore, in step (7), the definition of the evaluation metric is as follows:

[0055] BLEU - 4 is a variant of the BLEU metric, which evaluates the generation quality by calculating the n - gram matching degree (here \(n = 4\)) between the generated code and the reference code. Its core idea is to measure the similarity between the generated text and the target text in terms of vocabulary and local structure. Its calculation method is shown in Equation (9) below:

[0056]

[0057] where \(p\) n represents the proportion of the number of matching n - grams in the total n - grams in the generated text, is the weight of \(p\) n , and BP is the brevity penalty factor, which is used to penalize the case where the generated text is too short.

[0058] Edit Similarity (ES) measures the character - level modification cost between the generated code and the target code based on the Levenshtein distance and normalizes it to a similarity score. Its calculation method is shown in Equation (10) below, where Lev is the Levenshtein distance and \(Y\) represents the target sequence, represents the predicted sequence:

[0059]

[0060] The exact match accuracy (EM) is the ratio of the generated code that is exactly the same as the target code.

[0061] The latency is the average time required for the model to output a completion suggestion from receiving a single input (batch size is 1).

[0062] Compared with the prior art, the present invention has the following advantages:

[0063] Different from the traditional code completion method based on the autoregressive model, the present invention uses a non-autoregressive model, which has a smaller latency, that is, the inference speed is significantly improved. In addition, the present invention adopts a hybrid grammar-guided sampling strategy, which is a method similar to Teacher Forcing, and the sampling size can be dynamically adjusted according to the learning effect of the model. This method enables the model to better learn the grammar and semantic information therein to improve the accuracy of code completion. In summary, the present invention proposes a Verilog code completion method with high precision and significantly improved inference speed. Description of the Drawings

[0064] Figure 1 It is a schematic flowchart of the Verilog line-level code completion method based on the non-autoregressive model of the present invention.

[0065] Figure 2 It is a sub-flowchart of the process of collecting and cleaning data in the Verilog line-level code completion method based on the non-autoregressive model of the present invention.

[0066] Figure 3 It is a sub-flowchart of the process of splitting tokens in the Verilog line-level code completion method based on the non-autoregressive model of the present invention.

[0067] Figure 4 It is a sub-flowchart of the process of model training and inference in the Verilog line-level code completion method based on the non-autoregressive model of the present invention.

[0068] Figure 5 It is a sub-flowchart of the process of model evaluation in the Verilog line-level code completion method based on the non-autoregressive model of the present invention. Detailed Embodiments

[0069] The following will describe the embodiments of the present invention in detail in conjunction with the drawings and technical solutions.

[0070] Please refer to Figure 1, embodiments of the present invention disclose a Verilog line-level code completion method based on a non-autoregressive model. The Verilog code completed by the model generated according to this method has high precision, and the generation efficiency is significantly better than other similar methods. The specific steps are as follows:

[0071] (1) Collect and clean data from open-source hardware projects and industrial-level design cases. The specific sub-flow chart is as Figure 2 shown. This step requires collecting a large number of Verilog code files from open-source hardware projects and cleaning them to form a high-quality Verilog code dataset. The specific process is as follows:

[0072] 1.1. First, collect Verilog code files from open-source hardware projects and industrial-level design cases to form an original dataset D raw , which consists of a large number of code files.

[0073] 1.2. Conduct a first preliminary screening of the code files in the original dataset D raw , removing the overly long (more than 20,000 characters) code files and keeping only one of the completely duplicate code files.

[0074] 1.3. For the remaining code files after the above step, perform keyword matching and screen out the code files that do not contain any of the basic keywords. The basic keywords include: always, assign, always_ff, always_comb, always_latch.

[0075] 1.4. For the remaining code files after the above step, calculate the Jaccard similarity between each pair using the MinHash algorithm. If the similarity between two code files is higher than 0.8, only keep one of them. After completing this step, the remaining code files form a high-quality Verilog code dataset, namely the processed dataset D process .

[0076] Here is a sample code file in the embodiment, and its format in the processed dataset D process is as follows (block comments are omitted when showing):

[0077]

[0078]

[0079] (2) Split the processed dataset D process into the format of Token sequences and extract the syntax category corresponding to each Token. The specific sub-flow chart is as Figure 2 shown.

[0080] 2.1. First, perform regular expression matching on the processed dataset D process . The matching rule is: split by spaces and special characters, remove spaces but retain special characters and line breaks. After matching, each file in the processed dataset D process is split into Token sequences.

[0081] For the sample code file in the above embodiment, it is split in this step, and the obtained Token sequences are as follows:

[0083] ​'module', 'fetch', '(', 'clk', ',','stall', ',', 'busy', ',', 'pc', ',', 'rw', ',', 'access_size', ',', 'enable', ',', 'j_addr', ',', 'jump', ',', 'br_addr', ',', 'branch', ')', ';', '\n', 'parameter', 'START_ADDR', '=', '32\'h8002_0000', ';', '\n', 'input', 'clk', ';', '\n', 'input','stall', ';', '\n', 'input', 'busy', ';', '\n', 'input', '[', '31', ':', '0', ']', 'j_addr', ';', '\n', 'input', 'jump', ';', '\n', 'input', '[', '31', ':', '0', ']', 'br_addr', ';', '\n', 'input', 'branch', ';', '\n', 'output', '[', '31', ':', '0', ']', 'pc', ';', '\n', 'output', '[', '2', ':', '0', ']', 'access_size', ';', '\n', 'output', 'rw', ';', '\n', 'output', 'enable', ';', '\n','reg', '[', '31', ':', '0', ']', 'pc_reg', '=', '32\'h8001_FFFC', ';', '\n','reg', '[', '2', ':', '0', ']', 'access_size_reg', '=', '3\'b000', ';', '\n','reg', 'rw_reg', '=', '1\'b0', ';', '\n','reg', 'enable_reg', '=', '1\'b1', ';', '\n', 'assign', 'pc', '=', 'pc_reg', ';', '\n', 'assign', 'access_size', '=', 'access_size_reg', ';', '\n', 'assign', 'rw', '=', 'rw_reg', ';', '\n', 'assign', 'enable', '=', 'enable_reg', ';','\n','always','@','(','posedge','clk',')','\n','begin','\n','if','(','stall','!=','1 ','&','busy','!=','1',')','\n','begin','\n','if','(','jump','!=','1','&','branch','!=' ,'1',')','\n','begin','\n','pc_reg','=','pc_reg','+','32\'h0000_0004',';','\n','end',' \n','else','if','(','branch','==','1',')','\n','begin','\n','pc_reg','=','br_addr',';', '\n','end','\n','else','if','(','jump','==','1',')','\n','begin','\n','pc_reg','=','j_ addr',';','\n','end','\n','end','\n','else','if','(','branch','===','1',')','\n','begin' ,'\n','pc_reg','=','br_addr',';','\n','end','\n','else','if','(','jump','===','1',')',' \n','begin','\n','pc_reg','=','j_addr',';','\n','end','\n','end','\n','endmodule','\n'; ]

[0085] 2.2. The Verilog parsing tool Pyverilog is used to further parse the Token sequence. Comments are automatically filtered out during the process. After parsing, the corresponding syntax category of each Token can be obtained. process The segmentation is completed and converted into the format of Token sequence and its corresponding grammatical category.

[0086] For the example of the embodiment, a syntax type sequence corresponding to the Token sequence is obtained in this step, as follows: [

[0088] 'KEYWORD','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','NAME','OP','OP','OP','NEWLINE','KEYWORD','NAME','OP','NUMBER','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','OP','NUMBER','OP','NUMBER','OP','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','OP','NUMBER','OP','NUMBER','OP','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','OP','NUMBER','OP','NUMBER','OP','NAME','OP','NEWLINE','KEYWORD','OP','NUMBER','OP','NUMBER','OP','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NEWLINE','KEYWORD','OP','NUMBER','OP','NUMBER','OP','NAME','OP','NUMBER','OP','NEWLINE','KEYWORD','OP','NUMBER','OP','NUMBER','OP','NAME','OP','NUMBER','OP','NEWLINE','KEYWORD','NAME','OP','NUMBER','OP','NEWLINE','KEYWORD','NAME','OP','NUMBER','OP','NEWLINE','KEYWORD','NAME','OP','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NAME','OP','NEWLINE','KEYWORD','NAME','OP','NAME','OP','NEWLINE','KEYWORD','OP','OP','OP','KEYWORD','NAME','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','NAME','OP','OP','NAME','OP','OP','NUMBER','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','NAME','OP','NAME','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','NAME','OP','NAME','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','NAME','OP','NAME','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','NAME','OP','NAME','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','OP','OP','NAME','OP','OP','NUMBER','OP','OP','NEWLINE','KEYWORD','NEWLINE','NAME','OP','NAME','OP','OP','NEWLINE','KEYWORD','NEWLINE','KEYWORD','NEWLINE',

[0090] (3) Traverse the token sequence and generate the input token - target token pairs for each code file based on the sliding window method. Specifically, set the sliding window size to 10 lines, and determine the line numbers through the newline characters reserved in the token sequence. Each time, select the first 10 lines as the input tokens, and the next line as the target token. Then slide the window downward to generate the next input token and the corresponding target token, and so on until the window cannot be slid. Finally, combine them into a structured dataset D structure 。

[0091] Here are two sets of datasets generated by the first and second sliding windows for the examples of the embodiments. Each set of datasets consists of X (input tokens), Y (target tokens), and T (syntactic type of the target tokens), which is equivalent to the structured dataset D structure of 2 examples.

[0092] The input token - target token pairs of Example 1 are the first 10 lines and the 11th line of the implementation example code, as follows:

[0093] X:

[0095] ​​'module', 'fetch', '(', 'clk', ',','stall', ',', 'busy', ',', 'pc', ',', 'rw', ',', 'access_size', ',', 'enable', ',', 'j_addr', ',', 'jump', ',', 'br_addr', ',', 'branch', ')', ';', 'NEWLINE', 'parameter', 'START_ADDR', '=', '32\'h8002_0000', ';', 'NEWLINE', 'input', 'clk', ';', 'NEWLINE', 'input','stall', ';', 'NEWLINE', 'input', 'busy', ';', 'NEWLINE', 'input', '[', '31', ':', '0', ']', 'j_addr', ';', 'NEWLINE', 'input', 'jump', ';', 'NEWLINE', 'input', '[', '31', ':', '0', ']', 'br_addr', ';', 'NEWLINE', 'input', 'branch', ';', 'NEWLINE', 'output', '[', '31', ':', '0', ']', 'pc', ';', 'NEWLINE'

[0097] Y:

[0099] 'output', '[', '2', ':', '0', ']', 'access_size', ';', 'NEWLINE'

[0101] T:

[0103] 'KEYWORD', 'OP', 'NUMBER', 'OP', 'NUMBER', 'OP', 'NAME', 'OP', 'NEWLINE'

[0105] The input Token - target Token pairs for Example 2 are the 2nd to 11th lines and the 12th line of the implementation example code, as follows:

[0106] X:

[0108] ​​​​​​'parameter', 'START_ADDR', '=', '32\'h8002_0000', ';', 'NEWLINE', 'input', 'clk', ';', 'NEWLINE', 'input','stall', ';', 'NEWLINE', 'input', 'busy', ';', 'NEWLINE', 'input', '[', '31', ':', '0', ']', 'j_addr', ';', 'NEWLINE', 'input', 'jump', ';', 'NEWLINE', 'input', '[', '31', ':', '0', ']', 'br_addr', ';', 'NEWLINE', 'input', 'branch', ';', 'NEWLINE', 'output', '[', '31', ':', '0', ']', 'pc', ';', 'NEWLINE', 'output', '[', '2', ':', '0', ']', 'access_size', ';', 'NEWLINE'

[0110] Y:

[0112] 'output', 'rw', ';', 'NEWLINE'

[0114] T:

[0116] 'KEYWORD', 'NAME', 'OP', 'NEWLINE'

[0118] (4) Design a non-autoregressive encoder-decoder architecture. The architecture diagram, training phase, and inference phase of the model are all reflected in Figure 4 . The structure of this model includes:

[0119] Encoder: Add an additional Token and predict the target length based on the output of this Token on the encoder. In addition, there is a set of attention heads on the output of the encoder, similar to the Transformer model.

[0120] Soft copy process: According to the target Token length predicted by the output of the encoder, map the output S = {s1, s2,..., s m}} so that the length becomes the target Token to be used as the input H = {h1, h2,.., h m ​​​​​}, and its method is shown in the following formulas (1) and (2):

[0121]

[0122] W ji = W j W i (i + s i ) (2)

[0123] where W i , W j , W ji are weights. W ji depends on the distance relationship between the source position i and the target position j. These weights will be updated during the model training process, but only the parameters are updated in the second decoding process.

[0124] Decoder with shared parameters: There are two decoding processes during the model training process, but only one decoder with shared parameters is used, and there is another set of self-attention heads on the generated code. The way to generate the predicted Token in the first decoding process is as follows, as shown in formula (7) below:

[0125]

[0126] where X is the input sequence, is the predicted sequence, f decode (·) is the decoding process, f encode (·) is the encoding process, soft-copy(·) is the soft-copy process, θ1 is the encoder parameter, and θ2 is the decoder parameter.

[0127] In the inference stage, the predicted Token generated at this time is the predicted target sequence; in the training stage, it is only the initial predicted Token, and a second decoding process is required. The accuracy of the initial predicted Token is used to evaluate the difficulty of fitting the current target, which is used to determine the sampling quantity during the sampling process. The input Token for the second decoding comes from the input Token of the first decoding mixed with the true value sampling. The same decoder is used in the second decoding to enable the model to learn the remaining Tokens that have not been selected.

[0128] Sampling process: The sampling process is divided into two steps. First, determine the sampling quantity N, and then sample N Tokens from the true value according to the sampling strategy guided by the mixed grammar.

[0129] In the first step, the hyperparameter sampling rate λ can take a value of 0.3 during implementation, and based on this, the sampling quantity N can be calculated, as shown in the following formula (3):

[0130]

[0131] Among them, λ is the sampling rate, which is a hyperparameter. represents the Hamming distance between the target sequence Y and the predicted sequence and is the Token predicted at the current time step t, and y t is the corresponding true value.

[0132] In the second step, a random probability p is introduced into the sampling process. With a probability of 1 - p, the sampling is completely random; for the remaining probability of p, the sampling needs to consider the corresponding syntactic types, that is, the ratio of keywords, identifiers, and operators is 2:1:1, and a total of N true values are sampled according to this ratio.

[0133] (5) Model training, this part involves the entire non-autoregressive encoder-decoder architecture. The encoder part uses the self-attention mechanism, similar to the Transformer model, but the difference is that the decoder part is non-autoregressive and can generate all Tokens in parallel, thus significantly improving the generation speed. During the training process, a sampling strategy is introduced to optimize the training effect of the model by mixing true values and predicted values.

[0134] First, the encoder encodes the input Token sequence, and then maps the output of the encoder to the input of the decoder by means of soft copy. During the soft copy process, the weights are adjusted according to the distance relationship between the source position and the target position to correct the length of the target Token. After that, sampling is performed according to the output of the first decoder. The sampling strategy will dynamically adjust the sampling quantity according to the accuracy of the preliminary prediction. Through a certain ratio, the sampled true values can contain richer semantic and syntactic information, enhancing the training effect of the model. After sampling, the second decoding is performed and the weight parameters are updated. The ultimate goal is to maximize the following loss function, as shown in formula (4) below:

[0135]

[0136] Among them, X is the input sequence, Y is the target sequence, is the predicted sequence in the first decoding process, T is the syntactic type corresponding to the target sequence, θ is the model parameter, and y t is the Token generated at the current time step t, is the subset of Tokens not selected in the sampling strategy, is the posterior probability of generating the target Token under the condition that the input sequence is X, the model parameter is θ, and the sampling strategy is combined.

[0137] For example, in the initial stage of training, the prediction accuracy of the model is relatively low, and the sampling strategy will sample more real values to provide more information to guide the learning of the model. As the training progresses, the prediction ability of the model gradually improves, and the number of samples will gradually decrease until the model can predict the entire code snippet in one pass without any real value information. Through this dynamically adjusted sampling strategy, the model can gradually improve its generation ability for the code completion task during the training process.

[0138] During the training process of the embodiment, the hyperparameters used by the model are shown in Table (1) below:

[0139] Table 1 Hyperparameters of the model

[0140]

[0141]

[0142] After the training of the embodiment is completed, a series of parameters of the model will be obtained. For example, the tensor size of the word embedding matrix is (50000, 512). The encoder and decoder have a total of 6 layers. For each layer, its weight matrix includes: the Q, K, and V matrices in the self-attention mechanism and their corresponding offsets, with tensor sizes of (512, 512) and (512,); the self-attention layer normalization matrix and its corresponding offset, with tensor sizes of (512,) and (512,); two fully connected layers and their corresponding offsets, the former with tensor sizes of (2048, 512) and (2048,), and the latter with tensor sizes of (512, 2048) and (512,). The following shows some matrix weights (taking the first layer of the decoder as an example):

[0143] Word embedding matrix (50000, 512):

[0144] tensor([[-0.0605, 0.0561, 0.0446,..., 0.0609, -0.0517, -0.0072],

[0145] [0.0139, -0.0021, 0.0115,..., 0.0012, 0.01...166, 0.0407, 0.0195,...,-0.0364, 0.0371, -0.0377],

[0146] [0.0600, 0.0135, -0.0238,...,-0.0070, 0.0247, -0.0072]])

[0147] Q matrix (512, 512):

[0148] tensor([[0.0327, -0.0129, 0.0511, ..., 0.0191, 0.0244, -0.0111],

[0149] [0.0423, -0.0396, -0.0242, ..., 0.0287, 0.03...174, 0.0266, 0.0393, ..., 0.0603, -0.0094, 0.0198],

[0150] [-0.0147, 0.0075, -0.0217, ..., -0.0262, -0.0096, -0.0124]])

[0151] Q matrix offset (512,):

[0152] tensor([-0.1362, 0.1343, -0.0839, -0.0988, -0.0734, 0.0306, -0.1235, 0.0817, -0.0884, -0.1119, -0.1070, 0.0514, 0...,-0.1097, 0.1118, 0.1031, 0.0515, -0.0127,

[0154] 0.0967, 0.0406, -0.1086, 0.0717, 0.0639, -0.1184, -0.1235, 0.0497])

[0155] Self - attention layer normalization matrix (512,):

[0156] tensor([0.9561, 0.9634, 0.9751, 0.9844, 0.9604, 0.9771, 0.9897, 0.9790, 0.9824, 0.9805, 0.9692, 0.9985, 0.9707, 0.9595....9688, 0.9702, 0.9766, 0.9766, 0.9653, 0.9761, 0.9707,

[0158] 0.9692, 0.9629, 0.9590, 0.9746, 0.9673, 0.9648, 0.9570, 0.9629])

[0159] Self - attention layer normalization matrix offset (512,):

[0160] tensor([7.6962e - 04, -2.9507e - 03, -7.1526e - 03, -7.5569e - 03, 4.1237e - 03,

[0161] -7.2594e-03,-1.0422e-02,2.4452e-03,1.086...1e-03,-1.2970e-02,

[0162] 3.5381e-04,6.4240e-03,-4.9591e-03,6.5041e-04,-3.1528e-03,

[0163] -9.4299e-03,-2.0142e-03])

[0164] The first fully connected layer matrix (2048, 512):

[0165] tensor([[5.6648e-03,-1.4008e-02,3.0731e-02,...,-5.2185e-02,

[0166] 3.2318e-02,-2.0798e-02],

[0167] [2.2446e-02,...3453e-02,-2.1286e-02],

[0168] [2.7130e-02,-3.2120e-03,7.9651e-03,...,3.1921e-02,

[0169] 3.8849e-02,-6.5899e-04]])

[0170] The offset of the first fully connected layer matrix (2048,):

[0171] tensor([-0.0202,-0.0042,0.0077,...,-0.0403,-0.0345,-0.0077])

[0172] The second fully connected layer matrix (512, 2048):

[0173] tensor([[-1.3885e-02,-1.9836e-02,1.5001e-03,...,-1.1147e-02,

[0174] -2.5742e-02,2.2182e-03],

[0175] [1.6212e-05,...4610e-02,6.1722e-03],

[0176] [-7.7477e-03, -5.7907e-03, -2.6154e-02,..., 1.8539e-02,

[0177] 6.1941e-04, 2.6840e-02]])

[0178] The offset of the second fully connected layer matrix (512,):

[0179] tensor([7.9651e-03, -1.2253e-02, -2.3010e-02, -8.5831e-03, 4.0016e-03,

[0180] 1.9363e-02, 7.6637e-03, 1.6418e-02, -1.825...4e-02, -1.5625e-02,

[0181] 8.2626e-03, -1.8524e-02, -6.0081e-03, 4.4479e-03, 1.4191e-02,

[0182] 6.5956e-03, -1.1482e-02])

[0183] (6) Model inference refers to the process in which the trained model completes the process of recommending the next line of code based on the user's input code. The user's input is first converted into a Token sequence that the model can process. In the inference stage of the model, the weight parameters of the model have been trained, and the input sequence only needs to go through the processes of encoding, soft copying, and one decoding in sequence to obtain the predicted target Tokens, and these target Tokens are generated concurrently, thus greatly improving the code completion speed.

[0184] (7) Model evaluation is to evaluate the model from two dimensions: the accuracy and generation efficiency of the generated code to reflect the quality of the model, such as Figure 5 shown. Specifically, using the divided test set, the model makes recommendations on the test set and compares the obtained results with the true values. The following evaluation indicators are available:

[0185] BLEU-4 is a variant of the BLEU metric. It evaluates the generation quality by calculating the n-gram matching degree (here n = 4) between the generated code and the reference code, reflecting the similarity of the generated text to the target text in terms of vocabulary and local structure.

[0186] Edit Similarity (ES) measures the character-level modification cost between the generated code and the target code based on the Levenshtein distance and normalizes it into a similarity score.

[0187] The exact match accuracy (EM) is the proportion of the generated code that is exactly the same as the target code.

[0188] The latency is the average time required for the model to output a completion suggestion from receiving a single input (batch size is 1).

[0189] The three metrics of BLEU-4, edit similarity (ES), and exact match accuracy (EM) reflect the precision of the model in generating code, and the latency reflects the generation efficiency of the model. These four metrics together evaluate the quality of the model.

[0190] In the embodiment, the evaluation metrics of the trained model are shown in Table (2) below:

[0191] Table 2 Evaluation Metrics of the Model

[0192] Evaluation Index Result BLEU-4 31.72% EM 29.14% ES 65.93% Latency 35ms

[0193] It can be seen from this that the three metrics of BLEU-4, edit similarity (ES), and exact match accuracy (EM) have all reached a relatively high precision. From a theoretical perspective, the model has a probability of about 29.14% of completely completing the next line of code required by the user, and the correlation between the completed code and the next line of code required by the user is about 65.93%. Moreover, the completion efficiency is high and the required latency time is extremely low.

Claims

1. A Verilog line-level code completion method based on a non-autoregressive model, characterized in that: The steps include: Step (1) To build a high-quality dataset suitable for Verilog code completion, we first collect Verilog code files from open source hardware projects and industrial-grade design cases to obtain the original dataset D raw ; Filter the original data set D raw The long items and duplicate items are cleaned to remove the non-Verilog code data, and finally the MinHash algorithm is used to remove the code files with too similar content to obtain the processed data set D consisting of several code file data. process ; Step (2) Traverse the processed data set D process For all code files in the Verilog, regular expressions are used to split the code into token sequences for each file, and Pyverilog, an open source parsing tool dedicated to Verilog, is used to filter comments, extract keywords, operators, and identifiers contained in the code, and then obtain the syntax information of each token. Finally, the processed data set D is obtained process The corresponding Token sequence of each code file and the corresponding grammatical category of each Token; Step (3) traverses the Token sequence and identifies the line breaks in it, which allows the Token sequence to retain the line number information in the original code file, and then generates the input Token-target Token pair for each code file based on the sliding window method; specifically, the sliding window size is set to N w Lines, that is, for each code file, if the total number of code lines is N c Less than N w +1 row, then directly extract the first N c -1 and the last line as input Token-target Token pairs; otherwise, the beginning N of the code is selected w The Nth row is used as the input token. w +1 row as the target Token; then slide the window down to generate the next input Token and the corresponding target Token, where the next input Token is the 2nd to Nth row. w +1 line, target Token is Nth w +2 rows, and so on, until the window cannot slide any more; finally, we get the structured data set D structure ; Step (4) Design a non-autoregressive encoder-decoder model architecture; the encoder adopts a self-attention mechanism, and in order to solve the problem of token-by-token generation delay in the autoregressive model, a non-autoregressive decoder is designed, which generates tokens in parallel during the decoding process; introduce a mixed grammar-guided sampling strategy, for this purpose, design a shared decoder for the model to perform secondary decoding during training, that is, use the same decoder to perform an additional decoding during the training process; in this additional decoding, the input of the second decoding comes from the input of the first decoding mixed with part of the true value, relying on the mixed grammar-guided sampling strategy; Step (5) transforms the structured dataset D obtained in step (3) into structure Divide into training set D train , validation set D val and the test set D test , use the model architecture designed in step (4) to train, stop training after reaching the set training rounds, and obtain the trained model M; Step (6) Use the model M trained in step (5) to test the set D test Perform reasoning on the generated code; Step (7) evaluates the quality of the model-generated code based on the code results generated by reasoning in step (6) from four dimensions: BLEU-4, edit similarity ES, exact match accuracy EM, and latency.

2. The Verilog line-level code completion method based on a non-autoregressive model according to claim 1, characterized in that: Step (1) specifically includes the following steps: 1-1) Collect Verilog code files from open source hardware projects and industrial-grade design cases to obtain an original dataset D containing several Verilog code files. raw ; 1-2) The initially collected code files contain a large number of duplicates, so the original dataset D is filtered raw After the code files with more than 20,000 characters are compared, the code files with completely duplicate content are removed; 1-3) Cleaning to remove non-Verilog code data: Verilog code may contain one or more of the following basic keywords: always, assign, always_ff, always_comb, always_latch, so remove the code that does not contain any of the above basic keywords, thereby removing all non-Verilog code; 1-4) Use the MinHash algorithm to remove similar codes and set the Jaccard similarity threshold to 0.8; finally, we obtain a processed data set D consisting of several code file data. process .

3. The Verilog line-level code completion method based on a non-autoregressive model according to claim 1, characterized in that: Step (2) specifically includes the following steps: 2-1) Design a regular expression, treat spaces and special characters as split points, filter spaces and retain special characters; traverse the processed data set D process For all code files in the , the above-designed regular expression is used to split the code into Token sequences for each file; finally, the processed dataset D process For each code file in , a corresponding code Token sequence set consisting of several Token sequences can be obtained; 2-2) Use the Python open source library Pyverilog for parsing Verilog code to automatically filter comments and parse the tokens contained in the corresponding token sequence obtained in step 2-1) of each code file in turn, locate the identifiers and keywords, obtain the syntax information of each token and establish an index with the corresponding token; finally obtain the processed data set D process The corresponding Token sequence of each code file and the corresponding grammatical category of each Token.

4. The Verilog line-level code completion method based on a non-autoregressive model according to claim 1, characterized in that: The non-autoregressive encoder-decoder architecture in step (4) is specifically described as follows: Encoder: Add an extra token and use the output of the token on the encoder to predict the target length; in addition, there is a set of attention heads on the output of the encoder; Soft copy process: According to the target Token length predicted by the encoder output, the encoder output S = {s1, s2, ..., s m } is mapped so that the length becomes the target Token, which is used as the input of the decoder H = {h1,h2,..,h m }, the method is shown in the following formulas (1) and (2): W ji =W j W i (i+s i ) (2) Among them, W i , W j , W ji is the weight; W ji Depends on the distance relationship between source position i and target position j; these weights are updated during model training, but only the second decoding process updates the parameters; Shared parameter decoder: There are two decoding processes in the training process of the model, but only one shared parameter decoder is used, and there is another set of self-attention heads on the generated code; the way the first decoding process generates the predicted token is as follows, as shown in the following formula (3): Where X is the input sequence, is the prediction sequence, f decode (·) is the decoding process, f encode (·) is the encoding process, soft-copy (·) is the soft copy process, θ1 is the encoder parameter, and θ2 is the decoder parameter; In the inference phase, the prediction token generated at this time is the predicted target sequence; in the training phase, it is only the initial prediction token, and a second round of decoding is required; the accuracy of the initial prediction token is used to evaluate the difficulty of fitting the current target, which is used to determine the number of samples during the sampling process; the input token of the second round of decoding comes from the mixed real value sampling of the input token of the first round of decoding; the second round of decoding uses the same decoder to enable the model to learn the remaining tokens that are not selected; Hybrid grammar-guided sampling strategy: The sampling process is divided into two steps. First, the sampling number N is determined, and then N tokens are sampled from the true value according to the hybrid grammar-guided sampling strategy; The first step is to calculate the number of samples N, as shown in the following formula (4): Among them, λ is the sampling rate, which is a hyperparameter, Represents the target sequence Y and the predicted sequence The Hamming distance between is the token predicted at the current time step t, y t is the corresponding true value; In the second step, a random probability p is introduced into the sampling process. With a probability of 1-p, sampling is completely random. For the remaining probability of p, sampling needs to consider the corresponding grammatical type, that is, the ratio of keywords, identifiers, and operators is 2:1:

1. A total of N real values ​​are sampled according to this ratio.

5. The Verilog line-level code completion method based on a non-autoregressive model according to claim 1, characterized in that: The training process of step (5) is as follows: 5-1) In step (3), the data that meets the model input has been obtained and divided into training set D according to the proportion train , validation set D val and the test set D test ; 5-2) First, in the encoding process, the predicted target Token length is embedded into the first position of the encoder input, and in the soft copy process, the predicted value is used to correct the target Token length in a mapping manner; That is, the output of the encoder S = {s1, s2, ..., s m } is mapped to the decoder input H = {h1,h2,..,h m }, the mapping process is shown in the following formulas (5) and (6): W ji =W j W i (i+s i ) (6) Among them, W i , W j , W ji is the weight; W ji Depends on the distance relationship between source position i and target position j; all weights are trained during the model training process; 5-3) The result obtained from the soft copy in step 5-2) is used as the input of the first decoding process, and the output obtained is the preliminary prediction Token; 5-4) In step 5-3), a preliminary prediction Token has been obtained. Its accuracy indicates the difficulty of fitting the current target. As the training progresses, the sampling strategy gradually reduces the number of sampled true values, and eventually the model learns how to predict the entire code snippet in one traversal without seeing any true values. The number of samples N is dynamically adjusted by the Hamming distance between the preliminary prediction Token and the true sequence. The dynamic adjustment method is shown in the following formula (7): Among them, λ is the sampling rate, which is a hyperparameter, Represents the target sequence Y and the predicted sequence The Hamming distance between is the token predicted at the current time step t, y t is the corresponding true value; 5-5) After determining the number of samples N, a mixed grammar guided sampling strategy with a ratio of 2:1:1 for keywords, identifiers and operators is selected for sampling; a random probability p is introduced, stipulating that there is a probability of 1-p, and sampling is completely random; for the remaining probability p, the N real values ​​sampled must satisfy the ratio of keywords, identifiers and operators of 2:1:1; 5-6) After the sampling process is completed, the input sequence of the second decoding process is obtained by mixing. The parameters of the model will only be updated in the second decoding process. The ultimate goal is to maximize the following loss function, which is shown in the following formula (8): Where X is the input sequence, Y is the target sequence, is the predicted sequence in the first decoding process, T is the grammatical type corresponding to the target sequence, θ is the model parameter, and y t is the Token generated at the current time step t, is a subset of Tokens not selected in the sampling strategy. It is the posterior probability of generating the target Token under the condition that the input sequence is X and the model parameter is θ, combined with the sampling strategy.

6. The Verilog line-level code completion method based on a non-autoregressive model according to claim 1, characterized in that: Step (6) is the reasoning process of the model. Steps 5-2) and 5-3) are repeated again, and the output Token is the final reasoning result.

7. The Verilog line-level code completion method based on a non-autoregressive model according to claim 1, characterized in that: In step (7), the evaluation index is defined as follows: BLEU-4 is a variant of the BLEU indicator. It evaluates the generation quality by calculating the n-gram match between the generated code and the reference code. It measures the similarity between the generated text and the target text in vocabulary and local structure. The calculation method is shown in the following formula (9): Among them, p n Indicates the ratio of the number of matching n-grams to the total n-grams in the generated text, Yes n The weight of , BP is the short penalty factor, which is used to penalize the situation where the generated text is too short; The edit similarity ES measures the character-level modification cost between the generated code and the target code based on the Levenshtein distance and normalizes it into a similarity score; the calculation method is shown in the following formula (10), where Lev is the Levenshtein distance, Y represents the target sequence, Represents the prediction sequence: The exact match accuracy EM is the ratio of the generated code to the target code. Latency is the average time it takes for the model to take in a single input and output a suggested completion.