FPGA simulation verification tool fault report positioning method based on large language model

By combining the large language model and BUG report information, using the bidirectional adapter and cross-attention network model, the association between the BUG report and the source code in the FPGA simulation verification tool is established, which solves the problems of traditional methods in defect positioning and realizes high-precision defect positioning and automated positioning capabilities.

CN120179527APending Publication Date: 2025-06-20DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510240941.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

FPGA simulation verification tools face many challenges in defect location, including difficult to capture cross-language semantic associations, complex software and hardware interaction scenarios, multi-threaded and distributed computing environments, etc., making it difficult for traditional methods to achieve high-precision defect location.

Method used

Using a method based on a large language model, combining the context information and the structured characteristics of the source code in the BUG report, a vector representation containing semantic information is generated through a pre-trained large language model, and a dynamic association between the BUG report and the source code is established through a bidirectional adapter and a cross-attention network model, the similarity and dependency relationship are captured, and the file error probability is finally obtained through the activation function.

Benefits of technology

It realizes automatic positioning of defect code in the FPGA simulation verification tool software package without running test cases, significantly improving the accuracy and robustness of defect positioning, and is suitable for traditional software systems and heterogeneous code environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179527A_ABST
    Figure CN120179527A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of software engineering, and particularly relates to an FPGA simulation verification tool fault report positioning method based on a large language model. The method focuses on two widely used open source FPGA simulation verification tools, namely IVerilog and Verifier, and comprises the following steps: firstly, constructing a project data set, and carrying out preliminary screening on the data set to reduce the range of a positioning file; corresponding BUG reports and code files in the data set are input into a pre-trained large language model to obtain vector representation containing semantic information, and on the basis, a bidirectional adapter model is trained to obtain hidden state vectors containing complete context information in a suspicious mode; training a cross attention network model according to a BUG report vector and a code vector to establish a relation between the two different input sequences, capturing the similarity and dependency relationship of the two different input sequences, finally obtaining the file error probability through an activation function, and obtaining a suspicious file ranking list and the corresponding probability through sorting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software engineering, and particularly relates to a method for locating faults in FPGA simulation verification tool fault reports based on large language models, which is particularly suitable for automated defect location in FPGA simulation verification tool software packages. Background Art

[0002] In the entire life cycle of software development, software maintenance is the most time-consuming and costly stage, accounting for about 40% of the entire development workload. After the software is delivered, it enters the maintenance stage, and developers need to locate and fix defects according to the defect reports submitted by users. However, with the continuous increase in software scale and complexity, the number of defects has increased exponentially, and the defect tracking system needs to process a large number of defect reports every day. Such submitted defect reports are also called BUG reports, which are detailed records of software defects or abnormal behaviors by developers, testers, or users. It usually includes defect descriptions, reproduction steps, environmental information, error logs, etc., providing key clues for developers to locate problems. Through BUG reports, developers can clarify the nature of the defects, quickly reproduce the problems, and narrow down the location range. The traditional manual defect location method relies on developers to gradually search and locate the error files according to BUG reports, with low efficiency and difficult to meet the maintenance requirements of modern software systems. Therefore, how to automatically locate defect codes through BUG reports has always been a research hotspot in the field of software engineering.

[0003] Existing automated defect location technologies using BUG reports mainly rely on text similarity matching to locate defects by comparing the text similarity between BUG report content and source code files. The most commonly used method is to represent BUG reports and source code files as high-dimensional vectors through the Vector Space Model (VSM), and usually use feature extraction technologies such as TF-IDF (Term Frequency-Inverse Document Frequency) or BM25 to capture keyword information in the text. By calculating the cosine similarity between vectors, the similarity between the BUG report and the source code file is measured, so as to screen out the code files most relevant to the defect description as the candidate set. However, the traditional method has significant limitations: BUG reports usually describe problems in natural language, while source code is a structured programming language, and there is a semantic gap between the two, resulting in the difficulty of the traditional method to capture deep semantic associations; at the same time, word vectors or embedding representations only focus on shallow matching at the text level, ignoring the context information and structured features of the code (such as function call relationships, variable passing paths, and module dependencies), thus making it difficult to effectively model complex logics and resulting in low defect location accuracy.

[0004] In recent years, large language models (LLMs) have made remarkable progress in the fields of natural language processing and code understanding. These models are trained on large-scale corpora and possess powerful text understanding, generation, and reasoning capabilities, enabling them to capture complex patterns and semantic relationships in language. In the field of software engineering, LLMs can not only analyze the syntax and semantic structure of code but also assist developers in locating potential problems by learning the context information of code and error logs. This ability provides a new technical means for defect localization, especially showing significant advantages when dealing with complex software systems.

[0005] In the field of integrated circuit design, simulation verification is a crucial step in ensuring the correct functionality of chips. However, with the continuous increase in design complexity, defect localization in FPGA simulation verification tools faces numerous challenges. The most obvious difference from previous software packages is that the development code of FPGA simulation verification tools is usually written in high-level languages such as C++, while the compiled and executed code is a hardware description language such as Verilog. This heterogeneity of software and hardware codes makes it difficult for traditional text similarity matching methods to effectively capture cross-language semantic associations, increasing the difficulty of defect localization. Moreover, FPGA simulation verification tools need to handle a large number of software-hardware interaction scenarios, involving complex timing logic, signal transmission, and state transitions. The complexity of this interaction makes the manifestation forms of defects diverse, and it is difficult to accurately locate the root cause of problems through simple text matching. In addition, the operating environment of FPGA simulation verification tools usually involves complex mechanisms such as multi-threading and distributed computing, further increasing the difficulty of defect reproduction and localization. At the same time, there are significant differences in the implementation principles, code structures, and operating mechanisms among different FPGA simulation verification tools (such as IVerilog and Verilator), lacking a unified defect localization framework, which poses a huge challenge to the development of general solutions. Finally, the defects of FPGA simulation verification tools are often closely related to hardware design specifications, simulation accuracy, and the internal implementation details of the tools. Traditional methods based on text or code similarity are difficult to fully exploit these deep-level association information, resulting in insufficient localization accuracy. Summary of the Invention

[0006] To solve the problem that it is difficult to locate defects in FPGA simulation verification tools, the present invention proposes an efficient and accurate defect location method by leveraging the powerful semantic understanding ability of large language models, combining the context information in BUG reports and the structured features of source code. Specifically, it is a static defect location method for FPGA simulation verification tool software packages based on large language models combined with BUG report information, which can automatically locate the defective code in the software package without running test cases. Due to different simulation verification environments, it may be difficult to reproduce different test cases or additional dependencies need to be installed. The present invention combines the context information in BUG reports and the structured features of source code through large language models, does not require running test cases, effectively overcomes the limitations of traditional methods, realizes high-precision defect location, and provides strong technical support for the maintenance of FPGA simulation verification tools. The present invention focuses on two widely used open-source FPGA simulation verification tools, IVerilog and Verilator. First, a project dataset is constructed and the dataset is initially screened to narrow the scope of the location files. The corresponding BUG reports and code files in the dataset are input into a pre-trained large language model to obtain vector representations containing semantic information. On this basis, a bidirectional adapter model is trained to suspiciously obtain hidden state vectors containing complete context information. A cross-attention network model is trained according to the BUG report vector and the code vector to establish a connection between these two different input sequences, capture their similarities and dependencies, and finally obtain the file error probability through an activation function, and obtain a suspicious file ranking list and its corresponding probability through sorting.

[0007] This method is not only applicable to traditional software systems, but also can effectively handle the common heterogeneous code environments (such as C++ and Verilog) in FPGA simulation verification tools, providing strong technical support for software maintenance and defect repair.

[0008] The technical solution of the present invention:

[0009] A method for locating fault reports of FPGA simulation verification tools based on large language models, the specific steps are as follows:

[0010] Step (1): Collect BUG reports and source code files of FPGA simulation verification tools as a dataset and conduct preliminary screening. Download the source code files of two open-source tools, IVerilog and Verilator, from the official GitHub website, and screen the source code files to obtain their actual executable code files, removing description files and non-executable files that are irrelevant to the actual executed code. Then use the Faiss open-source tool to achieve efficient similarity search to complete the screening.

[0011] The collection of BUG reports is to crawl the reports labeled "BUG" on the Issues interface of the above Github project. The Faiss open-source tool is good at performing fast searches in high-dimensional vector spaces and can quickly find the "most similar" content to the given content in a large amount of data. It is mainly used to find other vectors that are most similar to the given vector. Therefore, it can quickly screen out the source code files related to BUG reports as the alternative set of error-prone files, improve the calculation efficiency, and the Faiss search can ensure that the screened results are likely to contain the source code files that actually cause errors, thus enhancing the pertinence and effectiveness of the analysis.

[0012] Step (2): Use the pre-trained large language model to convert the screened BUG reports and source code files into vectors, and then add a bidirectional adapter to increase the attention of the vectors to context information to obtain the final hidden state vectors. Among them, the bidirectional adapter is added after the large prediction model is processed. An encoder layer without a causal mask is designed in it, mainly to eliminate the influence of the representation that the vectors of the large language model only focus on the previous information and lack the subsequent information, so that the finally generated vectors can contain complete context semantic information, thus more comprehensively capturing the logical structure and dependency relationships of the code.

[0013] The large language model trained on a large-scale dataset can deeply understand the text semantics and convert it into vectors in a high-dimensional space. These vectors not only capture the surface features of the text (such as keywords, grammatical structures), but also contain deep semantic information (such as code logic, function call relationships, etc.), providing high-quality feature representations for subsequent defect localization. However, existing large language models add a "Causal Mask" during training - restricting the visible range of the model to prevent the model from seeing future data. The causal mask makes the model only focus on the current token and the tokens before it, while ignoring the subsequent context information. However, in the defect localization task, the subsequent information (such as function return values, subsequent logical judgments, etc.) is also crucial. The lack of context information may reduce the accuracy and interpretability of defect localization. In practical applications, debugging engineers often need to pay attention to the propagation relationships between suspicious statements and their contexts, and these relationships help to understand the cause of failure and accurately locate defects.. To solve the above problems, in this step, on the basis of the existing large language model, a new encoder layer is added, and the causal mask is removed in it. This design enables the model to simultaneously focus on the context information before and after the current token, thus generating hidden state vectors containing complete context information.

[0014] Step (3): Train a Cross-Attention Model based on the hidden state vectors that fuse the BUG reports and source code containing complete context information, and output the probability value of defect localization.

[0015] Cross-Attention is an extended attention mechanism in the Transformer architecture. It can establish dynamic associations between two different input sequences, capturing their semantic similarities and dependencies. In this method, the cross-attention mechanism realizes the deep fusion of the features of both by calculating the attention weights between the BUG report vector and the source code vector. Specifically, the cross-attention model assigns an attention score to each BUG report vector to measure its correlation with the source code vector, thereby highlighting the key code segments related to the defect. Next, the fused vector output by the cross-attention mechanism is input into a fully connected neural network to further extract high-level feature representations. The output of the model is transformed into an interpretable probability value through the Sigmoid activation function, and then the suspicious error file ranking list is obtained through sorting.

[0016] Furthermore, step (1) specifically includes the following steps:

[0017] 1-1) Find the source projects of IVerilog and Verilaotr from the official Github website for download as the initial set of source code files. Write a script to screen according to the file names to obtain the actual executable code files, only keep the executable code files such as ".c", ".cpp", ".h", etc., and exclude non-code-related description files and non-executable files such as README, license files, and test scripts.

[0018] 1-2) Write a crawler script to extract the issue reports marked as "BUG" from the Issues interface of the IVerilog and Verilator projects on GitHub. These reports usually contain key information such as specific problem descriptions encountered by users, reproduction steps, environment information, and error logs. The first problem description of each BUG report is extracted and used as the main input data for defect localization.

[0019] 1-3) Combine the code files actually modified by developers when fixing BUGs with the corresponding BUG reports to construct a positive sample dataset, and add the label "1" to these samples. These positive samples reflect the real defect-code correspondence and provide a reliable supervision signal for model training.

[0020] 1-4) Since the number of error-free files in the source code files is much larger than the number of actually error-prone files, directly combining them with BUG reports to construct negative samples will lead to an overly large dataset size and increase unnecessary computational overhead. To solve this problem, the Faiss tool is used to efficiently screen the source code files:

[0021] The specific implementation of Faiss screening is as follows:

[0022] a) Use Faiss to construct a vector index of the source code files, and calculate the similarity between files based on the hidden state vectors.

[0023] b) For each BUG report, use Faiss to perform k-Nearest Neighbors search, and filter out the top K source code files (K can be set to 30) that are most relevant to the BUG report.

[0024] c) Combine the filtered files that are different from the actual error-causing code file with the BUG report to construct a negative sample dataset, and add the label "0" to these samples.

[0025] Experiments show that the screening results are likely to include the source code files that actually cause errors, and at the same time significantly reduce the number of negative samples. Through the above steps, a high-quality and moderately sized defect localization dataset is constructed, providing a reliable data basis for subsequent model training and defect localization.

[0026] Furthermore, step (2) specifically includes the following steps:

[0027] 2-1) Input the BUG reports and corresponding source code files in the dataset obtained in step (1) into the same pre-trained large language model (such as the CodeGen model), and use the in-depth understanding of text semantics by the large language model to convert them into hidden state vectors in a high-dimensional space.

[0028] 2-2) Add a new encoder layer and remove the causal mask restriction, enabling the large language model to access the context information before and after the current token simultaneously. In the specific implementation, this encoder layer will recalculate the attention weights of each token, no longer shielding the influence of subsequent tokens, thereby generating a vector representation containing complete context information. After being processed by this encoder layer, the output vector will contain the complete context information (including the previous and subsequent texts) of the current token, thus more comprehensively reflecting the semantics and logical structure of the code.

[0029] Furthermore, step (3) specifically includes the following steps:

[0030] 3-1) Construct a cross-attention network to establish a dynamic semantic association between the BUG report sequence and the code sequence.

[0031] The implementation of the cross-attention network algorithm is as follows:

[0032] First, the correlation between the BUG report and the code sequence is measured by calculating the attention scores. Next, the attention scores are normalized to obtain the attention weights, which represent the correlation distribution between the BUG report and the code vectors. The code vectors are weighted and summed according to the attention weights to generate an output vector that fuses the semantic information of the BUG report and the code. Finally, the output vector is mapped to a probability value through the Sigmoid function for sorting to generate a list of suspicious file rankings.

[0033] The inputs to the cross-attention network are:

[0034] Query vector (Q): The hidden state vector from the BUG report, with a shape of (data batch size b, sequence length l1, model dimension d);

[0035] Key (K) and Value (V): The hidden state vectors from the source code file, with a shape of (data batch size b, sequence length l2, model dimension d);

[0036] a) Calculate the attention scores:

[0037]

[0038] The attention scores are calculated as shown in Equation (3.1), where Q×K T represents the dot product similarity between the BUG report and the code sequence, which is used to scale the scores to prevent gradient vanishing or explosion. The attention scores measure the correlation between each token in the BUG report and each token in the code.

[0039] b) Calculate the attention weights (α):

[0040] α = softmax(Attention Scores) (3.2)

[0041] The attention scores are Softmax-normalized to obtain the attention weights, which represent the correlation distribution between each BUG report token and the code tokens. The higher the weight, the stronger the correlation.

[0042] c) Weighted summation to generate the output:

[0043] Output = α×V (3.3)

[0044] The code Value vector (V) is weighted and summed according to the attention weights to obtain the interactive representation Output of the BUG report for the code. This output vector fuses the semantic information of the BUG report and the code sequence, providing a high-quality feature representation for subsequent probability calculations.

[0045] (3-2) The probability is obtained through mapping by the activation function. The Sigmoid function maps the value of the fusion vector to the interval [0, 1], and the output probability value Prediction represents the error probability of each source code file.

[0046] Prediction = Sigmod(Output) (3.4)

[0047] (3-3) According to the calculated probability values, the probability values of all suspicious source code files are sorted in descending order to generate a ranking list from high to low. The higher the ranking of the file, the higher the possibility of containing defects. According to actual requirements, select the top K files as the candidate set (such as Top-10). These files will be preferentially recommended to developers for further analysis and repair.

[0048] Compared with the prior art, the present invention has the following advantages and effects:

[0049] By combining the large language model and BUG report information, the present invention proposes a method for locating FPGA simulation verification tool fault reports based on the large language model, which has significant advantages compared with traditional methods. Traditional defect localization techniques usually rely on running test cases to reproduce problems. However, during the simulation verification process, different test environments may make it difficult to reproduce defects or require installing additional dependencies, increasing the complexity and cost of localization. The present invention realizes defect localization without running test cases by statically analyzing BUG reports and source code files, effectively avoiding the problem of environmental dependence. At the same time, by utilizing the powerful semantic understanding ability of the large language model, combining context information and code structural features, a hidden state vector containing complete context information is generated, significantly improving the accuracy and robustness of defect localization. In addition, by dynamically modeling the dependence relationship between the BUG report and the code sequence through the cross-attention network, and combining attention weights and probability values to generate an interpretable localization result, it provides an intuitive analysis basis for developers.

[0050] The present invention also quickly screens out source code files related to BUG reports through efficient similarity search techniques (such as Faiss), significantly reducing the computational overhead and improving the processing efficiency. Compared with traditional methods that treat BUG reports and source code files as plain text, the present invention can effectively capture the context information and structural features in the code, and is particularly suitable for heterogeneous code environments (such as C++ and Verilog) commonly found in FPGA simulation verification tools. In addition, the present invention focuses on two widely used open-source FPGA simulation verification tools, IVerilog and Verilator, and has carried out targeted optimizations for their code structures and defect characteristics, with broad applicability and practical application value. By combining the powerful semantic understanding ability of large language models and the dynamic dependency modeling of cross-attention networks, the present invention not only improves the accuracy and efficiency of defect localization, but also enhances the interpretability of the results, providing strong technical support for the maintenance of FPGA simulation verification tools.

[0051] This method also has strong scalability and flexibility, can adapt to the needs of different FPGA simulation verification tools, and can integrate multiple large language models to further improve performance. This characteristic not only reduces the threshold of technology application, but also enables more developers to quickly adopt and deploy this technology, thereby effectively improving the overall stability and reliability of FPGA simulation verification tools. Compared with traditional methods, the present invention has significant advantages in terms of positioning accuracy, computational efficiency, environmental dependence, interpretability, and applicability, providing a new solution for defect localization in FPGA simulation verification tools, and having important theoretical significance and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a schematic flowchart of the method for locating fault reports of an FPGA simulation verification tool based on a large language model according to the present invention.

[0053] Figure 2 is a sub-diagram of the process of constructing a data set and performing preliminary screening in the method for locating fault reports of an FPGA simulation verification tool based on a large language model according to the present invention.

[0054] Figure 3 is a sub-diagram of the process of feature extraction in the method for locating fault reports of an FPGA simulation verification tool based on a large language model according to the present invention.

[0055] Figure 4 is a sub-diagram of the process of cross-computing to obtain probabilities in the method for locating fault reports of an FPGA simulation verification tool based on a large language model according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0056] The method of the present invention will be described in detail below in conjunction with the drawings, technical solutions, and embodiments.

[0057] As Figure 1 shown, the present invention proposes a method for locating faults in FPGA simulation verification tool fault reports based on large language models. Its process is mainly divided into three steps: First, collect BUG reports and source code files of two open-source FPGA simulation verification tools, IVerilog and Verilator, from the official GitHub website, remove irrelevant files through screening, and use the Faiss tool for efficient similarity search to quickly screen out source code files related to the BUG report as a candidate set; Second, use a pre-trained large language model to convert the BUG report and source code file into high-dimensional vectors, and generate hidden state vectors containing complete context information by adding an encoder layer that removes causal masks to capture the context dependencies of the code; Finally, establish a dynamic association between the BUG report vector and the source code vector based on the cross-attention mechanism, calculate the attention weights and fuse the features, output the defect location probability value through the Sigmoid function, and generate a ranked list of suspicious error files.

[0058] Taking the IVerilog defect #765 as an example, the implementation details of each process are described in detail. The specific implementation is as follows:

[0059] (1) Collect BUG reports and source code files of the FPGA simulation verification tool as a data set and conduct preliminary screening. Download the source code files and BUG reports of the two open-source tools, IVerilog and Verilator, from the official GitHub website, and screen the source code files to obtain their actual executable code files, removing description files and non-executable files irrelevant to the actual executed code. Then use the open-source Faiss tool to achieve efficient similarity search and complete the screening.

[0060] 1.1. Crawl the BUG report of defect #765 with the label "BUG" on the Issues interface of the IVerilog project on the Github website as:

[0061] "Summary":"Error in determine width of the unbased unsized constant'1in port assignment",

[0062] "Description":"Example code:\nmodule mod0(input[3:0]x);\ninitialbegin\n$display(\"%b\",x);\n$finish;\nend\nendmodule\n\nmodule test;\n mod0m('1);\nendmodule\nOutput:\ntest.sv:9:warning:Port 1(x)of mod0 expects 4bits,got 1.\ntest.sv:9::Padding 3high bits of the port.\n0001\n\nOther simulators(Verilator,Modelsim)and synthesizers(Yosys,Quartus,Vivado)assigns value'b1111to port."

[0063] The corresponding actual error source code is in the iverilog / elaborate.cc file.

[0064] 1.2. For the source code library of the IVerilog tool, the total number of original files is 5,746, which includes not only the actually executable code files, but also non-code related description files and non-executable files such as README documents, license files, and test scripts. To build a high-quality dataset, we first screened the source code files and only retained the executable files directly related to code implementation, specifically including files with suffixes ".c", ".cc", ".cpp", ".h", ".hpp", ".py", ".ts", ".js", ".java". After screening, a total of 647 executable code files were obtained, and these files constitute the core dataset for subsequent defect localization.

[0065] 1.3. Next, use the Faiss tool for efficient similarity search to quickly screen out the source code files most relevant to the BUG report. As Figure 2 shown, the internal process of Faiss search is as follows:

[0066] a) First, input the BUG report and the screened source code files into the pre-trained CodeBERT model to generate high-dimensional vector representations. For each BUG report, the generated vector shape is (1, 768), representing the summary of its semantic information; while for all 647 screened source code files, the generated vector shape is (647, 768), and each vector corresponds to the semantic representation of a source code file.

[0067] b) Next, input the vector representation of the source code files into Faiss to construct an efficient vector index. Faiss organizes these vectors by optimizing data structures (such as IVF index or HNSW graph structure) to support fast similarity search. For the vector representation of each BUG report, use Faiss to perform k-Nearest Neighbors (k-NN) search to find the source code file vectors that are most similar to the BUG report vector. Faiss measures the similarity by calculating the Euclidean distance between vectors, and the smaller the distance, the higher the similarity. Finally, Faiss returns the 30 source code files that are most similar to each BUG report as the candidate set. The 30 code files most relevant to the IVerilog defect #765 are shown in Table 1:

[0068] Table 1 Ranking table of the 30 code files most relevant to the IVerilog defect #765

[0069] File Name Euclidean Distance iverilog / vvp / vpi_const.cc 7.7632 iverilog / PScope.cc 7.7984 iverilog / elaborate.cc 8.4328 iverilog / PNamedItem.h 9.0148 iverilog / PEvent.h 9.1445 iverilog / vpi / sys_vcd.c 9.3562 iverilog / PModport.h 9.7369 iverilog / version_base.h 9.8479 iverilog / netdarray.h 9.9113 iverilog / tgt-vvp / eval_object.c 10.3108 …

[0070] (2) Use a pre-trained large language model to convert the filtered BUG reports and source code files into high-dimensional vector representations, and then add a bidirectional adapter to increase the attention of the vectors to context information to obtain the final hidden state vectors. Among them, the core design of the bidirectional adapter is to remove the causal mask encoder layer, which solves the problem that the vectors generated by the large language model only focus on the previous information and ignore the subsequent information, enabling the finally generated vectors to contain complete context semantic information, so as to more comprehensively capture the logical structure and dependencies of the code. The specific implementation is as follows:

[0071] 2.1. Input the content of the BUG report of the IVerilog defect #765 into the CodeGen model to generate its vector representation. After tokenization, the BUG report content obtains 202 tokens, and each token generates a 1024-dimensional vector. Therefore, the vector shape of the BUG report is (202, 1024).

[0072] 2.2. Process the 30 candidate source code files (such as iverilog / elaborate.cc) returned by Faiss. Split the code of each file by line, and input each line of code into the CodeGen model to generate 1024-dimensional vectors. To extract the vector representation of each line, use the newline character as the delimiter, input each line of code as an independent input sequence into the CodeGen model, and extract the [CLS] token vector of each line as the representation of that line. For example, the iverilog / elaborate.cc file has m = 7724 lines of code, so its vector shape is (7724, 1024).

[0073] 2.3. To enhance the context information, an encoder layer without causal mask is added on top of the CodeGen model. This layer captures the context dependencies of the code through a bidirectional attention mechanism and generates a hidden state vector containing complete context information. Finally, the shape of the hidden state vector of the BUG report is (202, 1024), while the shape of the hidden state vectors of the 30 source code files are (m_i, 1024) respectively, where m_i is the number of lines of code in the i-th file.

[0074] (3) Based on the hidden state vectors generated in the second step, train a cross-attention network to establish dynamic dependencies between the BUG report and the source code files, and generate defect localization results through probability calculation.

[0075] The following are the detailed implementation processes and experimental results:

[0076] 3.1. For the two types of input vectors, the hidden state vector from the BUG report is used as the Query vector (Q), with a shape of (202, 1024), where 202 represents the number of tokens in the BUG report and 1024 represents the vector dimension of each token; the hidden state vectors of the 30 source code files are each used to form a Key vector (K) and a Value vector (V) for training, with a shape of (m_i, 1024), where m_i represents the number of lines of the i-th source code file and 1024 represents the vector dimension of each line of code. The vector output by weighted summation through cross-attention calculation fuses the semantic information of the BUG report and the source code files, with a shape of (202, 1024).

[0077] 3.2. Pass the output of the cross-attention network through a fully connected neural network to further extract high-level feature representations. Use the Sigmoid function to map the output to a probability value, as shown in formula (3.1), where x is the input fused vector:

[0078]

[0079] The mapped probability value represents the possibility that each source code file contains the target defect. The closer the probability value is to 1, the higher the probability that the file is incorrect.

[0080] 3.3. Sort the probability values of all source code files in descending order to generate a ranking list from high to low. Select the top K files as the candidate set (K = 10 in this experiment), and these files will be preferentially recommended to developers for further analysis and repair. Taking the IVerilog defect #765 as an example, the final list of suspicious files sorted according to the suspicious score of each file is shown in Table 2:

[0081] Table 2 List of Sorted Suspicious Files for IVerilog Defect #765

[0082] Rank File Name Probability Value Whether Defective 1 iverilog / elaborate.cc 0.92 Yes 2 iverilog / vvp / vpi_const.cc 0.87 No 3 iverilog / PScope.cc 0.85 No 4 iverilog / PNamedItem.h 0.81 No 5 iverilog / PEvent.h 0.78 No 6 iverilog / vpi / sys_vcd.c 0.75 No 7 iverilog / PModport.h 0.72 No 8 iverilog / version_base.h 0.69 No 9 iverilog / netdarray.h 0.65 No 10 iverilog / tgt-vvp / eval_object.c 0.62 No …

[0083] Among them, the probability value of the actual error file iverilog / elaborate.cc is 0.92, ranking first, indicating that it is very likely to be the source code file causing the defect. Through manual verification, it is confirmed that this file does contain error codes related to the BUG report description, verifying the effectiveness of this method.

Claims

1. A method for locating fault reports of FPGA simulation verification tools based on a large language model, characterized in that: The specific steps are as follows: Step (1): Collect the BUG reports and source code files of FPGA simulation verification tools as data sets and perform preliminary screening; download the source code files of two open source tools, IVerilog and Verilator, from the official GitHub website, and screen the source code files to obtain their actual executable code files, removing description files and non-executable files that are not related to the actual execution code; then use the Faiss open source tool to implement efficient similarity search and complete the screening; Step (2): Use the pre-trained large language model to convert the screened BUG reports and source code files to obtain vectors, and then add a bidirectional adapter to increase the vector's attention to contextual information to obtain the final hidden state vector; wherein the bidirectional adapter is added after the large prediction model is processed, and an encoder layer that removes the causal mask is designed in the encoder layer to eliminate the influence of the large language model's vector that only pays attention to the previous information but lacks the following information, so that the finally generated vector contains complete contextual semantic information; Step (3): Based on the hidden state vector of the bug report and source code that contains complete context information, a cross-attention model is trained and the probability value of defect location is output; Specifically, the cross-attention model assigns an attention score to each bug report vector to measure its correlation with the source code vector, thereby highlighting the key code snippets related to the defect; next, the fused vector output by the cross-attention mechanism is input into a fully connected neural network to further extract high-level feature representations; the output of the model is converted into an interpretable probability value through the Sigmoid activation function, and then a ranked list of suspicious error files is obtained by sorting.

2. The method for locating fault reports of FPGA simulation verification tools based on a large language model according to claim 1, characterized in that: Step (1) specifically includes the following steps: 1-1) Find the source projects of IVerilog and Verilaotr from the official Github website and download them as the initial source code file set. Write a script to filter the actual executable code files according to the file names, and only keep the executable code files, excluding non-code related description files and non-executable files; 1-2) Write a crawler script to extract problem reports marked as "BUG" from the Issues interface of IVerilog and Verilator projects on GitHub; the first problem description of each BUG report is extracted and used as the main input data for defect location; 1-3) Combine the code files actually modified by developers when fixing bugs with the corresponding bug reports, build a positive sample dataset, and add labels "1" to these samples; 1-4) Since the number of source code files without errors is much greater than the number of files with actual errors, the Faiss tool is used to efficiently screen the source code files: The specific implementation of Faiss screening is as follows: a) Use Faiss to build a vector index of source code files and calculate the similarity between files based on the hidden state vector; b) For each bug report, use Faiss to perform k-nearest neighbor search to filter out the top K source code files most relevant to the bug report; c) Combine the filtered files that are different from the actual error code files with the BUG report to construct a negative sample dataset and add a label "0" to these samples.

3. The method for locating fault reports of FPGA simulation verification tools based on a large language model according to claim 1, characterized in that: Step (2) specifically includes the following steps: 2-1) Input the BUG reports and corresponding source code files in the data set obtained in step (1) into the same pre-trained large language model, and use the large language model's in-depth understanding of text semantics to convert them into hidden state vectors in a high-dimensional space; 2-2) Add a new encoder layer to remove the limitation of causal mask, so that the large language model can access the context information of the current token at the same time; in the specific implementation, the encoder layer will recalculate the attention weight of each token and no longer mask the influence of subsequent tokens, thereby generating a vector representation containing complete context information; after being processed by this encoder layer, the output vector will contain the complete context information of the current token.

4. The method for locating fault reports of FPGA simulation verification tools based on a large language model according to claim 1, characterized in that: Step (3) specifically includes the following steps: 3-1) Construct a cross-attention network to establish dynamic semantic associations between bug report sequences and code sequences; The cross-attention network algorithm is implemented as follows: First, the correlation between the bug report and the code sequence is measured by calculating the attention score. Next, the attention score is normalized to obtain the attention weight, which represents the correlation distribution between the bug report and the code vector. The code vector is weighted and summed according to the attention weight to generate an output vector that combines the bug report and code semantic information. Finally, the output vector is mapped to a probability value through the Sigmoid function to generate a suspicious file ranking list. The input of the cross-attention network is: Query vector Q: hidden state vector from the bug report, with the shape of data batch size b, sequence length l1, model dimension d; Key value K and Value value V: hidden state vector from the source code file, with the shape of data batch size b, sequence length l2, model dimension d; a) Calculate the attention scores: Among them, Q×K T represents the dot product similarity between the bug report and the code sequence, Used to scale fractions; b) Calculate the attention weight α: α=softmax(Attention Scores)(3.2) c) Weighted summation generates output: Output = α × V (3.3) Perform weighted summation on the code Value vector V according to the attention weight to obtain the interactive representation of the bug report on the code. The output vector integrates the semantic information of the bug report and the code sequence. 3-2) Obtain the probability through activation function mapping. The Sigmoid function maps the value of the fusion vector to the interval [0,1]. The output probability value represents the error probability of each source code file. Prediction = Sigmod (Output) (3.4) 3-3) According to the calculated probability values, the probability values ​​of all suspicious source code files are sorted in descending order to generate a ranking list from high to low; the higher the ranking of the file, the higher the possibility of containing defects; according to actual needs, the top K ranked files are selected as candidate sets; these files will be recommended to developers for further analysis and repair.