Methods, apparatus, and devices for retrieving interface documents based on electronic design automation
By breaking down and filtering electronic design query information at multiple levels, and using the first and second models to select target interface document blocks, the problem of low API document retrieval performance in the electronic design automation toolchain is solved, achieving high-precision and high-quality retrieval results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN RES INST THE CHINESE UNIV OF HONG KONG
- Filing Date
- 2026-01-16
- Publication Date
- 2026-06-02
AI Technical Summary
In the field of electronic design automation, the API documentation retrieval performance of existing electronic design automation toolchains is low and the context integrity is poor, leading to retrieval difficulties.
By breaking down the electronic design question information and using the first and second models for multi-level filtering, the target interface document blocks that match the sub-task information are selected, the execution order is determined, and the corresponding answer information is generated.
It improves the accuracy of electronic design automation interface document retrieval and the quality of response information, ensuring that the selected interface document blocks are strongly related to the user's query intent and that the generated response information is comprehensive and accurate.
Smart Images

Figure CN122132524A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic design technology, and in particular to an interface document retrieval method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on electronic design automation. Background Technology
[0002] With the development of electronic design technology, Electronic Design Automation (EDA) technology has emerged. The EDA toolchain is extremely complex, and the massive amount of API (Application Programming Interface) descriptions in its documentation constitutes a huge retrieval barrier.
[0003] In traditional technologies, for complex EDA domain API document retrieval tasks, retrieval performance is usually low and context integrity is poor. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for retrieving interface documents based on electronic design automation, which can improve retrieval accuracy and enhance the quality of electronic design response information, in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for retrieving interface documents based on electronic design automation, including:
[0006] Obtain electronic design question information, which is composed of natural language; decompose the electronic design question information to obtain multiple sub-task information;
[0007] For each of the subtask information, the subtask information is input into the first model so as to retrieve the target candidate interface document block that matches the subtask information through the first model;
[0008] The target candidate interface document blocks are input into the second model. For each target candidate interface document block, the second model analyzes the semantic correlation between the target candidate interface document block and the subtask information in multiple target dimensions, and selects target interface document blocks from the target candidate interface document blocks. Based on the correlation between the subtask information, the execution order between the target interface document blocks is determined.
[0009] Based on the target interface document block and the corresponding execution order, answer information corresponding to the electronic design question information is generated.
[0010] Secondly, this application also provides an interface document retrieval device based on electronic design automation, comprising:
[0011] The acquisition module is used to acquire electronic design question information, which is composed of natural language.
[0012] The disassembly module is used to disassemble the electronic design query information to obtain multiple sub-task information;
[0013] The first filtering module is used to input the subtask information into the first model for each subtask information, so as to retrieve the target candidate interface document block that matches the subtask information through the first model;
[0014] The second filtering module is used to input the target candidate interface document blocks into the second model, and for each target candidate interface document block, analyze the semantic correlation between the target candidate interface document block and the subtask information through the second model, and filter out target interface document blocks from the target candidate interface document blocks; and determine the execution order between the target interface document blocks based on the correlation between the subtask information.
[0015] The answer module is used to generate answer information corresponding to the electronic design question information based on the target interface document block and the corresponding execution order.
[0016] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described embodiments of the method.
[0017] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the embodiments of the above methods.
[0018] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the embodiments of the above methods.
[0019] The aforementioned interface document retrieval method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on electronic design automation (EDA) decomposes EDA query information into multiple sub-task information. For each sub-task information, it filters out the target interface document block (i.e., target API) corresponding to that sub-task information. This comprehensive breakdown of the intent of complex API document retrieval tasks through user query intent reconstruction and task decomposition improves the comprehensiveness of the response information. Secondly, through multi-level filtering—inputting the candidate target interface document blocks into a second model, and then using the second model to filter out the final target interface document block—it ensures that the filtered target interface document block is strongly correlated with the EDA query information, thereby ensuring the quality of the response information generated based on the target interface document block. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is an application environment diagram of an interface document retrieval method based on electronic design automation in one embodiment;
[0022] Figure 2 This is a flowchart illustrating an interface document retrieval method based on electronic design automation in one embodiment.
[0023] Figure 3 This is a flowchart illustrating an interface document retrieval method based on electronic design automation in another embodiment;
[0024] Figure 4 This is a structural block diagram of an interface document retrieval device based on electronic design automation in one embodiment;
[0025] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0027] The interface document retrieval method based on electronic design automation provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. After terminal 102 receives electronic design query information input by the user through the corresponding input / output device, it sends the electronic design query information to server 104. Then, server 104 receives the electronic design query information, which is composed of natural language. The electronic design query information is decomposed into multiple sub-task information. For each sub-task information, the sub-task information is vectorized to obtain a sub-task vector. The sub-task vector is input into a first model to retrieve target candidate interface document blocks that match the sub-task vector through a target model. The target candidate interface document blocks are input into a second model. For each target candidate interface document block, the second model analyzes the semantic correlation between the target candidate interface document block and the sub-task information, and selects the target interface document block from the target candidate interface document blocks. Based on the correlation between the various sub-task information, the execution order between the target interface document blocks is determined. Based on the target interface document block and its corresponding execution order, generate answer information corresponding to the electronic design question.
[0028] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0029] In one exemplary embodiment, such as Figure 2 As shown, an interface document retrieval method based on electronic design automation is provided, which is then applied to... Figure 1 The following describes 104 server examples, including steps 202 to 208. Wherein:
[0030] Step 202: Obtain electronic design question information, which is composed of natural language; break down the electronic design question information to obtain multiple sub-task information.
[0031] Electronic Design Automation (EDA) refers to the engineering discipline and technical field that uses computer software tools to design, verify, simulate, and optimize integrated circuits (ICs), printed circuit boards (PCBs), and other electronic systems.
[0032] Electronic design questions refer to questions asked by users in natural language (such as Chinese and English) about how to use EDA tools to complete specific design tasks (such as circuit design, simulation, placement and routing).
[0033] Subtask information refers to several more specific and atomic task description units obtained after parsing and decomposing the original question information.
[0034] For example, electronic design questions can be input into a large model, which can then decompose the questions into several sub-tasks based on a hierarchical thinking chain.
[0035] Alternatively, the electronic design question information can be broken down based on a rule engine, such as by performing keyword matching and template filling, to decompose the electronic design question information into several sub-task information.
[0036] Step 204: For each subtask information, input the subtask information into the first model so that the first model can retrieve the target candidate interface document block that matches the subtask information.
[0037] The first model comprises a first sub-network and a second sub-network. The first sub-network is, for example, an embedding model, and the second sub-network is, for example, a reranking model. The embedding model filters candidate interface document blocks relevant to the electronic design query by calculating the similarity between the vector corresponding to the input information and the vectors corresponding to each interface document block in the electronic design interface document. The reranking model further filters target candidate interface document blocks relevant to the electronic design query by performing fine-grained semantic matching between the candidate interface document blocks input to the model and the electronic design query information.
[0038] The interface documentation block refers to the content related to an Application Programming Interface (API). An API is a set of predefined rules, protocols, and tools used to build software applications or enable different software systems to communicate and interact with each other. The interface documentation block can include a description of the API's functionality, parameter rules, sample code, etc.
[0039] Step 206: Input the target candidate interface document blocks into the second model. For each target candidate interface document block, analyze the semantic correlation between the target candidate interface document block and the subtask information in multiple target dimensions through the second model, and select the target interface document blocks from the target candidate interface document blocks. Based on the correlation between the information of each subtask, determine the execution order between the target interface document blocks.
[0040] The second model is an artificial intelligence model used to perform fine-grained analysis. Its function is to perform deep semantic understanding, logical reasoning and multi-dimensional evaluation on the candidate documents selected by the first model to determine their practical applicability.
[0041] Target Interface Document Block: After two levels of model screening and evaluation, the relevant content of the specific and usable electronic design interface that is finally confirmed to be able to solve the corresponding sub-task is finally identified.
[0042] Target dimensions include at least one of the following: functional matching, parameter compatibility, return type usability, and category relevance. Functional matching refers to the degree of alignment between the core functionality, main purpose, or specific problem solved by the candidate API documentation block and the fundamental intent and core requirements expressed in the user's question. Parameter compatibility refers to the degree of matching and adaptation between the input parameters, configuration options, and environment settings required or supported by the candidate API and the explicitly or implicit design constraints, boundary conditions, and operating environment in the user's question. Return type usability refers to the degree of alignment between the format, content, and form of the output results (such as data files, reports, database objects, and status codes) generated after the candidate API is executed and the expected outputs or downstream processing requirements explicitly or implicitly stated in the user's question. Category relevance refers to the degree of consistency between the tool category, applicable design stage, and type of abstract problem solved by the candidate API and the overall design process context, problem abstraction level, and professional domain of the user's question.
[0043] For example, after the first model filters out multiple candidate interface document blocks, the second model analyzes the correlation between each candidate interface document block and its corresponding subtask information from multiple target dimensions to filter out the target interface document blocks. This eliminates semantically inconsistent or unused noisy interface document blocks, and through multi-layered filtering, ensures that the selected target interface document blocks are strongly correlated with the electronic design questions input by the user.
[0044] Step 208: Generate answer information corresponding to the electronic design question information based on the target interface document block and the corresponding execution order.
[0045] The execution order refers to the sequence of specific operation steps represented by the target interface document block, determined by the logical dependencies (such as data flow and control flow) between subtasks. By determining the execution order, the dependencies (i.e., execution order) between various interface document blocks can be determined, thereby ensuring that the logic of the code in the generated answer information is correct.
[0046] In the aforementioned interface document retrieval method based on electronic design automation (EDA), the EDA query information is broken down into multiple sub-task information. For each sub-task information, the target interface document block (i.e., the target API) corresponding to that sub-task information is selected. This approach comprehensively deconstructs the intent of complex API document retrieval tasks by reconstructing the user's query intent and decomposing the task, thereby improving the comprehensiveness of the response information. Secondly, through multi-level filtering—inputting the candidate target interface document blocks into a second model, and then using the second model to filter out the final target interface document block—it is ensured that the selected target interface document block is strongly correlated with the EDA query information, thus ensuring the quality of the response information generated based on the target interface document block.
[0047] In some embodiments, the electronic design query information is decomposed into multiple sub-task information, including: inputting the electronic design query information into a large language model, extracting the corresponding design purpose, electronic object, and constraints from the electronic design query information through the large language model; parsing the electronic design query information based on the design purpose, electronic object, and constraints to obtain the corresponding task graph; the task graph represents each sub-task and the execution order between each sub-task; traversing each task node in the task graph, for each task node, identifying the corresponding functional intent, input / output, and constraint information, and decomposing the task node based on the functional intent, input / output, and constraint information to obtain at least one corresponding atomic task information and the dependencies between each atomic task information.
[0048] Among them, a large language model can be an artificial intelligence model based on the Transformer architecture, pre-trained on massive amounts of text data, and possessing powerful language understanding and generation capabilities, such as the GPT, LLaMA, Qwen and other series of models.
[0049] A task graph is a data structure used to represent complex tasks, typically presented in the form of a graph. Nodes in a task graph represent decomposed subtasks, and edges represent dependencies or execution order between subtasks, thus visually or formally describing the steps and logic for completing the overall task corresponding to the electronic design query.
[0050] Atomic task information is the smallest unit of task obtained by further breaking down task nodes. Its characteristics include single function, clear boundaries, and it can usually be completed directly by a specific operation or API call. It cannot be further divided, or further division would be meaningless. Atomic task information includes, for example, the specific atomic task, its identity description, execution specifications, input / output specifications, and exception handling strategies.
[0051] For example, firstly, leveraging the general knowledge understanding and reasoning capabilities of a large language model, the model performs in-depth analysis of the user's electronic design questions. Specifically, the model extracts three key elements from the question text: "design purpose" (the user's ultimate goal, such as "reducing power consumption"), "electronic object" (the target object to be manipulated, such as "clock tree network" or "a certain module"), and "constraints" (restrictions that must be met to achieve the goal, such as "using 7nm process" or "area increase not exceeding 5%"). Based on this structured information, the server constructs a "task graph" representing the complete task flow. This graph not only lists all necessary sub-tasks but also clarifies their execution order.
[0052] After obtaining a high-level task graph, each task node in the graph undergoes a more refined granular transformation. For each node, the system identifies its specific "functional intent" (what operation needs to be performed, such as "wiring" or "simulation"), "input / output" (what input data is needed and what output results are produced), and "constraint information." Based on this detailed information, the task nodes, which may still be quite complex, are further broken down into one or more "atomic task information" items, and the dependencies between these atomic tasks are clarified.
[0053] In this embodiment, a two-layer decomposition method based on task parsing and graph construction of a large language model, followed by atomic-level task splitting, enables a deep and accurate structured understanding of complex and ambiguous natural language design requests. This method refines grand problems layer by layer to the atomic task level that can be directly corresponding to a single API, thereby reducing the difficulty and ambiguity of subsequent interface retrieval.
[0054] In some embodiments, before obtaining electronic design query information, it is necessary to process the electronic design interface document in the EDA field to facilitate subsequent API retrieval. For example, the electronic design interface document is obtained; various document blocks of different types are identified in the electronic design interface document; the electronic design interface document is segmented based on these document blocks to obtain multiple interface document blocks; different interface document blocks correspond to different electronic design interfaces; and each interface document block is vectorized to obtain an embedding vector corresponding to each interface document block.
[0055] Among them, class document blocks are document fragments with the same or similar types identified based on the inherent structure and semantics of electronic design interface documents, such as "function declaration block", "parameter description block", "example code block", "notes block", etc.
[0056] Embedded vectors: Numerical representations of high-dimensional discrete data such as text and images, mapped to a relatively low-dimensional continuous vector space using neural network models. In this context, it refers to fixed-length real-valued vectors of interface document blocks processed by a vectorization model, used to calculate semantic similarity.
[0057] Specifically, the electronic design interface (EDI) document is obtained. The various class document blocks within the EDI document are identified, and the EDI document is segmented based on these class document blocks to obtain multiple interface document blocks. That is, the methods of the classes in the EDI document are segmented. Different interface document blocks correspond to different EDI interfaces. In other words, the EDI document is divided into multiple chunks. Each chunk contains only information about one API. Thus, retrieval can be performed on a chunk-by-chunk basis. Optionally, the size of each interface document block can be set to avoid excessively large chunks that could reduce retrieval efficiency. For example, the maximum number of tokens per chunk can be limited to 5000. Next, each interface document block is vectorized to obtain its corresponding embedding vector. Specifically, for each interface document block, each token is mapped to a target-dimensional vector and input into the target encoder layer to obtain the corresponding sequence representation. Then, the sequence representations corresponding to each token are normalized to obtain the final embedding vector for each interface document block. For example, for each interface document block... ,in, Indicates the first Each word is mapped to a d-dimensional vector, which is then fed into the Transformer encoder layer to obtain a sequence representation. Finally, L2 normalization is performed on the chunks according to the following formula to obtain the final embedding vectors:
[0058]
[0059] To facilitate retrieval, the format of electronic design documents can be preprocessed to the target format. For example, assuming the initial electronic design document is an HTML document, the original HTML document is first converted into a uniformly formatted structured JSON document. Subsequently, these structured JSON documents are converted into the more readable Markdown format.
[0060] In this embodiment, by preprocessing the electronic design document, the efficiency of subsequent retrieval can be improved, and by segmenting the electronic design document, it is easier to manage the relevant content of each interface involved in the electronic design document.
[0061] After preprocessing the electronic design interface (EDI) document, the efficiency of API retrieval can be improved based on the processed EDI document. For example, the first model includes a first sub-network and a second sub-network. Subtask vectors are input into the first model to retrieve target candidate interface document blocks that match the subtask vectors. This includes: inputting the subtask information into the first sub-network; vectorizing the subtask information through the first sub-network to obtain corresponding subtask vectors; calculating the similarity score between the subtask vectors and each embedded vector; determining a target similarity score from multiple similarity scores based on a first preset value; and selecting candidate similarity scores greater than or equal to the target similarity score from the multiple similarity scores. The similarity score is used to output the first candidate interface document block corresponding to each candidate similarity score. Each first candidate interface document block and the electronic design question information are input into the second sub-network to perform semantic matching between the electronic design question information and each first candidate interface document block, so as to obtain the relevance score corresponding to each first candidate interface document block. Based on the second preset value, the target relevance score is determined from multiple relevance scores. Each candidate relevance score that is greater than or equal to the target relevance score is selected from multiple relevance scores, and the target candidate interface document block corresponding to each candidate relevance score is output.
[0062] The first sub-network refers to the part of the first model responsible for encoding the query (subtask information) and the document (interface document block) into vectors respectively and calculating the similarity between them. It is usually a dual-tower encoder structure. The second sub-network refers to the part of the first model responsible for performing a deeper interactive matching between the initially selected candidate document blocks and the original query. It is usually a cross encoder or fine-ranking model structure used to calculate a more accurate relevance score.
[0063] For example, firstly, for each subtask information, the subtask information is vectorized to obtain the corresponding subtask vector. Then, the subtask vector is compared with the embedding vectors of all interface document blocks involved in the electronic design document (i.e., the embedding vectors corresponding to each API) using a similarity calculation (e.g., cosine similarity). Based on a preset similarity threshold (i.e., the first preset value), a batch of "first candidate interface document blocks" with higher scores is selected. For example, the interface document blocks with the top 15 similarities are selected. Next, the candidate document blocks obtained in the previous step are input along with the complete original electronic design query information (or subtask information). This sub-network can perform more complex semantic matching that considers contextual interaction, calculating a more accurate relevance score for each candidate document block. Again, based on a preset threshold (the second preset value), the final "target candidate interface document blocks" are selected from this batch of results. This two-level architecture, combining vector retrieval for coarse screening with interactive fine selection, balances retrieval efficiency and accuracy.
[0064] It is understandable that the process of vectorizing subtask information can be referred to in the above process of vectorizing interface document blocks.
[0065] In this embodiment, efficient and accurate preliminary screening in a large-scale EDA interface document library is achieved through document block preprocessing, vectorization, and two-level sub-network collaborative retrieval. By transforming the search problem of massive documents into efficient vector comparison, the retrieval speed is greatly improved. Through the two-stage design of "coarse screening + fine ranking", while ensuring the rapid recall of potentially relevant results from hundreds of thousands or even millions of documents (high recall rate), a more complex model is used to finely distinguish a small number of candidates, effectively improving the quality of the final candidate set (high precision rate) and avoiding sending a large number of irrelevant documents to subsequent stages with higher computational costs.
[0066] In practical applications, to ensure that the first model used is suitable for electronic design automation (EDA) scenarios, a dedicated training process is required for the general model. In this embodiment, fine-tuning the general model can both accelerate model convergence and ensure the correctness of the model's basic functions. For example, before obtaining electronic design query information, a first initial sub-network is obtained; a first training set corresponding to the first initial sub-network is obtained, the first training set including multiple first training pairs, each first training pair including a training task, and a positive training interface document block and a negative training interface document block corresponding to the training task; the network parameters of the target network layer in the first initial sub-network are adjusted to obtain a first intermediate sub-network; multiple training pairs are input into the first intermediate sub-network, and the training task is processed based on the parameters of the target network layer in the first intermediate sub-network to obtain a predicted interface document block; a contrastive loss is calculated based on the similarity between the predicted interface document block and the corresponding positive training interface document block and negative training interface document block; a regularization loss is calculated based on the network parameters of the first initial sub-network and the network parameters of the first intermediate sub-network; the network parameters of the target network layer in the first intermediate sub-network are adjusted based on the regularization loss and the contrastive loss; the above steps are repeated until the regularization loss and the contrastive loss reach the target conditions to obtain the first sub-network.
[0067] The first initial sub-network refers to the initial state of the neural network model used to perform the initial vector retrieval task. It is usually a base model (such as BERT, RoBERTa, etc.) pre-trained on a large-scale general text corpus, which has general language understanding and text representation capabilities, but has not yet specifically learned the professional semantic relationship between interface documents and queries in the field of electronic design automation (EDA).
[0068] The first training set is a dataset specifically used to train the first sub-network (vector retrieval model). Its core organizational form is "training pairs".
[0069] The first training pair constitutes the basic data unit of the first training set. Each training pair simulates one retrieval task and includes:
[0070] Training task: A text simulating a user's question, usually a natural language query collected or constructed from real-world problems in the EDA field;
[0071] Positive Training Interface Document Block: Documentation fragments of EDA interface APIs that are functionally strongly related to and correctly matched to the current training task;
[0072] Negative training interface document block: Document fragments of EDA interface APIs that are irrelevant, weakly relevant, or functionally inconsistent with the current training task, used to provide counterexamples and teach the model to distinguish between them.
[0073] The target network layers are those layers that are selected during the fine-tuning (training) process and whose parameters (weights and biases) are allowed to be updated; not all parameters of the entire model are adjusted.
[0074] Contrastive loss is a loss function used within the contrastive learning framework. Its core idea is to allow the model to learn to reduce the distance between positive sample pairs (related queries and documents) in the representation space, while increasing the distance between negative sample pairs (irrelevant queries and documents). A commonly used example is the InfoNCE loss.
[0075] Regularization loss is an additional constraint introduced during training to prevent the model from overfitting the training data and to encourage the model to learn more generalized or concise patterns. Here, it specifically refers to L1 regularization loss, which acts on the amount of change in model parameters and aims to encourage sparsity in parameter updates.
[0076] It is understandable that the core training process is an iterative optimization process. For example, in each iteration, the first initial sub-network is first initialized with "sparseness": only the parameters of a pre-selected subset of "target network layers" in the model can be adjusted, while the parameters of the remaining layers remain frozen, thus obtaining the "first intermediate sub-network". Then, a batch of training pairs is input into this first intermediate sub-network. The first intermediate sub-network retrieves the predicted interface document blocks corresponding to the training task; for example, it calculates the similarity between the various interface document blocks involved in the training task and the electronic design document, and selects the predicted interface document blocks. Subsequently, two key losses are calculated.
[0077] For contrastive loss: The similarity between the predicted interface document block and the positive training interface document block and the negative training interface document block is calculated separately to obtain the contrastive loss. In other words, the model is encouraged to assign higher similarity scores to positive documents and lower scores to negative documents.
[0078] Specifically, a pre-trained Qwen3-Embedding-8B model is loaded. Then, a SWIFT sparse update mechanism is introduced, where SWIFT only updates the parameters of key layers in the model, and L1 regularization ensures sparsity. The fine-tuned objective function is:
[0079]
[0080] in, is the regularization coefficient used to control sparsity. S is the set of selected layer indices. - (Parameter update amount). In other words, Characterizing the total loss, Indicates comparative loss, This represents the regularization loss.
[0081] The formula for contrast loss can be found below:
[0082]
[0083]
[0084]
[0085]
[0086] in, This is used to filter out false negatives. In the EDA field, due to the sheer number of APIs, many have similar functions. When related documentation for such APIs is used as a negative training interface document block, it should be considered a false negative. That is, during model training, we avoid penalizing relevant but suboptimal choices to better suit real-world engineering scenarios.
[0087] in, This represents the total number of samples in the training batch. This represents the index of the i-th training sample in the current training batch. This represents the user query (training task) in the i-th training sample. This indicates the corresponding positive training interface document block. express Compared with positive samples cosine similarity, This represents the temperature parameter (which controls the smoothness of the probability distribution). This represents the encoder function (used to map text to vectors). This indicates the number of negative training interface document blocks. This represents the Kth negative training interface document block. This corresponds to a false negative.
[0088] For regularization loss: it is calculated based on the changes in the model parameters themselves. It calculates the difference (L1 norm) between the parameter values of the target layer in the current intermediate sub-network and the corresponding parameter values of the initial sub-network before training, in each iteration. This aims to constrain the magnitude of such changes, prevent the model from completely "forgetting" its basic language understanding capabilities acquired during pre-training due to learning domain-specific knowledge, and encourage updates to focus on a few of the most critical network connections.
[0089] Finally, the contrastive loss and regularization loss are weighted and summed to obtain the total loss. Using the backpropagation algorithm, the gradient is calculated based on this total loss, and only the parameters of the "target network layers" are updated. This process is repeated, and the model gradually learns to accurately match EDA queries and relevant API documentation during iterations, while its overall parameter structure remains relatively stable. When the values of the two loss functions decrease to a preset threshold or no longer decrease significantly (the target condition is met), training stops, and the model at this point is the final, production-ready first sub-network.
[0090] During the training process of the first sub-network, the AdamW optimizer can be used to optimize the model parameters, including the learning rate. The training rounds are 5. The update rules are as follows:
[0091]
[0092] Only when , Update when the preset sparsity threshold is 0.01. (Parameters of the first initial sub-network), otherwise set It is 0.
[0093] The above describes the fine-tuning process for the first sub-network. The fine-tuning process for the second sub-network can be found below: Obtain the second initial sub-network; obtain the second training set corresponding to the second initial sub-network. The second training set includes multiple second training pairs, each containing training instructions, training interface queries, training interface documents, and training labels. Training instructions characterize the processing of training interface queries and training interface documents; training labels are divided into two categories, used to characterize whether training interface queries and training interface documents match; input the training instructions, training interface queries, and training interface documents into the second initial sub-network, and process these elements through the second initial sub-network to obtain predicted labels; calculate the model loss based on the predicted labels and training labels, and adjust the network parameters of the second initial sub-network based on the model loss until training is complete, thus obtaining the second sub-network.
[0094] For example, the training of the second sub-network employs a supervised learning paradigm, specifically a binary classification task. The training data consists of a four-tuple, including a "training instruction" (defining the task), a "training interface query," a "training interface document," and a "training label" ("match" or "not match"). During training, the instruction, query, and document are concatenated and input into the second initial sub-network. The model outputs a predicted label (or matching probability), and the network parameters are adjusted by calculating the "model loss" (such as cross-entropy loss) between the predicted label and the true label. This training enables the second sub-network to learn to make accurate binary judgments about the relevance between the query and the document based on detailed contextual information.
[0095] Specifically, for the second sub-network, the training objective is modeled as a binary classification task. Each second training sample contains an instruction I, a query q, a document d, and a label l ∈ {"yes", "no"}. The second sub-network model calculates the label probability through language modeling and minimizes the negative log-likelihood. The loss function for the second sub-network is as follows:
[0096]
[0097] in, Indicates training instructions. This indicates a query to the training interface. This refers to the training interface documentation. This represents the training labels. The model must output "yes" or "no" as the basis for result evaluation. The learning rate of the AdamW optimizer. The training round is 1.
[0098] In this embodiment, the first sub-network, by combining contrastive learning and sparse fine-tuning, can efficiently train a powerful semantic retrieval model on labeled data in the EDA domain, accurately grasping the semantic relationships of EDA terminology. Moreover, due to the protection provided by sparse fine-tuning, catastrophic forgetting is avoided, and the original general language understanding ability of the pre-trained model is preserved, thus making it more generalizable. The second sub-network, through direct binary classification supervised training, can obtain the ability to perform high-precision relevance discrimination on preliminary candidates.
[0099] In some embodiments, to improve the robustness of the model, multiple checkpoints can be saved during the training of the first and second sub-networks, and the final first and second sub-networks can be obtained by merging the models. For example, during the training of the first initial model, multiple checkpoints are obtained; different checkpoints correspond to the model parameters of the first model at different stages during training; the first initial model corresponds to either the first initial sub-network or the second initial sub-network; an interpolation angle is obtained from a preset range; the weights corresponding to each checkpoint are obtained based on the interpolation angle; the checkpoints are fused based on the weights to obtain fused model parameters; and the first model is obtained based on the first initial model corresponding to the fused model parameters.
[0100] Checkpoints are periodic snapshots of files containing information such as all current model parameters and optimizer state, saved during model training. They represent the "knowledge state" of the model at a specific training moment.
[0101] In spherical linear interpolation, the interpolation angle is a parameter used to control the position of two interpolated vectors (in this case, model parameters) along the interpolation path. Different angles correspond to different weight combinations.
[0102] The fusion model parameters are a new set of model parameters obtained by weighting and combining the model parameters saved at multiple checkpoints according to a specific algorithm (such as spherical linear interpolation).
[0103] For example, in this embodiment, the last checkpoint at the end of training is not directly used as the final model. Instead, multiple checkpoints saved at different stages of training are utilized. These checkpoints represent the "knowledge state" of the model at different training periods, such as the model in the early rapid learning stage, the model in the mid-term stable convergence stage, and the model before potential overfitting in the later stage. The specific fusion method uses "spherical linear interpolation". First, a specific "interpolation angle" is selected from a preset range (e.g., 0 to 90 degrees). Then, the corresponding weights for fusing the parameters of the two checkpoints at that angle are calculated based on trigonometric functions (e.g., sine and cosine). Finally, the calculated weights are used to weight and combine the parameters of the two (or more, recursively) checkpoints to generate a new set of "fusion model parameters". The model initialized with this set of parameters is the first model (or its sub-network) finally used for deployment.
[0104] Specifically, let's assume two checkpoints. and The parameter matrices are respectively and The parameters after fusion are:
[0105]
[0106] in, α∈[0,π / 2] It is the interpolation angle. It is the Frobenius norm. This fusion method can effectively integrate the model characteristics of different training stages, and it performs particularly well in the EDA field for embedding and ranking long documents.
[0107] In this embodiment, by performing spherical linear interpolation fusion on multiple checkpoints on the training trajectory, random fluctuations during the training process can be effectively smoothed, reducing the dependence of the final model performance on the randomness of the training stopping point. By fusing the generalization tendency of the early model and the fitting ability of the later model, an enhanced model that is more balanced and stable in terms of generalization and accuracy is obtained. In complex tasks such as EDA long document understanding, this integration strategy helps the model to comprehensively grasp features of different granularities and aspects, thereby exhibiting more robust and outstanding performance.
[0108] In some embodiments, the second model is used to analyze the semantic correlation between target candidate interface document blocks and subtask information, and target interface document blocks are selected from the target candidate interface document blocks. This includes: obtaining target evaluation rules, whereby the target evaluation criteria include at least one of functional matching, parameter compatibility, return type usability, and category relevance for interface documents; for each target candidate interface document block, the second model is used to process the target candidate interface document block and subtask information to obtain a reference value corresponding to the target evaluation rule; a reference score is calculated for each target candidate interface document block based on the reference value corresponding to the target evaluation rule; and target interface document blocks are selected based on the reference scores.
[0109] Among them, the target evaluation rules are a set of predefined, multi-dimensional quantitative standards used to comprehensively evaluate whether an interface document is suitable for a specific task from an engineering practice perspective.
[0110] For example, each initially retrieved "target candidate interface document block" is processed independently. The current "subtask information" and a "target candidate interface document block" are submitted as input to the second model (e.g., Qwen3 32B). Based on its built-in domain knowledge and reasoning capabilities, the second model performs in-depth analysis and comparison of these two elements according to preset target evaluation rules. For instance, the model determines whether the interface's "optimization" function matches the subtask's "repair" intent (functionality matching), checks whether the interface supports the "28nm" process mentioned in the subtask (parameter compatibility), analyzes whether the interface's output "database file" facilitates the user's acquisition of the required "optimization report" (return type applicability), and confirms whether the "physical design" tool to which the interface belongs is suitable for the user's current "post-layout" stage (category relevance). For each rule, the second model calculates a specific "reference value."
[0111] After obtaining reference values for all dimensions, these reference values are integrated into a single "reference score" through a predetermined calculation logic (e.g., weighted summation after assigning different weights to different dimensions). This score comprehensively reflects the overall suitability of the candidate interface in multiple key aspects. Finally, the system sorts and compares the reference scores of all candidate interface document blocks, and selects the final set of "target interface document blocks" based on preset thresholds or selection principles (e.g., selecting Top-N or scores above a certain standard). This process converges a large number of possible candidates to a few technically reliable and practically usable optimal options.
[0112] In this embodiment, by using a large model to perform in-depth analysis on each candidate interface document block obtained from the initial retrieval, the accuracy and practicality of API recommendations can be greatly improved through four-fold filtering of function, parameters, output, and context, ensuring that the final results are not only literally relevant, but also truly solve user problems and integrate into their workflow.
[0113] In some embodiments, to improve the accuracy of the response information, the response information can be validated. For example, duplicate APIs can be removed from the filtered list, or similar APIs can be prioritized. Specifically, see the following: If similar API document blocks with the same function are identified in the target API document block, the priority information corresponding to the similar API document blocks is determined based on the semantic correlation between the similar API document blocks and the subtask information, and the similar API document blocks are sorted based on the priority information.
[0114] Among them, similar interface document blocks with consistent functionality refer to two or more interface documents in the selected set of "target interface document blocks" that describe the same core functions and solve the same or very similar fundamental problems. These interfaces are usually different implementations of the same function. For example, for the function of "performing static timing analysis", there may be tool APIs from different vendors, or APIs from new and old versions of the same tool.
[0115] Semantic relevance refers to the degree of detail in which each similar API fits the current specific subtask information (including its unique constraints, context, and objectives) under the premise of functional consistency.
[0116] Priority information is a quantitative or hierarchical metric used to rank multiple similar interfaces based on their merits. It is generated based on a multi-faceted evaluation of the "degree of semantic relevance" and is used to identify which similar interface is the most suitable and recommended choice in the current specific task scenario.
[0117] For example, for similar interface document blocks with identical functions, the corresponding priority information can be determined based on the semantic relevance calculated by the second model. This priority information can then be used to determine the API (e.g., the highest priority API) that the example code in the response information should use. It is understood that the output response information may include all the final filtered APIs (including similar APIs) to allow users to understand all available APIs in the API documentation.
[0118] In this embodiment, the system intelligently selects the most suitable and most likely to efficiently complete the task from multiple feasible solutions based on the specific details of the current task, thus demonstrating the "intelligence" and "personalization" of the recommendation system.
[0119] In some embodiments, before generating the answer information, the sample code generated based on the retrieved API is validated to ensure the high quality of the answer information. For example, generating answer information corresponding to the electronic design question based on the target interface document block and its corresponding execution order includes: generating sample code based on the target interface document block and its corresponding execution order; the sample code includes multiple sample code blocks; for each sample code block, identifying non-built-in functions within the sample code block; if no interface document block matching a non-built-in function can be found in the electronic design interface document, the answer information is regenerated; and / or for each sample code block, if Chinese characters are identified within the sample code block, the answer information is regenerated.
[0120] The example code is a program code snippet generated based on a defined target interface document block and its execution order, demonstrating how to call these interfaces to complete the task. It typically uses the scripting language of a specific EDA tool.
[0121] A sample code block is a logical paragraph or a specific function / command call unit in the generated sample code.
[0122] Non-built-in functions are functions or command names that appear in the generated sample code, are not part of the scripting language's standard library or core syntax, and require references to external libraries or specific tool environments.
[0123] For example, to ensure the seriousness and usability of the generated code, this method introduces a dual verification mechanism. The first verification targets "non-built-in functions": the system checks every non-built-in function or command appearing in the code to confirm whether a matching documentation block exists in the pre-processed electronic design interface documentation library. If a called function is found to lack corresponding official documentation support (i.e., it may be a fictitious API generated by a model "illusion"), the code is considered unreliable, triggering a regeneration process. The second verification targets the "purity" of the code: EDA tool scripts typically require the use of English characters and specific syntax. The system checks for Chinese characters in the generated code blocks (which may be meaningless mixing or explanatory text mistakenly entering the code area). If found, a regeneration is also triggered. These two verifications can be used individually or in combination.
[0124] In this embodiment, by concretizing the answer information into example code and introducing a strict code verification mechanism, the direct executableness and high reliability of the solution deliverables are achieved.
[0125] In one specific embodiment, a knowledge base is constructed by acquiring the original HTML documents related to electronic design (i.e., the aforementioned electronic design documents). Fine-tuning training data is then built to fine-tune the Embedding model (i.e., the aforementioned first initial sub-network) and the Reranker model (i.e., the aforementioned second initial sub-network), resulting in the trained Embedding model (i.e., the aforementioned first sub-network) and the trained Reranker model (i.e., the aforementioned second sub-network). With this preparation completed, the user's question information can then be processed and corresponding answers generated based on the pre-built knowledge base and the pre-tuned models.
[0126] Specifically, refer to Figure 3 The process involves acquiring user-generated question information (i.e., the aforementioned electronic design question information), rewriting / splitting this information, and efficiently breaking down the entire task into several atomic, ordered sub-questions (corresponding to the aforementioned decomposition of the electronic design question information into multiple sub-task information). Next, it receives a set of sub-question lists with atomic tags produced by the rewriting / splitting module, and uses a multi-stage funnel-style filtering mechanism to construct a preliminary set of effective API document content (i.e., the aforementioned target candidate interface document blocks) as preliminary search results using the trained Embedding and Reranker models. Specifically, for search filtering, the sub-questions are first embedded, and a hierarchical coarse search is performed according to certain priority tags to extract the Top 15 document blocks (i.e., the aforementioned first candidate interface document blocks). Subsequently, these candidate document blocks are sent to the Reranker for deep relevance fine-tuning to obtain the Top 3 results (i.e., the aforementioned target candidate interface document blocks). Finally, the top 3 document blocks are input into the LLM filter (i.e., the second model mentioned above) to perform semantic consistency checks to eliminate logical noise and inconsistencies, and finally output the most suitable set of valid API document content (i.e., the target candidate interface document blocks mentioned above).
[0127] Next, Qwen3 32B is used as the Generator to generate an initial answer based on the candidate chunks obtained from the multi-level filtering module. The core function of this Generator is to generate an initial answer containing sample code and API descriptions, including the API interface name, API interface description, and sample code. First, it determines whether the user's question requires reference materials. If not, the user's question is answered directly; if so, relevant information is extracted from the reference materials, and a step-by-step answer is generated. During the generation process, the Generator must strictly adhere to priority rules: multiple interfaces implementing the same function are selected and ordered according to a certain priority order. If no relevant information is found in the reference materials, the model will provide a standardized rejection answer to avoid fabricating non-existent API names.
[0128] After generating an initial response, the system performs error checking: First, it extracts all example code blocks from the Generator's output text. Then, it performs two core hard checks on each code block: API dependency verification, which checks if all non-built-in functions used in the code are present in the system's supported API list; and language specification verification, which checks if the code snippet contains Chinese characters. If any code block fails either check (e.g., using an unsupported API or containing Chinese comments), the entire output is deemed invalid, and a detailed analysis of all errors is generated. This error analysis is then used to guide the Generator in precise code logic and syntax fixes until the example code in the Generator's generated response passes the checks.
[0129] Finally, the sample answer is optimized using the preliminary search results obtained from the multi-level filtering module to ensure consistency and completeness between the final output answer and the retrieved chunk. The module receives the preliminary search results and the sample answer as input. Its main functions include: optimization alignment, i.e., using the preliminary search results to perform final optimization and formatting on the sample answer; noise reduction, i.e., removing API documents or context blocks not referenced in the final answer; evidence supplementation, i.e., identifying APIs that appear in the answer but were not captured by the preliminary search and triggering supplementary searches to obtain their complete documents; finally, the system integrates and encapsulates the sample answer with the proofread complete documents as the final output of the API retrieval task.
[0130] In this embodiment, by utilizing a large-model-driven query intent reconstruction and task decomposition module, the user's macro-level subjective intent is transformed into atomic, high-precision EDA terminology. This solves the granularity mismatch problem caused by traditional RAG's reliance on single semantic matching, significantly improving the retrieval recall and contextual completeness of complex queries. Secondly, by fine-tuning the Embedding and Reranker pre-trained models using EDA domain documents and combining them with a multi-level filtering mechanism, retrieval performance is greatly improved, enhancing system retrieval performance from the source while reducing the illusionary tendency of LLM. By integrating a rule-driven closed-loop verification mechanism, the generated sample code undergoes syntactic and logical verification, and feedback-based automatic correction is implemented, ultimately ensuring that the generated API descriptions and sample code are accurate and meet user needs. Comparative experiments show that for complex EDA domain API document retrieval tasks, the recall rate increased from 78.53% to 86.67%, and the average reciprocal ranking score increased from 0.5794 to 0.7074, demonstrating significant improvement.
[0131] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0132] Based on the same inventive concept, this application also provides an interface document retrieval device for implementing the interface document retrieval method based on electronic design automation as described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the interface document retrieval device based on electronic design automation provided below can be found in the limitations of the interface document retrieval method based on electronic design automation described above, and will not be repeated here.
[0133] In one exemplary embodiment, such as Figure 4 As shown, an interface document retrieval device 400 based on electronic design automation is provided, including: an acquisition module 401, a disassembly module 402, a first filtering module 403, a second filtering module 404, and a response module 405, wherein:
[0134] The acquisition module 401 is used to acquire electronic design question information, which is composed of natural language.
[0135] The disassembly module 402 is used to disassemble the electronic design question information to obtain multiple sub-task information.
[0136] The first filtering module 403 is used to input the subtask information into the first model for each subtask information, so as to retrieve the target candidate interface document block that matches the subtask information through the first model.
[0137] The second filtering module 404 is used to input the target candidate interface document blocks into the second model. For each target candidate interface document block, the second model analyzes the semantic correlation between the target candidate interface document block and the subtask information, and filters out the target interface document blocks from the target candidate interface document blocks. Based on the correlation between the subtask information, the execution order between the target interface document blocks is determined.
[0138] The answer module 405 is used to generate answer information corresponding to the electronic design question information based on the target interface document block and the corresponding execution order.
[0139] In some embodiments, the electronic design query information is decomposed into multiple sub-task information. Specifically, the decomposition module 402 is used to input the electronic design query information into a large language model, extract the corresponding design purpose, electronic object, and constraints from the electronic design query information through the large language model; parse the electronic design query information based on the design purpose, electronic object, and constraints to obtain the corresponding task graph; the task graph represents each sub-task and the execution order between each sub-task; traverse each task node in the task graph, and for each task node, identify the corresponding functional intent, input / output, and constraint information; decompose the task node based on the functional intent, input / output, and constraint information to obtain at least one corresponding atomic task information and the dependencies between each atomic task information.
[0140] In some embodiments, before acquiring electronic design query information, the interface document retrieval device 400 based on electronic design automation further includes a construction module, which is specifically used for: acquiring electronic design interface documents; identifying various types of document blocks in the electronic design interface documents; segmenting the electronic design interface documents based on the various types of document blocks to obtain multiple interface document blocks; different interface document blocks correspond to different electronic design interfaces; and performing vectorization processing on each interface document block to obtain the embedding vector corresponding to each interface document block.
[0141] In some embodiments, in inputting a subtask vector into a first model to retrieve target candidate interface document blocks that match the subtask vector, the first filtering module 403 is specifically configured to: input the subtask information into the first sub-network; vectorize the subtask information through the first sub-network to obtain corresponding subtask vectors; calculate the similarity score between the subtask vector and each of the embedded vectors; determine a target similarity score from multiple similarity scores based on a first preset value; and filter out each similarity score greater than or equal to the target similarity score from the multiple similarity scores. The candidate similarity scores are selected, and the first candidate interface document blocks corresponding to each candidate similarity score are output. Each first candidate interface document block and the electronic design question information are input into the second sub-network to perform semantic matching between the electronic design question information and each first candidate interface document block, so as to obtain the relevance score corresponding to each first candidate interface document block. Based on the second preset value, the target relevance score is determined from multiple relevance scores, and each candidate relevance score that is greater than or equal to the target relevance score is selected from multiple relevance scores. The target candidate interface document blocks corresponding to each candidate relevance score are output.
[0142] In some embodiments, before acquiring electronic design query information, the interface document retrieval device 400 based on electronic design automation further includes a training module, which is specifically used for: acquiring a first initial sub-network; acquiring a first training set corresponding to the first initial sub-network, the first training set including multiple first training pairs, each first training pair including a training task, and a positive training interface document block and a negative training interface document block corresponding to the training task; adjusting the network parameters of the target network layer in the first initial sub-network to obtain a first intermediate sub-network; inputting multiple training pairs into the first intermediate sub-network, processing the training task based on the parameters of the target network layer in the first intermediate sub-network to obtain a predicted interface document block, calculating a contrastive loss based on the similarity between the predicted interface document block and the corresponding positive training interface document block and negative training interface document block; calculating a regularization loss based on the network parameters of the first initial sub-network and the network parameters of the first intermediate sub-network; adjusting the network parameters of the target network layer in the first intermediate sub-network based on the regularization loss and the contrastive loss; repeating the above steps until the regularization loss and the contrastive loss reach the target conditions to obtain the first sub-network.
[0143] In some embodiments, the training module is further specifically used for: obtaining a second initial sub-network; obtaining a second training set corresponding to the second initial sub-network, the second training set including multiple second training pairs, the second training pairs including training instructions, training interface queries, training interface documents, and training labels, the training instructions being used to characterize the processing of training interface queries and training interface documents; training labels being divided into two categories, used to characterize whether training interface queries and training interface documents match; inputting the training instructions, training interface queries, and training interface documents into the second initial sub-network, processing the training instructions, training interface queries, and training interface documents through the second initial sub-network to obtain predicted labels; calculating the model loss based on the predicted labels and training labels, adjusting the network parameters of the second initial sub-network based on the model loss until training is completed, and obtaining the second sub-network.
[0144] In some embodiments, the training module is further specifically used for: obtaining multiple checkpoints during the training of the first initial model; different checkpoints correspond to model parameters of the first model at different stages during the training process; the first initial model corresponds to a first initial sub-network or a second initial sub-network; obtaining an interpolation angle from a preset range; obtaining the weights corresponding to each checkpoint based on the interpolation angle; fusing the checkpoints based on each weight to obtain fused model parameters; and obtaining the first model based on the first initial model corresponding to the fused model parameters.
[0145] In some embodiments, in analyzing the semantic correlation between target candidate interface document blocks and subtask information using a second model, and selecting target interface document blocks from the target candidate interface document blocks, the second filtering module 404 is specifically used to: obtain target evaluation rules, wherein the target evaluation criteria include at least one of functional matching, parameter compatibility, return type usability, and category relevance for interface documents; for each target candidate interface document block, process the target candidate interface document block and subtask information using the second model to obtain a reference value corresponding to the target evaluation rule; calculate a reference score corresponding to each target candidate interface document block based on the reference value corresponding to the target evaluation rule; and select target interface document blocks based on each reference score.
[0146] In some embodiments, in generating answer information corresponding to electronic design questions based on target interface document blocks and their corresponding execution order, the answer module 405 is specifically configured to: generate sample code based on target interface document blocks and their corresponding execution order; the sample code includes multiple sample code blocks; for each sample code block, identify non-built-in functions in the sample code block; if an interface document block matching a non-built-in function cannot be found in the electronic design interface document, then regenerate the answer information; and / or for each sample code block, if Chinese characters are identified in the sample code block, then regenerate the answer information.
[0147] The modules in the aforementioned interface document retrieval device based on electronic design automation can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0148] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data used for electronic design interface documents and other data for electronic design automation. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an interface document retrieval method based on electronic design automation.
[0149] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0150] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0152] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0153] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0156] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for retrieving interface documents based on electronic design automation, characterized in that, The method includes: Obtain electronic design question information, which is composed of natural language; decompose the electronic design question information to obtain multiple sub-task information; For each of the subtask information, the subtask information is input into the first model so as to retrieve the target candidate interface document block that matches the subtask information through the first model; The target candidate interface document blocks are input into the second model. For each target candidate interface document block, the second model analyzes the semantic correlation between the target candidate interface document block and the subtask information in multiple target dimensions, and selects target interface document blocks from the target candidate interface document blocks. Based on the correlation between the subtask information, the execution order between the target interface document blocks is determined. Based on the target interface document block and the corresponding execution order, answer information corresponding to the electronic design question information is generated.
2. The method according to claim 1, characterized in that, The electronic design query information is broken down into multiple sub-task information, including: The electronic design question information is input into a large language model, which extracts the corresponding design purpose, electronic object, and constraints from the electronic design question information. Based on the design purpose, electronic object, and constraints, the electronic design question information is parsed to obtain a corresponding task graph. The task graph represents each subtask and the execution order between each subtask. Traverse each task node in the task graph. For each task node, identify the corresponding functional intent, input / output, and constraint information. Based on the functional intent, input / output, and constraint information, split the task node to obtain at least one corresponding atomic task information and the dependencies between each atomic task information.
3. The method according to claim 1, characterized in that, Before obtaining electronic design query information, the following are included: Obtain the electronic design interface document; identify the various document blocks in the electronic design interface document; The electronic design interface document is segmented based on each of the aforementioned class document blocks to obtain multiple interface document blocks; different interface document blocks correspond to different electronic design interfaces. Each of the interface document blocks is vectorized to obtain the embedding vector corresponding to each interface document block; The first model includes a first sub-network and a second sub-network. The step of inputting the subtask information into the first model to retrieve target candidate interface document blocks matching the subtask information includes: The subtask information is input into the first sub-network, and the subtask information is vectorized through the first sub-network to obtain the corresponding subtask vector. The similarity score between the subtask vector and each of the embedding vectors is calculated. Based on a first preset value, a target similarity score is determined from multiple similarity scores. Each candidate similarity score that is greater than or equal to the target similarity score is selected from multiple similarity scores. The first candidate interface document block corresponding to each candidate similarity score is output. Each of the first candidate interface document blocks and the electronic design question information are input into the second sub-network, so that the electronic design question information and each of the first candidate interface document blocks are semantically matched through the second sub-network to obtain the relevance score corresponding to each of the first candidate interface document blocks; a target relevance score is determined from multiple relevance scores based on a second preset value, and each candidate relevance score that is greater than or equal to the target relevance score is selected from the multiple relevance scores, and the target candidate interface document block corresponding to each candidate relevance score is output.
4. The method according to claim 3, characterized in that, Before obtaining the electronic design query information, the following is also included: Obtain the first initial subnetwork; Obtain a first training set corresponding to the first initial sub-network. The first training set includes multiple first training pairs. Each first training pair includes a training task and a positive training interface document block and a negative training interface document block corresponding to the training task. The network parameters of the target network layer in the first initial sub-network are adjusted to obtain a first intermediate sub-network. Multiple training pairs are input into the first intermediate sub-network, and the training task is processed based on the parameters of the target network layer in the first intermediate sub-network to obtain predicted interface document blocks. A contrastive loss is calculated based on the similarity between the predicted interface document blocks and the corresponding positive and negative training interface document blocks. A regularization loss is calculated based on the network parameters of the first initial sub-network and the network parameters of the first intermediate sub-network. The network parameters of the target network layer in the first intermediate sub-network are adjusted based on the regularization loss and the contrastive loss. The above steps are repeated until the regularization loss and the contrastive loss meet the target conditions to obtain the first sub-network. Before obtaining the electronic design query information, the following is also included: Obtain the second initial subnetwork; Obtain a second training set corresponding to the second initial sub-network. The second training set includes multiple second training pairs. Each second training pair includes a training instruction, a training interface query, a training interface document, and training labels. The training instruction is used to characterize the processing of the training interface query and the training interface document. The training labels are divided into two categories and are used to characterize whether the training interface query and the training interface document match. The training instructions, the training interface query, and the training interface document are input into the second initial sub-network. The second initial sub-network processes the training instructions, the training interface query, and the training interface document to obtain predicted labels. The model loss is calculated based on the predicted labels and the training labels. The network parameters of the second initial sub-network are adjusted based on the model loss until training is completed, thus obtaining the second sub-network.
5. The method according to claim 4, characterized in that, The method further includes: During the training of the first initial model, multiple checkpoints are obtained; different checkpoints correspond to the model parameters of the first model at different stages of the training process; the first initial model corresponds to the first initial sub-network or the second initial sub-network. Obtain the interpolation angle from the preset range; The weights corresponding to each checkpoint are obtained based on the interpolation angle. The checkpoints are fused based on each of the weights to obtain the fusion model parameters; The first model is obtained based on the first initial model corresponding to the parameters of the fusion model.
6. The method according to claim 1, characterized in that, The step of analyzing the semantic correlation between the target candidate interface document block and the subtask information using the second model, and selecting the target interface document block from the target candidate interface document block, includes: Obtain target evaluation rules, wherein the target evaluation criteria include at least one of the following: functional matching of interface documentation, parameter compatibility, return type usability, and category relevance; For each of the target candidate interface document blocks, the second model is used to process the target candidate interface document block and the subtask information to obtain a reference value corresponding to the target evaluation rule; a reference score is calculated for each of the target candidate interface document blocks based on the reference value corresponding to the target evaluation rule; and the target interface document blocks are selected based on each of the reference scores.
7. The method according to claim 1, characterized in that, The step of generating answer information corresponding to the electronic design question information based on the target interface document block and the corresponding execution order includes: Example code is generated based on the target interface document block and the corresponding execution order; the example code includes multiple example code blocks. For each example code block, identify the non-built-in functions within the example code block. If no interface document block matching the non-built-in function can be found in the electronic design interface document, then regenerate the response information; and / or For each example code block, if Chinese characters are detected in the example code block, the answer information is regenerated.
8. An interface document retrieval device based on electronic design automation, characterized in that, The device includes: The acquisition module is used to acquire electronic design question information, which is composed of natural language. The disassembly module is used to disassemble the electronic design query information to obtain multiple sub-task information; The first filtering module is used to input the subtask information into the first model for each subtask information, so as to retrieve the target candidate interface document block that matches the subtask information through the first model; The second filtering module is used to input the target candidate interface document blocks into the second model, and for each target candidate interface document block, analyze the semantic correlation between the target candidate interface document block and the subtask information through the second model, and filter out target interface document blocks from the target candidate interface document blocks; and determine the execution order between the target interface document blocks based on the correlation between the subtask information. The answer module is used to generate answer information corresponding to the electronic design question information based on the target interface document block and the corresponding execution order.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.