Knowledge reasoning method and device for integrated circuit wafer process and medium
Through a multimodal chain reasoning framework and adaptive sorting method, the problems of insufficient understanding of professional knowledge and document redundancy in wafer manufacturing are solved, efficient and accurate defect detection and root cause analysis are achieved, and the knowledge reasoning ability of wafer processes is improved.
Patent Information
- Application Number
- CN202511204253.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing large-scale multimodal language models in the wafer manufacturing field suffer from insufficient understanding of professional knowledge, missing details in complex queries, and interference from redundant documents, resulting in inaccurate generation and excessive computational burden.
A multimodal chain reasoning framework (CoT) combined with an adaptive ranking method is used to dynamically decompose complex queries through logical decomposition and cross-modal fusion. The basic language model is fine-tuned using a hybrid retrieval strategy and LoRA technology to enhance the understanding of wafer processes and the accuracy of document sorting.
It significantly improves defect detection, root cause analysis, and question-answering capabilities in the wafer manufacturing field, reduces the computational burden of the model, and improves the accuracy and efficiency of answer generation.
Smart Images

Figure CN120744140A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wafer manufacturing, and in particular to a knowledge reasoning method, device and medium for integrated circuit wafer technology. Background Art
[0002] Wafer manufacturing is the critical process of processing high-purity silicon wafers into semiconductor chips. As the core link in integrated circuit (IC) production, wafer manufacturing directly affects chip performance, power consumption, and reliability. During this process, wafers undergo precision processes such as photolithography, ion implantation, thin film deposition, and etching to construct complex microcircuits. However, even tiny surface defects such as cracks, particle contamination, or uneven oxide layers can significantly affect the performance and yield of integrated circuits, and even cause device failure. As chip manufacturing processes evolve to smaller nodes (such as 3nm and below), quality control in the wafer manufacturing process faces greater technical challenges. Integrated and automated wafer defect detection, root cause analysis, and reasoning query technologies are crucial to improving the yield of semiconductor manufacturing and promoting technological advancement.
[0003] In recent years, breakthroughs in large language models (LLMs) in natural language processing (NLP) have significantly enhanced the capabilities of artificial intelligence (AI) in language understanding and generation. Multimodal large language models (MLLMs) offer new solutions for defect detection and root cause analysis in wafer analysis by integrating multimodal information such as language, vision, and audio. For example, by combining optical microscopy images and inspection log data, MLLMs can help engineers quickly locate the root cause of wafer defects. However, when applied to the integrated circuit field, these models still face the challenge of "hallucination" (generating inaccurate or false information) and have significant limitations in understanding the highly specialized domain knowledge of wafer manufacturing.
[0004] To address these issues, methods based on the Retrieval-Augmented Generation (RAG) framework have been widely used. The RAG framework improves generation accuracy and domain adaptability by retrieving relevant information from external knowledge bases and integrating it into model input. In the wafer manufacturing domain, RAG technology can leverage structured documents (such as manufacturing process standards and defect analysis reports) to optimize the model's understanding of process steps and terminology. Retrieval strategies are generally categorized into three categories: sparse retrieval, dense retrieval, and hybrid retrieval. Sparse retrieval (e.g., keyword search, BM25, TF-IDF, and inverted index) relies on an exact match between the query and document vocabulary and is computationally efficient. It works well for structured wafer knowledge documents with fixed terminology, but struggles to capture semantic similarity between terms, limiting understanding of the rich terminology of wafer processes and defect analysis. Dense retrieval (semantic search) utilizes embedding models to convert text into high-dimensional vectors and identifies relevant documents based on vector similarity. This can capture deep semantic relationships within the wafer manufacturing process, but it is computationally expensive. Hybrid retrieval combines and deduplicates the results of both retrieval methods, balancing efficiency and semantic understanding. These strategies enhance the information retrieval capability of RAG and the reasoning ability of large models in professional fields to a certain extent.
[0005] Although these strategies are effective, they still have the following limitations: 1) Insufficient understanding of specialized wafer knowledge: Wafer manufacturing involves complex processes and expertise, but existing models have difficulty understanding some uncommon terms (such as "doping" and "oxidation") in general databases, which exacerbates the illusion phenomenon; 2) Omission of details in complex queries: Direct document retrieval of complex queries (such as defect root cause analysis or surface treatment processes) is difficult to cover all key information, especially when multiple technical details are involved, which may omit important content and affect the integrity of the reasoning; 3) Interference from redundant or irrelevant documents: The documents returned by existing retrieval strategies are often lengthy and contain a large amount of irrelevant information, which may introduce text noise and weaken the reliability of the generated model.
[0006] Therefore, the present invention aims to solve the following 3 technical problems:
[0007] 1) How to assist in optimizing retrieval enhancement generation technology, enhance comprehensive understanding of complex professional tasks, and effectively alleviate the emergence of hallucination problems;
[0008] 2) How to efficiently conduct in-depth analysis of query intent and refined retrieval of wafer knowledge to improve the authenticity and accuracy of language model responses;
[0009] 3) How to provide the model with more reliable document knowledge and reduce the interference of irrelevant knowledge while reducing the computational burden of the model. Summary of the Invention
[0010] The purpose of the present invention is to provide a knowledge reasoning method, device and medium for integrated circuit wafer process to solve the problems raised in the above background technology.
[0011] To achieve the above object, the present invention provides the following technical solutions:
[0012] A knowledge reasoning method for integrated circuit wafer technology, comprising:
[0013] Step 1: Extract multimodal features from the input query wafer image to detect image surface defects, and finally generate a mask image that can intuitively represent the defect area;
[0014] Step 2: Perform wafer knowledge retrieval. In wafer knowledge retrieval, the retrieval task is broken down into multiple sub-steps based on the complexity of the user's query. In each sub-step, wafer images, defect mask information, and wafer IDs are dynamically integrated to gradually complete the search for relevant documents.
[0015] Step 3: sort the documents. During the document sorting process, the wafer knowledge search results are weighted and scored. The retrieved documents are prioritized using a scientific scoring mechanism, and the top several documents most relevant to the query are selected.
[0016] Step 4: Generate query answers. During the query answer generation process, the wafer image, defect mask, and sorted high-relevance documents are input as prompts into the basic language model. The basic language model generates accurate and explanatory answers by combining multimodal data and professional document information.
[0017] Furthermore, step 1 includes: using a frozen pre-trained encoder to encode the input wafer image and the extracted wafer ID information, and fusing visual features and text features into a multimodal vector representation.
[0018] Furthermore, the step 2 includes:
[0019] Step 2.1, logical decomposition of queries: Design a query reasoner that gradually decomposes user queries into multiple subqueries and forms a solution chain to capture key details. In this stage, hint learning is fine-tuned on a dataset of queries of varying complexity, enabling the base language model to identify whether a query needs to be decomposed and generate the corresponding subqueries.
[0020] For simple queries, the final answer is directly output; for complex queries, the length of the solution chain is dynamically adjusted to further optimize the logical plan;
[0021] Step 2.2, cross-modal fusion: Design a shared gating unit to evaluate the relevance of the wafer image, generated mask, and extracted wafer ID information to each sub-query, and select the most relevant modality information to be fused with the query vector.
[0022] Step 2.3, multi-round retrieval: In the multi-round retrieval stage, based on logical decomposition and cross-modal fusion, for each user query and its sub-queries and multimodal vectors, relevant candidate documents are gradually retrieved from the wafer knowledge base; each round of retrieval provides the necessary external knowledge hints for generating intermediate answers; a hybrid retrieval method is adopted to combine keyword-based retrieval with semantic-based retrieval. In operation, sub-queries use keyword matching to ensure accurate alignment of terms, and multimodal vectors capture complex contextual relationships through semantic retrieval.
[0023] Furthermore, the step 2.2 specifically includes:
[0024] The frozen pre-trained encoder generates an image vector, a mask vector, and a text vector. These three vectors are concatenated and mapped into a smaller latent space through a linear transformation, thereby achieving a comprehensive evaluation of multimodal information.
[0025] Subsequently, the correlation between modalities is calculated through the self-attention mechanism, and Softmax is used to assign weights to filter out the most relevant modal information for subsequent reasoning. Finally, a multimodal query vector is obtained.
[0026] Furthermore, the step 3 specifically includes:
[0027] Step 3.1, multi-dimensional scoring strategy for candidate documents: For each subquery and its multimodal query vector, three search engines (TF-IDF, BM25, and Faiss) are used to extract a set of candidate documents from the knowledge base, and each document is assigned a corresponding score.
[0028] Step 3.2, semantic density calculation and adaptive adjustment of retriever weights: Calculate the semantic density of each subquery;
[0029] Based on semantic density, the TF-IDF, BM25, and Faiss weight vectors are adaptively adjusted to achieve a dynamic balance in the scores. The final score is calculated based on the TF-IDF, BM25, and Faiss scores and weight vectors, semantic density, and the balance coefficient.
[0030] Step 3.3, accurate re-ranking of candidate documents: In the final ranking stage, the candidate documents are re-ranked according to the final scores to select the most relevant documents.
[0031] Furthermore, the step 4 includes:
[0032] The basic language model is fine-tuned using LoRA technology. Using the multimodal wafer corpus as the training basis, the basic language model is customized into a dedicated multimodal large language model. During the fine-tuning process, the input mode is expanded based on the characteristics of wafers in the integrated circuit field:
[0033] First, the mask vector extracted from the query image is used as visual input to enhance the model's accurate understanding of the defect area;
[0034] Secondly, the first several retrieved knowledge documents that are highly relevant to the query are appended to the user query as additional text input to provide rich contextual support for the model;
[0035] During the model adaptation process, LoRA technology adjusts model parameters by introducing a low-rank matrix.
[0036] Furthermore, the basic language model is a Qwen2VL-7B model.
[0037] The present invention also provides a knowledge reasoning device for integrated circuit wafer technology, comprising one or more processors for implementing the above-mentioned knowledge reasoning method for integrated circuit wafer technology.
[0038] The present invention also provides a readable storage medium having a program stored thereon, which, when executed by a processor, implements the above-mentioned knowledge reasoning method for integrated circuit wafer process.
[0039] Compared to existing technologies, the present invention offers the following advantages: It proposes a novel multimodal large language model for detection, retrieval, and question-answering tasks in the integrated circuit (IC) field. This model progressively achieves dynamic decomposition and high-precision solution of complex queries through three core stages. The introduced multimodal Chained Reasoning (CoT) framework enables the model to combine logical reasoning with multimodal information to construct a coherent solution chain for complex tasks. Furthermore, an enhanced hybrid retrieval strategy combined with an adaptive ranking method effectively improves document relevance and ranking accuracy. The present invention also utilizes LoRa technology to fine-tune a dedicated wafer corpus, enhancing the model's capabilities in defect area detection, root cause analysis, and question-answering regarding defect type, location, and quantity, while significantly reducing the model's information forgetting during multi-turn conversations. Results demonstrate that the proposed model approach performs exceptionally well in detection, retrieval, and question-answering tasks on internal wafer datasets and multiple test benchmarks (including 100 defect-related questions and general IC questions), significantly outperforming existing multimodal large language models. This invention not only provides strong support for knowledge analysis in the field of wafer manufacturing and integrated circuits, but also provides new insights and important references for the application and research of multimodal large models in the field of industrial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flow chart of a knowledge reasoning method for integrated circuit wafer technology of the present invention.
[0041] Figure 2 This is a schematic diagram of the structure of a knowledge reasoning device for integrated circuit wafer technology according to the present invention.
[0042] Figure 3 This is an example diagram of the wafer knowledge dialogue of the present invention. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0044] like Figure 1 As shown, a knowledge reasoning method for integrated circuit wafer processes consists of four core steps: wafer defect detection, wafer knowledge retrieval, document sorting, and query answer generation. The Qwen2VL-7B model is used as the basic language model. Each step plays an indispensable role in the operation, working together to achieve efficient and accurate wafer manufacturing defect detection and knowledge query. The simple invention framework steps are described as follows:
[0045] Step 1: Wafer defect detection
[0046] The wafer defect detection step is the foundation of the entire invention. In this step, multimodal features are extracted from the input query wafer image to detect surface defects, ultimately generating a mask image that can intuitively represent the defect area.
[0047] In the wafer defect detection step, the present invention uses a frozen pre-trained encoder to encode the input wafer image and the wafer ID information extracted by OCR, and fuses the visual features and text features into a multimodal vector representation. This multimodal representation method can effectively combine the information advantages of images and texts, so that the present invention can capture the subtle features of wafer surface defects. For example, for common defects such as particle contamination, scratches and film unevenness, the present invention can generate accurate defect area masks and provide quantitative defect indicators in wafer surface quality assessment. This information can not only be used for real-time quality control, but also provide data support for subsequent process optimization. Subsequently, the present invention uses a mask decoder containing four layers of upsampling to further decode the encoded information to generate an accurate defect area mask, thereby providing high-quality input data for subsequent retrieval and analysis tasks.
[0048] Step 2: Wafer knowledge retrieval
[0049] In the wafer knowledge retrieval step, the retrieval task is decomposed into multiple sub-steps according to the complexity of the user's query. In each sub-step, the wafer image, defect mask information and wafer ID are dynamically integrated to gradually complete the search for relevant documents. For example, when a user asks the query "How to reduce particle contamination in the lithography process", the present invention will first locate the historical defect data related to the lithography step, and determine the typical location and characteristics of particle contamination in combination with the mask information, and then further retrieve the best process practices related to contamination control. This step-by-step dynamic retrieval strategy can not only effectively cover all the key information of complex queries, but also optimize the retrieval accuracy in each step, ensuring that the search results are highly relevant to user needs, especially when facing highly specialized and technical fields.
[0050] In the wafer knowledge retrieval step, due to the high complexity and specialization of the wafer knowledge base, traditional retrieval methods usually rely on directly matching the query content, but this makes it difficult to effectively capture the user's potential intentions. Especially in the field of integrated circuit manufacturing, when conducting root cause analysis, it is necessary to comprehensively utilize wafer images, process data, and technical documents to locate problem points that affect yield. To address this bottleneck, we proposed a chain retrieval method of "logical decomposition-cross-modal fusion-multi-round retrieval". By refining the query, integrating multimodal information, and constructing contextual semantics, it accurately meets the user's personalized needs. The steps of the chain retrieval method are as follows:
[0051] Step 2.1, logical decomposition of the query: In the query of wafer manufacturing field, the user's needs often contain multiple intentions. Directly processing these complex queries can easily lead to information omission. To this end, the present invention designs a dedicated query reasoner, which uses Qwen2.5-14B to gradually decompose the user query Q into multiple sub-queries and form a solution chain to capture key details. At this stage, the present invention fine-tunes the data set of queries of different complexity through prompt learning, so that the Qwen2.5-14B model can identify whether the query needs to be decomposed and generate the corresponding sub-query q i ∈{q1,q2,...,q k},q i represents the i-th subquery, k represents the number of all subqueries, i.e. the length of the solution chain, and each subquery q i The generation of is determined by the following formula:
[0052]
[0053] Among them, arg max represents the parameter value that makes the objective function reach the maximum value, R i-1 For the previous subquery q i-1 The answer obtained is , where P is the conditional probability distribution. For simple queries, this method directly outputs the final answer. For complex queries, it further optimizes the logical plan by dynamically adjusting the length k of the solution chain. This strategy effectively reduces the computational cost of simple queries while ensuring detailed coverage.
[0054] Step 2.2, cross-modal fusion: In scenarios involving visual reasoning, understanding the user's query requires not only analyzing its logic but also combining relevant multimodal information. To this end, the present invention designs a shared gating unit to evaluate the relevance of the wafer image, the generated mask, and the extracted wafer ID information to each sub-query, and selects the most relevant modal information and query vector. Specifically, the frozen pre-trained encoder generates the image vector v image , mask vector v mask and text vector v id Finally, these three vectors are concatenated and mapped to a smaller hidden space through linear transformation, thereby achieving a comprehensive evaluation of multimodal information. The expression is:
[0055]
[0056] Among them, concat represents the operation of splicing data along the specified dimension, W and b represent the weight matrix and bias term respectively. F represents v image 、v mask and v id The concatenated vector, Fh Represents the final vector after mapping by linear transformation.
[0057] Then, the correlation between modalities is calculated through the self-attention mechanism, and Softmax is used to assign weights to filter out the most relevant modal information for subsequent reasoning. Finally, the multimodal query vector representation v mq The calculation is as follows:
[0058]
[0059] in, The i-th subquery vector generated by the encoder of Qwen2VL-7B, T represents the transpose, d Fh Represents vector F h Through this process, the model can effectively combine user queries with multimodal features to improve its understanding of the query context.
[0060] Step 2.3, multi-round retrieval: In the multi-round retrieval stage, the present invention uses logical decomposition and cross-modal fusion to perform a multi-round retrieval on each user query Q and its subquery q i and the multimodal vector v mq , gradually retrieve relevant candidate documents from the wafer knowledge base. Each round of retrieval provides necessary external knowledge hints for generating intermediate answers. To improve retrieval performance, the present invention adopts a hybrid retrieval method, combining keyword-based retrieval (such as TF-IDF, BM25) with semantic-based retrieval (such as Faiss). In the operation, the subquery q i Keyword matching is used to ensure accurate alignment of terms, and the multimodal vector v mq Capturing complex contextual relationships through semantic retrieval. This iterative strategy can optimize the model's global retrieval capabilities and improve the accuracy of obtaining wafer-related documents.
[0061] Step 3: Document sorting
[0062] In the document sorting step, the search results are further weighted and scored based on the knowledge retrieval. Through a scientific scoring mechanism (such as keyword matching weight based on BM25, semantic embedding similarity, and wafer defect category relevance), the present invention is able to prioritize the retrieved documents and select the top three documents most relevant to the query. For example, when processing queries involving "ion implantation doping concentration optimization", the present invention will give priority to documents containing detailed doping concentration curves and relevant experimental results. This process effectively reduces information redundancy and interference from irrelevant documents, providing highly relevant input materials for the subsequent answer generation step, thereby improving the overall reliability and accuracy of the present invention.
[0063] In the document sorting step, in the root cause diagnosis process of wafer manufacturing, quickly locating the most relevant process conditions or previous failure analysis documents is of great significance for improving problem troubleshooting efficiency and yield optimization. Wafer manufacturing is the core link in the production of integrated circuits, involving complex processes such as lithography, etching, ion implantation and thin film deposition. Any tiny defects or parameter deviations may cause chip performance degradation or even failure. In order to achieve this goal, the present invention designs an adaptive weighted sorting method, which scores the relevance of all candidate documents and selects the document with the highest score as an external prompt, thereby reducing the inference burden of the basic language model while ensuring retrieval accuracy. This method generates the final sorting result by dynamically adjusting the weight of the retriever and combining multiple scoring strategies. The specific implementation steps are as follows:
[0064] Step 3.1, multi-dimensional scoring strategy for candidate documents: For each subquery q i and its multimodal query vector v mq , the present invention will simultaneously use three search engines (TF-IDF, BM25 and Faiss) to extract the candidate document set D={d1,d2,...,d n}, where D represents the candidate document set, d1...d n Represents the candidate documents, n represents the total number of candidate documents, and finally assigns a corresponding score to each document. The following are the scoring methods of each search engine: TF-IDF score evaluates the relevance of documents based on the importance of terms, and the calculation formula is:
[0065]
[0066] Among them, S t represents the TF-IDF score, and t represents the subquery q i A term in , IDE(t) is the inverse document frequency of the term, f(t,d i ) is the term t in document d i This method is suitable for exact matching of queries containing clear terms.
[0067] The BM25 score optimizes the relevance between the query and the document by combining term frequency and document length. Its formula is as follows:
[0068]
[0069] Among them, S b represents the BM25 score, avgdl is the average length of documents in the corpus, and k1 and b1 are tuning parameters, typically set to 1.5 and 0.75, respectively. The BM25 method smoothes high-frequency terms in long documents and is suitable for scenarios with long or unstructured corpora.
[0070] The Faiss score is based on the cosine similarity of vectors to evaluate the semantic relevance between multimodal vector representations and documents. The formula is as follows:
[0071]
[0072] Among them, S f represents the score of Faiss, v mq represents the multimodal query vector, and For document d i The Faiss score is particularly suitable for capturing complex contextual relationships and semantic information.
[0073] Step 3.2, semantic density calculation and adaptive adjustment of retriever weight: In order to dynamically evaluate the semantic complexity of the query, the present invention calculates each subquery q i The semantic density f sem (q i ), which is defined as:
[0074]
[0075] Among them, cosine represents cosine similarity, and is the vector representation of terms t1 and t2. Semantic density reflects the semantic similarity between terms in subqueries. For complex queries, its value is usually higher. sem (q i ), the present invention adaptively adjusts the weight vectors W of the three retrievers t (TF-IDF), W b (BM25) and W f (Faiss), thus achieving a dynamic balance of scores. The final score is calculated as follows:
[0076] Where β is the balancing coefficient, set to 0.6. For semantically complex queries, the semantic retrieval weight of Faiss will be amplified, while for simpler queries, the weights of TF-IDF and BM25 will dominate.
[0077] Step 3.3: Accurately rerank candidate documents: In the final ranking stage, the present invention reranks candidate documents based on the final score S, selecting the most relevant documents. This method eliminates irrelevant or redundant content, retaining only high-quality documents that are relevant to the query. This not only improves the accuracy of knowledge retrieval but also significantly reduces the inference pressure on the underlying language model.
[0078] Step 4: Query answer generation
[0079] The query answer generation step is the final link in the interaction between the user and the system. In this step, the present invention inputs the wafer image, defect mask, and sorted high-relevance documents as prompts into the basic language model. The basic language model generates accurate and explanatory answers by combining multimodal data and professional document information to meet the user's professional query needs in the field of wafer manufacturing. For example, when a user asks "how to solve the problem of increased wafer scratches during chemical mechanical polishing", the present invention can not only output a detailed cause analysis (such as uneven pressure distribution or excessively large polishing liquid particles), but also provide specific improvement suggestions (such as adjusting the hardness of the polishing pad or optimizing the polishing liquid formula). In addition, the design of this step can generate detailed reasoning logic through an additional reasoning process, such as the prediction of the simulation effect of adjusting different process parameters, so that the generated answers are excellent in both technical and practical aspects.
[0080] In the query answer generation step, to meet the highly specialized needs of the wafer manufacturing field, this paper fine-tunes the Qwen2VL-7B model using LoRA technology. Using a multimodal wafer corpus as a training foundation, the model is customized into a specialized multimodal large language model (MLLM). During the fine-tuning process, the paper expands the input modality to address the specific characteristics of wafers in the integrated circuit field. First, mask vectors extracted from the query image are used as visual input to enhance the model's accurate understanding of defect regions. Second, the top three retrieved knowledge documents highly relevant to the query are appended to the user query as additional textual input, providing rich contextual support for the model. This expanded input design enables the model to better integrate information from both visual and textual modalities, thereby improving its ability to answer complex technical questions. During the model adaptation process, LoRA technology introduces a low-rank matrix to adjust model parameters, preserving the robust capabilities of the original model while significantly reducing the computational overhead of fine-tuning. This optimization results in a customized MLLM that demonstrates superior performance in wafer manufacturing-related tasks, including high-precision defect detection, complex technical Q&A, and root cause analysis, while also flexibly addressing diverse user needs. This domain-specific model provides strong support for solving technical problems and significantly improves knowledge management and problem diagnosis capabilities in the wafer manufacturing field.
[0081] Through the seamless collaboration of these four core steps, the present invention can efficiently solve complex query tasks in the wafer manufacturing field. Its technical advantages lie in accurate defect detection, comprehensive knowledge retrieval, reliable document sorting, and highly specialized answer generation capabilities.
[0082] See also Figure 2An embodiment of the present invention provides a knowledge reasoning device for an integrated circuit wafer process, comprising one or more processors for implementing a knowledge reasoning method for an integrated circuit wafer process in the above embodiment.
[0083] An embodiment of a knowledge reasoning device for integrated circuit wafer process of the present invention can be applied to any device with data processing capability, and the device with data processing capability can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capability in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, if Figure 2 As shown, it is a hardware structure diagram of any device with data processing capability where the knowledge reasoning device of integrated circuit wafer process of the present invention is located. Figure 2 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in the embodiment may also include other hardware according to the actual function of the device with data processing capabilities, which will not be described in detail.
[0084] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0085] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0086] An embodiment of the present invention further provides a readable storage medium having a program stored thereon. When the program is executed by a processor, the knowledge reasoning method for an integrated circuit wafer process in the above embodiment is implemented.
[0087] The readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0088] Example
[0089] 1. Experimental data set: The present invention randomly divides the internal wafer data set into a training set and a test set in a ratio of 7:3 for training and performance evaluation of supervised defect detection. At the same time, the basic language model is trained using a multimodal wafer corpus. In order to test the performance of the model in retrieval and question-answering tasks, the present invention uses Qwen2-72B to randomly generate 200 questions related to wafer defects and 200 general questions related to integrated circuits (ICs) based on the knowledge database and corpus. Subsequently, 100 questions from each category are selected to form a test set. In the retrieval test set, each question is associated with three related documents; in the question-answering test set, each question is accompanied by a reference answer generated by a domain expert for performance comparison and verification.
[0090] 2. Experimental details: This paper uses frozen CLIP-L / 14@336px as the embedding model. Through its powerful multimodal feature extraction capabilities, it maps image embedding to a high-dimensional feature space, providing a solid foundation for subsequent multimodal information processing. At the same time, the fine-tuned Qwen2vl-7B is selected as the language generation model. Combined with its excellent performance in text generation and multimodal tasks, it ensures that the model has high accuracy and flexibility when processing complex queries, technical questions and answers, and specialized reasoning tasks. During the model training process, the efficient AdamW optimizer (β1=0.9,β2=0.999) was used, and the initial learning rate was set to 1e -4 , and then the cosine annealing scheduling strategy is used to gradually decay the learning rate to 1e -6. This optimization method can smoothly adjust the learning rate to avoid overfitting while accelerating the convergence speed. In addition, combined with the LoRA fine-tuning strategy, the number of parameters and computational costs during model fine-tuning are greatly reduced, while retaining the original performance of the model, it effectively adapts to the professional task requirements in the wafer manufacturing field. Model training is performed on 4 A100 GPUs, with a total of 50 rounds of training, and the batch size of each batch is set to 24. Through this high-performance hardware support and refined training process, the model of the present invention can efficiently capture key information in multimodal data, and demonstrate excellent capabilities in tasks such as defect detection, knowledge retrieval, and question answering in wafer manufacturing, providing important technical support for promoting the intelligence of integrated circuit manufacturing.
[0091] 3. Experimental evaluation indicators: A multi-dimensional and multi-metric approach is used to comprehensively evaluate the performance of the model to ensure that its performance in detection, retrieval, and generation tasks meets high standards.
[0092] In inspection tasks, models are evaluated using five key metrics: image-level AUC (Image-AUC), pixel-level AUC (Pixel-AUC), proportion of defect regions (PRO), average precision (AP), and accuracy. Specifically, ① Image-level AUC evaluates the binary classification capability of the entire image in defect detection tasks. AUC (Area Under the Curve) is the area under the receiver operating characteristic (ROC) curve, indicating the model's ability to distinguish between positive and negative samples under all possible classification thresholds. Image-level AUC measures the model's accuracy in identifying whether an image contains defects. ② Pixel-level AUC measures the model's ability to detect defects at a fine-grained (i.e., pixel-level) level. Unlike image-level AUC, pixel-level AUC focuses on the classification accuracy of each pixel within the defect region. It assesses the model's classification performance at each pixel and is particularly useful for boundary detection and precise localization of defect regions. ③ Proportion of defect regions measures the ratio of the defect region to the total image area when the model detects a defect. This indicator reflects the model's coverage ability and positioning accuracy of the defect area. Generally, the higher the better, indicating that the model can better identify defects and avoid missed detections. ④ Average precision refers to the average accuracy of the model under multiple thresholds, which is often used to evaluate the performance of target detection tasks. It takes into account the performance of the model under different confidence levels, calculates the accuracy at each recall rate, and finally obtains a comprehensive evaluation value. The higher the AP, the better the overall performance of the model in the detection process. ⑤ Accuracy refers to the ratio of correctly classified samples to the total number of samples. In defect detection tasks, accuracy indicates the proportion of defects or non-defects correctly identified by the model.
[0093] In retrieval tasks, the model's retrieval capabilities are evaluated using three metrics: Recall@n, Normalized Discounted Cumulative Gain@n (NDCG@n), and Mean Rank Response Rate@3 (MRR@3). Specifically, ① Recall@n indicates the proportion of relevant documents included in the nth result returned by the model in the retrieval task. In other words, when the model only returns n results, the proportion of results that are relevant. ② Normalized Discounted Cumulative Gain@n measures the relevance of the model in the nth search result returned. This normalized score is used to compare the performance of different models. This metric takes into account the relevance and ranking of the results, with larger values indicating more accurate ranking. ③ Mean Rank Response Rate@3 refers to the average reciprocal of the position of the first relevant document in the top three search results. This metric reflects how quickly the model retrieves relevant documents, with larger values indicating earlier relevant documents.
[0094] For the evaluation of generation tasks, this paper introduces five metrics: Bilingual Evaluation Understanding Score (BLEU), Recall Benchmark-1 (ROUGE-1), Recall Benchmark-Longest Common Subsequence (ROUGE-L), Word Error Rate (Wer), Semantic Textual Similarity (STS), and Expert Accuracy, to comprehensively measure the quality of generated answers. Specifically, ① Bilingual Evaluation Understanding Score is used to assess machine translation quality, measuring the n-gram overlap between the machine-generated text and the reference text. A higher score indicates a closer match between the machine-generated text and the reference text, making it suitable for text generation tasks. ② Recall Benchmark-1 is used to evaluate metrics for text summarization or generation tasks, calculating the word-level recall between the generated text and the reference text. ROUGE-1 focuses on word recall, with a higher score indicating that the generated text contains more words from the reference text. ③ Recall Benchmark-Longest Common Subsequence measures the degree of matching of the longest common subsequence (LCS) between the generated text and the reference text. Unlike ROUGE-1, ROUGE-L takes the sequential structure of the generated text into account. A higher score indicates that the generated text is better at preserving order and structure. ④ Word error rate is used to assess the accuracy of automatic speech recognition or text generation tasks. It calculates the edit distance between the generated text and the reference text, including the minimum number of insertions, deletions, and substitutions. A lower word error rate indicates that the generated text is closer to the reference text. ⑤ Semantic text similarity measures the semantic similarity between two text segments. This metric is suitable for text question answering or text matching tasks. A higher score indicates greater semantic similarity between the two segments. ⑥ Expert accuracy rating provides a final professional assessment of the model's generated results through manual evaluation of the correctness and logic of the answers, ensuring the reliability and practicality of the generated content in the technical field. Through the comprehensive evaluation of these multi-dimensional indicators, we can fully understand the model's performance in different task scenarios, providing a scientific basis and direction for further optimization and improvement of the model.
[0095] 4. Experimental results: (1) Test results: Table 1 lists in detail the quantitative performance comparison results of the model method proposed in the present invention and other comparison models in five core indicators under the same training and testing conditions. As can be seen from the table, the model method of the present invention shows obvious advantages in all indicators, especially in the three key indicators of image-level AUC, average precision and accuracy. Among them, the image-level AUC is improved by 1.62 compared with the latest model AnomalyGPT, indicating that the model method of the present invention is more accurate in overall detection effect; the average precision is improved by 2.23, indicating that it is more superior in balancing precision and recall; and the accuracy is increased by 1.27, reflecting that the proportion of correct classification in all predictions has increased significantly, further verifying the reliability and adaptability of the model.
[0096] Table 1. Detection metric results of fully supervised experiments on an internal wafer dataset
[0097] (2) Search results: Using the same wafer knowledge database and generated test set, this paper comprehensively evaluates the performance of four retrieval-enhanced generation methods in retrieval tasks, focusing on comparing their performance across different metrics. Questions related to wafer defects typically focus on analyzing the causes of defects and developing prevention strategies. Here are two typical examples: 1) Please explain the cause of this defect; 2) How can such defects be prevented? Meanwhile, general questions related to integrated circuits (ICs) cover key aspects of process principles and manufacturing processes. Examples include: 1) Briefly describe the principles of ion diffusion and how it differs from thermal diffusion; 2) Why do wafers need to be cleaned after polishing? 3) Why is wafer planarization necessary? In a quantitative evaluation of retrieval performance, as shown in Table 2, the proposed model method achieved a 0.107 improvement in Recall@3 over the BGE-M3 method. This result demonstrates that the proposed model method achieves higher coverage when retrieving relevant documents, more effectively extracting the information required for user queries, and significantly improving the breadth and accuracy of results. By retrieving a more comprehensive set of relevant documents, the proposed model method provides users with richer background knowledge support, performing particularly well in handling complex technical queries and scenarios that incorporate multimodal data. Furthermore, as shown in Table 3, the proposed model method also demonstrates significant advantages in ranking performance metrics. In terms of the Normalized Loss Cumulative Gain@3 metric, the proposed model method achieved a 0.127 improvement over BGE-Reranker, demonstrating greater precision in ranking relevant documents and prioritizing the most relevant documents at the top of the search results. Furthermore, in terms of the Uniform Rank Recall@3 metric, the proposed model method achieved a 0.108 improvement over the BGE-M3 method. This demonstrates that the proposed model method offers significant advantages in optimizing the ranking of relevant documents and improving the overall quality of query results. The model approach presented in this paper achieves comprehensive optimization in retrieval tasks, from coverage to ranking accuracy. It ensures the completeness of the information required by the user's query while accurately identifying and ranking the most relevant documents. This performance advantage not only highlights the technological advancement of the model approach presented in this paper, but also provides an efficient and reliable solution for knowledge retrieval and complex query tasks in the wafer manufacturing field.
[0098] Table 2. Retrieval metric results on the wafer defect-related question-answering dataset
[0099] Table 3. Retrieval index results on general integrated circuit (IC) related question answering dataset
[0100] (3) Question and answer results: As shown in Tables 4 and 5, the present invention tested different methods on two test sets for question answering. The results show that the proposed model method demonstrates significant advantages across multiple metrics. For wafer-related questions, the proposed model method outperformed existing multimodal large language models across most evaluation metrics. For example, on the Recall Benchmark-1 metric, the proposed model method achieved a 0.093 improvement over Vicuna-7B, demonstrating higher accuracy in the similarity between generated content and reference answers and stronger text generation capabilities. This demonstrates that the model can more effectively capture key information and generate high-quality answers in specialized knowledge question answering tasks in the wafer manufacturing field. For general questions about integrated circuits (ICs), the proposed model method achieved a 0.410-point improvement in the key metric of expert accuracy over Qwen2VL-7B, further demonstrating its superior performance on technical and general tasks. This improvement in expert accuracy indicates that the proposed model method not only answers questions more accurately but also meets expert standards in logical reasoning and processing of technical details. This capability is particularly important for complex tasks related to IC process flow, design principles, and technical analysis.
[0101] Figure 3 The model method of the present invention further demonstrates the dialogue performance in the wafer defect knowledge question-answering scenario. As can be seen from the figure, the model is able to accurately extract and apply knowledge in the context of questions with technical depth, demonstrating strong understanding and flexibility. This adaptability makes it more prominent and reliable in question-answering tasks in the field of integrated circuits, especially when dealing with complex technical problems. It can be seen that the model method of the present invention not only performs well in professional questions related to wafer manufacturing, but also demonstrates a high degree of accuracy and versatility in general technical questions and answers about integrated circuits, providing a powerful solution for intelligent technical support in integrated circuit manufacturing and related fields.
[0102] Table 4 Question answering metric results on the wafer defect-related question answering dataset
[0103] Table 5 Question answering metric results on general integrated circuit (IC) related question answering dataset
[0104] (4) Ablation experiment:
[0105] We conducted in-depth ablation studies on two different test sets (wafer defects and general integrated circuit problems) to fully validate the effectiveness of the proposed modules. In the experiments, we gradually removed or replaced key modules in the model to evaluate the specific contribution of each module to the overall performance. For the wafer defect test set, the focus was on the module's impact on multimodal information processing and accurate defect identification; for the general integrated circuit problem test set, the module's performance in tasks such as logical reasoning, knowledge retrieval, and question-and-answer generation was evaluated. Through ablation experiments, we were able to clearly quantify the role and impact of each module in different tasks, providing strong theoretical support and data evidence for model design and optimization.
[0106] ① Retrieval Method Design: This paper conducted in-depth research under three different experimental conditions: using a fixed three-step query decomposition, not performing query decomposition, and not performing multimodal information fusion at all. By comparing these experimental conditions, the paper aimed to evaluate the specific impact of each key step on the quality of answer generation. As shown in Table 6, the experimental results clearly demonstrate that dynamic query decomposition and the integration of additional information play a crucial role in improving answer quality.
[0107] While the model can generate answers of reasonable quality using a fixed three-step query decomposition, its flexibility is significantly limited, making it difficult to handle the multiple layers of intent or underlying details in complex queries. When query decomposition is completely eliminated, the model's ability to handle complex questions declines significantly, manifesting as missing important information and incomplete answers. Furthermore, without multimodal information fusion, the model lacks a comprehensive understanding of visual and textual data, failing to fully leverage the advantages of multimodal input. This directly leads to significantly reduced accuracy and relevance of answers.
[0108] In contrast, when dynamic query decomposition and multimodal information fusion are employed, the model intelligently decomposes the query based on its complexity, gradually extracting and integrating key information relevant to the question. This dynamic adaptability enables the model to perform better when handling complex questions, while the effective fusion of multimodal information further enhances the model's understanding of context and the accuracy of answer generation.
[0109] Overall, dynamic query decomposition and multimodal information fusion not only enhance the model's flexibility and adaptability, but also provide key support for generating more comprehensive and accurate answers. These experimental results fully demonstrate the technical advantages of this method in complex task processing and provide an important reference for further optimizing the design of multimodal models.
[0110] Table 6 Ablation results of the Chain of Thought (CoT) retrieval method on wafer defect-related and general integrated circuit question answering datasets
[0111] ② Ranking Method Design: This paper systematically validates the effectiveness of three sorters and ranking calculation methods through carefully designed experiments. The experiments aim to evaluate the performance of different sorters in search results and analyze the impact of the ranking calculation methods on overall efficiency and accuracy. As shown in Table 7, the experimental results demonstrate that the proposed ranking method outperforms the average level of existing methods in many aspects, demonstrating higher efficiency.
[0112] Specifically, the experiment compared the performance of three rankers when processing queries of different complexity, including the accuracy of document relevance ranking, ranking efficiency, and support for retrieval tasks. The experimental results show that the proposed method can not only complete the priority ranking of candidate documents more quickly, but also dynamically adjust the weights during the sorting process to ensure that the most relevant documents are ranked at the top of the retrieval results. This ability is particularly significant when dealing with complex queries, significantly reducing the interference caused by irrelevant or low-quality documents. In addition, in the verification of the ranking calculation method, the comprehensive ranking mechanism proposed in the present invention achieves performance optimization by combining the results of multiple rankers. Compared with a single ranker, this method can more comprehensively utilize the multi-dimensional features of documents, such as keyword matching, semantic similarity, and multimodal information fusion features. This multi-level ranking calculation method makes the final ranking result more accurate and significantly improves efficiency.
[0113] It can be seen that the ranking method proposed in this paper achieves a good balance between efficiency and accuracy, overcoming the shortcomings of traditional ranking methods in highly complex tasks, and providing strong technical support for improving retrieval performance and optimizing user experience. Experimental results fully demonstrate the superiority and practical application value of this method, especially when handling complex technical document ranking tasks.
[0114] Table 7 Ablation results of retriever and ranking methods on wafer defect related and general integrated circuit question answering datasets
[0115] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A knowledge reasoning method for integrated circuit wafer technology, characterized in that: include: Step 1: Extract multimodal features from the input query wafer image to detect image surface defects, and finally generate a mask image that can intuitively represent the defect area; Step 2: Perform wafer knowledge retrieval. In wafer knowledge retrieval, the retrieval task is broken down into multiple sub-steps based on the complexity of the user's query. In each sub-step, wafer images, defect mask information, and wafer IDs are dynamically integrated to gradually complete the search for relevant documents. Step 3: sort the documents. During the document sorting process, the wafer knowledge search results are weighted and scored. The retrieved documents are prioritized using a scientific scoring mechanism, and the top several documents most relevant to the query are selected. Step 4: Generate query answers. During the query answer generation process, the wafer image, defect mask, and sorted high-relevance documents are input as prompts into the basic language model. The basic language model generates accurate and explanatory answers by combining multimodal data and professional document information.
2. The knowledge reasoning method for integrated circuit wafer process according to claim 1, characterized in that: The step 1 includes: using a frozen pre-trained encoder to encode the input wafer image and the extracted wafer ID information, and fusing visual features and text features into a multimodal vector representation.
3. The knowledge reasoning method for integrated circuit wafer process according to claim 1, characterized in that: The step 2 includes: Step 2.1, logical decomposition of queries: Design a query reasoner that gradually decomposes user queries into multiple subqueries and forms a solution chain to capture key details. In this stage, hint learning is fine-tuned on a dataset of queries of varying complexity, enabling the base language model to identify whether a query needs to be decomposed and generate the corresponding subqueries. For simple queries, the final answer is directly output; for complex queries, the length of the solution chain is dynamically adjusted to further optimize the logical plan; Step 2.2, cross-modal fusion: Design a shared gating unit to evaluate the relevance of the wafer image, generated mask, and extracted wafer ID information to each sub-query, and select the most relevant modality information to be fused with the query vector; Step 2.3, multi-round retrieval: In the multi-round retrieval stage, based on logical decomposition and cross-modal fusion, for each user query and its sub-queries and multimodal vectors, relevant candidate documents are gradually retrieved from the wafer knowledge base; each round of retrieval provides the necessary external knowledge hints for generating intermediate answers; a hybrid retrieval method is adopted to combine keyword-based retrieval with semantic-based retrieval. In operation, sub-queries use keyword matching to ensure accurate alignment of terms, and multimodal vectors capture complex contextual relationships through semantic retrieval.
4. The knowledge reasoning method for integrated circuit wafer process according to claim 3, characterized in that: The step 2.2 specifically includes: The frozen pre-trained encoder generates an image vector, a mask vector, and a text vector. These three vectors are concatenated and mapped into a smaller latent space through a linear transformation, thereby achieving a comprehensive evaluation of multimodal information. Subsequently, the correlation between modalities is calculated through the self-attention mechanism, and Softmax is used to assign weights to filter out the most relevant modal information for subsequent reasoning. Finally, a multimodal query vector is obtained.
5. The knowledge reasoning method for integrated circuit wafer process according to claim 1, characterized in that: The step 3 specifically includes: Step 3.1, multi-dimensional scoring strategy for candidate documents: For each subquery and its multimodal query vector, three search engines (TF-IDF, BM25, and Faiss) are used to extract a set of candidate documents from the knowledge base, and each document is assigned a corresponding score. Step 3.2, semantic density calculation and adaptive adjustment of retriever weights: Calculate the semantic density of each subquery; Based on semantic density, the TF-IDF, BM25, and Faiss weight vectors are adaptively adjusted to achieve a dynamic balance in the scores. The final score is calculated based on the TF-IDF, BM25, and Faiss scores and weight vectors, semantic density, and the balance coefficient. Step 3.3, accurate re-ranking of candidate documents: In the final ranking stage, the candidate documents are re-ranked according to the final scores to select the most relevant documents.
6. The knowledge reasoning method for integrated circuit wafer process according to claim 1, characterized in that: The step 4 comprises: The basic language model is fine-tuned using LoRA technology. Using the multimodal wafer corpus as the training basis, the basic language model is customized into a dedicated multimodal large language model. During the fine-tuning process, the input mode is expanded based on the characteristics of wafers in the integrated circuit field: First, the mask vector extracted from the query image is used as visual input to enhance the model's accurate understanding of the defect area; Secondly, the first several retrieved knowledge documents that are highly relevant to the query are appended to the user query as additional text input to provide rich contextual support for the model; During the model adaptation process, LoRA technology adjusts model parameters by introducing a low-rank matrix.
7. The knowledge reasoning method for integrated circuit wafer process according to claim 1, characterized in that: The basic language model is the Qwen2VL-7B model.
8. A knowledge reasoning device for integrated circuit wafer technology, characterized in that: The method comprises one or more processors for implementing a knowledge reasoning method for an integrated circuit wafer process according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, the knowledge reasoning method for integrated circuit wafer process according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for generating wafer defect description based on multi-modal fusion
CN118037706A
Wafer manufacturing process evaluation and anomaly detection method assisted by machine learning
CN118887208A
Integrated circuit process defect diagnosis and analysis method and device and medium
CN119027411A
Multi-modal big language model attribute prediction method based on multi-modal thinking chain
CN119693768A
Complex problem decomposition and multi-modal knowledge retrieval method combined with large model
CN119938832A
Cited By
Wafer graph defect semantic reasoning method driven by retrieval enhancement large model
CN121787596A