Chip defect detection method and system

Through neural architecture search and knowledge distillation algorithm optimization, a lightweight chip defect detection model is built, which solves the problems of large manual intervention and model redundancy in the existing technology, and realizes efficient chip defect detection and edge device deployment.

WO2025139322A1PCT designated stage expired Publication Date: 2025-07-03JIANGNAN UNIV +1

Patent Information

Application Number
PCT/CN2024/128016
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-10-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing deep learning models have high manual intervention in chip defect detection, fuzzy engineering application orientation, and redundant and bloated models, which are inconvenient for edge equipment deployment.

Method used

The neural architecture search algorithm model ASNDARTS and the knowledge distillation algorithm model DPSKD are used to optimize the neural architecture search algorithm NAS to build a lightweight knowledge distillation algorithm model DPSKD, and use probability space decoupling and hyperparameter optimization to realize chip defect detection.

Benefits of technology

Improves the accuracy and stability of chip defect detection, reduces hardware resource consumption, and is suitable for lightweight deployment of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128016_03072025_PF_FP_ABST
    Figure CN2024128016_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a chip defect detection method and system. The method comprises: acquiring a one-dimensional vibration signal of a chip and converting same into two-dimensional image data; training an algorithm NAS by means of the two-dimensional image data, and optimizing the algorithm NAS to obtain a neural architecture search algorithm model ASNDARTS; constructing a knowledge distillation algorithm model DPSKD on the basis of the neural architecture search algorithm model ASNDARTS; and acquiring a one-dimensional vibration signal of a chip under test and converting sane into two-dimensional image data, and detecting the two-dimensional image data of the chip under test by means of the knowledge distillation algorithm model DPSKD, so as to determine whether the chip under test has a defect. In the present invention, the algorithm NAS is optimized, the knowledge distillation transfer method is optimized, and the obtained model DPSKD effectively improves the accuracy of chip defect detection while having reduced model redundancy.
Need to check novelty before this filing date? Find Prior Art

Description

Chip defect detection method and system Technical Field

[0001] The present invention relates to the technical field of chip quality detection, and in particular to a chip defect detection method and system. Background Art

[0002] The rapid development of electronic products is driving the trend toward miniaturization and high integration of electronic devices, making traditional wire bonding technology difficult to meet. Flip-chip packaging technology is widely used in the microelectronics packaging industry due to its excellent manufacturability and high reliability. However, the mismatch in the thermal expansion coefficients of the chip and substrate can easily lead to internal defects in the chip, such as cracks, missing solder joints, and cold solder joints. Therefore, to ensure high performance and reliability of products, chip defect detection technology is gaining increasing attention in the electronic packaging field.

[0003] Deep learning technology has recently gained widespread attention among engineers, driven by its successful application in fault detection. First, deep learning can extract useful information from massive amounts of data and possesses strong big data processing capabilities. Second, industrial data collection and storage technologies continue to advance, and the scale of industrial data continues to expand. This massive amount of data has laid the foundation for the application of deep learning. Third, deep learning's powerful feature extraction and fitting capabilities enable end-to-end fault diagnosis, eliminating significant manual data processing and feature extraction steps. Fourth, deep learning can solve a wide range of problems, including classification, prediction, and decision-making.

[0004] Existing deep fault diagnosis models require extensive trial-and-error debugging by professionals when addressing different application scenarios, resulting in high development costs and long development cycles. Furthermore, to achieve improved prediction accuracy, many existing deep models with poor application performance will have additional modules added, making them increasingly bloated. This will undoubtedly slow down the model's prediction speed in practical applications. The design and adjustment of deep models requires a certain level of expertise. This is because different engineering application scenarios, different research objects, and different signal distributions require different model structures for specific tasks to accurately identify fault patterns. Designing algorithmic models that are suitable for various fault diagnosis tasks and have excellent performance is a significant challenge. Therefore, research on automatic, lightweight deep models for fault diagnosis is urgently needed.

[0005] Summary of the Invention

[0006] To this end, the technical problem to be solved by the present invention is to overcome the problems in the existing deep technology, such as high degree of manual intervention, vague engineering application direction, redundant and bloated models, and inconvenience in edge device deployment.

[0007] To solve the above technical problems, the present invention provides a chip defect detection method, comprising:

[0008] Acquire the one-dimensional vibration signal of the chip and convert it into two-dimensional image data;

[0009] Construct a neural architecture search algorithm model ASNDARTS, and construct a knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS;

[0010] The neural architecture search algorithm model ASNDARTS is constructed by inputting two-dimensional image data of a chip into a neural architecture search algorithm NAS for training, and optimizing the neural architecture search algorithm NAS during the training process, and obtaining the neural architecture search algorithm model ASNDARTS when the neural architecture search algorithm NAS reaches a Nash equilibrium;

[0011] The method for constructing the knowledge distillation algorithm model DPSKD is as follows: using the neural architecture search algorithm model ASNDARTS as a teacher network for knowledge distillation, and taking a single cell in the neural architecture search algorithm model ASNDARTS as a student network for knowledge distillation, wherein the neural architecture search algorithm model ASNDARTS includes a plurality of cells connected in series, each of which is a result obtained by the neural architecture search algorithm model ASNDARTS in a search space;

[0012] Based on the probability space decoupling method of knowledge transfer between the teacher network and the student network, and optimizing the hyperparameters of the student network, when the student network converges, the knowledge distillation algorithm model DPSKD is obtained;

[0013] The one-dimensional vibration signal of the chip to be detected is obtained and converted into two-dimensional image data. The two-dimensional image data of the chip to be detected is detected by the knowledge distillation algorithm model DPSKD to determine whether the chip to be detected has defects.

[0014] In one embodiment of the present invention, the optimization of the neural architecture search algorithm NAS during the training process includes a first optimization method, specifically:

[0015] During the training process, the probability of selecting the original skip_connect operation of the neural architecture search algorithm NAS is band-limited.

[0016] In one embodiment of the present invention, the optimization of the neural architecture search algorithm NAS during the training process includes a second optimization method, specifically:

[0017] During the training process, several operation modes OP are randomly selected from the search space of the algorithm NAS to form an initial search space Q1, and then the remaining operation modes OP in the search space of the algorithm NAS are formed into a spare search space Q2. According to the preset criteria, it is determined whether the operation modes OP in the initial search space Q1 and the operation modes OP in the spare search space Q2 need to be adaptively adjusted. If the operation mode OP needs to be adaptively adjusted, the operation mode OP with the lowest probability of being selected in the initial search space Q1 is replaced by the first operation mode OP in the spare search space Q2, and the operation mode OP with the lowest probability of being selected is used as the last operation mode OP in the spare search space Q2.

[0018] In one embodiment of the present invention, the optimization of the neural architecture search algorithm NAS during the training process includes a third optimization method, specifically:

[0019] During the training process, the number of operation mode OPs of the node is adaptively expanded according to the model contribution rate of each operation mode OP.

[0020] In one embodiment of the present invention, the method of performing band-limiting correction on the probability of selection of the original skip_connect operation of the neural architecture search algorithm NAS during training includes:

[0021] Real-time recording of the selected probability value α of the skip_connect operation during the neural architecture search algorithm NAS training process skip , α skip The probability distribution of being selected is P skip , define P skip The point where the gradient rise phenomenon occurs is the algorithm NAS crash starting point epoch_break; introduce the parameter k skip α for skip_connect operation skip Correction is performed to obtain the correction value αcorrect_skip=α skip *k skip , the α of the algorithm NAS collapse starting point skip The difference with αcorrect_skip is defined as △, according to α skip , αcorrect_skip and △ get the band-limited correction probability The formula is as follows:

[0022] Among them, * represents the multiplication sign, k skip Indicates the constraint on the probability of skip_connect operation during the NAS search process; k skipThe first half, kskip(tanh), is based on the tanh activation function to implement the number of early skip_connect operations at the epoch level. skip The second half, kskip(sigmoid), is to limit the number of skip_connect operations at the edge level based on the sigmoid activation function; Indicates α before the crash starting point epoch_break skip Probability of value; Indicates α after the crash starting point epoch_break skip Probability of taking value.

[0023] In one embodiment of the present invention, the preset criteria include:

[0024] The preset criteria include, in order of judgment, the algorithm accuracy trend under equal training epoch intervals, whether there is an abnormal accuracy value under equal training epoch intervals, the OP contribution value trend of each operation mode under equal training epoch intervals, and the OP frequency of each node selected operation mode under equal training epoch intervals, wherein:

[0025] The algorithm accuracy trend under the equal training epoch interval is as follows: a preset number of equal training epoch intervals are obtained, the algorithm accuracy value of each equal training epoch interval is counted, and the algorithm accuracy values ​​of all equal training epoch intervals are linearly fitted to obtain a first linear fitting value. If the first linear fitting value is greater than 0, it indicates that the algorithm accuracy is on an upward trend; if the first linear fitting value is less than 0, it indicates that the algorithm accuracy is on a downward trend.

[0026] Whether there is an abnormal accuracy value under the equal training epoch interval is as follows: a preset number of equal training epoch intervals are obtained, and the algorithm accuracy value of each equal training epoch interval is counted. If the algorithm accuracy value does not increase with the increase of the equal training epoch interval, it indicates that the algorithm accuracy value of the equal training epoch interval is abnormal; otherwise, it is normal.

[0027] The contribution value trend of each operation mode OP under the equal training epoch interval is as follows: a preset number of equal training epoch intervals are obtained, and the selection probability of each operation mode OP in each equal training epoch interval is obtained, and the selection probability of OP in all equal training epoch intervals is linearly fitted to obtain a second linear fitting value. If the second linear fitting value is greater than 0, it indicates that the contribution value of the operation mode OP is on an upward trend; if the second linear fitting value is less than 0, it indicates that the contribution value of the operation mode OP is on a downward trend.

[0028] The frequency of the operation mode OP selected by each node under the equal training epoch interval is: obtaining a preset number of equal training epoch intervals, and at the same time obtaining the number of times each operation mode OP is selected in all equal training epoch intervals, and according to the preset number of selections a, judging the relationship between the number of times each operation mode OP is selected and the preset number of selections a.

[0029] In one embodiment of the present invention, the method for adaptively expanding the number of operation modes OP of a node according to the model contribution rate of each operation mode OP includes:

[0030] In the NAS algorithm, an operation mode OP contribution discrimination mechanism is introduced into the predecessor connection edge of each node. A preset number of equal training epoch intervals are obtained, and the selection probability of each operation mode OP in each equal training epoch interval is obtained. The selection probability of OP in all equal training epoch intervals is linearly fitted to obtain a second linear fitting value. If the second linear fitting value is greater than 0, it indicates that the contribution value of the operation mode OP is on an upward trend; if the second linear fitting value is less than 0, it indicates that the contribution value of the operation mode OP is on a downward trend.

[0031] During the training process, the current node selects the operation mode OP whose contribution value is an upward trend and ignores the operation mode OP whose contribution value is a downward trend, thereby completing the adaptive expansion of the number of operation mode OPs of the current node.

[0032] In one embodiment of the present invention, the search space of the algorithm NAS includes several optional operation modes OP, and the search space of the algorithm NAS includes several optional operation modes OP, including several max_pool_nxn, several avg_pool_nxn, skip_connect, several sep_conv_nxn, several dil_conv_nxn, and none, where max_pool represents maximum pooling, avg_pool represents average pooling, skip_connect represents skip connection, sep_conv represents depthwise separable convolution, dil_conv represents void convolution, none represents no operation, and n takes an odd number in the value range ∈ [1,9].

[0033] In one embodiment of the present invention, the method for decoupling the knowledge transfer between the teacher network and the student network based on the probability space includes:

[0034] The model ASNDARTS is reassigned with probability and divided into three hierarchical category probability spaces, namely the target space t, the target class space O and the overall space C. In order to separate the prediction of the three hierarchical category probability spaces, the symbol P = [p t ,p O\t ,p\O ], p t represents the target space t probability, p O\t Indicates the probability of removing target t in the target class space O, p \O Indicates the exclusion of O probability in the overall space C, where

[0035] Among them, z t represents the soft probability of the target class, z j represents the soft probability of all classes, z k represents the soft profile of the non-target class, k represents the in-class index of the non-target class, and j represents the out-class index of the non-target class;

[0036] The various probabilities in O\t space and \O are defined as and Among them, ∧ represents the final correction symbol of skip_connect, Indicates the probability of excluding other independent similar targets in the target space t within the target space O; similarly, Represents the probability of other independent similar objects in the overall space C after excluding the target class space O, where

[0037] The knowledge distillation KD is decoupled into three-level category probability space, and the formula is:

[0038] make ⊙ represents dot product;

[0039] in, Represents the KL divergence TKD of the teacher-student model in the three defined spaces; Represents the KL divergence CKD of the teacher-student model on the non-target; represents the KL divergence OKD of the teacher-student model between other classes;

[0040] The KL divergence TKD, KL divergence CKD and KL divergence OKD of the three-level category probability space are reorganized, and the knowledge transfer between the teacher network and the student network is realized based on the reorganization results.

[0041] To solve the above technical problems, the present invention provides a chip defect detection system, comprising:

[0042] Acquisition module: used to acquire the one-dimensional vibration signal of the chip and convert it into two-dimensional image data;

[0043] Construction module: used to construct a neural architecture search algorithm model ASNDARTS, and to construct a knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS;

[0044] The neural architecture search algorithm model ASNDARTS is constructed by inputting two-dimensional image data of a chip into a neural architecture search algorithm NAS for training, and optimizing the neural architecture search algorithm NAS during the training process, and obtaining the neural architecture search algorithm model ASNDARTS when the neural architecture search algorithm NAS reaches a Nash equilibrium;

[0045] The method for constructing the knowledge distillation algorithm model DPSKD is as follows: using the neural architecture search algorithm model ASNDARTS as a teacher network for knowledge distillation, and taking a single cell in the neural architecture search algorithm model ASNDARTS as a student network for knowledge distillation, wherein the neural architecture search algorithm model ASNDARTS includes a plurality of cells connected in series, each of which is a result obtained by the neural architecture search algorithm model ASNDARTS in a search space;

[0046] Based on the probability space decoupling method of knowledge transfer between the teacher network and the student network, and optimizing the hyperparameters of the student network, when the student network converges, the knowledge distillation algorithm model DPSKD is obtained;

[0047] Detection module: used to obtain the one-dimensional vibration signal of the chip to be detected and convert it into two-dimensional image data, and detect the two-dimensional image data of the chip to be detected through the knowledge distillation algorithm model DPSKD to determine whether the chip to be detected has defects.

[0048] The above technical solution of the present invention has the following advantages over the prior art:

[0049] The present invention selects the neural architecture search algorithm NAS as the basic network model, effectively solving the time-consuming problem of manually designing models, and designs a lightweight network model with excellent performance that can effectively target specific engineering application backgrounds;

[0050] This paper implements a band-limited correction on the skip_connect operation in the NAS algorithm to resolve the NAS algorithm crash phenomenon. It also proposes the concept of an alternative search space to address the excessive GPU usage of the NAS algorithm by expanding the diversity of the search space. It also introduces a contribution discrimination mechanism for the predecessor connection edges of each node in the NAS algorithm to obtain a variable number of node structures to form diverse cells. This paper significantly improves the training stability and detection accuracy of the neural architecture search algorithm NAS, while reducing hardware resource consumption.

[0051] This paper uses the network model ASNDARTS obtained from neural architecture search as the basic teacher model, further lightweights the network structure through knowledge distillation technology to obtain the student model, and decouples the knowledge transfer logic between the teacher and student models, effectively improving the ability to obtain information from the teacher model and reducing the model's computing resource consumption without reducing prediction accuracy.

[0052] The present invention can adaptively design deep neural networks for different industrial backgrounds, improve the prediction accuracy of model data while lightweighting the model structure, making it easy to deploy on edge devices to detect chip defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings.

[0054] FIG1 is a flow chart of a chip defect detection method according to the present invention;

[0055] FIG2 is a flow chart of a skip_connect operation band limit correction according to an embodiment of the present invention;

[0056] FIG3 is a logic diagram of a skip_connect operation band limit correction according to an embodiment of the present invention;

[0057] FIG4 is a logic diagram of the decoupling of the teacher-student knowledge transfer logic in the DPSKD model according to an embodiment of the present invention;

[0058] FIG5 is a schematic diagram of a result structure (ie, a cell) searched by the ASNDARTS model in an embodiment of the present invention. DETAILED DESCRIPTION

[0059] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0060] Example 1

[0061] 1 , the present invention relates to a chip defect detection method, comprising:

[0062] Acquire the one-dimensional vibration signal of the chip and convert it into two-dimensional image data;

[0063] Construct a neural architecture search algorithm model ASNDARTS, and construct a knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS;

[0064] The neural architecture search algorithm model ASNDARTS is constructed by inputting two-dimensional image data of a chip into a neural architecture search algorithm NAS for training, and optimizing the neural architecture search algorithm NAS during the training process, and obtaining the neural architecture search algorithm model ASNDARTS when the neural architecture search algorithm NAS reaches a Nash equilibrium;

[0065] The method for constructing the knowledge distillation algorithm model DPSKD is as follows: using the neural architecture search algorithm model ASNDARTS as a teacher network for knowledge distillation, and taking a single cell in the neural architecture search algorithm model ASNDARTS as a student network for knowledge distillation, wherein the neural architecture search algorithm model ASNDARTS includes a plurality of cells connected in series / parallel, each of which is a result obtained by the neural architecture search algorithm model ASNDARTS in a search space, and the search space is used to store an operation mode OP;

[0066] Based on the probability space decoupling method of knowledge transfer between the teacher network and the student network, and optimizing the hyperparameters of the student network, when the student network converges, the knowledge distillation algorithm model DPSKD is obtained;

[0067] The one-dimensional vibration signal of the chip to be detected is obtained and converted into two-dimensional image data. The two-dimensional image data of the chip to be detected is detected by the knowledge distillation algorithm model DPSKD to determine whether the chip to be detected has defects.

[0068] Furthermore, the optimization of the neural architecture search algorithm NAS during the training process includes a first optimization method, specifically: performing a band-limited correction on the original skip_connect operation selection probability of the neural architecture search algorithm NAS during the training process.

[0069] Specifically, during the training process, the original skip_connect operation of the neural architecture search algorithm NAS is subjected to a band-limited correction, and the method includes:

[0070] Real-time recording of the selected probability value α of the skip_connect operation during the neural architecture search algorithm NAS training process skip , α skip The probability distribution of being selected is P skip , define P skip The point where the gradient rise phenomenon occurs is the algorithm NAS crash starting point epoch_break; introduce the parameter k skip α for skip_connect operation skip Correction is performed to obtain the correction value αcorrect_skip=α skip *kskip , the α of the algorithm NAS collapse starting point skip The difference with αcorrect_skip is defined as △, according to α skip , αcorrect_skip and △ get the band-limited correction probability The formula is as follows:

[0071] Among them, * represents the multiplication sign, k skip Indicates the restriction on the probability of skip_connect operation in the NAS search process; in the early stage of the algorithm search, the data features can be fully extracted without using skip_connect operation for skip connection; k skip The first half, kskip(tanh), is based on the tanh activation function to implement the number of early skip_connect operations at the epoch level. skip The second half, kskip(sigmoid), is based on the sigmoid activation function to implement the number of skip_connect operations on the edge layer (edge ​​is built into the NAS algorithm, for example, the line between 0 and 2 in Figure 5 is an edge). (The first and second half of the operation have a greater impact on the epoch / edge in the first and second half of the process.) Indicates α before the crash starting point epoch_break skip Probability of value; Indicates α after the crash starting point epoch_break skip Probability of taking value.

[0072] This embodiment corrects the probability of the skip_connect operation being selected, which can effectively solve the algorithm crash phenomenon.

[0073] Furthermore, the neural architecture search algorithm NAS is optimized during the training process, including a second optimization method, specifically: during the training process, a number of operation modes OP are randomly selected from the search space of the algorithm NAS to form an initial search space Q1, and the remaining operation modes OP in the search space of the algorithm NAS are then combined to form a spare search space Q2. According to a preset criterion, it is determined whether the operation modes OP in the initial search space Q1 and the operation modes OP in the spare search space Q2 need to be adaptively adjusted. If the operation modes OP need to be adaptively adjusted, the operation mode OP with the lowest probability of being selected in the initial search space Q1 is replaced by the first operation mode OP in the spare search space Q2, and the operation mode OP with the lowest probability of being selected is used as the last operation mode OP in the spare search space Q2. It should be noted that the initial order of the operation modes OP in the spare search space Q2 is random.

[0074] In this embodiment, the search space of the algorithm NAS includes several optional operation modes OP, including several max_pool_nxn, several avg_pool_nxn, skip_connect, several sep_conv_nxn, several dil_conv_nxn, and none. Among them, max_pool represents maximum pooling, avg_pool represents average pooling, skip_connect represents skip connection, sep_conv represents depthwise separable convolution, dil_conv represents dilated convolution, none represents no operation, and n is an odd number in the range ∈ [1, 9]. The search strategy of the algorithm NAS in this embodiment adopts the gradient descent method.

[0075] Specifically, the preset criteria include, in order of judgment, (1) the algorithm accuracy trend under equal training epoch intervals, (2) whether there is an abnormality in the accuracy value under equal training epoch intervals, (3) the OP contribution value trend of each operation mode under equal training epoch intervals, and (4) the OP frequency of each node selected operation mode under equal training epoch intervals, wherein,

[0076] (1) The algorithm accuracy trend under the equal training epoch interval is as follows: a preset number of equal training epoch intervals are obtained, the algorithm accuracy value of each equal training epoch interval is counted, and the algorithm accuracy values ​​of all equal training epoch intervals are linearly fitted to obtain a first linear fitting value. If the first linear fitting value is greater than 0, it indicates that the algorithm accuracy is on an upward trend; if the first linear fitting value is less than 0, it indicates that the algorithm accuracy is on a downward trend.

[0077] (2) Whether there is an abnormal accuracy value under the equal training epoch interval is as follows: a preset number of equal training epoch intervals are obtained, and the algorithm accuracy value of each equal training epoch interval is counted. If the algorithm accuracy value does not increase with the increase of the equal training epoch interval, it indicates that the algorithm accuracy value of the equal training epoch interval is abnormal; otherwise, it is normal. For example, the algorithm accuracy values ​​of four equal training epoch intervals are obtained: 0.3, 0.7, 0.4, and 0.8. Then 0.4 is an abnormal accuracy value because the accuracy of the model training will generally increase. However, a sudden drop of 0.4 is judged to be an abnormal accuracy value.

[0078] (3) The contribution value trend of each operation mode OP under the equal training epoch interval is as follows: a preset number of equal training epoch intervals are obtained, and at the same time, the selection probability of each operation mode OP in each equal training epoch interval is obtained, and the selection probability of OP in all equal training epoch intervals is linearly fitted to obtain a second linear fitting value. If the second linear fitting value is greater than 0, it indicates that the contribution value of the operation mode OP is on an upward trend; if the second linear fitting value is less than 0, it indicates that the contribution value of the operation mode OP is on a downward trend.

[0079] (4) The frequency of the operation mode OP selected by each node under the equal training epoch interval is as follows: a preset number of equal training epoch intervals is obtained, and the number of times each operation mode OP is selected in all equal training epoch intervals is obtained at the same time; and according to the preset number of selections a, the relationship between the number of selections of each operation mode OP and the preset number of selections a is determined.

[0080] Furthermore, the neural architecture search algorithm NAS is optimized during the training process, including a third optimization method, specifically: during the training process, the number of operation modes OP of the node is adaptively expanded according to the model contribution rate of each operation mode OP.

[0081] Specifically, the method for adaptively expanding the number of operation modes OP of a node according to the model contribution rate of each operation mode OP includes:

[0082] In the NAS algorithm, an operation mode OP contribution discrimination mechanism is introduced into the predecessor connection edge of each node. A preset number of equal training epoch intervals are obtained, and the selection probability of each operation mode OP in each equal training epoch interval is obtained. The selection probability of OP in all equal training epoch intervals is linearly fitted to obtain a second linear fitting value. If the second linear fitting value is greater than 0, it indicates that the contribution value of the operation mode OP is on an upward trend (equivalent to a positive significance); if the second linear fitting value is less than 0, it indicates that the contribution value of the operation mode OP is on a downward trend.

[0083] During the training process, the current node selects the operation mode OP whose contribution value is an upward trend and ignores the operation mode OP whose contribution value is a downward trend, thereby completing the adaptive expansion of the number of operation mode OPs of the current node.

[0084] The method of knowledge transfer between the teacher network and the student network based on probabilistic space decoupling includes:

[0085] The deep learning soft probability model ASNDARTS is probability reassigned and divided into three hierarchical category probability spaces, namely the target space t, the target class space O and the overall space C. In order to separate the prediction of the three hierarchical category probability spaces, the symbol P = [p t ,p O\t ,p \O ], p t represents the target space t probability, p O\t Indicates the probability of removing target t in the target class space O, p \O Indicates the exclusion of O probability in the overall space C, where

[0086] Among them, z t represents the soft probability of the target class, z j represents the soft probability of all classes, z k represents the soft profile of the non-target class, k represents the in-class index of the non-target class, and j represents the out-class index of the non-target class;

[0087] The various probabilities in O\t space and \O are defined as and Among them, ∧ represents the final correction symbol of skip_connect, Indicates the probability of excluding other independent similar targets in the target space t within the target space O; similarly, It represents the probability of other independent similar defects after excluding the target class space O in the overall space C (other independent similar defects represent different defect degrees of the same type of defects), where

[0088] The knowledge distillation KD is decoupled into three-level category probability space, and the formula is:

[0089] make ⊙ represents dot product;

[0090] in, Represents the KL divergence TKD of the teacher-student model in the three defined spaces; Represents the KL divergence CKD of the teacher-student model on the non-target; represents the KL divergence OKD of the teacher-student model between other classes;

[0091] The KL divergence TKD, KL divergence CKD and KL divergence OKD of the three-level category probability space are reorganized (DPSKD = ​​α*TKD + β*CKD + γ*OKD), and the knowledge transfer between the teacher network and the student network is realized based on the reorganization results.

[0092] Figure 4 shows the decoupling logic of knowledge transfer logic using three-level category probability space. A1, A2, A3, A4, and A5 in Figure 4 are as follows:

[0093] A1:

[0094] A2:

[0095] A3:

[0096] A4:

[0097] A5:

[0098] The technical solution of the present invention is further described below with reference to specific embodiments:

[0099] Step S1: Obtain a one-dimensional vibration signal from the faulty chip and sample the data with equal steps using a window width of n*n length. Reconstruct the sampled signal into an n*n matrix after Fourier transformation, and splice three layers of n*n matrices to form a three-channel image data. Each sample has a uniform size of n*n, and a one-hot type label of the corresponding category is prepared for each sample (since this is a prior art, it will not be described in detail in this embodiment). All samples are divided into training sets and test sets according to the ratio.

[0100] Step S2: Construct a neural architecture search algorithm model ASNDARTS, and construct a knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS.

[0101] Build an improved neural architecture search algorithm ASNDARTS and obtain the lightweight model model_V1.0.

[0102] Step S3: Based on the original NAS model, a correction value for the skip_connect operation is introduced and combined with the original skip_connect operation selection probability value. Based on the accuracy convergence trend and model structure convergence trend during algorithm training, the skip_connect operation selection probability is corrected to solve the algorithm crash. The specific workflow and usage logic are shown in Figure 2.

[0103] Real-time record of the selected probability value α of the skip_connect operation in its search space skip and correction value αcorrect_skip, α skip The distribution of P skip ; Define P skip The point where the gradient rise occurs is the algorithm crash starting point epoch_break. Introducing parameter k skip Correct the a value of the skip operation to obtain αcorrect_skip=α skip *k skip , the corrected coefficient αcorrect_skip will tend to be stable. The difference between the algorithm crash point and is defined as △, for α after the crash point skip The probability of taking the value is Get the final band-limited corrected probability

[0104] Among them, k skip Indicates the restriction on the probability of skip_connect operation in the algorithm search process. In the early stage of the algorithm search, data features can be fully extracted without using skip_connect for skip connection, k skip The first half uses the tanh activation function to limit the number of skip_connect operations at the epoch level. The second half uses the sigmoid activation function to limit the number of skip_connect operations at the edge level.

[0105] Step S4: For example, define the initialization search space of the neural architecture search algorithm: Q1 = {max_pool_3x3, sep_conv_3x3, sep_conv_7x7, dil_conv_3x3, dil_conv_7x7, skip_connect, none}, and the backup search space: Q2 = {avg_pool_3x3, avg_pool_5x5, sep_conv_5x5, dil_conv_5x5, max_pool_5x5}. According to the set criteria, the optional operations within the search space are replaced and adjusted. The criteria include: (1) the trend of model accuracy under equal training epoch intervals, (2) the abnormality of the model accuracy value under equal training epoch intervals, (3) the trend of the contribution value of each operation mode under equal training epoch intervals, and (4) the frequency of the operation mode selected by each node under equal training epoch intervals. The specific judgment logic is shown in Table 1 below.

[0106] Table 1 Search space adaptation criteria

[0107] Step S5: introducing an operation mode OP contribution discrimination mechanism for each node's predecessor connection edge, adaptively expanding the number of node operations according to the OP model contribution rate, and obtaining a variable number of node structures;

[0108] Step S6: Repeat steps S3 to S5 with equal epochs. In this embodiment, after approximately 150 iterations, the model accuracy converges to the optimal value and the model structure obtained by the neural architecture search stabilizes, obtaining the neural architecture search algorithm model ASNDARTS. The test set is then input, completing the construction of the lightweight chip defect detection model model_V1.0.

[0109] Step S7: Use model_V1.0 (i.e., model ASNDARTS) as the teacher network in the knowledge distillation model and take a single cell as the student network.

[0110] Figure 5 shows the structure of the ASNDARTS search result, which is a cell. Blocks 1 through 4 in Figure 5 represent four nodes, and each node's input predecessor edge corresponds to an operation mode (OP). Model_V1.0 is formed by multiple cells connected in series or parallel, and it also reflects the structure of model_V2.0 (i.e., model DPSKD).

[0111] Step S8: Construct a new information transfer logic for the teacher-student network in the knowledge distillation algorithm. Combined with the chip fault category and fault severity, the target class, target family class, and all classes are decoupled into three probability spaces, and the knowledge transfer logic formula under different probability spaces is obtained. The specific workflow and usage logic are shown in Figure 3.

[0112] The model probability is reallocated and divided into three levels of class probability space, namely target space t, target class space O and overall space C. In order to separate the three levels of prediction, the following symbol P = [p t ,p O\t ,p \O ]. t represents the target t probability, p O\t Indicates the probability of removing the target t in O space, p \O represents the probability of excluding O in the overall space C.

[0113] Among them, z t represents the soft probability of the target class, z j represents the soft probability of all classes, z krepresents the soft profile of the non-target class, k represents the in-class index of the non-target class, and j represents the out-class index of the non-target class;

[0114] The various probabilities in O\t space and \O are defined as and Among them, ∧ represents the correction symbol of the final skip_connect, Indicates the probability of excluding other independent similar cases of t in the class O space. Similarly, Represents other independent probabilities in C space after excluding O space.

[0115] Decouple the traditional KD into three probability spaces:

[0116] according to and ⊙ represents dot product:

[0117] Among them, ① represents the KL divergence TKD of the teacher-student model in the three defined spaces; ② represents the KL divergence CKD of the teacher-student model in non-target; ③ represents the KL divergence OKD of the teacher-student model between other classes.

[0118] The KL divergence TKD, KL divergence CKD and KL divergence OKD of the three-level category probability space are reorganized, and the knowledge transfer between the teacher network and the student network is realized based on the reorganization results.

[0119] Step S9: Optimize the hyperparameters of the student model using a hyperparameter optimization algorithm;

[0120] Step S10: Repeat steps S8 and S9 for approximately 50 iterations, until the student model's accuracy converges to the optimal value. The knowledge distillation algorithm model DPSKD is obtained. The test set is input to complete the construction of the minimalist self-search model model_V2.0. Model model_V2.0 can be deployed on edge machines for chip defect detection applications on edge devices.

[0121] Example 2

[0122] This embodiment provides a chip defect detection system, including:

[0123] Acquisition module: used to acquire the one-dimensional vibration signal of the chip and convert it into two-dimensional image data;

[0124] Construction module: used to construct a neural architecture search algorithm model ASNDARTS, and to construct a knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS;

[0125] The neural architecture search algorithm model ASNDARTS is constructed by inputting two-dimensional image data of a chip into a neural architecture search algorithm NAS for training, and optimizing the neural architecture search algorithm NAS during the training process, and obtaining the neural architecture search algorithm model ASNDARTS when the neural architecture search algorithm NAS reaches a Nash equilibrium;

[0126] The method for constructing the knowledge distillation algorithm model DPSKD is as follows: using the neural architecture search algorithm model ASNDARTS as a teacher network for knowledge distillation, and taking a single cell in the neural architecture search algorithm model ASNDARTS as a student network for knowledge distillation, wherein the neural architecture search algorithm model ASNDARTS includes a plurality of cells connected in series, each of which is a result obtained by the neural architecture search algorithm model ASNDARTS in a search space;

[0127] Based on the probability space decoupling method of knowledge transfer between the teacher network and the student network, and optimizing the hyperparameters of the student network, when the student network converges, the knowledge distillation algorithm model DPSKD is obtained;

[0128] Detection module: used to obtain the one-dimensional vibration signal of the chip to be detected and convert it into two-dimensional image data, and detect the two-dimensional image data of the chip to be detected through the knowledge distillation algorithm model DPSKD to determine whether the chip to be detected has defects.

[0129] Example 3

[0130] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the chip defect detection method described in the first embodiment are implemented.

[0131] Example 4

[0132] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the chip defect detection method described in the first embodiment are implemented.

[0133] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.

[0134] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0135] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0137] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0138] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A method for detecting chip defects, characterized in that: Including: Obtain the one-dimensional vibration signal of the chip and convert it into two-dimensional image data; Construct a neural architecture search algorithm model ASNDARTS, and construct a knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS; The construction method of the neural architecture search algorithm model ASNDARTS is: input the two-dimensional image data of the chip into the neural architecture search algorithm NAS for training, and optimize the neural architecture search algorithm NAS during the training process. When the neural architecture search algorithm NAS reaches Nash equilibrium, the neural architecture search algorithm model ASNDARTS is obtained; The construction method of the knowledge distillation algorithm model DPSKD is: use the neural architecture search algorithm model ASNDARTS as the teacher network for knowledge distillation, and take a single cell in the neural architecture search algorithm model ASNDARTS as the student network for knowledge distillation. Among them, the neural architecture search algorithm model ASNDARTS includes several serially connected cells, and each cell is the result obtained by the neural architecture search algorithm model ASNDARTS searching in the search space; Decouple the knowledge transfer method between the teacher network and the student network based on the probability space, and perform hyperparameter optimization on the student network. When the student network converges, the knowledge distillation algorithm model DPSKD is obtained; Obtain the one-dimensional vibration signal of the chip to be detected and convert it into two-dimensional image data, and detect the two-dimensional image data of the chip to be detected through the knowledge distillation algorithm model DPSKD to determine whether there are defects in the chip to be detected.

2. The chip defect detection method according to claim 1, characterized in that: The optimization of the neural architecture search algorithm NAS during the training process includes a first optimization method, specifically: During the training process, perform band-limited correction on the selected probability of the original skip_connect operation of the neural architecture search algorithm NAS.

3. The chip defect detection method according to claim 1, wherein: The optimization of the neural architecture search algorithm NAS during the training process includes a second optimization method, specifically: During the training process, randomly select several operation methods OP from the search space of the algorithm NAS to form an initial search space Q1, and then form a spare search space Q2 with the remaining operation methods OP in the search space of the algorithm NAS. Determine whether the operation method OP in the initial search space Q1 needs to be adaptively adjusted with the operation method OP in the spare search space Q2 according to a preset criterion. Among them, if the operation method OP needs to be adaptively adjusted, replace the operation method OP with the lowest selected probability in the initial search space Q1 with the first operation method OP in the spare search space Q2, and use the operation method OP with the lowest selected probability as the last operation method OP in the spare search space Q2.

4. The chip defect detection method according to claim 1, characterized in that: The optimization of the neural architecture search algorithm NAS during the training process includes a third optimization method, specifically: During the training process, perform adaptive expansion of the number of operation methods OP of the nodes according to the model contribution rate of each operation method OP.

5. The chip defect detection method according to claim 2, characterized in that: During the training process, a bandwidth-limited correction is performed on the selection probability of the original skip_connect operation of the neural architecture search algorithm NAS. The method includes: Record the selected probability value α of the skip_connect operation during the training process of the neural architecture search algorithm NAS in real time skip , α skip The selected probability distribution is P skip , define P skip The point where the gradient ascent phenomenon occurs is defined as the starting point of the algorithm NAS crash epoch_break; introduce the parameter k skip For the α of the skip_connect operation skip Perform correction to obtain the corrected value αcorrect_skip = α skip *k skip , define the difference between the α at the starting point of the algorithm NAS crash and αcorrect_skip as △, and obtain the band-limited corrected probability according to α skip , αcorrect_skip and △ skip ​ The formula is as follows: where * represents the multiplication sign, and k skip represents the constraint on the occurrence probability of the skip_connect operation during the NAS search process of the algorithm; k skip The first half kskip(tanh) is the limit on the number of early skip_connect operations at the epoch level implemented based on the tanh activation function, and k skip The second half kskip(sigmoid) is the limit on the number of skip_connect operations at the edge level implemented based on the sigmoid activation function; Indicates α before the epoch_break at the crash starting point skip Probability of taking values; Indicates α after the epoch_break of the crash starting point skip Value probability.

6. The chip defect detection method according to claim 3, wherein: The preset criteria include: The preset criteria sequentially include, according to the discrimination order, the accuracy trend of the algorithm at equal training epoch intervals, whether there is an abnormal accuracy value at equal training epoch intervals, the contribution value trend of each operation mode OP at equal training epoch intervals, and the selection frequency of the operation mode OP selected by each node at equal training epoch intervals. Among them, The accuracy trend of the algorithm at equal training epoch intervals is as follows: Obtain a preset number of equal training epoch intervals, count the accuracy values of the algorithm for each equal training epoch interval, and perform a linear fit on the accuracy values of the algorithm for all equal training epoch intervals to obtain a first linear fit value. If the first linear fit value is greater than 0, it indicates that the accuracy of the algorithm is on the rise; if the first linear fit value is less than 0, then It indicates that the accuracy of the algorithm is on the decline; Whether there is an abnormal accuracy value at equal training epoch intervals is as follows: Obtain a preset number of equal training epoch intervals, count the accuracy values of the algorithm for each equal training epoch interval. If the accuracy value of the algorithm does not increase with the increase of the equal training epoch interval, it indicates that the accuracy value of the algorithm for this equal training epoch interval is abnormal, otherwise it is normal; The contribution value trend of each operation mode OP at equal training epoch intervals is as follows: Obtain a preset number of equal training epoch intervals, and at the same time obtain the selection probability of each operation mode OP in each equal training epoch interval, and perform a linear fit on the selection probabilities of all OPs for all equal training epoch intervals to obtain a second linear fit value. If the second linear fit value is greater than 0, it indicates that the contribution value of the operation mode OP is on the rise; if the second linear fit value is less than 0, it indicates that the contribution value of the operation mode OP is on the decline; The selection frequency of the operation mode OP selected by each node at equal training epoch intervals is as follows: Obtain a preset number of equal training epoch intervals, and at the same time obtain the number of times each operation mode OP is selected in all equal training epoch intervals, and judge the relationship between the number of times each operation mode OP is selected and the preset number of selected times a according to the preset number of selected times a.

7. The chip defect detection method according to claim 4, characterized in that: The method for adaptively expanding the number of operation modes OP of nodes according to the model contribution rate of each operation mode OP includes: Introduce an operation mode OP contribution discrimination mechanism for the precursor connection edges of each node in the algorithm NAS. Obtain a preset number of equal training epoch intervals, and at the same time obtain the selection probability of each operation mode OP in each equal training epoch interval, and perform a linear fit on the selection probabilities of all OPs for all equal training epoch intervals to obtain a second linear fit value. If the second linear fit value is greater than 0, it indicates that the contribution value of the operation mode OP is on the rise; if the second linear fit value is less than 0, it indicates that the contribution value of the operation mode OP is on the decline; During the training process, the current node selects the operation mode OP with an increasing contribution value, ignores the operation mode OP with a decreasing contribution value, and completes the adaptive expansion of the number of operation modes OP of the current node.

8. The chip defect detection method according to claim 3, characterized in that: The search space of the algorithm NAS includes several optional operation modes OP, and the operation mode OP includes several max_pool_nxn, several avg_pool_nxn, skip_connect, several sep_conv_nxn, several dil_conv_nxn, none, where max_pool represents max pooling, avg_pool represents average pooling, skip_connect represents skip connection, sep_conv represents depthwise separable convolution, dil_conv represents dilated convolution, none represents no operation, and n takes odd values in the range ∈[1,9].

9. The chip defect detection method according to claim 1, wherein: The method for decoupling the knowledge transfer method between the teacher network and the student network based on the probability space includes: The model ASNDARTS is probabilistically reassigned and divided into three hierarchical category probability spaces, namely the target space t, the within-target-class space O, and the overall space C; to separate the prediction situations of the three hierarchical category probability spaces, the symbol P = [p t , p Ot , p O is defined, where p t represents the probability of the target space t, p Ot represents the probability excluding the target t within the within-target-class space O, and p O represents the probability excluding O in the overall space C. Among them, Among them, z t represents the soft probability of the target class, z j represents the soft probability of all classes, z k represents the soft profile of non-target classes, k represents the within-class index of non-target classes, and j represents the out-of-class index of non-target classes; Define the O space and various probabilities below O respectively as and Among them, ∧ represents the correction symbol of the final skip_connect, Indicates the probability of other independent same classes excluding the target space t within the target class inner space O; similarly, Indicates other independent same-class probabilities after excluding the within-class space O of the target class within the overall space C, where Decouple the knowledge distillation KD into three levels of category probability spaces, and the formula is: Let ⊙ represents the dot product; Among them, Denotes the KL divergence TKD of the teacher-student model on the three defined spaces; Denote the teacher-student model in the non-target KL divergence CKD; Denote the KL divergence OKD between the teacher-student model in other classes; Recombine the KL divergence TKD, KL divergence CKD, and KL divergence OKD of the three-level category probability space, and realize the knowledge transfer between the teacher network and the student network according to the recombination result.

10. A chip defect detection system, characterized in that: Including: Acquisition module: used to acquire the one-dimensional vibration signal of the chip and convert it into two-dimensional image data; Construction module: used to construct the neural architecture search algorithm model ASNDARTS, and construct the knowledge distillation algorithm model DPSKD based on the neural architecture search algorithm model ASNDARTS; The construction method of the neural architecture search algorithm model ASNDARTS is: input the two-dimensional image data of the chip into the neural architecture search algorithm NAS for training, and optimize the neural architecture search algorithm NAS during the training process. When the neural architecture search algorithm NAS reaches the Nash equilibrium, the neural architecture search algorithm model ASNDARTS is obtained; The construction method of the knowledge distillation algorithm model DPSKD is: use the neural architecture search algorithm model ASNDARTS as the teacher network for knowledge distillation, and take a single cell in the neural architecture search algorithm model ASNDARTS as the student network for knowledge distillation, where the neural architecture search algorithm model ASNDARTS includes several cascaded cells, and each cell is the result obtained by the neural architecture search algorithm model ASNDARTS in the search space; Decouple the knowledge transfer method between the teacher network and the student network based on the probability space, and perform hyperparameter optimization on the student network. When the student network converges, the knowledge distillation algorithm model DPSKD is obtained; Detection module: used to acquire the one-dimensional vibration signal of the chip to be detected and convert it into two-dimensional image data, and detect the two-dimensional image data of the chip to be detected through the knowledge distillation algorithm model DPSKD to determine whether there are defects in the chip to be detected.

Citation Information

Patent Citations

  • Neural network search method and system based on knowledge distillation

    CN111445008A

  • Surface defect detection method based on differentiable neural architecture search

    CN117173091A

  • Chip defect detection method and system

    CN117788427A

  • Mask apparatus

    KR1020220001755A

  • Neural architecture search method based on knowledge distillation

    US20220156596A1

Cited By

  • Flow measurement data verification method based on abnormal water consumption analysis

    CN121256505A