Machine learning-based tool for characterizing individual oxide defects
By applying Bayesian machine learning and Markov chain modeling and other technical means in deep-scaling CMOS devices, the charging/discharge behavior of individual oxide defects has been successfully detected and characterized, and the problem that the existing technology is difficult to effectively detect and characterize these defects is solved, and more accurate device behavior modeling and reliability prediction are achieved.
Patent Information
- Application Number
- CN202411558787.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-05
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-06
AI Technical Summary
In deep-scaling complementary metal oxide semiconductor (CMOS) devices, individual oxide defects have a significant impact on device reliability and current characteristics, and it is difficult for the prior art to effectively detect and characterize these defects.
The charging/discharge behavior of individual oxide defects in deep-scaling semiconductor devices is used to detect and characterize their characteristics, and the Bayesian machine learning algorithm and Markov chain model are used, combined with clustering models and maximum likelihood estimator, to achieve automated quantization of measurement data and determination of defect configuration.
This method can effectively detect and characterize complex random telegraph noise (RTN) characteristics of individual oxide defects in deep-scaling CMOS devices in a fully automated manner, providing more accurate device behavior modeling and reliability prediction.
Smart Images

Figure CN119939377A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of semiconductor device characterization. More specifically, it relates to methods of identifying and characterizing individual oxide defects in deeply scaled complementary metal oxide semiconductor (CMOS) devices. Background Art
[0002] It is widely recognized that degradation mechanisms involving field effect transistor (FET) gate current, such as stress induced leakage current (SILC) and time-dependent dielectric breakdown (TDDB), are mainly affected by a small number of material defects. These defects can lead to significant leakage current and potential bridging ("percolation") of the gate oxide. In addition, the continued reduction of VLSI devices has led to a reduction in lateral dimensions to the 10nm range and a reduction in gate oxide thickness to ~1nm. In this case, some randomly behaving defects in the FET gate oxide have a considerable impact on both the drive current and the gate leakage current. This manifests itself as time-dependent variability and random telegraph noise (RTN). As described in this article, in large devices, the (random) behavior of each defect is averaged together, which produces a well-defined lifetime. However, in deeply scaled devices, only a small number of defects exist, and their random nature based on Poisson statistics leads to large variations in device behavior.
[0003] Therefore, in this new paradigm of reliability physics, it is necessary to understand device degradation at the level of individual defects. Due to scaling at nanometer and smaller dimensions, reliability studies using current statistical methods and simulations are becoming increasingly complex. The present disclosure addresses the need for defect analysis methods to enable state-of-the-art reliability studies. Summary of the invention
[0004] This disclosure describes tools and methods for detecting and characterizing the charge / discharge behavior of individual oxide defects in semiconductor devices through deep scaling. These tools and associated methods operate in a fully automated manner to analyze the complex RTN signatures of these defects. While conventional methods can only provide signatures when one or two defects cause RTN, the methods of this disclosure can handle more complex signals caused by multiple defects.
[0005] In one aspect of the invention, a method is provided. The method includes receiving measurement data. The measurement data indicates a random telegraph noise (RTN) signal. The method further includes extracting a set of extracted discrete levels from the measurement data using a Bayesian machine learning algorithm. The method also includes quantizing the measurement data by assigning each data point of the measurement data to a corresponding extracted discrete level. The method further includes determining a plurality of possible defect configurations based on the quantized measurement data. The method includes applying a Markov chain model to the time evolution of the extracted discrete levels to provide a state transition matrix. The state transition matrix includes a transition probability for each extracted discrete level to a different extracted discrete level. The method further includes determining a corresponding transition probability for each extracted discrete level. The method also includes forming an amplitude / transition probability space based on the amplitude difference and transition probability between the respective discrete levels. The method further includes applying a clustering model to the amplitude / transition probability space to reduce the total number of plausible level combinations to a filtered set of plausible level combinations. The method includes determining a base leakage value using a maximum likelihood estimator. The method also includes determining a best fit data reconstruction of the measured data and corresponding predicted levels by mapping the extracted discrete levels to a linear superposition of each level in a set selected from the filtered set of reasonable level combinations.
[0006] Particular aspects of the embodiments are set out in the accompanying independent and dependent claims. Features from the dependent claims may be combined with features of the independent claims and with features of other dependent claims as appropriate and not just as explicitly set out in the claims.
[0007] These and other aspects of the invention will be apparent from the embodiment(s) described hereinafter and will be elucidated with reference to these embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and additional features may be better understood from the following illustrative and non-limiting detailed description of example embodiments with reference to the accompanying drawings.
[0009] Figure 1 A method according to an example embodiment is shown.
[0010] Figure 2 A semiconductor device according to example embodiments is shown.
[0011] Figure 3 Electrical measurement data of two semiconductor devices of different sizes are shown according to example embodiments.
[0012] Figure 4 Electrical measurement data of a semiconductor device according to example embodiments is shown.
[0013] Figure 5 According to an example embodiment, Figure 1 part of each step of the method.
[0014] Figure 6 A graph showing absolute current magnitude differences between adjacent current levels from measurement data according to an example embodiment is shown.
[0015] Figure 7 A portion of a Bayesian learning loop is shown according to an example embodiment.
[0016] Figure 8 Finding an optimal noise level according to an example embodiment is shown.
[0017] Fig. 9 A graph of probability density versus gate current is shown according to an example embodiment.
[0018] Fig.10 A graph of noise level versus current level is shown according to an example embodiment.
[0019] Fig.11 A best fit data reconstruction graph is shown according to an example embodiment.
[0020] Fig.12 A state transition matrix according to an example embodiment is shown.
[0021] Fig.13 A graph showing transition probability versus current level variation according to an example embodiment is shown.
[0022] Fig.14 The first six transition probabilities and current level variations are shown along with error bars according to an example embodiment.
[0023] Fig.15 The first six transition probabilities and current level variations are shown along with normalized error bars according to an example embodiment.
[0024] Fig.16 The first six transition probabilities and current level variations are shown according to an example embodiment.
[0025] Fig.17 A graph showing the first six transition probabilities versus current level according to an example embodiment is shown.
[0026] Fig.18 The first six transition probabilities after applying the clustering model according to an example embodiment are shown.
[0027] Fig.19 The distribution of Bayesian based levels and probability density functions of individual discrete levels according to an example embodiment are shown.
[0028] Fig. 20 A Bayesian based distribution of levels and optimal current configuration according to an example embodiment is shown.
[0029] Fig.21 An optimal solution and corresponding data reconstruction according to an example embodiment are shown.
[0030] Fig. 22 Deconvolution after finding an optimal solution is shown according to an example embodiment.
[0031] Any reference signs in the claims should not be construed as limiting the scope.
[0032] The same reference numbers in different drawings refer to the same or similar elements.
[0033] All figures are schematic, not necessarily to scale, and generally show only parts which are necessary to elucidate the example embodiments, wherein other parts may be omitted or merely indicated. DETAILED DESCRIPTION
[0034] Example embodiments will now be described more fully below with reference to the accompanying drawings. However, the content covered by the claims can be embodied in many different forms and should not be construed as limited to the embodiments described herein; rather, these embodiments are provided by way of example. In addition, the same reference numerals always refer to the same or similar elements or components.
[0035] The drawings described are merely schematic and non-limiting. In the drawings, for illustrative purposes, the size of some of the elements may be exaggerated and not drawn to scale. The dimensions and relative dimensions do not correspond to actual reductions in practice of the embodiments.
[0036] Furthermore, the terms top, bottom, etc. in the specification and claims are used for descriptive purposes and not necessarily for describing relative positions. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments described herein are capable of operation in orientations other than those described or illustrated herein.
[0037] It should be noted that the term "comprising" used in the claims should not be interpreted as being limited to the means listed thereafter; it does not exclude other elements or steps. Thus, the term should be interpreted as specifying the presence of the stated features, integers, steps or components as mentioned, but does not exclude the presence or addition of one or more other features, integers, steps or components, or groups thereof. Thus, the scope of expressing a "device comprising means A and B" should not be limited to devices consisting only of components A and B. It means that for embodiments of the present invention, the only relevant components of the device are A and B.
[0038] Throughout this specification, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, but may refer to different embodiments. Furthermore, in one or more embodiments, as would be apparent to one of ordinary skill in the art from this disclosure, the particular features, structures, or characteristics may be combined in any suitable manner.
[0039] Similarly, it should be appreciated that in the description of the exemplary embodiments, for the purpose of streamlining the disclosure and assisting in the understanding of one or more of the various aspects, various features are sometimes grouped together in a single embodiment, figure, or description thereof. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than those expressly recited in each claim. On the contrary, as reflected in the appended claims, various aspects consist of fewer features than all of the features of a single previously disclosed embodiment. Thus, the claims appended to the Detailed Description are hereby expressly incorporated into the Detailed Description, with each claim itself representing a separate embodiment.
[0040] In addition, although some embodiments described herein include some features included in other embodiments but do not include other features included in these other embodiments, the combination of features of different embodiments is intended to fall within the scope of the present invention and form different embodiments as will be understood by those skilled in the art. For example, in the appended claims, any of the embodiments claimed for protection may be used in any combination.
[0041] In the description provided herein, numerous specific details are set forth. However, it should be understood that various embodiments may be practiced without these specific details. In other instances, well-known methods, structures, and techniques are not shown in detail to avoid obscuring the understanding of this specification.
[0042] Overview
[0043] This disclosure describes a physically-informed machine learning (PIML) approach for detecting and modeling the behavior of individual defects in thin gate oxides of deeply scaled CMOS devices.
[0044] The presence of a small number of individual defects in the oxide results in random telegraph noise (RTN), which can be measured by monitoring the gate leakage current as a function of time. The generation of new oxide defects is reflected by a sudden increase in the time-resolved gate leakage current measured at a constant stress voltage. This generation of a single defect is called a 'discrete' stress-induced leakage current (SILC) event, which can be observed due to the presence of only a few oxide defects.
[0045] For example, the device of interest may include a scaled high-k metal gate (HKMG) MOSFET formed based on a technology node such as 28 nm. In such a scenario, the gate oxide may be a hafnium oxide layer having a thickness of 1.8 nm.
[0046] In some embodiments, SILC can be measured as a function of time at a constant gate voltage bias of 1.2V. It will be appreciated that other gate voltage bias values are also possible and contemplated. In various examples, the gate voltage bias can reflect standard device operating conditions, rather than accelerated operating conditions. In other words, the device under test can have a voltage bias according to the gate bias recommended by the manufacturer, rather than a higher voltage that may be applied in an accelerated test scenario.
[0047] In other words, the measurement data described herein may be obtained under normal operating conditions of the semiconductor device. Among other elements, the normal operating conditions of the semiconductor device may include a standard operating bias voltage recommended by the semiconductor device manufacturer. In various embodiments, the measurement data described herein are not obtained under accelerated operating conditions of the semiconductor device. For example, the accelerated operating conditions may include an elevated bias voltage.
[0048] In an example embodiment, data may be collected for 104 seconds at a sampling rate of 50 Hz (20 ms period). It will be appreciated that other sampling rates and time periods may be utilized. The experimental data has inherent fluctuations in measurement noise and discrete changes between current levels corresponding to the charging / discharging of individual defects.
[0049] The experimental data are analyzed and the total number of observed current levels is extracted, where each current level is represented by a Gaussian distribution with two parameters (mean and standard deviation). A Bayesian-inspired algorithm is developed to extract the different current levels, which clearly explains the observed experimental data.
[0050] The measurement noise is then filtered out by quantizing the data, where each data point is assigned to its corresponding extracted current level based on the distribution parameters. Since the defects in the thin gate oxide are independent and the defect physics has no internal memory, the evolution of the extracted current levels corresponds to a Markov chain. An algorithm is used to extract the transition probabilities from the observed switching between the extracted current levels.
[0051] Each extracted current level corresponds to a defect configuration. In an N-defect system, a total of 2N configurations are possible, i.e., each defect is in a charged state or a discharged state, corresponding to high or low conductivity in the leakage path corresponding to the defect. The capture (charging) and emission (discharging) times of oxide defects can vary from a few microseconds to a few days. When there are only a small number of defects, the probability of two defects switching at the same time is extremely low. This mathematical constraint simplifies the observed transfer between two current levels to the switching of a single defect. In addition, the absolute amplitude difference between the two current levels corresponds to the defect current, i.e., the leakage current through the path involved by the defect. The available information related to the transition probability and the current level difference is used to extract the individual defects (defect currents) observed in the data. Machine learning algorithms such as nearest neighbor propagation can be used for this purpose.
[0052] In the next step, a maximum likelihood estimator is used to determine the base leakage current, i.e., the residual leakage current when all leakage current paths interposed by defects are non-conductive. The base leakage current may include a state where all defects are non-conductive. It should be noted that in some cases, the charged state corresponds to the conductive / active state, while the discharged / inactive state corresponds to the non-conductive state. However, in other cases, the discharged state may correspond to the conductive state, and vice versa.
[0053] Finally, the individual defect currents are analyzed to accurately reconstruct the experimental data. The reconstructed data are deconvoluted into the conductive / non-conductive cycles of each extracted current level / defect, which in turn provides the time constant associated with each defect.
[0054] The systems and methods described in this disclosure provide several advantages over conventional solutions. For example, the methods herein can be used to detect the generation and charging / discharging of individual defects in the gate oxide of FET-type devices. The successful detection of individual defects in the gate oxide opens the way to a variety of unprecedented applications:
[0055] Oxide degradation monitoring tool (at low voltage);
[0056] Predict gate oxide breakdown of individual transistors;
[0057] Mapping and characterizing each defect signature as an oxide defect type in new materials, especially in multilayer materials, which can provide unprecedented insights into material structure and degradation; and
[0058] Conventional device degradation models (developed under accelerated conditions) are validated by obtaining stress induced leakage current (SILC) data under normal device operating conditions.
[0059] Example Method
[0060] Figure 1 A method 100 according to an example embodiment is shown. By way of illustrative example, the steps or blocks of the method 100 will be described with reference to other figures. Although the example embodiments may be described as performing certain blocks or steps in a particular order, it will be understood that these blocks or steps may be performed in a different order. Furthermore, various steps or blocks may be repeated and / or omitted within the scope of the present disclosure.
[0061] The method 100 may include three modules: (i) first, using a Bayesian heuristic algorithm to convert the measured I g,leak The proposed method is described as a set of discrete constant observed current levels. (ii) These discrete current levels are then mapped to possible defect sets. Each set generates a list of linear superpositions of individual leakage paths, called a "configuration". This mapping is achieved using Markov process modeling combined with the proximity propagation (AP) clustering algorithm. (iii) Finally, a cost function is defined to select the optimal solution after data reconstruction, i.e., the optimal defect set that can explain the observed data. These modules are now described in detail.
[0062] The method 100 includes receiving measurement data 110. In an example embodiment, the measurement data indicates a random telegraph noise (RTN) signal. RTN is a type of electronic noise that occurs in semiconductors and ultra-thin gate oxide films. Specifically, the RTN signal may be caused by defects in the gate oxide of the semiconductor device. It is also called burst noise, popcorn noise, impulse noise, bistable noise, or random telegraph signal (RTS) noise.
[0063] RTNs are typically caused by the trapping and de-trapping of charge carriers at defects in the gate dielectric of a semiconductor device. These defects may be caused by manufacturing processes (such as heavy ion implantation) or unintentional side effects (such as surface contamination), or may be inherently present due to the amorphous structure of the material. RTNs manifest themselves as discrete jumps in the current or voltage of a semiconductor device. The jumps are random in time and amplitude, but they can be characterized by two characteristic time constants: one for the trapping of charge carriers and one for the de-trapping of charge carriers.
[0064] RTN can be a problem for semiconductor devices because it causes errors in analog circuits and reduces the reliability of digital circuits. RTN is particularly important in advanced semiconductor devices, such as those used in smartphones and other mobile devices, where the gate dielectric is very thin.
[0065] In some embodiments, the measurement data 110 may include information indicating a stress induced leakage current (SILC) measurement. As such, the measurement data may include SILC data 112. In some scenarios, receiving the measurement data may include receiving an electrical signal from a semiconductor device. For example, the semiconductor device may include a high-k gate dielectric layer of a metal oxide semiconductor field effect transistor (MOSFET). In this case, at least a portion of the RTN signal indicates a time-dependent dielectric breakdown (TDDB) of the gate dielectric layer.
[0066] Figure 2 A semiconductor device 200 according to an example embodiment is shown. In an example embodiment, the semiconductor device 200 can be a MOSFET having a gate dielectric stack 202. The gate dielectric stack 202 can be a high-k gate stack including a SiO2 layer (~0.7nm thick), a HfO2 layer (1.8nm thick), a TiN layer (1mm thick), a TiAl layer (3nm thick), and a TiN layer (3nn thick), among other possibilities. In some examples, the gate dielectric stack 202 can have a thickness of 300×300nm. 2 Additionally or alternatively, the gate dielectric stack 202 may be similar to or identical to a gate dielectric stack of a 28 nm technology node.
[0067] It will be understood that other dielectric stacks and discrete dielectric layers are possible and contemplated.
[0068] Figure 3 A relatively large area gate dielectric 300 (eg, 1×1 μm) is shown according to an example embodiment. 2 ) and a relatively small area gate dielectric 310 (e.g., 0.1×0.1 μm 2 ). It can be observed from the electrical measurement data 302 that when a static gate bias voltage is applied, the gate leakage current increases relatively smoothly over time. In such a scenario, the gate oxide may include traps on the order of 800.
[0069] As shown in the electrical measurement data 312, the gate leakage current changes abruptly over time in a step-like manner. This behavior indicates an RTN signal. In such a scenario, the gate oxide may include traps on the order of 8.
[0070] Figure 4 Electrical measurement data 400 of a semiconductor device according to an example embodiment is shown. The measurement data 400 shows the gate leakage current as a function of time and is obtained with a SiO2 / HfO2 gate dielectric (0.7 / 1.8 nm thick) at a gate voltage of 1.2 volts. As shown, the measurement data 400 indicates multiple sudden jumps in the leakage current. These sudden changes generally indicate the creation and / or destruction of a single trap conduction path, which can be viewed as a discrete SILC event.
[0071] The method 100 includes extracting a set of extracted discrete levels (eg, different current levels 122) from the measurement data using a Bayesian machine learning algorithm 120. It will be appreciated that the set of extracted discrete levels may include a set of current values or a set of voltage values.
[0072] The method 100 also includes performing data quantization 130 by assigning each data point of the measured data to a corresponding extracted discrete level.
[0073] The method 100 also includes determining a plurality of possible defect configurations based on the quantified measurement data.
[0074] Figure 5 According to an example embodiment, Figure 1 5. For example, after receiving the measurement data 110, the SILC data 112 can be quantized into different current levels 122 obtained using the Bayesian inspired algorithm 120. In addition, based on N defects, 2N defect configurations are possible due to independent charge / discharge trapping events.
[0075] For example, assuming device area = 300 × 300 nm 2 , oxide thickness = 1.8 nm, defect density = ~[10 14 ,5×10 16 ]cm -3 , we can estimate that in the time period t[0,10 4 ] seconds. In this case, the number of defects per device = D ox ×A×t ox =~[0.02,8.1]. It will be appreciated that other ranges for the number of defects per device may vary based on, for example, process variables and device variations, among other possibilities.
[0076] Figure 6 A graph 600 of current amplitude differences between adjacent current levels from measurement data according to an example embodiment is shown. Based on this information, the median of the sorted current increments is approximately σ=0.2 pA. It will be appreciated that other measurement data may have other medians. In various examples, extracting the current increments in this manner serves as an initial setting for the Bayesian algorithm. In this case, determining the median of the sorted current increments may include initially determining an optimal noise level based on a median 2x2 range standard deviation estimator of the level transitions in the measurement data.
[0077] Figure 7A portion 700 of a Bayesian learning loop is shown according to an example embodiment. Bayes' theorem assigns a degree of confidence to a hypothesis and rationally updates the probability based on new evidence. Mathematically, this can be defined as the likelihood of event E occurring given that hypothesis H is true. In other words, the Bayesian learning loop described herein is a learning algorithm that interprets some given data based on prior hypotheses and then adapts the hypotheses based on new data.
[0078] In various example embodiments, assumption H may include one current level in the presence of Gaussian noise, and event E may include a measured data point (eg, a gate leakage current value).
[0079]
[0080] Where P(H|E) = the probability of a current level existing at a given data point;
[0081] P(H) = prior probability;
[0082] P(E|H) = the probability that the data point belongs to this current level;
[0083] P(E) = probability of finding this data point = P(H).P(E|H)+P(-H).P(E|-H).
[0084] Under this model, the probability density function of each current level can be expressed by a mean μ and a standard deviation σ. In this case, μ can be initialized to the measured data points, and σ can be estimated by a median moving range (such as (median|Ii+1-Ii|)), as Figure 6 shown.
[0085] By iteratively applying the Bayesian learning loop, the hypothesis can be adapted to include multiple Gaussian peaks along the PDF versus current axis.
[0086] In the Bayesian learning loop, the probability of a current level in the presence of Gaussian noise is initialized to P(H) = 0.5. For example, the current level is set to μ = I0 = 0.33 nA, σ initial = 0.2pA. In this case, μ and σ provide Figure 7 Then, for each discrete level (e.g., i=1, 2, 3, ...), the loop consists of calculating the likelihood P(e|H) to find the I given the hypothesis H. i .
[0087] Next, calculate the posterior likelihood:
[0088]
[0089] Based on the predetermined threshold, if it is inappropriate, update H, P(H) = 0.5, μ = I i , σ=σ initial .
[0090] If appropriate, and Only when the extracted current level I n σ is updated only when a statistically significant amount of data points are assigned.
[0091] In this case, the posterior likelihood becomes the new prior likelihood P(H)→P(H|E).
[0092] like Figure 7 As shown, P(E|H) = the probability of finding I1 in the interval [μ+Δ, +∞]U[-∞, μ-Δ], where P = 1-F(μ+Δ, μ, σ) or P = F(μ-Δ, μ, σ).
[0093] Thus, for I1 = μ: P = 1, and for I1 very far from μ: P → 0. Increment i until all current values have been evaluated.
[0094] In the second loop portion, for the "L" current levels found in the measured data, set i = 1 and find the current level closest to I i Next, find the current level closest to I i+1 The current level. If I i and I i+1 The level matches, then I i+1 Assigned to I i The same level. If I i and I i+1 If the levels do not match, use the continuity equation to verify the certainty of the measurement and set I i+1 Assigned to I i Then, i is incremented until all data points of the measured data have been evaluated. The continuity equation provides a continuity condition under which a data point I is considered to be a valid current level if at least N consecutively observed data points can be assigned to the extracted current level. i can only be assigned to the extracted level. This ensures that outliers, single point measurement errors and single point transitions between levels are not incorrectly assigned to the level of the extracted current. In addition, the continuity condition improves the separation of adjacent current levels with current level differences.
[0095] In an example embodiment, in order to accurately estimate the number of current levels, method 100 may include extracting an optimal noise level. As an example, Figure 8800 for finding the optimal noise level 802. In such a scenario, the Bayesian algorithm 120 may use different initial noise levels (σ initial ) is run several times to find the optimal noise level 802. For example, the Bayesian loop can be run at σ initial =[0.2pA,1.0pA]. In this case, when the estimated σ average (σ average) is approximately equal to σ initial When the corresponding starting value of (σ initial value) is set, the optimal noise level 802 can be determined. In the illustrated case, the optimal noise level 802 can be determined to be approximately 0.477 pA. In some examples, the noise level can correspond to multiple discrete levels, N level =32. It will be appreciated that the optimum noise level 802 and N level Other values of are possible and contemplated.
[0096] Fig. 9 A graph 900 of probability density versus gate current according to an example embodiment is shown. In this case, the graph 900 corresponds to N above. level =32. In this case, the graph 900 includes a plurality of Gaussian curves corresponding to each of the 32 levels determined using the Bayesian algorithm 120.
[0097] Fig.10 A graph 1000 of noise level versus current level is shown according to an example embodiment. In graph 1000, each data point represents the noise level for each of the 32 current levels extracted from the Bayesian algorithm 120. Thus, the average value of the individual data points = σ average =0.477pA, corresponding to the optimal noise level 802.
[0098] Fig.11 A best fit data graph 1100 is shown according to an example embodiment. In this case, the fit curve is provided by the second loop portion of the Bayesian algorithm 120, where each data point is assigned to a discrete current level. In other words, the data points can be quantized into discrete extracted current levels.
[0099] In general, a Markov chain can be used as a mathematical model to describe a system that transitions between different states over time, with the probability of transitioning from one state to another depending only on the current state and not on any past states. In other words, a Markov chain can represent a random walk through a set of states. At each step, the walker moves from one state to another according to a certain probability distribution. The probability distribution depends only on the current state and not on the path the walker took to get there.
[0100] Applied to the present disclosure, the evolution of the current level in the measured data corresponds to a Markov chain. In this case, defect physics does not deal with long-term history - there is no internal memory. In addition, the defects in the thin gate oxide are independent of each other.
[0101] As such, the method 100 additionally includes applying the Markov chain model 140 to the extracted discrete levels to provide a state transition matrix 142. In such a scenario, for each extracted discrete level, the state transition matrix 142 may include a transition probability from a given extracted discrete level to a different extracted discrete level.
[0102] Fig.12 1 is a diagram showing a state transfer matrix 1200 according to an example embodiment. In some embodiments, the state transfer matrix 1200 or the state transfer matrix 142 can be visualized as an NxN matrix A ij In such scenarios, A ij It can be expressed from the current level I i Transfer to current level I j For clarity of illustration, the diagonal matrix elements of the state transfer matrix 1200 are left blank.
[0103] Fig.13 A graph 1300 of transition probability versus current level according to an example embodiment is shown. The 31 data points of graph 1300 correspond to an example where N = 32. In such a scenario, the 31 data points represent the probability of transitioning from one state to another (y-axis) versus the difference in current level (x-axis).
[0104] ΔI level It can form a sudden change in the RTN signal and the measured data, and can be regarded as a defect current {I defect}.
[0105] As described herein, boundary conditions include that the probability of simultaneous charging / discharging of defects is extremely low, i.e., a given configuration "i" can only be transferred to at most 5 different configurations. It will be understood that a 5-defect system will have 5 different configurations, 6 different configurations will be available for a 6-defect system, and so on. The rest of this article uses a 6-defect system as an example. Therefore, the corresponding description refers to the "first 6" transition probabilities, current levels, etc.
[0106] The method 100 further includes forming an amplitude / transition probability space based on the amplitude differences and transition probabilities between the respective discrete levels.
[0107] Clustering is only valid for a 2-variable system if the mean error bars of the system are on the same scale. Therefore, applying a clustering model may include initially normalizing the state transition matrix so that the corresponding error bars of the transition probabilities and level differences are similar.
[0108] If the current level increment is given as ΔI level =I x –I y ,but
[0109] ΔI level The error on can be written as
[0110] Through ΔI level,norm =ΔI level / median(σ ΔI ) to normalize the current level increment.
[0111] Similarly, by σ ΔI,norm =σ ΔI / median(σ ΔI ) to normalize the corresponding error.
[0112] If the transition probability is written as but
[0113] P transition The error can be given as (For Poisson process). Normalizing the transition probability can be done by normalizing log(P transition )=log(P transition ) / median(log(E transition )) and the normalized log(E transition )=log(E transition ) / median(log(E transition )) to execute,
[0114] Where log(E transition )=log(P transition +E transition )-log(P transition ).
[0115] Fig.14 A graph 1400 showing the first six transition probabilities versus current level changes along with error bars (before normalization) according to an example embodiment is shown.
[0116] Fig.15 A graph 1500 showing transition probability versus current level after normalization according to an example embodiment is shown.
[0117] After normalization, the state transition data points can be clustered. In some embodiments, the clustering model may include a neighbor propagation algorithm. Near neighbor propagation (AP) is an unsupervised machine learning algorithm for clustering data points into multiple groups based on the similarity of data points. In AP, compared with k-means and hierarchical clustering models, the number of clusters does not need to be specified in advance. AP works by passing messages between data points. Each data point sends a message to all other data points, indicating the degree of suitability as each other's exemplars. An exemplar is a data point representing a cluster. Messages are iteratively exchanged, and the algorithm converges when a stable set of exemplars is found. Data points belonging to the same cluster are then assigned to the exemplar of the changed cluster.
[0118] As such, the method 100 additionally includes applying a clustering model 150 to the amplitude transition probability space to reduce the total number of reasonable level combinations to a filtered set of reasonable level combinations and defect currents 152 .
[0119] To determine the most likely configuration, formation clustering may be performed based on different time constants of different defects. For example, method 100 may include determining 182 a corresponding time constant for each extracted discrete level.
[0120] Fig.16 A graph 1600 showing the first six transition probabilities versus current level before normalization according to an example embodiment is shown.
[0121] Fig.17 A graph 1700 of the first six transition probabilities versus current level according to an example embodiment is shown. Based on the neighbor propagation clustering model, the data points can be viewed as "messaging" each other. In addition, a model data point can be extracted, i.e., a cluster leader selected by all data points. In addition, the similarity between two points can be expressed by -|ΔI i -ΔI j | 2 Uneven cluster sizes are allowed, and cluster leaders may correspond to input data points. In some embodiments, clustering parameters (eg, preference parameters and damping parameters) may be optimized.
[0122] Fig.18 A graph 1800 is shown of the first six transition probabilities after applying a clustering model according to an example embodiment.
[0123] The method 100 includes determining a base leakage value 162 using a maximum likelihood estimator 160. In such a scenario, the maximum likelihood estimator 160 may be configured to fit a set of Gaussian peaks associated with the extracted discrete levels to a superposition of a filtered set of reasonable level combinations. base , 64 linear configurations {6 I defect}+I base should correspond to the Bayesian current level {32 I level}.
[0124] Fig.19 A graph 1900 is shown showing the distribution of the Bayesian-based levels and the probability density function of the individual discrete levels according to an example embodiment. As shown in the graph 1900, the individual (unweighted) I level PDF to provide I level The income distribution.
[0125] Fig. 20 Graph 2000 shows a distribution of Bayesian based levels and an optimal current configuration according to an example embodiment.
[0126] I base The range is assumed to be ~I level The extreme value of .
[0127] Interpolation of PDF values → 64 configurations {each I defect +I base}
[0128] Furthermore, in some embodiments, a log-likelihood estimator may be provided as In such scenarios, maximizing the log-likelihood estimator can provide the optimal I base .
[0129] The method 100 also includes data reconstruction using the defect current 170. The data reconstruction includes determining a best fit data reconstruction of the measured data and corresponding predicted levels by mapping the extracted discrete levels to a linear superposition of each level in a set selected from the filtered set of reasonable level combinations.
[0130] Fig.21 An optimal solution and corresponding data reconstruction 2100 according to an example embodiment is shown.
[0131] In various examples, the optimal solution may include an optimal set of defects that can explain the observed data. In such scenarios, the optimal solution may be selected by defining a cost function after data reconstruction. In some embodiments, for each given extracted discrete level, reconstructing the data includes determining whether the amplitude difference relative to an adjacent predicted level is within a threshold deviation range of an independent probability density function (PDF) corresponding to the adjacent predicted level.
[0132] In such a scenario, if the amplitude difference relative to the adjacent predicted level is greater than a threshold deviation range (e.g., 3σ) of the independent PDF corresponding to the given predicted level, then a no-solution identifier is assigned to the adjacent predicted level. It is understood that other threshold deviation ranges (e.g., 2σ or less; or greater than 3σ) are possible and contemplated.
[0133] As an initial boundary condition, it is assumed that transfers involving simultaneous switching of multiple defects are impossible.
[0134] Data reconstruction includes:
[0135] - Initialize i=1.
[0136] - For the quantified SILC data (using Bayesian I level ): Find the closest i {I defects+base} configuration and find the closest I i+1 {I deffets+base ] configuration.
[0137] A) If from {I db} i The configured transfer is not mapped to {I db} i+1 Configuration, then BC violence +=1.
[0138] B) If from {I db} i The configured transfer is mapped to {I db} i+1 Configuration, then BC valid +=1.
[0139] C) If {I db} i Configuration is not in I i Within the 3σ range, BC no-solution +=1.
[0140] Next, increment i and loop to find more configurations and evaluate them as described above for A, B, and C.
[0141] For example, for a 6-I defect (pA): {16.7,15.5,9.4,5.0,3.3,2.0} and I base (pA): 315.7, a data reconstruction process can be performed to produce an optimal solution, where the valid transition count = 24280, the violation count = 1892, and the no-solution count = 0. It will be understood that other optimal solutions (and corresponding valid transition counts, violation counts, and no-solution counts) are possible and contemplated.
[0142] In various embodiments, the method 100 also includes deconvolving the reconstructed data 180 to determine a time constant 182 for each individual defect 184 .
[0143] Table 1 shows the equations used in the present method 100 .
[0144] Table 1:
[0145]
[0146]
[0147] To find the optimal solution, the temporal evolution (transition) of discrete current levels is reconstructed using each possible set of defects. A "cost" function (Equation 7 in Table 1) is defined to track the goodness of each solution. This function considers two criteria: (1) all observed levels must be explained within a preset error bar, and (2) more than one transient leakage path switch is penalized. The reconstruction is performed by finding the base (tunneling) leakage current (I base ), which is the current level when no defect path is active. Note that this current level is not necessarily observed. Fig. 20 As shown, I base It can be easily found by matching all configurations of a single defect set with the observed current level using the maximum likelihood method. Finally, a cost function is calculated for each possible defect set, and the set with the minimum cost function is selected as the optimal solution. In some examples, the optimal solution is Fig.21 The 6-defect set with the most accurate reconstruction is shown. This means that only 32 of the 64 configurations were actually observed during the measurement time. Fig. 22 Deconvolution 2200 is shown after finding the optimal solution according to an example embodiment. In addition, Fig. 22 The individual activity of each separate leak path is shown.
[0148] Based on the methods described herein, information related to defects in the gate oxide layer can be used for various applications, including predicting an estimated time to failure (ETTF) of a semiconductor device based on a best fit data reconstruction. Alternatively or additionally, based on a set of possible transitions, various applications may include mapping a set of possible defects in a gate dielectric layer of a semiconductor device or a defect density of a gate dielectric layer of a semiconductor device. It will be appreciated that other practical and useful applications of the disclosed methods are possible and contemplated.
[0149] Listed Example Embodiments
[0150] The present disclosure is described in further detail in the following enumerated example embodiments (EEE). These embodiments are not intended to be limiting, and other embodiments are possible.
[0151] EEE 1 is a method that includes:
[0152] receiving measurement data, wherein the measurement data is indicative of a random telegraph noise (RTN) signal;
[0153] extracting a set of extracted discrete levels from the measurement data using a Bayesian machine learning algorithm;
[0154] quantizing the measurement data by assigning each data point of the measurement data to a corresponding extracted discrete level;
[0155] determining a plurality of possible defect configurations based on the quantified measurement data;
[0156] applying a Markov chain model to the extracted discrete levels to provide a state transition matrix, wherein the state transition matrix includes, for each extracted discrete level, a transition probability to a different extracted discrete level;
[0157] Based on the amplitude difference and transition probability between each discrete level, an amplitude transition probability space is formed;
[0158] applying a clustering model to the amplitude transition probability space to reduce the total number of reasonable level combinations to a filtered set of reasonable level combinations;
[0159] determining a base leakage value using a maximum likelihood estimator; and
[0160] A best fit data reconstruction of the measured data and corresponding predicted levels is determined by mapping the extracted discrete levels to a linear superposition of each level in a set selected from a filtered set of reasonable level combinations.
[0161] EEE 2 is a method according to EEE 1, wherein the measurement data includes information indicative of a stress induced leakage current (SILC) measurement.
[0162] EEE 3 is a method according to EEE 1, wherein receiving the measurement data includes receiving an electrical signal from a semiconductor device, wherein the semiconductor device includes a high-k gate dielectric layer of a metal oxide semiconductor field effect transistor (MOSFET), wherein at least a portion of the RTN signal indicates a time-dependent dielectric breakdown (TDDB) of the gate dielectric layer.
[0163] EEE 4 is a method according to EEE 3, wherein the measurement data is obtained under normal operating conditions of the semiconductor device, wherein the normal operating conditions of the semiconductor device include a standard operating bias voltage recommended by a manufacturer of the semiconductor device.
[0164] EEE 5 is a method according to EEE 3, wherein the measurement data is not obtained under accelerated operating conditions of the semiconductor device, wherein the accelerated operating conditions include an elevated bias voltage.
[0165] EEE 6 is a method according to EEE 3, wherein the RTN signal is based on the presence of defects in a gate oxide of the semiconductor device.
[0166] EEE 7 is a method according to EEE 1, wherein applying the Bayesian machine learning algorithm includes initially determining an optimal noise level.
[0167] EEE 8 is a method according to EEE 1, wherein applying the clustering model includes initially normalizing the state transition matrix so that corresponding error bars of transition probabilities and level differences are similar.
[0168] EEE 9 is a method according to EEE 1, wherein the clustering model includes an affinity propagation algorithm.
[0169] EEE 10 is a method according to EEE 1, wherein the maximum likelihood estimator is configured to fit a set of Gaussian peaks associated with the extracted discrete levels to a superposition of a filtered set of reasonable level combinations.
[0170] EEE 11 is a method according to EEE 1, further comprising:
[0171] For each given extracted discrete level, it is determined whether the amplitude difference relative to an adjacent predicted level is within a threshold deviation of an independent probability density function (PDF) corresponding to the adjacent predicted level.
[0172] EEE 12 is a method according to EEE 11, wherein a no-solution identifier is assigned to the adjacent predicted level if the amplitude difference relative to the adjacent predicted level is greater than a threshold deviation range of an independent PDF corresponding to a given predicted level.
[0173] EEE 12 is a method according to EEE 11, wherein the threshold deviation range includes 3σ.
[0174] EEE 14 is a method according to EEE 3, further comprising:
[0175] Predict the estimated time to failure (ETTF) of semiconductor devices based on best fit data reconstruction.
[0176] EEE 15 is a method according to EEE 3, further comprising:
[0177] Based on the set of possible transitions, a set of possible defects in a gate dielectric layer of a semiconductor device or a defect density of a gate dielectric layer of a semiconductor device is mapped.
[0178] Although some embodiments have been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be regarded as illustrative and not restrictive. Other variations of the disclosed embodiments may be understood and effected from a study of the drawings, the present disclosure, and the appended claims when implementing the claims. The mere fact that certain measures or features are recited in mutually different dependent claims does not indicate that a combination of these measures or features cannot be used. Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method comprising: receiving measurement data (110), wherein the measurement data is indicative of a random telegraph noise (RTN) signal; extracting a set of extracted discrete levels (122) from the measurement data using a Bayesian machine learning algorithm (120); quantizing the measurement data by assigning each data point of the measurement data to a corresponding extracted discrete level (130); determining a plurality of possible defect configurations based on the quantified measurement data; applying a Markov chain model to the extracted discrete levels (140) to provide a state transition matrix (142), wherein the state transition matrix includes, for each extracted discrete level, a transition probability to a different extracted discrete level; Based on the amplitude difference between each discrete level and the transition probability, an amplitude transition probability space is formed; Applying a clustering model (150) to the amplitude transition probability space to reduce the total number of reasonable level combinations to a filtered set of reasonable level combinations; determining a base leakage value using a maximum likelihood estimator (160); as well as A best fit data reconstruction of the measured data and corresponding predicted levels is determined by mapping the extracted discrete levels to a linear superposition of each level in a set selected from the filtered set of reasonable level combinations.
2. The method according to claim 1, characterized in that The measurement data includes information indicative of a stress induced leakage current (SILC) measurement (112).
3. The method according to any one of claims 1 or 2, characterized in that Receiving measurement data (110) includes receiving an electrical signal from a semiconductor device, wherein the semiconductor device includes a high-k gate dielectric layer of a metal oxide semiconductor field effect transistor (MOSFET) (200, 202), wherein at least a portion of the RTN signal indicates a time-dependent dielectric breakdown (TDDB) of the gate dielectric layer.
4. The method according to claim 3, characterized in that The measurement data is obtained under normal operating conditions of the semiconductor device, wherein the normal operating conditions of the semiconductor device include a standard operating bias voltage recommended by a manufacturer of the semiconductor device.
5. The method according to claim 3, characterized in that: The measurement data is not obtained under accelerated operating conditions of the semiconductor device, wherein the accelerated operating conditions include an elevated bias voltage.
6. The method according to any one of claims 3 to 5, characterized in that The RTN signal is based on the presence of defects in a gate oxide of the semiconductor device.
7. The method according to any one of the preceding claims, characterized in that Applying the Bayesian machine learning algorithm (120) includes initially determining an optimal noise level.
8. The method according to any one of the preceding claims, characterized in that Applying the clustering model (150) includes initially normalizing the state transition matrix so that corresponding error bars of transition probabilities and level differences are similar.
9. The method according to any one of the preceding claims, characterized in that The clustering model includes an affinitized neighbor propagation algorithm.
10. The method according to any one of the preceding claims, characterized in that The maximum likelihood estimator (160) is configured to fit a set of Gaussian peaks associated with the extracted discrete levels to a superposition of the filtered set of reasonable level combinations.
11. The method according to any one of the preceding claims, characterized in that Also includes: For each given extracted discrete level, it is determined whether the amplitude difference relative to an adjacent predicted level is within a threshold deviation of an independent probability density function (PDF) corresponding to the adjacent predicted level.
12. The method according to claim 11, characterized in that If the magnitude difference relative to the neighboring predicted level is greater than a threshold deviation range for an independent PDF corresponding to a given predicted level, a no solution identifier is assigned to the neighboring predicted level.
13. The method according to claim 11 or 12, characterized in that: The threshold deviation range includes three σ.
14. The method according to any one of claims 3 to 5, characterized in that Also includes: An estimated time to failure (ETTF) of the semiconductor device is predicted based on the best fit data reconstruction.
15. The method according to any one of claims 3 to 5, characterized in that Also includes: Based on the set of possible transitions, a set of possible defects in a gate dielectric layer of the semiconductor device or a defect density of a gate dielectric layer of the semiconductor device is mapped.