Neural network training method and device, equipment and storage medium
By decomposing or transforming the original Heczone matrix, the target Heczone matrix is obtained, which is used to reduce the memory usage and computational complexity in the neural network training process, solving the problems of slow training speed and large memory usage in the existing technology, and achieving more efficient neural network training.
Patent Information
- Application Number
- CN202510645222.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When using the Levinberg-Marquard algorithm to train neural networks, the sample data and model parameters are large, resulting in a large amount of memory required for the training process and slow speed.
By decomposing or converting the original Hellscrew matrix, the calculation complexity is reduced and the target Hellscrew matrix is obtained. Then use the target optimization algorithm to train the neural network to reduce the need for data storage during the training process.
It reduces the storage amount and calculation complexity required during neural network training, and improves training speed and accuracy.
Smart Images

Figure CN120197663A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a neural network training method, apparatus, device, and storage medium. Background Art
[0002] In the manufacture of high-density integrated circuits, the size measurement of various semiconductor materials is crucial for ensuring device performance.
[0003] Currently, considering the measurement accuracy requirements for the size of semiconductor materials, the Levenberg-Marquardt (LM) optimization algorithm can be used to train a neural network, and then the size of the semiconductor material can be obtained through the trained neural network.
[0004] However, in the above method, there are a large number of sample data participating in the neural network training, a large number of parameters in the neural network model, the training process requires extremely large training memory, and the training speed is slow. Summary of the Invention
[0005] Embodiments of this application provide a neural network training method, apparatus, device, and storage medium, which can reduce the storage amount required in the training process and improve the training speed.
[0006] In a first aspect, embodiments of this application provide a neural network training method, including: Inputting N sample data into a first original neural network to obtain N predicted spectra. Each sample data in the N sample data includes the size of a first device, the first device includes a wafer, and each sample data is associated with a true spectrum corresponding to the size of the first device. N is a positive integer greater than or equal to 2; obtaining a loss value according to the N predicted spectra and the N true spectra; training the first original neural network according to the loss value through a target optimization algorithm to obtain a first neural network. The target optimization algorithm is obtained through a first process, and the target optimization algorithm includes a target Hessian matrix. The first process includes decomposing the original Hessian matrix to obtain the target Hessian matrix, or transforming the original optimization algorithm to obtain the target optimization algorithm. The computational complexity of the target optimization algorithm after the first process is lower than that before the first process. The target optimization algorithm is used to adjust the network parameters of the first original neural network.
[0007] In some embodiments, the target optimization algorithm is the Levenberg-Marquardt LM algorithm.
[0008] In a possible implementation, the decomposition process includes performing matrix decomposition or first-order approximation on the original Hessian matrix. Matrix decomposition includes eigenvalue decomposition or singular value decomposition. Eigenvalue decomposition decomposes the original Hessian matrix into eigenvectors and eigenvalues, and singular value decomposition decomposes the original Hessian matrix into an orthogonal matrix and a diagonal matrix. The first-order approximation process reconstructs the original Hessian matrix by first-order derivative approximation. The conversion process includes converting the first formula from taking the inverse of the target Hessian matrix to solving a linear equation to obtain the target optimization algorithm. The first formula is the formula of the original optimization algorithm, and the first formula includes the target Hessian matrix.
[0009] In a possible implementation, the method further includes: During the training of the first original neural network, monitoring the loss value, the gradient of the first original neural network, and the curvature of the target Hessian matrix; when the first condition is satisfied, stopping the training of the first original neural network. The first condition includes at least one of the following: the decrease amplitude between adjacent loss values among the multiple loss values corresponding to multiple iterative trainings is less than the first threshold, the gradient is less than or equal to the second threshold, and the curvature is positive; inputting the N sample data into the first original neural network again.
[0010] In a possible implementation, before inputting the N sample data into the first original neural network to obtain N predicted spectra, or before inputting the N sample data into the first original neural network again, the method further includes: Training the second original neural network through a first-order optimizer to obtain the second neural network; determining the first network parameters of the second neural network; determining the first network parameters among all the network parameters of the first original neural network as the initial values of the network parameters of the first original neural network.
[0011] In a possible implementation, inputting the N sample data into the first original neural network to obtain N predicted spectra includes: obtaining N data, where each of the N data includes the size of the first device; preprocessing the N data to obtain N sample data, and the preprocessing includes dimensionality reduction processing; inputting the N sample data into the first original neural network to obtain N predicted spectra.
[0012] In a possible implementation, after training the first original neural network according to the loss value through a target optimization algorithm to obtain the first neural network, the method further includes: determining the network quality of the first neural network; when the network quality meets the second condition, performing a second process; where the second condition includes that the goodness of fit of the first neural network is less than a third threshold, and the goodness of fit is used to indicate the matching degree between the predicted spectrum corresponding to the first neural network and the true spectrum, and the second process includes at least one of adjusting the number of N sample data, performing preprocessing again, and adjusting the structure of the first original neural network.
[0013] In a possible implementation, training the first original neural network according to the loss value through a target optimization algorithm to obtain the first neural network includes: inputting each loss value in the loss value into the target optimization algorithm to obtain a model parameter update amount; updating the model parameters of the first original neural network according to the model parameter update amount to obtain the first neural network.
[0014] In a possible implementation, the optimization algorithm is the Levenberg-Marquardt LM algorithm.
[0015] In a second aspect, an embodiment of the present application provides a neural network training device, including: An obtaining module, configured to input N sample data into the first original neural network to obtain N predicted spectra. Each sample data in the N sample data includes the size of a first device. The first device includes a wafer. Each sample data is associated with a true spectrum corresponding to the size of the first device. N is a positive integer greater than or equal to 2; The obtaining module is further configured to obtain a loss value according to the N predicted spectra and the N true spectra; A training module, configured to train the first original neural network according to the loss value through a target optimization algorithm to obtain the first neural network. The target optimization algorithm is obtained through a first process. The target optimization algorithm includes a target Hessian matrix. The first process includes decomposing the original Hessian matrix to obtain the target Hessian matrix, or performing a transformation process on the original optimization algorithm to obtain the target optimization algorithm. The computational complexity of the target optimization algorithm after the first process is lower than that before the first process. The target optimization algorithm is used to adjust the network parameters of the first original neural network.
[0016] In a third aspect, an embodiment of the present application provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the neural network training method according to any one of the first aspects above.
[0017] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the neural network training method described in any one of the above first aspects.
[0018] Fifthly, an embodiment of the present application provides a computer program product, which when running on a terminal device, causes the terminal device to execute the neural network training method described in any one of the above first aspects.
[0019] It can be understood that the beneficial effects of the above second to fifth aspects can be referred to the relevant descriptions in the above first aspect, and will not be repeated here.
[0020] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: Input N sample data into the first original neural network to obtain N predicted spectra. According to the N predicted spectra and N true spectra, obtain a loss value, so as to facilitate training the first neural network according to the loss value. Specifically, through a target optimization algorithm, train the first original neural network according to the loss value to obtain the first neural network. Among them, the target optimization algorithm can perform a first process in advance to reduce the computational complexity of the original optimization algorithm and obtain the target optimization algorithm. In this way, when the target optimization algorithm trains the first original neural network, the data storage amount in the memory during the calculation process can be reduced, thereby reducing the data storage amount in the overall processing process of the optimization algorithm and ensuring the smoothness of the optimization algorithm for training the first original neural network, and improving the training speed of the optimization algorithm for the first original neural network. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 is a schematic flowchart of a neural network training method provided by an embodiment of the present application; Figure 2 is a schematic flowchart of a neural network training method provided by another embodiment of the present application; Figure 3 is a schematic flowchart of a neural network training method provided by another embodiment of the present application; Figure 4 is a schematic structural diagram of a neural network training device provided by an embodiment of the present application. Detailed Embodiments
[0023] In the following description, for purposes of illustration and not limitation, specific details such as specific system architectures, technologies, etc. are set forth in order to provide a thorough understanding of embodiments of the present application. However, those skilled in the art should appreciate that the present application may be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary details.
[0024] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their combinations.
[0025] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0026] As used in the specification of the present application and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be construed, depending on the context, as meaning "once determined" or "in response to determining" or "once detected [the described condition or event]" or "in response to detecting [the described condition or event]".
[0027] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0028] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0029] For ease of understanding, examples are given to illustrate some concepts related to the embodiments of the present application for reference.
[0030] 1. LM Algorithm An optimization algorithm for nonlinear least squares problems, commonly used in curve fitting and neural network training. It combines the advantages of gradient descent and Gauss-Newton methods, achieving a balance between convergence speed and stability.
[0031] 2. Hessian Matrix The Hessian matrix, also known as the Hessian matrix, Hesse matrix, or H matrix, is a square matrix composed of the second partial derivatives of a multivariate function, describing the local curvature of the function.
[0032] 3. First-Order Optimizer A first-order optimizer (first-order optimizer) is an optimization algorithm that updates parameters based on the first derivative (gradient) of the loss function. Its core goal is to iteratively adjust the model parameters to minimize the loss function, specifically by using the first derivative (gradient) of the loss value function to adjust the model parameters to minimize the loss value function. Different from second-order optimizers (such as Newton's method and the LM algorithm), first-order optimizers do not calculate the Hessian matrix (second derivative).
[0033] 4. Wafer A wafer refers to the silicon wafer used to fabricate silicon semiconductor circuits, and its raw material is silicon. High-purity polysilicon is dissolved and doped with silicon crystal seeds, and then slowly pulled out to form a cylindrical single-crystal silicon. After the silicon ingot is ground, polished, and sliced, a silicon wafer is formed, which is the wafer.
[0034] In the manufacturing of high-density integrated circuits, the dimensional measurement of various semiconductor materials is crucial for ensuring the performance of semiconductor devices. Ellipsometry, as an important nanoscale metrology technique, is widely used to measure the critical dimensions and optical constants of semiconductor material structures due to its surface sensitivity, non-destructiveness, non-invasiveness, etc. Ellipsometry analyzes the change in the polarization state of polarized light after reflection from the sample, combines with theoretical models, uses the Fourier modal method (FMM) and similar algorithms to calculate and simulate spectra, and then obtains the final dimensional parameters of the semiconductor material by fitting the simulated spectra with the experimental spectra. For traditional complex nanoscale wafer structures, the measurement methods for their width, length, and shape are based on complex electromagnetic solution calculations, which are accurate but have high computational costs and slow speeds. To accelerate the measurement process, neural networks are used to replace traditional calculations.
[0035] Currently, considering the measurement accuracy requirements for the size of semiconductor materials, the Levenberg-Marquardt (LM) optimization algorithm can be used to train a neural network, and then the size of the semiconductor material can be obtained through the trained neural network.
[0036] However, in the above method, there are a large number of sample data participating in the neural network training, and the number of parameters in the neural network model is large. The training process requires a large amount of training memory and the training speed is slow.
[0037] In the face of the above problems, the present application proposes a neural network training method, device, server, and computer-readable storage medium. In this method, the original Hessian matrix in the optimization algorithm can be decomposed or transformed in advance to reduce the complexity of the original Hessian matrix, thereby obtaining a target Hessian matrix. As the complexity of the target Hessian matrix decreases, when the optimization algorithm trains the first original neural network, the computational complexity of the target Hessian matrix decreases, and the amount of data to be stored in memory decreases, so that the training speed and accuracy of the optimization algorithm for training the first original neural network can be improved, and the memory occupancy can also be reduced.
[0038] The neural network training method provided by the embodiments of the present application can be applied to a server, and can also be applied to terminal devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of terminal devices.
[0039] Figure 1 The schematic flowchart of the neural network training method provided by the present application is shown. By way of example and not limitation, this method can be applied to the above-mentioned server.
[0040] Please refer to Figure 1 , Figure 1 which shows the schematic flowchart of the neural network training method provided by an embodiment of the present application.
[0041] As Figure 1 shown, the neural network training method provided by the present application may include: S101. Input N sample data into the first original neural network to obtain N predicted spectra.
[0042] Among them, each of the N sample data includes the size of the first device. The first device includes a wafer, where N is a positive integer greater than or equal to 2. Each sample data is associated with a true spectrum corresponding to the size of the first device.
[0043] Among them, each of the N sample data includes the size of the first device. The size here may include, but is not limited to, length, width, film thickness, trench depth, and shape. The present application does not limit the specific type of the size of the first device.
[0044] The first device may be a wafer or other devices. The present application does not limit the specific type of the first device.
[0045] If the first original neural network is trained to measure the size of a wafer, then each of the N sample data includes the wafer size. If the first original neural network is trained to measure the size of other devices, then each of the N sample data includes the size of other devices. That is to say, the data type included in each of the N sample data depends on the data type for which the first original neural network is trained to measure.
[0046] For example, each of the N sample data includes the size of a semiconductor chip.
[0047] Among them, each sample data is associated with a true spectrum corresponding to the size of the first device. As the size of the first device is different, the true spectrum corresponding to this size is different.
[0048] For example, when the first device is a wafer, different thicknesses of the silicon dioxide film on the wafer will cause changes in the reflection light interference conditions, and the positions of the interference peaks in the spectrum will shift accordingly. That is, wafers of different sizes can correspond to different true spectra.
[0049] Another example is that when the first device is a wafer, the shape of the wafer can affect the spectrum. The shape of the periodic microstructure (such as a rectangle or a trapezoid) will affect the distribution of the diffraction spectrum.
[0050] Based on the above description, since each of the N sample data is associated with a true spectrum corresponding to the size of the first device, therefore, by inputting the N sample data into the first original neural network, N predicted spectra can be obtained, and the N true spectra and the N predicted spectra correspond one by one.
[0051] S102. Obtain a loss value according to the N predicted spectra and the N true spectra.
[0052] The loss value can be used to represent the difference between the predicted spectrum and the true spectrum.
[0053] S103. Train the first original neural network according to the loss value through a target optimization algorithm to obtain the first neural network.
[0054] Among them, the target optimization algorithm is obtained through a first process. The target optimization algorithm includes a target Hessian matrix. The first process includes decomposing the original Hessian matrix to obtain the target Hessian matrix, or performing a transformation process on the original optimization algorithm to obtain the target optimization algorithm. The computational complexity of the target optimization algorithm after the first process is lower than that before the first process. The target optimization algorithm is used to adjust the network parameters of the first original neural network.
[0055] It should be understood that the essence of the optimization algorithm is to solve the equation containing the inverse of the second-order original Hessian matrix to obtain the updated network parameters. However, the sample data volume is large, and the number of parameters of the first original neural network is large, with high complexity, resulting in a high computational complexity of the original Hessian matrix, and the calculation is extremely time-consuming and memory-consuming.
[0056] That is to say, if directly using the optimization algorithm including the original Hessian matrix to train the first original neural network, problems such as low accuracy, long training time, and high memory occupancy may occur.
[0057] Based on this, before the server trains the first original neural network, some approximation processing methods can be adopted to simplify the expression of the original target optimization algorithm. This processing method can be the first process, and the first process includes a decomposition process or a transformation process.
[0058] In some embodiments, the decomposition process includes performing matrix decomposition or first-order approximation processing on the original Hessian matrix. Matrix decomposition includes eigenvalue decomposition or singular value decomposition. Eigenvalue decomposition is to decompose the original Hessian matrix into eigenvectors and eigenvalues. Singular value decomposition is to decompose the original Hessian matrix into an orthogonal matrix and a diagonal matrix. First-order approximation processing is to reconstruct the original Hessian matrix through first-order derivative approximation.
[0059] Exemplarily, the eigenvalue decomposition can refer to the following formula: The original Hessian matrix H ∈ R d×d , can be decomposed into: H = QΛQ T .
[0060] Among them, Q is an orthogonal matrix (QQ T = I), and its column vectors are eigenvectors. Λ = diag(λ 1 , λ 2 ,..., λ d ) is a diagonal matrix, and the elements are eigenvalues.
[0061] Thus, through the eigenvalue decomposition in the above formula, the inverse of the Hessian matrix can be quickly calculated.
[0062] Exemplarily, the singular value decomposition can be seen in the following formula: The original Hessian matrix H ∈ R d×d , can be decomposed as: H = UΣV T .
[0063] Where U and V are orthogonal matrices, and Σ = diag(σ 1 , σ 2 ,..., σ d ) is a diagonal matrix.
[0064] Thus, low-rank approximation is achieved by retaining the first k principal components through the singular value decomposition in the above formula.
[0065] Among them, the first-order approximation process can be understood as ignoring the second-order derivative terms in H and approximating H only with first-order information.
[0066] In one possible implementation, the original Hessian matrix can be reconstructed by first-order derivative approximation in the form of Gauss-Newton approximation, and the formula can be: H ≈ J T J, where J is the Jacobian matrix of the residual function with respect to the parameters.
[0067] In another possible implementation, the original Hessian matrix can be reconstructed by first-order derivative approximation in the form of diagonal approximation: only retain the diagonal elements of H, that is, Hdiag = diag(H 11 , H 22 , …, H nn ), and the inversion complexity is reduced from O(n 3 ) to O(n).
[0068] Based on the above description, the server can decompose the original Hessian matrix in the optimization algorithm, reduce the computational complexity of the original Hessian matrix, and obtain the target Hessian matrix. In this way, when the target optimization algorithm including the target Hessian matrix trains the first original neural network, the data storage amount in memory during the calculation process can be reduced, thereby reducing the data storage amount of the overall processing process of the target optimization algorithm, ensuring the smoothness of the optimization algorithm for training the first original neural network, and improving the training speed of the optimization algorithm for the first original neural network.
[0069] In some embodiments, the conversion process includes converting the first formula from inverting the target Hessian matrix to solving a linear equation to obtain the target optimization algorithm. The first formula is the formula of the original optimization algorithm, and the first formula includes the target Hessian matrix.
[0070] Among them, the conversion process can be understood as converting matrix inversion into solving a system of linear equations and using the iterative method to avoid directly storing H.
[0071] Exemplarily, the calculation formula of the optimization algorithm is dw = -(H + μI) -1 g, where H + μI is the damping term, which is used to ensure the positive definiteness of the optimization algorithm.
[0072] The conversion process can be: converting dw = -(H + μI) -1 g to (H + μI)dw = -g.
[0073] Based on this, the server can perform conversion processing on the first formula corresponding to the original optimization algorithm to avoid directly storing H. In this way, when the target optimization algorithm including the target Hessian matrix trains the first original neural network, the data storage amount in memory during the calculation process can be reduced, thereby reducing the data storage amount in the overall processing process of the target optimization algorithm, ensuring the smoothness of the optimization algorithm for training the first original neural network, and improving the training speed of the optimization algorithm for the first original neural network.
[0074] In some embodiments, the optimization algorithm is the Levenberg-Marquardt LM algorithm, and the LM algorithm is a second-order gradient optimization algorithm. The optimization algorithm can also be the quasi-Newton method (limited-memory BFGS, L-BFGS), the natural gradient method, etc. The specific type of the optimization algorithm in this application is not limited.
[0075] In some embodiments, through the target optimization algorithm, the first original neural network is trained according to the loss value to obtain the first neural network, including: inputting each loss value in the loss value into the target optimization algorithm to obtain the model parameter update amount; updating the model parameters of the first original neural network according to the model parameter update amount to obtain the first neural network.
[0076] It should be noted that when the loss value is less than or equal to the error threshold, the network stops training, and the network at the time of stopping training is the first neural network.
[0077] In the embodiments of this application, the server inputs N sample data into the first original neural network to obtain N predicted spectra, and obtains the loss value according to the N predicted spectra and the N true spectra, so as to facilitate training the first neural network according to the loss value. Specifically, through the target optimization algorithm, the first original neural network is trained according to the loss value to obtain the first neural network. Among them, the target optimization algorithm can perform the first processing in advance to reduce the computational complexity of the original optimization algorithm to obtain the target optimization algorithm. In this way, when the target optimization algorithm trains the first original neural network, the data storage amount in memory during the calculation process can be reduced, thereby reducing the data storage amount in the overall processing process of the optimization algorithm, ensuring the smoothness of the optimization algorithm for training the first original neural network, and improving the training speed of the optimization algorithm for the first original neural network.
[0078] Based on Figure 1 According to the description of S103, during the training of the first original neural network, the server can also monitor the training process. Moreover, after the training of the first original neural network is completed, the server can also determine the network quality of the first neural network and the database quality of the database where the N sample data are located, so as to facilitate relevant adjustments and retraining when the network quality is poor.
[0079] As Figure 2 shown, the server can monitor the training process in real time during the training of the first original network, so as to facilitate stopping the training when falling into a local minimum; after the training is completed, it can also perform result evaluation and make relevant adjustments and retraining when the network quality of the first neural network is poor.
[0080] Please refer to Figure 3 , Figure 3 which shows a schematic flowchart of a neural network training method provided by an embodiment of the present application.
[0081] As Figure 3 shown, the neural network training method provided by the present application may include: S201. Obtain N data, and each of the N data includes the size of the first device.
[0082] Among them, there are various methods for obtaining the N data.
[0083] In some embodiments, the N data are obtained by the user manually measuring the size of the first device with a micrometer, vernier caliper, laser rangefinder and other manual measuring tools.
[0084] In other embodiments, the N data are obtained by measuring the size of the first device with an automated measuring device such as a 3D scanner, machine vision system, coordinate measuring machine, etc.
[0085] In other embodiments, the N data are obtained by simulating the size of the first device with drawing software such as computer-aided design (CAD) software.
[0086] In other embodiments, the N data include the sizes of the first devices with existing production records.
[0087] S202. Preprocess the N data to obtain N sample data, and the preprocessing includes dimensionality reduction processing.
[0088] The dimensionality reduction process can adopt the principal component analysis (PCA) method. The dimensionality reduction process is used to remove redundant or irrelevant features by reducing the dimensions of data features (i.e., the number of variables), while retaining as much important information of the original data as possible. For example, any one of the N data includes 5 features of the length, width, height, volume, and surface area of the first device. Since the volume and surface area can be calculated from the length, width, and height (redundant), the data can be reduced to 3 principal components (length, width, and height) without losing information.
[0089] In addition, the preprocessing can also include denoising processing, data augmentation processing (augmenting the sample size based on the N data), outlier processing (which can be deletion), etc. This application does not make specific limitations on the type of preprocessing.
[0090] S203. Input the N sample data into the first original neural network to obtain N predicted spectra.
[0091] S204. Obtain a loss value based on the N predicted spectra and the N true spectra.
[0092] S205. Through a target optimization algorithm, train the first original neural network according to the loss value to obtain a first neural network. The target optimization algorithm is obtained through a first process. The target optimization algorithm includes a target Hessian matrix. The first process includes decomposing the original Hessian matrix to obtain the target Hessian matrix, or performing a transformation process on the original optimization algorithm to obtain the target optimization algorithm. The computational complexity of the target optimization algorithm after the first process is lower than that before the first process. The target optimization algorithm is used to adjust the network parameters of the first original neural network.
[0093] Among them, S203, S204, and S205 are respectively similar to Figure 1 the implementation manners of S101, S102, and S103 in the illustrated embodiment, and will not be specifically described here.
[0094] S206. During the training of the first original neural network, monitor the loss value, the gradient of the first original neural network, and the curvature of the target Hessian matrix.
[0095] It should be understood that the optimization algorithm is a trust region method, which is prone to falling into a local minimum, resulting in training stagnation, and thus the final training effect is poor: specifically, the optimization algorithm iterates the network parameters multiple times, and the reduced loss value is very small, and the gradient is close to 0, indicating that the training is close to completion, but the current overall loss value level is still at a relatively high position, indicating that the current training has fallen into a local minimum and it is difficult to iterate out of it.
[0096] Based on this, during the process of the server training the first original neural network, it is necessary to conduct monitoring. Specifically, the loss value, the gradient of the first original neural network, and the curvature of the target Hessian matrix can be monitored.
[0097] S207. Determine whether the first condition is satisfied. The first condition includes at least one of the following: the reduction amplitude between adjacent loss values among the multiple loss values corresponding to multiple iterative trainings is less than the first threshold, the gradient is less than or equal to the second threshold, and the curvature is a positive value.
[0098] Generally, if the reduction amplitudes of consecutive loss values among the multiple loss values corresponding to multiple iterative trainings are all less than the first threshold, it indicates that the first original neural network is not continuously optimized, and the first original neural network may fall into a local minimum.
[0099] For example, if the preset number of times is 10 and the loss value does not decrease in 10 consecutive iterations, it indicates that the training of the network may fall into a local minimum.
[0100] The gradient can be understood as the partial derivative of the loss function with respect to the network parameters, indicating the update direction of the network parameters. During the training process of the first original neural network, if the gradient is less than or equal to the second threshold, it indicates that the gradient approaches 0, and the first original neural network may fall into a local minimum.
[0101] The curvature is used to determine whether the current network parameters are in a convex region (the target Hessian matrix is positive definite) or a saddle point / non-convex region (including negative eigenvalues). If the curvature is a positive value, it indicates that the target Hessian matrix is positive definite, indicating that it has fallen into a relatively wide flat saddle point neighborhood, that is, the first original neural network corresponding to the current network parameters may fall into a local minimum.
[0102] To prevent misjudgment, the first condition can include at least two of the above three conditions. That is to say, it is necessary to comprehensively consider at least two indicators to more accurately determine whether the current training has fallen into a local minimum.
[0103] When the first condition is satisfied, the server can execute S208; when the first condition is not satisfied, the server can execute S209.
[0104] S208. When the first condition is satisfied, stop training the first original neural network, and input the N sample data into the first original neural network again.
[0105] If the first condition is satisfied, it indicates that the first original neural network may fall into a local minimum and the training has stagnated. It is necessary to stop training the first original neural network and re-input the N sample data into the first original neural network for training.
[0106] In some embodiments, before inputting N sample data into the first original neural network to obtain N predicted spectra, or before inputting N sample data into the first original neural network again, the method further includes: The second original neural network is trained by a first-order optimizer to obtain a second neural network; a first network parameter of the second neural network is determined; and a first network parameter among all network parameters of the first original neural network is determined as an initial value of the network parameter of the first original neural network.
[0107] Among them, the first-order optimizer can be an adaptive moment estimation (Adam) algorithm or a stochastic gradient descent (stochastic gradient descent) algorithm. This application does not limit the specific type of the first-order optimizer.
[0108] The second original neural network can be trained multiple times, and the first-order optimizer used for training the second original neural network each time can be different. The second original neural network is trained multiple times to obtain multiple network parameters, and the optimal network parameter can be determined from the multiple network parameters as the first network parameter. For example, the optimal network parameter can be the network parameter that is the same the most times among the multiple network parameters.
[0109] Usually, the initial value of the network parameter of the first original neural network is any one of all the network parameters of the first original neural network. Training the first original neural network based on the initial value is more likely to fall into a local minimum during the training of the first original neural network.
[0110] Based on this, the server can first train another neural network, that is, the second original neural network, through a first-order optimizer to obtain the second neural network. At this time, the first network parameters are the best network parameters for the second neural network. Therefore, the server can directly apply the first network parameters to the first original neural network as the initial values of the network parameters of the first original neural network. In this way, the first network parameters are better neural network parameters for the first original neural network.
[0111] Therefore, using the first network parameters as the initial values of the network parameters of the first original neural network can avoid falling into the local minimum during the training of the first original neural network as much as possible. In addition, the process of determining the initial values of the network parameters of the first original neural network is simple and efficient, and takes a short time, so that the training process can reach the global minimum more quickly, effectively improving the training speed and improving the network accuracy.
[0112] S209: When the first condition is not met, no operation is performed.
[0113] When the first condition is not met, it indicates that the first original neural network is in the normal training process and no operation needs to be performed, which can ensure the normal training of the first original neural network.
[0114] S210. Determine the network quality of the first neural network.
[0115] The network quality of the first neural network determines the accuracy of the first neural network during subsequent applications. Therefore, it is necessary to determine the network quality of the first neural network.
[0116] S211. When the network quality meets the second condition, perform a second process.
[0117] Among them, the second condition includes that the goodness of fit of the first neural network is less than a third threshold, and the goodness of fit is used to indicate the matching degree between the predicted spectrum and the true spectrum corresponding to the first neural network. The second process includes at least one of adjusting the number of N sample data, performing preprocessing again, and adjusting the structure of the first original neural network.
[0118] After obtaining the first neural network, at least one device size can be selected to test the first neural network, and a predicted spectrum can be obtained. By comparing the difference between the predicted spectrum and the true spectrum corresponding to the foregoing device size, the goodness of fit can be obtained. Thus, the network quality can be determined through the goodness of fit. The greater the goodness of fit, the higher the network quality of the first neural network. Similarly, the smaller the goodness of fit, the lower the network quality of the first neural network.
[0119] The specific network quality can be measured by the third threshold. If the goodness of fit of the first neural network is less than the third threshold, the network quality meets the second condition and the network quality is poor.
[0120] Based on this, if the goodness of fit is less than the third threshold, the server can perform the second process. The second process includes at least one of adjusting the number of N sample data (expanding the sample size), performing preprocessing again, and adjusting the structure of the first original neural network (such as adjusting the number of convolutional layers, stride, etc.). Thus, the server can improve the network quality of the first neural network through the second process.
[0121] In the embodiment of the present application, during the training process of the first original neural network, the loss value, the gradient of the first original neural network, and the curvature of the target Hessian matrix can be monitored. If the first condition is met (the first condition includes at least one of that the decrease amplitude between adjacent loss values among the multiple loss values corresponding to multiple iterative trainings is less than the first threshold, the gradient is less than or equal to the second threshold, and the curvature is positive), it indicates that the first original neural network may fall into a local minimum and the training stagnates. It is necessary to stop training the original neural network and re-enter the N sample data into the first original neural network for training.
[0122] Moreover, before inputting N sample data into the first original neural network to obtain N predicted spectra, or before inputting the N sample data into the first original neural network again, the server can obtain the initial network parameters of the first original neural network through a first-order optimizer, which can avoid falling into local minima as much as possible during the training process of the first original neural network. Moreover, the process of determining the initial network parameters of the first original neural network is simple and efficient, takes a short time, and can enable the training process to reach the global minimum more quickly, effectively improving the training speed and the model accuracy.
[0123] Moreover, the server can also determine the network quality of the first neural network. If the network quality meets the second condition, the server can perform a second process, and the second process includes at least one of adjusting the number of N sample data, performing preprocessing again, and adjusting the structure of the first original neural network. Thus, the server can improve the network quality of the first neural network through the second process and ensure the prediction quality when the first neural network is finally applied.
[0124] In addition, after obtaining the first neural network, the size of the target device can be predicted by using the first neural network, and the target device and the first device can be devices of the same type.
[0125] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0126] Corresponding to the neural network training method described in the above embodiments, Figure 4 FIG. shows a structural block diagram of a neural network training apparatus 300 provided by an embodiment of the present application. For the sake of simplicity, only the parts related to the embodiment of the present application are shown.
[0127] Referring to Figure 4 , the neural network training apparatus 300 includes: An obtaining module 301, configured to input N sample data into the first original neural network to obtain N predicted spectra. Each sample data in the N sample data includes the size of the first device. The first device includes a wafer. Each sample data is associated with a true spectrum corresponding to the size of the first device. N is a positive integer greater than or equal to 2.
[0128] The obtaining module 301 is further configured to obtain a loss value according to the N predicted spectra and the N true spectra.
[0129] A training module 302, configured to train a first original neural network according to a loss value by a target optimization algorithm to obtain a first neural network. The target optimization algorithm is obtained through a first process. The target optimization algorithm includes a target Hessian matrix. The first process includes decomposing an original Hessian matrix to obtain the target Hessian matrix, or transforming an original optimization algorithm to obtain the target optimization algorithm. The computational complexity of the target optimization algorithm after the first process is lower than that before the first process. The target optimization algorithm is used to adjust the network parameters of the first original neural network.
[0130] In some embodiments, the decomposition process includes performing matrix decomposition or first-order approximation on the original Hessian matrix. Matrix decomposition includes eigenvalue decomposition or singular value decomposition. Eigenvalue decomposition decomposes the original Hessian matrix into eigenvectors and eigenvalues. Singular value decomposition decomposes the original Hessian matrix into an orthogonal matrix and a diagonal matrix. First-order approximation reconstructs the original Hessian matrix through first-order derivative approximation. The transformation process includes converting a first formula from inverting the target Hessian matrix to solving a linear equation to obtain the target optimization algorithm. The first formula is the formula of the original optimization algorithm, and the first formula includes the target Hessian matrix.
[0131] In some embodiments, the neural network training device 300 may further include a monitoring module, configured to: monitor the loss value, the gradient of the first original neural network, and the curvature of the target Hessian matrix during the training of the first original neural network; stop training the first original neural network when a first condition is met. The first condition includes at least one of that the reduction amplitude between adjacent loss values among the multiple loss values corresponding to multiple iterative trainings is less than a first threshold, the gradient is less than or equal to a second threshold, and the curvature is a positive value; input the N sample data into the first original neural network again.
[0132] In some embodiments, the training module 302 is specifically configured to: train a second original neural network through a first-order optimizer to obtain a second neural network; determine the first network parameters of the second neural network; determine the first network parameters among all the network parameters of the first original neural network as the initial values of the network parameters of the first original neural network.
[0133] In some embodiments, the obtaining module 301 is specifically configured to: obtain N data, where each of the N data includes the size of a first device; preprocess the N data to obtain N sample data, and the preprocessing includes dimensionality reduction processing; input the N sample data into the first original neural network to obtain N predicted spectra.
[0134] In some embodiments, the neural network training device 300 may further include an evaluation module, and the evaluation module is configured to: determine the network quality of the first neural network; when the network quality meets the second condition, perform a second process; wherein, the second condition includes that the goodness of fit of the first neural network is less than a third threshold, and the goodness of fit is used to indicate the matching degree between the predicted spectrum corresponding to the first neural network and the true spectrum, and the second process includes at least one of adjusting the number of N sample data, performing preprocessing again, and adjusting the structure of the first original neural network.
[0135] In some embodiments, the training module 302 is specifically configured to: input the loss value into a target optimization algorithm to obtain a model parameter update amount; update the model parameters of the first original neural network according to the model parameter update amount to obtain the first neural network.
[0136] In some embodiments, the optimization algorithm is the Levenberg-Marquardt LM algorithm.
[0137] It should be noted that the information interaction, execution process, etc. between the above-mentioned device / units, due to being based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0138] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details are not described herein again.
[0139] The embodiment of the present application further provides a server, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and when the processor executes the computer program, the steps in any of the foregoing method embodiments are implemented.
[0140] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0141] An embodiment of the present application provides a computer program product. When the computer program product runs on a server, the server is caused to execute the steps in the above-mentioned method embodiments.
[0142] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0143] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0144] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0145] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0146] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0147] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A neural network training method, characterized in that: include: Inputting N sample data into a first original neural network to obtain N predicted spectra, wherein each sample data in the N sample data includes a size of a first device, the first device includes a wafer, and each sample data is associated with a real spectrum corresponding to the size of the first device, and N is a positive integer greater than or equal to 2; Obtaining a loss value according to the N predicted spectra and the N real spectra; The first original neural network is trained according to the loss value through a target optimization algorithm to obtain a first neural network. The target optimization algorithm is obtained through a first process. The target optimization algorithm includes a target Hessian matrix. The first process includes decomposing the original Hessian matrix to obtain the target Hessian matrix, or converting the original optimization algorithm to obtain the target optimization algorithm. After the first process, the computational complexity of the target optimization algorithm is lower than the computational complexity before the first process. The target optimization algorithm is used to adjust the network parameters of the first original neural network.
2. The method according to claim 1, characterized in that The decomposition process includes performing matrix decomposition or first-order approximation process on the original Hessian matrix, the matrix decomposition includes eigenvalue decomposition or singular value decomposition, the eigenvalue decomposition is to decompose the original Hessian matrix into eigenvectors and eigenvalues, the singular value decomposition is to decompose the original Hessian matrix into an orthogonal matrix and a diagonal matrix, and the first-order approximation process is to reconstruct the original Hessian matrix through first-order derivative approximation; The conversion process includes converting a first formula from inverting the target Hessian matrix to solving a linear equation to obtain the target optimization algorithm, the first formula is the formula of the original optimization algorithm, and the first formula includes the target Hessian matrix.
3. The method according to claim 1, characterized in that The method further comprises: During the training of the first original neural network, monitoring the loss value, the gradient of the first original neural network, and the curvature of the target Hessian matrix; When a first condition is met, the training of the first original neural network is stopped, wherein the first condition includes at least one of the following: the magnitude of decrease between adjacent loss values in the multiple loss values corresponding to the multiple iterative trainings is less than a first threshold, the gradient is less than or equal to a second threshold, and the curvature is a positive value; The N sample data are input into the first original neural network again.
4. The method according to claim 3, characterized in that Before inputting the N sample data into the first original neural network to obtain N predicted spectra, or before inputting the N sample data into the first original neural network again, the method further includes: Training the second original neural network through a first-order optimizer to obtain a second neural network; determining first network parameters of the second neural network; The first network parameter among all network parameters of the first original neural network is determined as an initial value of the network parameter of the first original neural network.
5. The method according to any one of claims 1 to 4, characterized in that The step of inputting N sample data into the first original neural network to obtain N predicted spectra includes: Acquire N data, each of the N data including the size of the first device; Preprocessing the N data to obtain the N sample data, wherein the preprocessing includes dimensionality reduction processing; The N sample data are input into the first original neural network to obtain the N predicted spectra.
6. The method according to claim 5, characterized in that After the first original neural network is trained according to the loss value by the target optimization algorithm to obtain the first neural network, the method further includes: determining a network quality of the first neural network; When the network quality meets the second condition, performing a second process; Among them, the second condition includes that the goodness of fit of the first neural network is less than a third threshold, and the goodness of fit is used to indicate the degree of match between the predicted spectrum corresponding to the first neural network and the true spectrum, and the second processing includes adjusting the number of the N sample data, performing the preprocessing again, and adjusting at least one of the structure of the first original neural network.
7. The method according to any one of claims 1 to 4, characterized in that The step of training the first original neural network according to the loss value by using a target optimization algorithm to obtain a first neural network includes: Input each of the loss values into the target optimization algorithm to obtain a model parameter update amount; The model parameters of the first original neural network are updated according to the model parameter update amount to obtain the first neural network.
8. A neural network training device, characterized in that: include: An obtaining module, used for inputting N sample data into a first original neural network to obtain N predicted spectra, wherein each sample data in the N sample data includes the size of a first device, the first device includes a wafer, and each sample data is associated with a real spectrum corresponding to the size of the first device, and N is a positive integer greater than or equal to 2; The obtaining module is further used to obtain a loss value according to the N predicted spectra and the N real spectra; A training module is used to train the first original neural network according to the loss value through a target optimization algorithm to obtain a first neural network, the target optimization algorithm is obtained through a first process, the target optimization algorithm includes a target Hessian matrix, the first process includes decomposing the original Hessian matrix to obtain the target Hessian matrix, or converting the original optimization algorithm to obtain the target optimization algorithm, the computational complexity of the target optimization algorithm after the first process is lower than the computational complexity before the first process, and the target optimization algorithm is used to adjust the network parameters of the first original neural network.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data processing method, system and equipment based on distributed cluster and storage medium
CN116070720A
Structural parameter determination method
CN118999351A
Feature size parameter extraction method and system, terminal equipment and storage medium
CN119227481A
Automated Accuracy-Oriented Model Optimization System for Critical Dimension Metrology
US20180232630A1