Chip visual design defect detection method and system based on deep learning
By optimizing the neural network structure and precision configuration, the problem of inefficient model deployment in chip defect detection is solved, and efficient detection on a heterogeneous computing platform is achieved to meet the real-time detection requirements of chip production lines.
Patent Information
- Application Number
- CN202510456594.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-12
AI Technical Summary
The existing deep learning models have high computational complexity and large parameters in chip defect detection, making it difficult to deploy efficiently in real-time detection systems in production line, and have poor hardware adaptability, making it impossible to achieve consistent and efficient performance on different hardware platforms.
By optimizing the neural network structure and accuracy configuration, a dedicated search space and hardware performance evaluation model is built, differentiable structure search, mixed precision quantization and knowledge distillation technologies are implemented, and the deployment of the model on a heterogeneous computing platform is optimized.
It realizes efficient deployment of chip defect detection, significantly improves detection speed, reduces model size, reduces hardware resource requirements, adapts to a variety of hardware platforms, and improves the reliability and economicality of detection.
Smart Images

Figure CN120298388A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and particularly to a method and system for detecting chip visualization design defects based on deep learning. Background Art
[0002] With the continuous development of integrated circuit manufacturing technology, the integration and complexity of chips are constantly increasing, and the requirements for the accuracy and efficiency of chip appearance defect detection are also getting higher and higher. The traditional manual detection method can no longer meet the needs of modern chip production lines, and the automatic defect detection technology based on machine vision has gradually become the mainstream solution. As a cutting-edge technology in the field of machine vision, deep learning performs excellently in image recognition and target detection, and is therefore widely used in chip defect detection.
[0003] Currently, the chip defect detection methods based on deep learning mainly use models such as convolutional neural network (CNN) and region convolutional neural network (R-CNN) to detect and classify defects. However, the existing technologies have the following problems:
[0004] The existing deep learning models have high computational complexity and large number of parameters, making it difficult to be efficiently deployed in the real-time detection system of the production line, resulting in a long detection process and affecting production efficiency;
[0005] The traditional deep learning models occupy a large amount of storage space and computing resources, have high requirements for the hardware resources of the detection equipment, and increase the deployment cost;
[0006] The existing methods do not fully consider the diversity of hardware platforms and resource constraints in the production environment, resulting in limited model applicability and difficulty in achieving consistent performance on different hardware platforms;
[0007] The existing models run inefficiently on specific hardware, cannot exert the best performance, and cause waste of hardware resources.
[0008] In the existing technology, there is a chip defect detection method based on deep learning, which uses an improved U-Net network structure for defect segmentation. Although the detection accuracy is high, the computational complexity is large, and it is difficult to run in real time on resource-constrained hardware. There is a chip surface defect detection method based on convolutional neural network. Although this method simplifies the network structure, it does not consider model quantization and hardware adaptation problems, resulting in large differences in deployment effects on different hardware platforms.
[0009] Therefore, there is an urgent need for a chip defect detection method that can significantly improve the detection speed, reduce the model size, and adapt to multiple hardware platforms while ensuring high accuracy, so as to meet the needs of real-time detection in the chip production line. Summary of the Invention
[0010] The present invention provides a method and system for detecting chip visualization design defects based on deep learning. By optimizing the neural network structure and precision configuration, the method realizes the efficient deployment of the detection model on a heterogeneous computing platform, and solves the technical problems of high computational complexity, large model size, and poor hardware adaptability in the prior art.
[0011] To achieve the above object, the technical solution provided by the present invention is:
[0012] A method for detecting chip visualization design defects based on deep learning, comprising the following steps:
[0013] Construct a dedicated search space and a hardware performance evaluation model. The search space includes neural network structure components suitable for chip defect feature extraction, and the hardware performance evaluation model is used to predict the execution time of the neural network on the target hardware;
[0014] Implement differentiable architecture search to optimize the neural network structure for chip defect detection;
[0015] Implement mixed-precision quantization to adaptively allocate the optimal bit width for different layers of the neural network, and optimize the model storage and computing efficiency;
[0016] Perform multi-objective joint optimization to optimize both the neural network structure and the quantization bit width configuration;
[0017] Apply knowledge distillation technology to transfer the knowledge of a large teacher model to the optimized small student model to improve the chip defect detection accuracy.
[0018] Further, the step of constructing the dedicated search space and the hardware performance evaluation model includes:
[0019] According to the characteristics of chip defect images, construct a neural network search space suitable for defect feature extraction, including convolutional layers, pooling layers, attention modules, etc.;
[0020] Collect the performance data of the target hardware platform to form a hardware characteristic data set;
[0021] Train a hardware execution time prediction model. This model is implemented using a multi-layer perceptron structure, including an input layer, a feature extraction layer, a fusion layer, and an output layer. It receives neural network operation parameters and hardware specifications as inputs and is trained by minimizing the mean square error between the predicted time and the actual time.
[0022] Further, the step of implementing differentiable architecture search includes:
[0023] Based on the search space, establish a differentiable network architecture representation, and represent network architecture selection as continuous variables;
[0024] Adopt the alternating optimization algorithm to optimize the weight parameters while fixing the structural parameters on the training set, and optimize the structural parameters while fixing the weight parameters on the validation set. Use the second-order approximation method to accelerate the structural optimization;
[0025] During the training process, according to the feedback of the hardware performance evaluation model, add a latency-aware regularization term.
[0026] Further, the steps of implementing mixed-precision quantization include:
[0027] Establish a quantization representation method to convert full-precision parameters into low-bitwidth representations;
[0028] Sequentially apply different bitwidth quantizations to each layer in the neural network, calculate the precision loss, generate a sensitivity map, and record the precision changes of each layer under different quantization bitwidths;
[0029] Adopt an iterative greedy algorithm. On the premise of meeting the overall precision requirements, preferentially apply lower bitwidths to layers with low sensitivity;
[0030] Implement hardware-aware quantization. According to the characteristics of the target hardware platform, select the most suitable quantization bitwidth and computing mode;
[0031] Construct a quantization-aware training loop, including applying quantization configurations, quantizing weights and activation values in forward propagation, using a straight-through estimator to process gradients in backward propagation, and updating full-precision weights.
[0032] Further, the steps of performing multi-objective joint optimization include:
[0033] Establish a joint optimization objective function, considering detection precision, inference latency, and model size simultaneously;
[0034] Construct a joint search space, including two dimensions: network structure and quantization bitwidth;
[0035] Adopt a differential evolution algorithm to optimize the joint search space. The differential evolution algorithm includes population initialization, mutation operation, crossover operation, and selection operation. Use the fast non-dominated sorting algorithm to sort candidate solutions, and adopt a crowding distance mechanism to maintain the diversity of solutions;
[0036] Train the corresponding neural network for each candidate solution and evaluate its performance;
[0037] According to specific application requirements, select the most suitable solution from the Pareto front.
[0038] Further, the steps of applying knowledge distillation technology include:
[0039] Construct a multi-level knowledge distillation framework and define a teacher model and a student model;
[0040] Design a multi-task distillation loss function for chip defect detection, including hard label loss, soft label loss, feature matching loss, and boundary awareness loss;
[0041] Implement an attention-guided knowledge distillation method to generate the teacher model's attention map and construct an attention-guided loss;
[0042] Implement a progressive knowledge distillation training process, including an initialization stage, a feature matching stage, a classification optimization stage, and a quantization fine-tuning stage.
[0043] Furthermore, the multi-task distillation loss function further includes:
[0044] Hard label loss, which calculates the cross-entropy loss using the true labels;
[0045] Soft label loss, which uses the output of the teacher model as soft labels to guide the student model to learn the class probability distribution;
[0046] Feature matching loss, which aligns the feature representations of the teacher and student models in the intermediate layers;
[0047] Boundary awareness loss, which focuses on the feature learning of the chip defect boundary region.
[0048] Furthermore, the attention-guided knowledge distillation method includes:
[0049] Generate the attention map of the teacher model to reveal the key regions that the model focuses on;
[0050] Construct an attention-guided loss to guide the student model to focus on the same key regions as the teacher model;
[0051] Design a selective distillation strategy for different types of defects, including focusing on shallow features for edge-type defects, middle-layer features for texture-type defects, and deep-layer features for structure-type defects.
[0052] Furthermore, the progressive knowledge distillation training process includes:
[0053] Initialization stage, using a pre-trained teacher model to randomly initialize the student model;
[0054] Feature matching stage, freezing the classification head of the student model and only optimizing the feature extraction part;
[0055] Classification optimization stage, unfreezing the classification head of the student model and training using the complete distillation loss function;
[0056] Quantization fine-tuning stage, applying quantization-aware training to further reduce the precision loss caused by quantization.
[0057] The second invention of the present invention provides a chip visualization design defect detection system based on deep learning for implementing the above-mentioned chip visualization design defect detection method based on deep learning. The system includes:
[0058] A search space construction module for constructing a dedicated search space and a hardware performance evaluation model;
[0059] A structure search module for implementing differentiable structure search to optimize the neural network structure for chip defect detection;
[0060] An accuracy quantization module for implementing mixed-precision quantization to adaptively allocate the optimal bit width for different layers of the neural network;
[0061] A joint optimization module for performing multi-objective joint optimization to simultaneously optimize the network structure and the quantization bit width configuration;
[0062] A knowledge distillation module for applying knowledge distillation technology to transfer the knowledge of a large teacher model to an optimized small student model.
[0063] The beneficial effects of the present invention are as follows: By jointly optimizing the network structure and the accuracy configuration, the detection speed of the present invention is significantly improved, meeting the real-time detection requirements of the chip production line; the model size is greatly reduced, saving a large amount of storage space and transmission bandwidth, reducing the demand for hardware resources, and ensuring the reliability and stability of chip defect detection; the model can be efficiently deployed on various computing platforms (including edge devices and heterogeneous acceleration platforms) to adapt to the real-time detection requirements of different production scenarios; and the hardware cost is reduced, and high-efficiency detection can be achieved using devices with lower configurations, improving the universality and economy of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a flowchart of the steps of the chip visualization design defect detection method based on deep learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0066] A chip visualization design defect detection method based on deep learning provided in this embodiment, as Figure 1 shown, includes the following steps:
[0067] Step 1: Construct a dedicated search space and a hardware performance evaluation model. The search space includes neural network structure components suitable for chip defect feature extraction, and the hardware performance evaluation model is used to predict the execution time of the neural network on the target hardware;
[0068] Specifically include:
[0069] According to the characteristics of the chip defect image, construct a neural network search space S, which contains network components such as convolutional layers, pooling layers, and attention modules suitable for defect feature extraction. The parameter range of each component is set through pre-analysis, expressed as:
[0070]
[0071] Among them, O i represents the i-th type of network operation, is the set of all optional operations, P i represents the parameter set of the network operation type O i , represents the value range of the parameter. For example, for the convolution operation, the parameters include the convolution kernel size, the number of channels, the stride, etc.
[0072] Collect the performance data of the target hardware platform, including the execution efficiency of various neural network operations by different computing units such as CPU, GPU, and dedicated chips, to form a hardware characteristic data set D hw :
[0073] D hw ={(op j , hw k , t j,k )|j∈[1,J],k∈[1,K]};
[0074] Among them, op j represents the j-th neural network operation, hw k represents the k-th hardware device, t j,k represents the execution time of op j on the hardware device hw k . J represents the total number of network operations, K represents the total number of hardware devices, ∈[1,J] represents the operation index range from 1 to J, and k∈[1,K] represents the hardware device index range from 1 to K. This data set contains the execution time information of all operations on all target hardware platforms, providing basic data support for subsequent optimization.
[0075] Train a hardware execution time prediction model T. The input of this model is network operation parameters and hardware specifications, and the output is the predicted execution time:
[0076] T(op,hw) = f φ (opparams , hw specs );
[0077] Among them, f φ represents the prediction function of the set of learnable parameters φ of the model, φ represents the set of learnable parameters of the model; op represents a neural network operation; hw represents a hardware platform; T(op, hw) represents the predicted execution time of the neural network operation op on the hardware platform hw; op params represents the set of parameters of the neural network operation, including operation type, input and output sizes, number of channels, convolution kernel size, etc.; hw specs represents the specification parameters of the hardware platform, including processor type, memory size, number of computing units, etc. The hardware execution time prediction model is implemented using a specific multi-layer perceptron structure, which includes:
[0078] Input layer: Receives the operation parameter vector (including operation type, input and output sizes, number of channels, convolution kernel size, etc.) and the hardware specification vector (including processor type, memory size, number of computing units, etc.);
[0079] Feature extraction layer: Consists of 3 fully connected layers, each layer has multiple neurons, uses the ReLU activation function, and extracts the association patterns of operation and hardware characteristics;
[0080] Fusion layer: After concatenating the operation features and hardware features, performs feature fusion through 2 fully connected layers;
[0081] Output layer: A single neuron that outputs the predicted execution time.
[0082] This prediction model is trained by minimizing the mean square error between the predicted time and the actual time:
[0083]
[0084] Among them, min φ represents optimizing the set of learnable parameters φ of the model to make the objective function reach the minimum value, ∑ represents summing over all samples; (op, hw, t) ∈ D hw represents a sample taken from the hardware characteristics dataset D hw which contains the neural network operation op, the hardware platform hw, and the actual execution time t; T(op, hw) represents the execution time of the operation op predicted by the model on the hardware platform hw; (T(op, hw) - t) 2 represents the square of the difference between the predicted time and the actual time, that is, the mean square error.
[0085] In the actual application of the chip production line, the hardware execution time prediction model can be applied to evaluate the performance of different defect detection network structures on the edge computing devices of the production line. For example, when a new type of image acquisition and processing device is introduced into the production line, the model can predict the execution efficiency of the existing defect detection network without time-consuming actual deployment tests, so as to quickly decide whether it is necessary to adjust and optimize the model.
[0086] Through this step, a neural network search space suitable for chip defect detection tasks and a reliable hardware performance evaluation model are formed, providing a basis for subsequent network optimization.
[0087] Step 2: Implement differentiable architecture search to optimize the network architecture and optimize the network architecture for chip defect detection;
[0088] Specifically include:
[0089] Based on the search space constructed in Step 1, establish a differentiable network architecture representation α, representing network architecture selection as continuous variables:
[0090]
[0091] where E represents the set of edges in the network, (i, j) represents the connection from node i to node j, represents the weight of a single optional operation O on this connection.
[0092] Construct a network representation method so that the selection of each neural network operation is determined by a set of weights α (i,j) and convert the discrete selection into a continuous variable through the softmax function:
[0093]
[0094] where, represents the output of the mixed operation, x represents the input feature, o(x) represents the result of processing the input X by operation o, exp represents the exponential function, O′ represents another network operation, represents the probability weight of operation O calculated through the softmax function.
[0095] Adopt an alternating optimization algorithm to iteratively optimize the weight parameter w and the architecture parameter α:
[0096] Fix the architecture parameter α and optimize the weight parameter W on the training set D train :
[0097]
[0098] where W * represents the optimized weight parameter, argmin wDenote finding the weight parameters that minimize the objective function, denote the loss function on the training set;
[0099] Fix the optimized weight parameter W * , and optimize the structure parameter α on the validation set D val :
[0100]
[0101] where α * denote the optimized structure parameter, and argmin α denote finding the structure parameter that minimizes the objective function, denote the loss function on the validation set, including the classification loss and the regularization term:
[0102]
[0103] where denote the total loss function, is the cross-entropy loss, R(α) is the sparsity regularization term, and λ is the trade-off coefficient.
[0104] Specific implementation method for realizing differentiable architecture search:
[0105] Use the second-order approximation method to accelerate the architecture optimization process:
[0106]
[0107] where denote the gradient of the structure parameter α, denote the gradient of the validation set loss function with respect to the structure parameter α, denote the gradient of the weight parameter W, denote the gradient of the training set loss function with respect to the weight parameter w, denote the result after updating the weight parameter w one step along the training gradient direction, and ξ is the approximate learning rate.
[0108] Adopt a specific network unit design, including:
[0109] Edge detection unit: composed of 3×3 and 5×5 depthwise separable convolutions, suitable for detecting chip edge defects;
[0110] Texture analysis unit: composed of 3×3 convolution and self-attention module, suitable for fine-grained texture defect detection;
[0111] Multi-scale feature fusion unit: combine feature maps of different scales and adjust the feature channels through 1×1 convolution.
[0112] During the training process, according to the feedback of the hardware execution time prediction model, a latency-aware regularization term is added to reduce the inference latency while ensuring the detection accuracy of the network structure:
[0113] R delay (α) = T pred (N α , hw target )
[0114] where R delay (α) represents the latency-aware regularization term, T pred is the execution time prediction model in Step 1, N α is the network under the current structure parameters, and hw target is the target hardware platform. The final structure optimization goal is:
[0115]
[0116] where β is the weight coefficient of latency regularization.
[0117] In a practical application example, when this differentiable architecture search algorithm is applied to chip surface scratch detection, by designing a specific search space, it can discover a network architecture that is particularly suitable for identifying linear scratch features. For example, for the micro-scratches on the surface of a certain type of semiconductor chip, this algorithm automatically selects a structure that includes cascaded depthwise separable convolutions and a direction-aware attention module. Compared with the manually designed network, it not only improves the detection accuracy to a certain extent but also enhances the inference speed. This optimized structure can effectively capture the direction features and edge characteristics of micro-scratches, significantly improving the detection performance.
[0118] Through this step, for the characteristics of chip defect detection, an optimized network structure is automatically searched and determined, improving the ability and efficiency of feature extraction.
[0119] Step 3: Implement mixed-precision quantization to adaptively allocate the optimal bitwidth for different layers of the neural network and optimize the model storage and computational efficiency;
[0120] In this step, through the mixed-precision quantization technology, the optimal bitwidth is adaptively allocated for different layers of the network, reducing the model size and computational complexity while maintaining the detection accuracy. Specifically, it includes:
[0121] Establish a quantization representation method to convert full-precision parameters into low-bitwidth representations:
[0122] w q = Int(w f · s);
[0123] where w f is the full-precision weight, wq W is the quantized integer weight, s is the quantization scale factor, and Int is the rounding operation.
[0124] Construct a set B of quantization configurations with different precisions, including multiple different bit-width options:
[0125] B = {b1, b2,..., b n};
[0126] where b1, b2,..., b n represent 1, 2,..., n quantization bit-widths respectively, and n represents the total number of quantization bit-widths.
[0127] Establish a quantization sensitivity analysis method to evaluate the sensitivity of each layer of the network to quantization:
[0128] For each layer l in the network, sequentially apply different bit-width quantizations b ∈ B and calculate the accuracy loss on the validation set:
[0129] ΔAcc(l, b) = Acc(N) - Acc(N 1,b );
[0130] where ΔAcc(l, b) represents the accuracy loss of layer l at bit-width b, Acc(N) represents the accuracy of the original network, and Acc(N l,b ) represents the accuracy of the network after quantizing layer l to bit-width b.
[0131] Generate a sensitivity map S to record the accuracy changes of each layer at different quantization bit-widths:
[0132] S = {s l,b = ΔAcc(l, b)|l ∈ [1, L], b ∈ B};
[0133] where L is the number of network layers, and s l,b represents the accuracy loss when layer l uses bit-width b.
[0134] Implement a specific method for mixed-precision quantization based on sensitivity analysis:
[0135] Adopt an iterative greedy algorithm. On the premise of meeting the overall accuracy requirements, preferentially apply lower bit-widths to layers with lower sensitivity:
[0136] Initial state: All layers use the highest bit-width;
[0137] Iterative process: Each time, select the layer-bit-width combination (l, b) with the smallest accuracy loss and quantize layer l to a lower bit-width b;
[0138] Termination condition: Reach the target model size or the accuracy loss exceeds the preset threshold;
[0139] Build a quantization-aware training loop and fine-tune the network after the quantization configuration is determined:
[0140] Initialization: Apply the selected mixed-precision quantization configuration;
[0141] Forward propagation: Quantize weights → Perform calculations → Quantize activation values;
[0142] Backward propagation: Process gradients using the straight-through estimator (STE);
[0143] Update: Update the full-precision weights;
[0144] Specific optimizations for the chip defect detection scenario:
[0145] Use a higher bitwidth for the first and last channels of the convolutional layer to retain the precision of the input and output features;
[0146] Use a higher bitwidth for the feature extraction layers that detect key regions (such as edges and corners);
[0147] Use a lower bitwidth for the texture feature representation layer because texture features are insensitive to quantization;
[0148] Implement hardware-aware quantization and select the most suitable quantization bitwidth and calculation mode according to the characteristics of the target hardware platform:
[0149] Utilize the hardware execution time prediction model in Step 1 to evaluate the actual performance of different quantization configurations on the target hardware:
[0150] T quant (N,hw,Q)=f θ (N arch ,hw specs ,Q config );
[0151] where T quant represents the execution time of the quantized model, N represents the neural network model, Q represents the quantization configuration scheme, and f θ represents the hardware execution time prediction function, θ represents the prediction model parameters, N arch is the network architecture, and Q config is the quantization configuration.
[0152] Combine the hardware characteristics and accuracy requirements to select the optimal quantization configuration:
[0153] Q * =argmin Q T quant (N,hw,Q),s.t.ΔAcc(NQ)≤∈;
[0154] where Q *Denotes the optimal quantization configuration scheme, argmin Q Denotes the quantization configuration scheme that minimizes the objective function, T quant (N, hw, Q) represents the model execution time under the neural network model N, hardware platform hw, and quantization configuration scheme Q, ΔAcc(N Q ) is the precision loss after quantization, ΔAcc represents the precision loss, ∈ is the acceptable precision loss threshold, N Q Denotes the network model after applying the quantization configuration scheme Q, s.t. represents subject to the condition.
[0155] In a practical application example, when this mixed-precision quantization method is applied to chip package defect recognition on a certain embedded vision detection device, the original multi-bit floating-point model is successfully compressed to several bits / parameters, and the model size is reduced to a certain extent, while the precision decreases to a certain extent. In this application, it is found that the input layer and the feature extraction layer related to defect edge detection are more sensitive to quantization. Therefore, a higher bit width is allocated to these layers, while a lower bit width is used for the intermediate feature representation layer. This optimized configuration enables the detection model to run in real time on resource-constrained industrial edge computing devices and process high-resolution chip images per second.
[0156] Through this step, a significant reduction in model size and computational complexity is achieved, while maintaining high precision in chip defect detection, providing a basis for subsequent joint optimization.
[0157] Step 4: Perform multi-objective joint optimization to optimize both the network structure and bit width configuration simultaneously;
[0158] This step realizes the joint optimization of the network structure and quantization bit width, and simultaneously finds the optimal solution in multiple dimensions such as precision, latency, and model size through a global multi-objective optimization algorithm. Specifically, it includes:
[0159] Establish a joint optimization objective function, considering detection precision, inference latency, and model size simultaneously:
[0160]
[0161] Among them, Q represents the quantization configuration scheme, represents the detection loss function, T represents the inference latency on the hardware platform hw, and M represents the model size.
[0162] Construct a joint search space S joint , including two dimensions of network structure and quantization bit width:
[0163] S joint = S arch × S quant ;
[0164] Among them, S arch represents the network structure search space, and S quant represents the quantization configuration search space.
[0165] The differential evolution algorithm is used to optimize the joint search space of bit width and structure:
[0166] Population initialization: Generate multiple candidate solutions, and each candidate solution contains network structure parameters α and quantization bit width configuration Q:
[0167] P = {(α i , Q i ) | i = 1, 2,..., NP};
[0168] Among them, NP represents the population size, that is, the number of candidate solutions; α i represents the network structure parameters of the i-th candidate solution, which determines the topological structure and connection method of the network; Q i represents the quantization bit width configuration of the i-th candidate solution, which determines the precision of each layer of the network; P represents the entire population, which is a set containing all candidate solutions;
[0169] Mutation operation: For each individual (α i , Q i ), select three different individuals r1, r2, r3 to generate a mutation vector:
[0170]
[0171] Among them, v i represents the generated mutation vector, and respectively represent three different individuals randomly selected from the population, respectively represent the network structure parameters of the r1-th, r2-th, and r3-th candidate solutions, respectively represent the quantization bit width configurations of the r1-th, r2-th, and r3-th candidate solutions, and F is a scaling factor that controls the amplification degree of the differential vector, usually with a value range of [0.4, 1.0].
[0172] Crossover operation: Perform crossover on the current individual (α i , Q i ) and the mutation vector v i to generate a trial vector u i :
[0173]
[0174] Among them, represents the j-th component of the trial vector u i , represents the mutation vector v iThe j-th component of, (α i , Q j ) j represents the j-th component of the current individual, j represents the index of the decision variable, rand j represents a random number in the interval [0, 1], CR is the crossover probability (controlling the proportion of information inherited from the mutant vector), j rand is a randomly selected index (ensuring that at least one variable comes from the mutant vector), otherwise means otherwise.
[0175] Selection operation: Evaluate the performance of the trial vector u i . If it is better than the current individual, replace it:
[0176]
[0177] where, (α i , Q i ) t+1 represents the individual in the (t + 1)-th generation, u i represents the trial vector, (α i , Q i ) t represents the current individual in the t-th generation, F(u i ) represents the objective function value vector of the trial vector, F((α i , Q i ) t ) represents the objective function value vector of the current individual, ≤ represents the Pareto dominance relationship in multi-objective optimization (indicating that one solution is not inferior to another solution in all objectives and is superior to another solution in at least one objective).
[0178] The specific optimization process of implementing the differential evolution algorithm:
[0179] Evolution parameter configuration:
[0180] Population size: NP = 50;
[0181] Maximum number of generations: MaxGen = 100;
[0182] Scaling factor: F = 0.5;
[0183] Crossover probability: CR = 0.7;
[0184] Multi-objective evaluation and Pareto front selection:
[0185] Use the fast non-dominated sorting algorithm (NSGA-II) to sort the candidate solutions;
[0186] Adopt the crowding distance mechanism to maintain the diversity of solutions;
[0187] Retain the set of non-dominated solutions on the Pareto front;
[0188] Optimization strategies for the chip defect detection scenario:
[0189] Adopt different mutation and crossover strategies for the network structure parameter α and the quantization bit-width configuration Q;
[0190] The structure parameters adopt continuous real-value coding, and the bit-width configuration adopts discrete integer coding;
[0191] After the mutation and crossover operations, normalize the structure parameters and perform rounding operations on the bit-width configuration;
[0192] Implement candidate solution evaluation and final model selection:
[0193] For each candidate solution (α, Q), train the corresponding network and evaluate its performance:
[0194] Detection accuracy: Evaluate the detection accuracy, recall, and precision on the test set;
[0195] Latency performance: Use the prediction model in Step 1 to estimate the inference latency;
[0196] Model size: Calculate the number of parameters and the storage size;
[0197] According to specific application requirements, select the most suitable solution from the Pareto front:
[0198] Performance-first scenario: Select the model with the highest accuracy under the latency and size constraints;
[0199] Real-time-first scenario: Select the model with the minimum latency under the accuracy requirements;
[0200] Resource-constrained scenario: Select the model with the minimum size under the accuracy and latency constraints;
[0201] In an actual application example, when this multi-objective joint optimization algorithm is applied to the defect detection system of an automotive chip production line, by simultaneously optimizing the network structure and quantization configuration, three model versions with different trade-offs are obtained:
[0202] High-precision version: The detection accuracy is improved to some extent, suitable for the detection of high-end chips with strict quality control;
[0203] High-performance version: The detection accuracy decreases to some extent, suitable for real-time detection on the production line;
[0204] Ultra-lightweight version: The detection accuracy remains stable, suitable for embedded detection devices.
[0205] Compared with optimizing the network structure or quantization bitwidth alone, the joint optimization method can improve the detection performance to a certain extent and reduce the inference latency to a certain extent. Especially on resource-constrained edge devices, this method can enable the detection system to meet the real-time requirements while maintaining high detection accuracy by discovering the synergistic effect between the structure and quantization.
[0206] Through this step, the joint optimization of the network structure and bitwidth configuration is achieved, obtaining a series of Pareto optimal solutions with different trade-offs in terms of accuracy, latency, and model size, providing a basis for selecting the most suitable model for a specific scenario.
[0207] Step 5: Apply the knowledge distillation technique to transfer the knowledge of the large teacher model to the optimized small student model to improve the chip defect detection accuracy;
[0208] In this step, a knowledge distillation method for chip defect detection is constructed to further improve the model detection accuracy.
[0209] Specifically include:
[0210] Construct a multi-level knowledge distillation framework and define the teacher model and the student model:
[0211] Teacher model T: Use a large high-precision network, such as ResNet-101 or EfficientNet-B7 using multi-bit floating-point numbers, with strong feature extraction ability but high computational complexity;
[0212] Student model S: Use the lightweight model optimized in Step 4, with a compact structure and quantized to a low bitwidth, but there may be a loss of accuracy;
[0213] Establish a multi-level knowledge transfer channel:
[0214]
[0215] Among them, and respectively represent the feature maps of the teacher model and the student model at the l-th layer, and L KD represents the set of layers that need to perform knowledge distillation.
[0216] Design a multi-task distillation loss function for chip defect detection:
[0217]
[0218] Among them: Total loss function; λ1, λ2, λ3, λ4 represent the weight coefficients of each loss term, used to balance the contributions of different loss terms; is the hard label loss; represents the soft label loss; is the feature matching loss; represents the boundary awareness loss;
[0219] Hard label loss Calculate the cross-entropy loss using the true label to ensure basic classification ability:
[0220]
[0221] where S(x) represents the prediction of the student model for the input x, y represents the true label, and CE is the cross-entropy function.
[0222] Soft label loss Use the output of the teacher model as the soft label to guide the student model to learn the class probability distribution:
[0223]
[0224] where KL represents the KL divergence, τ is the temperature parameter that controls the smoothness of the soft label, and T(x) represents the prediction output of the teacher model for the input x.
[0225] Feature matching loss Align the feature representations of the teacher and student models in the intermediate layer:
[0226]
[0227] where φ1 is the feature transformation function used to adjust the feature dimension of the teacher model to match the student model, is the feature map of the teacher model at the l-th layer, is the feature map of the student model at the l-th layer, L KD is the set of layers that need to perform knowledge distillation, represents the square of the L2 norm.
[0228] Boundary awareness loss Focus on the feature learning of the chip defect boundary region:
[0229]
[0230] where M b is the boundary region mask, and ⊙ represents element-wise multiplication, focusing on the feature differences at the defect boundary.
[0231] Implement the knowledge distillation method based on attention guidance:
[0232] Generate the attention map of the teacher model Reveal the key regions that the model focuses on:
[0233]
[0234] Among them, C is the number of feature channels, represents the feature map of the c-th channel in the l-th layer, and is the attention map of the teacher model in the first layer.
[0235] Construct an attention-guided loss to guide the student model to focus on the same key regions:
[0236]
[0237] Among them, is the attention-guided loss, is the attention map of the student model in the first layer;
[0238] Design a selective distillation strategy for defect type perception:
[0239] For edge class defects, focus on the distillation of shallow features to enhance edge detection ability;
[0240] For texture class defects, focus on the distillation of middle-layer features to improve texture recognition ability;
[0241] For structure class defects, focus on the distillation of deep features to enhance semantic understanding ability;
[0242] Implement a progressive knowledge distillation training process:
[0243] Initialization stage: Use a pre-trained teacher model to randomly initialize the student model;
[0244] Feature matching stage: Freeze the classification head of the student model, only optimize the feature extraction part, and train using the feature matching loss:
[0245]
[0246] Classification optimization stage: Unfreeze the classification head of the student model and train using the complete distillation loss function:
[0247]
[0248] Quantization fine-tuning stage: After distillation, apply quantization-aware training to further reduce the accuracy loss caused by quantization:
[0249]
[0250] Among them, S Q represents the quantized student model, Q(S) represents the quantization operation on the student model, and γ is the trade-off coefficient.
[0251] Establish a method to evaluate the model performance and quantitatively compare the model performance before and after optimization:
[0252] Detection accuracy metrics: calculation accuracy, recall rate, F1-score, and mean average precision (mAP);
[0253] Efficiency metrics: measure the inference time, model size, and memory footprint on actual hardware;
[0254] Comprehensive performance metrics: calculate the precision-latency product (PLP) and precision-size product (PSP):
[0255] PLP = Accuracy × (1 / Latency);
[0256] PSP = Accuracy × (1 / Size);
[0257] Among them, Accuracy represents the detection accuracy of the model, Latency is the inference latency time of the model, Size represents the storage size of the model, (1 / Latency) represents the reciprocal of latency, which represents the speed of the model, and (1 / Size) is the reciprocal of size, which represents the storage efficiency of the model.
[0258] In an actual application example, when this knowledge distillation method is applied to the etching defect detection scenario, by transferring knowledge from the ResNet-101 teacher model with stable parameters to the lightweight MobileNet-V3 student model with only stable parameters, the detection performance is significantly improved. On a dataset containing multiple types of chip etching defects, compared with an equivalent-sized model without using knowledge distillation, the distilled student model has a certain improvement in the average detection accuracy, especially for difficult-to-identify subtle bridging defects and incomplete etching defects, with the accuracy being improved to a certain extent respectively.
[0259] This knowledge distillation method is particularly suitable for the case of unbalanced defect samples. By extracting knowledge about rare defect categories from the teacher model, the detection ability of the student model for these defects is significantly improved. In the actual deployment of a certain chip production line, the lightweight model after applying this method can process high-resolution chip images in real time on edge computing devices, the detection speed reaches stability, and the accuracy is comparable to that of a large model using a central server.
[0260] Through this step, the knowledge of the large high-precision model is effectively transferred to the optimized lightweight model, while maintaining high detection accuracy, meeting the deployment requirements on resource-constrained hardware.
[0261] 3. Technical effects of this embodiment:
[0262] The chip visualization design defect detection method based on deep learning provided by this embodiment realizes the following technical effects by optimizing the neural network structure and precision configuration:
[0263] The detection speed has been significantly improved, showing a certain increase compared to traditional deep learning models, meeting the requirements of real-time detection in the chip production line;
[0264] The model size has been greatly reduced, with the compression rate reaching stability, greatly saving storage space and transmission bandwidth, and reducing the demand for hardware resources;
[0265] The detection accuracy remains at a stable high level, ensuring the reliability and stability of chip defect detection;
[0266] The model can be efficiently deployed on various computing platforms (including edge devices and heterogeneous acceleration platforms), adapting to the real-time detection needs of different production scenarios;
[0267] The hardware cost has been reduced. High-efficiency detection can be achieved using devices with lower configurations, improving the universality and economy of the method.
[0268] Through the joint optimization of the network structure and bit-width configuration, and combined with the knowledge distillation technology, this embodiment successfully solves the problem of low model deployment efficiency in chip visualization design defect detection, providing an efficient and feasible technical solution for real-time defect detection in the chip manufacturing process.
[0269] This section details the application process and effects of the present invention in an actual chip production line through specific application examples.
[0270] The chip production line of an integrated circuit manufacturing enterprise needs to perform appearance defect detection on the produced large-scale integrated circuit chips. The main objectives are to identify the following types of defects:
[0271] Surface scratches: including fine linear scratches, deep scratches, etc.;
[0272] Packaging defects: including pin deformation, solder joint missing, packaging cracks, etc.;
[0273] Pattern defects: including incomplete lithography, pattern missing, bridging, etc.;
[0274] Material anomalies: including oxidation, corrosion, pollution, etc.
[0275] The detection system faces the following challenges: The production line speed is fast, and multiple chip images need to be processed per second; Resources are limited, and the detection device is configured with an Intel Core i5 processor and 8GB of memory, without a dedicated GPU; There is a need for multi-scenario deployment, including real-time detection on the production line and spot checks on portable devices; The types of defects are diverse, and new types of defects continue to emerge, requiring the system to have good generalization ability.
[0276] Based on the method of the present invention, we deployed an optimized defect detection system on this chip production line. The specific application process is as follows:
[0277] Data collection: Collect multiple high-resolution chip images from the production line, including various defect samples and normal samples;
[0278] Data annotation: Professional quality inspection personnel annotate the defect areas in the images, including information such as defect type, location, and size;
[0279] Data augmentation: Expand the dataset through methods such as rotation, flipping, and brightness adjustment, especially focusing on enhancing rare defect types;
[0280] Data division: Divide the dataset into a training set, a validation set, and a test set according to a certain ratio to ensure the balanced distribution of various defects in each subset.
[0281] Select the initial network structure: Based on EfficientDet-D0 as the base network, which has a good balance of accuracy and efficiency in object detection tasks;
[0282] Model training: Use the standard training process to train the initial model, including conventional optimization techniques such as learning rate scheduling and weight decay;
[0283] Performance evaluation: Evaluate the performance of the base model on the test set, and the results are as follows:
[0284] Average detection accuracy (mAP): Remains stable;
[0285] Processing time per image: 420ms;
[0286] Model size: 32.6MB;
[0287] Memory occupancy: 1.2GB.
[0288] Step 1: Construct a dedicated search space and a hardware performance evaluation model. The search space includes neural network structure components suitable for chip defect feature extraction, and the hardware performance evaluation model is used to predict the execution time of the neural network on the target hardware;
[0289] Based on the characteristics of chip defects, a search space containing various convolutional operations, various pooling operations, and various attention modules is constructed;
[0290] Test the execution time of each operation on the target Intel processor to construct a hardware performance prediction model;
[0291] The test results show that the prediction error of the prediction model for the operation execution time is within a certain range.
[0292] Step 2: Implement differentiable architecture search to optimize the neural network structure for chip defect detection;
[0293] Apply the search space in Step 1 and use multiple images for architecture search;
[0294] The architecture search process took a certain amount of time and finally a network architecture more suitable for chip defect detection was found;
[0295] The optimized architecture has a certain reduction in the number of parameters while maintaining the detection accuracy.
[0296] Step 3: Implement mixed-precision quantization, adaptively assign the optimal bitwidth to different layers of the neural network, and optimize the model storage and computational efficiency;
[0297] Perform sensitivity analysis on each layer of the network and find that the input layer and the last three layers are the most sensitive to quantization;
[0298] Retain stable accuracy for the input layer and output layer, and assign a certain accuracy to the intermediate layers according to the sensitivity;
[0299] The size of the quantized model is reduced to a certain extent.
[0300] Step 4: Perform multi-objective joint optimization to optimize both the neural network architecture and the quantization bitwidth configuration simultaneously;
[0301] Set the population size to be stable and finally obtain multiple candidate models on the Pareto front;
[0302] Select several representative models from the candidate models, which are respectively suitable for high-precision scenarios, real-time detection scenarios, and portable device scenarios.
[0303] Step 5: Apply the knowledge distillation technique to transfer the knowledge of the large teacher model to the optimized small student model to improve the chip defect detection accuracy;
[0304] Use ResNet101 as the teacher model and transfer the knowledge to the optimized lightweight student model;
[0305] Adopt a differentiated distillation strategy for different types of defects, especially strengthening the recognition ability of difficult-to-detect defects;
[0306] After distillation, the detection accuracy of the student model for some defect types even exceeds that of the teacher model.
[0307] The performance of the three finally obtained optimized models is shown in Table 1:
[0308] Table 1. Example of performance data of the three finally obtained optimized models:
[0309]
[0310] Deployment scenarios and effects:
[0311] Deployment on the main production line:
[0312] Deploy the standard version model and connect to 12 detection sites;
[0313] Achieve a detection speed of processing 25 chip images per second;
[0314] The missed detection rate has been reduced to a certain extent from the original, and the false detection rate has been decreased to a certain extent;
[0315] The system has been running stably for 6 months without performance degradation.
[0316] Deployment in the quality inspection laboratory:
[0317] Deploy the high-precision version model for accurate identification of difficult-to-detect defects;
[0318] Conduct secondary detection on suspicious samples on the production line with an accuracy rate of 99.1%;
[0319] Assist quality inspection personnel in establishing a more accurate defect classification standard.
[0320] Deployment of portable detection equipment:
[0321] Deploy the lightweight version model to the engineer's handheld device;
[0322] Achieve rapid on-site spot checks, with each single detection only taking 0.5 seconds;
[0323] Provide timely feedback for production line adjustment and reduce the output of defective products.
[0324] Economic benefits:
[0325] The detection speed has been improved to a certain extent, and the detection efficiency of the production line has been increased to a certain extent;
[0326] The detection accuracy rate has been improved to a certain extent, and the missed detection rate of defective products has been increased to a certain extent;
[0327] Reduce the workload of manual re-inspection and lower the labor cost to a certain extent;
[0328] The hardware cost of the detection equipment has been reduced by 60%, and there is no need to configure high-end GPUs.
[0329] Technological advancement:
[0330] For the first time, jointly optimize the network structure and quantization bit width in the field of chip defect detection;
[0331] Innovatively incorporate defect features and hardware characteristics into the optimization consideration;
[0332] The developed distillation method effectively solves the problem of sample imbalance.
[0333] Promotion value:
[0334] This method has been successfully extended to the company's eight chip production lines of different models;
[0335] Strong adaptability, only a small number of new samples are needed to migrate to new chip models for testing;
[0336] The method is universal and can be extended to other precision electronic component detection fields.
[0337] To verify the superiority of the method of the present invention, we conducted a comparative experiment with the four current mainstream deep learning optimization methods and tested the performance of the optimized models of each method on the same chip defect dataset. The comparison results are shown in Table 2:
[0338] Table 2: Examples of model performance data after testing each method on the same chip defect dataset:
[0339] Optimization method Detection accuracy (mAP) Processing speed (FPS) Model size (MB) Adapt to different hardware capabilities The method of the present invention 97.6% 31.2 3.2 High Network pruning method 95.8% 24.5 8.6 Medium Knowledge distillation method 96.7% 18.3 12.1 Low Quantization compression method 94.2% 27.1 4.8 Medium NAS method 97.1% 15.6 9.2 Low
[0340] From the comparison results, it can be seen that the method of the present invention achieves the fastest processing speed and the smallest model size while maintaining high detection accuracy, and has the best hardware adaptability, and its comprehensive performance is significantly better than other methods.
[0341] This application example fully verifies the effectiveness and advancement of the method of the present invention in an actual chip production line. Through the integrated application of technologies such as dedicated search space construction, differentiable structure search, mixed precision quantization, multi-objective joint optimization, and knowledge distillation, the problem of balancing efficiency and accuracy in chip defect detection has been successfully solved, providing chip manufacturing companies with an efficient and reliable defect detection solution. This method is not only applicable to the specific chip production line in this example, but also has a wide range of promotion and application value.
Claims
1. A method for detecting visual design defects of a chip based on deep learning, characterized in that, The steps include: Construct a dedicated search space and a hardware performance evaluation model. The search space includes neural network structure components suitable for chip defect feature extraction, and the hardware performance evaluation model is used to predict the execution time of the neural network on the target hardware; Implement differentiable architecture search to optimize the neural network structure for chip defect detection; Implement mixed-precision quantization to adaptively allocate the optimal bit width for different layers of the neural network, optimizing model storage and computing efficiency; Perform multi-objective joint optimization to simultaneously optimize the neural network structure and quantization bit width configuration; Apply knowledge distillation technology to transfer the knowledge of the large teacher model to the optimized small student model to improve the chip defect detection accuracy.
2. The method for detecting design defects in chip visualization based on deep learning according to claim 1, wherein, The steps of constructing the dedicated search space and the hardware performance evaluation model include: According to the characteristics of chip defect images, construct a neural network search space containing convolutional layers, pooling layers, and attention modules suitable for defect feature extraction; Collect the performance data of the target hardware platform to form a hardware characteristics dataset; Train a hardware execution time prediction model. This hardware execution time prediction model is implemented using a multi-layer perceptron structure, including an input layer, a feature extraction layer, a fusion layer, and an output layer. It receives neural network operation parameters and hardware specifications as inputs and is trained by minimizing the mean square error between the predicted time and the actual time.
3. A method for detecting design defects in chip visualization based on deep learning according to claim 1, characterized in that, The steps of implementing differentiable architecture search include: Based on the search space, establish a differentiable network architecture representation, representing network architecture selection as continuous variables; Adopt an alternating optimization algorithm to optimize the weight parameters while fixing the architecture parameters on the training set, and optimize the architecture parameters while fixing the weight parameters on the validation set, using a second-order approximation method to accelerate architecture optimization; During the training process, according to the feedback of the hardware performance evaluation model, add a latency-aware regularization term.
4. A method for detecting design defects in chip visualization based on deep learning according to claim 1, characterized in that The steps of implementing mixed-precision quantization include: Establish a quantization representation method to convert full-precision parameters into low-bit-width representations; Sequentially apply different bit-width quantizations to each layer of the neural network, calculate the precision loss, generate a sensitivity map, and record the precision changes of each layer under different quantization bit widths; Adopt an iterative greedy algorithm. On the premise of meeting the overall precision requirements, preferentially apply lower bit widths to layers with lower sensitivity; Implement hardware-aware quantization and select the most suitable quantization bit width and computing mode according to the characteristics of the target hardware platform; Construct a quantization-aware training loop, including applying quantization configurations, quantizing weights and activation values during forward propagation, using a straight-through estimator to process gradients during backward propagation, and updating full-precision weights.
5. The method for detecting design defects in chip visualization based on deep learning according to claim 1, wherein, The steps of performing multi-objective joint optimization include: Establish a joint optimization objective function, considering detection accuracy, inference latency, and model size simultaneously; Construct a joint search space, including two dimensions of network architecture and quantization bit width; Adopt a differential evolution algorithm to optimize the joint search space. The differential evolution algorithm includes population initialization, mutation operation, crossover operation, and selection operation. Use a fast non-dominated sorting algorithm to sort candidate solutions, and adopt a crowding distance mechanism to maintain the diversity of solutions; Train the corresponding neural network for each candidate solution and evaluate its performance; According to specific application requirements, select the most suitable solution from the Pareto front.
6. The method for detecting design defects in chip visualization based on deep learning according to claim 1, wherein, The steps of applying the knowledge distillation technology include: Construct a multi-level knowledge distillation framework and define the teacher model and the student model; Design a multi-task distillation loss function for chip defect detection, including hard label loss, soft label loss, feature matching loss, and boundary awareness loss; Implement an attention-guided knowledge distillation method to generate the attention map of the teacher model and construct the attention-guided loss; Implement a progressive knowledge distillation training process, including an initialization phase, a feature matching phase, a classification optimization phase, and a quantization fine-tuning phase.
7. A method for detecting design defects in chip visualization based on deep learning according to claim 6, characterized in that The multi-task distillation loss function further includes: Hard label loss, which calculates the cross-entropy loss using the true labels; Soft label loss, which uses the output of the teacher model as the soft labels to guide the student model to learn the class probability distribution; Feature matching loss, which aligns the feature representations of the teacher and student models in the intermediate layers; Boundary awareness loss, which focuses on the feature learning of the chip defect boundary region.
8. A method for detecting visualization design defects of a chip based on deep learning according to claim 6, characterized in that, The attention-guided knowledge distillation method includes: Generate the attention map of the teacher model to reveal the key regions that the model focuses on; Construct the attention-guided loss to guide the student model to focus on the same key regions as the teacher model; Design a selective distillation strategy for different types of defects, including focusing on the shallow features for edge-type defects, the middle features for texture-type defects, and the deep features for structure-type defects.
9. The method for detecting design defects in chip visualization based on deep learning according to claim 6, wherein, The progressive knowledge distillation training process includes: Initialization phase, using a pre-trained teacher model to randomly initialize the student model; Feature matching phase, freezing the classification head of the student model and only optimizing the feature extraction part; Classification optimization phase, unfreezing the classification head of the student model and training it using the complete distillation loss function; Quantization fine-tuning phase, applying quantization-aware training to further reduce the accuracy loss caused by quantization.
10. A chip visualization design defect detection system based on deep learning, characterized in that, For executing the deep learning-based chip visualization design defect detection method as described in any one of claims 1-9, the system includes: A search space construction module for constructing a dedicated search space and a hardware performance evaluation model; A structure search module for implementing differentiable structure search to optimize the neural network structure for chip defect detection; An accuracy quantization module for implementing mixed-precision quantization to adaptively allocate the optimal bitwidth for different layers of the neural network; A joint optimization module for performing multi-objective joint optimization to optimize both the network structure and the quantization bitwidth configuration; A knowledge distillation module for applying the knowledge distillation technology to transfer the knowledge of a large teacher model to the optimized small student model.
Citation Information
Patent Citations
Software and hardware joint search method, device and equipment oriented to storage and calculation integrated architecture
CN115293341A
Strip steel surface defect identification method based on soft optimization knowledge distillation
CN116468686A
Surface defect detection method based on differentiable neural architecture search
CN117173091A
Chip defect detection method and system
CN117788427A
Knowledge distillation method and system based on attention correction features and boundary constraints
CN117973500A