Chip multi-station parallel consistency test method and device and storage medium
By constructing a chip collaborative damage model based on the electric domain perturbation potential energy factor and thermomechanical coupling dispersion, the consistency problem in asynchronous parallel testing was solved, achieving efficient multi-station parallel testing and improving the utilization rate of testing equipment and chip yield.
Patent Information
- Application Number
- CN202610047677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-05-15
AI Technical Summary
In the semiconductor back-end testing process, the asynchronous parallel testing architecture causes voltage drops, ground bounce noise, dynamic spatial thermal gradients, and mechanical warping issues, leading to contact instability and false failures, which affect test consistency and efficiency.
By collecting electrical and thermomechanical parameters, an electric domain perturbation potential energy factor and thermomechanical coupling dispersion are constructed to generate a chip collaborative damage index. An automated collaborative damage model is used for risk assessment, and an elastic phase shift mechanism is triggered to ensure the consistency and stability of the test environment.
It effectively eliminates transient voltage drop and ground bounce noise interference from shared power rails, ensures the stability of probe contact with pads, improves the first pass rate and yield of mass production testing, and reduces the invalid retest rate.
Smart Images

Figure CN122045761A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of technology, specifically to a method, equipment, and storage medium for multi-station parallel consistency testing of chips. Background Technology
[0002] In the semiconductor back-end testing stage, multi-station parallel testing has become the industry standard to reduce the time cost of expensive automated testing equipment and maximize hourly output. However, in the traditional rigid synchronous testing architecture, all stations under test must follow a uniform timing sequence, and the system is waiting for all various synchronization points. Limited by the frequency lock-in time or internal self-test speed differences caused by chip process variations, the testing process is often forced to be constrained by the slowest responding chip, the so-called "bottleneck effect." This rigid coupling not only leads to the idle resources of high-speed stations and significantly reduces the effective utilization rate of expensive equipment, but also becomes a macro bottleneck restricting the improvement of mass production throughput.
[0003] To improve efficiency, existing technologies are gradually evolving towards a fully asynchronous parallel testing architecture. The core lies in decoupling the timing control of each workstation, allowing each test workstation to independently schedule test threads based on the actual response of the device under test (DUT), implementing an independent unloading strategy after completion. Existing technologies typically attempt to achieve workstation isolation at both the physical and logical levels by configuring independent power channels and multi-threaded concurrent control software. Theoretically, this eliminates the bottleneck effect, aiming to maximize the utilization of machine resources and thus maximize the throughput of mass production testing.
[0004] However, while asynchronous architecture solves the timing efficiency problem, it artificially disrupts the steady-state characteristics of the test environment, introducing more complex nondeterministic dynamic disturbances and leading to deep-seated technical contradictions in conformance testing. First, timing decoupling causes high-load transient actions at adjacent workstations to randomly overlap on the time axis, inducing unpredictable voltage drops and ground bounce noise on the shared power distribution network or ground plane, resulting in unreproducible false failures. Second, for high-density strip-level testing, asynchronous operation causes severe dynamic spatial thermal gradients on the strip. The non-uniform temperature fields of adjacent workstations, one hot and one cold, combined with the mismatch in thermal expansion coefficients between the packaging material and the test socket, induce nonlinear mechanical warping of the strip. This dynamic deformation causes drastic fluctuations in the contact impedance between the test probes and the pads, leading to contact instability failures.
[0005] Therefore, the present invention provides a method, device and storage medium for multi-station parallel consistency testing of chips. Summary of the Invention
[0006] The purpose of this invention is to provide a chip multi-station parallel consistency testing method, device, and storage medium to solve the existing problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a chip multi-station parallel consistency testing method, comprising the following steps: S1. Collect electrical and thermomechanical parameters of this workstation and adjacent workstations, perform time-domain alignment, and generate a time-synchronized original feature vector containing electrical interference characteristics and thermomechanical deformation characteristics. S2. Based on the original feature vector, calculate the electric domain perturbation potential energy factor and the thermo-mechanical coupling dispersion, respectively. S3. Construct an automated collaborative damage model, taking the electric domain perturbation potential energy factor and the thermomechanical coupling dispersion as inputs, and generating a chip collaborative damage index through multi-dimensional feature mapping; S4. Compare the chip collaborative damage index with a preset risk threshold, and determine whether to trigger the elastic phase shift mechanism based on the comparison result; S5. Combine the test results with the chip collaborative damage index to generate an evaluation score, and update the parameters of the chip collaborative damage index based on the test results.
[0008] A further improvement of this invention is that the process of constructing the time-synchronized original feature vector includes: establishing a dual-channel data buffer pool, wherein the first channel is configured to operate at a first sampling frequency. The system collects voltage data from the power supply pins of this workstation and current data from all neighboring workstations sharing the same power supply. The second channel is configured to use a second sampling frequency. Collect temperature and contact resistance data between workstations, and Subsequently, using a sliding time window algorithm, downsampling features are extracted from the data in the first channel, and the root mean square value and peak-to-peak value within the unit window are calculated to align the time granularity with the data in the second channel. The original feature vector is then generated, including the neighborhood aggregated load rate Nal based on the intensity of the electrical interference source, the power rail transient noise Ptn based on the degree of disturbance at this workstation, the spatial thermal gradient Stg based on the difference in thermal distribution, and the contact resistance dispersion Crd based on the mechanical contact stability.
[0009] A further improvement of this invention is that the electric domain perturbation potential energy factor is obtained by logarithmically smoothing the neighborhood aggregated load rate Nal; then, the percentage of the power rail transient noise Ptn relative to the chip's preset voltage safety margin is calculated to obtain the relative noise intensity; finally, the smoothed neighborhood aggregated load rate and the relative noise intensity are convolved and weighted to calculate the electric domain perturbation potential energy factor Epe. The thermomechanical coupling dispersion is determined based on the thermal expansion coefficients of the test socket and packaging material supporting the chip, and the thermal coupling weighting coefficient is determined; the spatial thermal gradient Stg is used as a gain factor and applied to the square term of the contact resistance dispersion Crd; based on the weighted calculation results, the thermomechanical coupling dispersion Tec is obtained.
[0010] A further improvement of this invention is that the construction process of the automated collaborative damage model includes: based on a dual-stream long short-term memory network, the time synchronization sequence of the electric domain perturbation potential energy factor and the thermo-mechanical coupling discreteness is used as input and fed into the automated collaborative damage model, including a first input stream configured to process the electric domain feature sequence and a second input stream configured to process the thermo-mechanical feature sequence; then, the forget gate in the model is adjusted through a heat dissipation mapping control strategy to simulate the rapid insertion process of transient electric shock on the system state; then, the sensitivity of the model to the recognition of the electrothermal environment combination is adjusted through a sample failure control strategy; and finally, the chip collaborative damage index is output.
[0011] A further improvement of this invention is that the heat dissipation mapping control strategy includes: obtaining the physical and thermal characteristic parameters of the test socket supporting the chip and the surrounding fixture materials, and calculating the thermal time constant of the system. Based on the thermal time constant, the forget gate parameters of the LSTM units in the network are initialized and adjusted, and the bias term of the forget gate is set so that its weight decay rate matches the thermal dissipation time constant of the test seat. At the same time, the input gate weights of the LSTM units in the network are adjusted by increasing the input weight coefficient for the electric perturbation potential factor Epe, thereby reducing the excitation threshold of the electric perturbation potential factor activating the input gate.
[0012] A further improvement of this invention is that the sample failure control strategy includes: extracting historical test data, marking samples that initially failed but passed the retest as false failure positive samples, and marking the remaining failure samples as ordinary failure negative samples; and counting the total number of ordinary failure negative samples in the historical test data. Total number of positive samples with false failures And calculate the category balance coefficient. The network is trained using a weighted binary cross-entropy loss function, and the class balance coefficients are... The loss weights assigned to the false failure positive samples are set to be greater than those assigned to ordinary failure samples.
[0013] A further improvement of this invention is that the triggering determination process of the elastic phase shift mechanism includes: comparing the chip collaborative damage index with a preset dynamic risk threshold; when the chip collaborative damage index is greater than or equal to the dynamic risk threshold, determining that the current test station is in a high-risk interference window, suspending the current test station and inserting a microsecond-level random backoff delay; after the delay ends, triggering the re-execution of steps S1 to S3 to update the chip collaborative damage index, until the chip collaborative damage index is less than the dynamic risk threshold, and then triggering the test engine to execute the test.
[0014] A further improvement of the present invention is that step S5 includes: When the test result is passed, the confidence penalty logic is executed: a nonlinear decay function negatively correlated with the chip co-damage index is constructed to calculate the environmental confidence score; if the environmental confidence score is greater than or equal to the first preset threshold, the chip is determined to be a high-quality product and is assigned to a high-grade material warehouse; if the environmental confidence score is less than the first preset threshold but greater than or equal to the second preset threshold, the chip is determined to be an edge product and is assigned to a downgraded material warehouse or a safety retest is triggered. When the test result is a failure, the environmental attribution logic is executed: a mapping function positively correlated with the chip co-damage index is constructed, and the environmental confidence score is calculated; if the environmental confidence score is greater than or equal to a third preset threshold, the current failure is determined to be a false failure, and the yield loss is not included and the cleaning and retesting process is automatically triggered; if the environmental confidence score is less than the third preset threshold, the current failure is determined to be a hard failure and is sorted into the scrap silo. Finally, the feature sequences, test results, and retest results generated in this test are stored in the historical database and configured to periodically update the model parameters of the automated collaborative damage model.
[0015] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described chip multi-station parallel consistency testing method.
[0016] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned chip multi-station parallel consistency testing method.
[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention first solves the problem of transient voltage drop and ground bounce noise interference on the shared power rail caused by the superposition of random high load actions of adjacent workstations in asynchronous parallel testing by constructing an electric domain perturbation potential energy factor and combining it with an elastic phase shift mechanism. It achieves the avoidance of high-risk electrical conflict windows without reducing the global test parallelism, effectively eliminates logic deadlock and false failure caused by power integrity damage, and significantly improves the pass rate of mass production testing.
[0018] 2. By introducing an automated collaborative damage model, the nonlinear mechanical warping and contact impedance fluctuation problems caused by the mismatch between dynamic temperature differences between workstations and the thermal expansion coefficient of materials in high-density strip-level testing are solved. This enables early warning and proactive thermal hysteresis management of ghost contact failures, ensuring the stability of the test probe and chip pad contact in an extremely uneven dynamic thermal field environment and avoiding false kills caused by mechanical deformation.
[0019] 3. By comprehensively evaluating the chip collaborative damage index and generating environmental confidence scores, the technical contradiction of traditional testing solutions being unable to balance the consistency of individual test environments while pursuing high hourly output is resolved. This significantly reduces the invalid retest rate and provides a third dimension of data support for chip reliability classification, in addition to electrical performance. Attached Figure Description
[0020] Figure 1 This is a flowchart of a chip multi-station parallel consistency testing method according to the present invention. Detailed Implementation
[0021] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0022] The term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone.
[0023] Example 1 Figure 1 This embodiment presents a flowchart of a chip multi-station parallel consistency testing method, the steps of which are as follows: Step S1: Collect electrical and thermomechanical parameters from this workstation and adjacent workstations, and perform time-domain alignment to generate a time-synchronized original feature vector containing electrical interference characteristics and thermomechanical deformation characteristics; wherein the process of generating the time-synchronized original feature vector includes: establishing a dual-channel data buffer pool, wherein the first channel is configured to use a first sampling frequency. The system collects voltage data from the power supply pins of this workstation and current data from all neighboring workstations sharing the same power supply. The second channel is configured to use a second sampling frequency. Collect temperature and contact resistance data between workstations, and Subsequently, a sliding time window algorithm is used to downsample and extract features from the data in the first channel, calculating the root mean square value and peak-to-peak value within a unit window to align its time granularity with the data in the second channel; the original feature vector is then generated, including: Read the real-time current values of all N stations other than this station on the ATE power module (DPS). The neighborhood aggregated load factor Nal, which characterizes the intensity of the electric domain interference source, is obtained by weighted summation. The set of power supply pin voltage waveforms at this workstation is acquired using a digitizer at a high frequency of 1MHz. By filtering out the DC component and calculating the root mean square or peak-to-peak value of the AC component within the time window, the power rail transient noise Ptn, which characterizes the degree of disturbance at this workstation, is obtained. The temperature at this workstation is collected from the temperature sensor. and the temperature of adjacent workstations The difference between the two is calculated to obtain the spatial thermal gradient characterizing the difference in heat distribution. ; The set of contact resistance values obtained by continuous measurement using a PMU within a small time period. The standard deviation is calculated to obtain the contact resistance dispersion Crd, which characterizes the mechanical contact stability.
[0024] Step S2: Based on the original feature vector, calculate the electric domain disturbance potential energy factor that characterizes the intensity of external electrical environment interference, and the thermomechanical coupling dispersion that characterizes mechanical contact instability. Monitoring environmental consistency typically employs a single-parameter static threshold method. Specifically, existing technologies usually monitor independently whether the power supply voltage of a certain workstation is lower than the minimum value, or whether the chip surface temperature exceeds the maximum value.
[0025] However, in asynchronous testing scenarios, current interference from a single neighboring workstation may not be sufficient to trigger a threshold alarm. But when multiple neighboring workstations operate under high load simultaneously, the transient noise superimposed on the shared power plane can generate a nonlinear cumulative effect. Existing technologies only monitor the absolute voltage value of the workstation itself and cannot detect this potential risk caused by congestion, often resulting in false failures where the voltage appears normal but the chip logic is deadlocked. Secondly, existing technologies often treat temperature and contact resistance as two independent variables. However, in high-density strip-level testing, the essential impact of temperature is physical warping caused by the mismatch of thermal expansion coefficients, which in turn worsens contact resistance. Usually, temperature difference is the driving force, and material properties are merely a lever. The shortcomings of existing technologies make it impossible to identify those critical states where the temperature changes only slightly, but the contact is already on the verge of failure due to material differences, leading to false negatives in yield.
[0026] To address the aforementioned technical problems, this invention introduces the electric domain perturbation potential energy factor Epc and the thermo-mechanical coupling dispersion Crd; specifically including: S21. The electric domain disturbance potential energy factor is configured to quantify the potential for damage to the signal integrity of this workstation by the external electrical environment at the current moment, reflecting both the noise level and the number of interference sources.
[0027] S211. Logarithmically smooth the neighborhood aggregated load rate. In physical circuits, interference does not increase linearly and infinitely with the number of interference sources. For example, when the number of neighbors increases from 0 to 1, the impact on the decoupling capacitors of the power rail is greatest; however, when the number of neighbors increases from 10 to 11, due to the impedance characteristics of the power distribution network and the charge pool effect of the capacitors, the marginal impact of the new interference decreases. Therefore, firstly, the system collects the instantaneous current values of all adjacent workstations of the shared power supply and sums them to obtain the neighborhood aggregated load rate. Subsequently, according to Perform the transformation, where, The smoothed load index represents a potential source of interference. It is the natural logarithm function. The normalization coefficient can be set to 0.1 in this embodiment. A logarithmic function can be used to simulate the aforementioned diminishing marginal returns effect. This avoids unnecessary test suspensions caused by the algorithm overestimating interference risks under extremely high concurrent loads, ensuring the robustness of the algorithm under full load conditions.
[0028] S212, Normalized Relative Noise Intensity. For example, 50mV of noise might be safe for a 3.3V I / O supply, but fatal for a 0.8V Core supply; simply observing... The absolute value is meaningless. Therefore, the system uses a high-frequency digitizer to collect the transient noise of the power rail at this workstation. This is typically the peak-to-peak value, combined with the chip's preset voltage safety margin. ,Right now Calculate the relative noise intensity : By dividing by Dimensionless normalization is performed, making the metric universally applicable to chip testing across different voltage domains and process nodes, thus achieving the algorithm's versatility. At this point, it can be indicated that the noise level has exceeded the safety threshold.
[0029] S213. Finally, the indicators of potential interference sources... Actual interference performance Perform non-linear convolution: ;in, Let be the system sensitivity constant. The noise penalty index is usually... In this embodiment, 1.5 can be selected. This embodiment introduces an exponential term. It can demonstrate high sensitivity to actual noise levels, even if the neighboring unit has a high load. The noise level is relatively high, but as long as the power supply performs well, the actual noise level will be low. A lower value will result in a lower final electric field perturbation potential factor Epe; conversely, if the actual noise is higher... Slightly elevated, due to The existence of exponentiation dramatically amplifies the risk. It differentiates between busy and harsh environments, thereby maximizing test parallelism while ensuring yield and safety.
[0030] S22. Calculate the Tec quantification of thermomechanical coupling dispersion. This addresses the risk of mechanical contact degradation due to uneven temperature distribution, demonstrating how temperature differences lead to deformation, and deformation leads to contact failure. Specifically, this includes: S221. Determine the thermal coupling weighting coefficient. The root cause of warpage lies in the mismatch of the thermal expansion coefficients of the two bonding materials. If their thermal expansion coefficients are identical, they will expand and contract synchronously even at high temperatures, preventing warping. Therefore, the system pre-stores the thermal expansion coefficients of the chip packaging materials. and the coefficient of thermal expansion of the test seat material Calculate the thermal coupling weighting coefficient : , where k is the structure factor constant.
[0031] S222. Using the spatial thermal gradient as a gain factor, calculate the absolute value of the temperature difference between this workstation and adjacent workstations as the spatial thermal gradient Stg, and construct the thermal stress gain factor. : thermal expansion coefficient mismatch The temperature difference (Stg) is the underlying factor, while the temperature difference itself is the driving force. The greater the temperature difference, the greater the internal stress caused by uneven expansion, resulting in a linear increase in the magnitude of physical deformation. Simultaneously, using Stg as a gain factor, the lever amplification effect of thermal stress on the mechanical structure was simulated.
[0032] S223. Because the contact quality between the probe and the pad does not deteriorate linearly, the resistance change is small when the warpage is slight; however, once the warpage exceeds a certain critical point, the probe may experience a momentary suspension on the order of microseconds, causing the resistance to tend towards infinity. Therefore, it is necessary to collect multiple measurements of the contact resistance dispersion. Combining the above gain factors, the thermo-mechanical coupling dispersion Tec is obtained as follows: This embodiment is... By employing a squared term, the characteristics of critical abrupt changes can be mathematically simulated. This means that under low temperature differences, slight resistance fluctuations are considered tolerable; however, under high temperature differences with a high risk of warping, the same resistance fluctuations will be amplified by both the squared term and the gain factor, thereby providing early warning of potential open-circuit risks and preventing false alarms.
[0033] In one possible embodiment, Site 5 has just finished a high-temperature test, such as 85°C, while the adjacent Site 6 has just started a room-temperature test, such as 25°C; at this time, the contact resistance dispersion Crd of Site 6 fluctuates slightly, for example, from 2mΩ to 5mΩ, which is usually within the tolerance range.
[0034] Due to the significant temperature difference between the two workstations, the calculated spatial thermal gradient (Stg) reaches as high as 60°C; simultaneously, because the chip uses plastic packaging and a metal base, the CTE mismatch coefficient is high. Relatively large.
[0035] According to the formula The huge temperature difference gain factor multiplied by the square of the resistance fluctuation caused Tec to spike instantly, exceeding the safety threshold.
[0036] This embodiment of the system identifies that although the resistance reading appears to be still on the edge, the physical structure is already in a highly unstable, warped state, and could break at any time. Therefore, the system triggers a phase shift and waits for heat dissipation, thus successfully intercepting a false failure that is highly likely to occur and avoiding the subsequent cumbersome retesting process.
[0037] Step S3: Construct an automated collaborative damage model, taking the electric domain perturbation potential energy factor and the thermomechanical coupling dispersion as inputs, and generating a chip collaborative damage index that characterizes the risk level of the current test environment through multi-dimensional feature mapping; The construction process of the automated collaborative damage model includes: based on a dual-stream long short-term memory network, using a dual-stream fusion architecture, including: The first input stream consists of several layers of LSTM units, used to extract the temporal characteristics of the electric domain perturbation potential energy factor; The second input stream consists of several layers of LSTM units, used to extract the temporal features of the heat-engine coupling discreteness; Feature fusion layer: The hidden layer state vectors of the first and second input streams at the last time step are concatenated to form a joint feature vector.
[0038] Output layer: The joint feature vector is input into the fully connected layer and processed by the Sigmoid activation function. The final output is the chip collaborative damage index in the range of [0,1].
[0039] In conventional parallel chip testing, the thermal state of the test environment is typically considered memoryless, meaning the heat is assumed to dissipate immediately upon completion of the test. However, in high-density strip test sockets, heat exhibits a significant hysteresis. The conventional approach is to use standard time-series models to learn this hysteresis.
[0040] However, standard LSTM weights and biases are typically initialized randomly, such as with Xavier or He initialization. This means the model must learn the pattern of heat decay over time using thousands of training datasets. This leads to a severe cold start problem. In situations with insufficient data or during the initial introduction of a new product, the model cannot accurately predict heat accumulation, easily resulting in misjudgments. Furthermore, purely data-driven models lack physical constraints and may learn incorrect spurious correlations.
[0041] Therefore, the time synchronization sequence of the electric domain perturbation potential energy factor and the thermomechanical coupling discreteness is used as input and fed into the automated collaborative damage model, including a first input stream configured to process the electric domain feature sequence and a second input stream configured to process the thermomechanical feature sequence; then, the forget gate in the model is adjusted through a heat dissipation mapping control strategy to simulate the rapid insertion process of transient electric shocks into the system state; subsequently, the model's sensitivity to the identification of electrothermal environment combinations is adjusted through a sample failure control strategy; the chip collaborative damage index (CSDI) is output, representing the probability of the current test environment causing damage to the test results. Specifically, it includes: S31. Before constructing the model, the thermal time constant of the test environment should first be determined through experiments or thermal simulation. : ,in The thermal resistance of the test socket and PCB. This is the heat capacity. This constant determines how much the local temperature decreases from its peak value after the heat source is removed. The required time, in this embodiment, is approximately 36.8%. For example, the measurement of a certain high-density Strip test socket... .
[0042] S32. According to Newton's law of cooling, the decay of temperature difference follows an exponential law: Cellular state of LSTM In the update formula, the retention of the old state is determined by the output value of the forget gate. Decide If no new input is received, the process continues. After time step, the state decays to To enable the LSTM to remember heat and forget it according to physical laws, this invention sets the value of the forget gate... The mandatory constraint is: ,in The time step size for LSTM processing of the sequence can be selected as 0.1 seconds in this implementation.
[0043] S33. Specifically, the heat dissipation mapping control strategy is executed during the parameter initialization phase before model training. This is due to the LSTM forget gate output... To control the retention ratio of cell states and ensure that the model possesses memory capabilities that conform to physical heat dissipation laws in its initial state, this embodiment does not employ random initialization but instead calculates the target bias value. The forget gate bias term is initialized in a fixed manner.
[0044] Set the time step of LSTM to In this embodiment, 0.1s can be selected, based on the measured thermal time constant. According to the physical attenuation formula And combined with the Sigmoid activation function Based on the characteristics under steady-state input, the initialization formula is derived: In this way, the LSTM already possesses the ability to "hot memorize" the test environment before training, and will not fail to learn the thermal hysteresis effect due to the sparsity of training data, but will instead be infused with prior physical knowledge. This is achieved by initializing the forget gate bias term to the calculated value mentioned above. This ensures that the natural decay rate of the network's internal state is synchronized with the physical heat dissipation rate of the test bench before training begins. This parameter can be used to freeze fixed prior knowledge or as an initial point for subsequent fine-tuning.
[0045] S34. In contrast to the slow decay of the forget gate, the input gate is responsible for simulating the conversion of electrical energy into heat energy. Electric domain perturbations are typically on the order of microseconds. The input gate formula for LSTM is: .in The Sigmoid activation function has an output range of... .
[0046] when When the door is closed, input is ignored; when... At that time, the door is fully open, and all input is recorded.
[0047] The electric perturbation potential factor Epe typically manifests as a Dirac pulse with extremely short duration but dramatic amplitude, representing a momentary energy injection. To ensure that this short pulse can be captured, this embodiment allows... The product quickly enters the saturation region on the right side of the sigmoid function. Specifically, this includes: For the input vector For the dimension corresponding to Epe, the initial weight values are manually increased. For example, the initial weight of this dimension is set to k times the average weight of other dimensions. In this embodiment, k=5 is chosen. Simultaneously, the bias term of the input gate is appropriately reduced. The degree of negation, or setting it to a slightly positive value, puts the input gate in a subcritical state.
[0048] Through the above adjustments, when the electric domain perturbation potential energy factor Epe experiences a small abrupt change or rising edge, the input variable of the Sigmoid function will increase instantaneously due to the amplification of the weighting coefficients, leading to... It rapidly jumps from 0 to near 1. At this point, the transient electrical shock is no longer smoothed out by the model, but is written into and stored. Then, the heat is slowly dissipated by the forgetting gate.
[0049] By increasing the initial weight of the input gate for the electric perturbation potential factor Epe, it is ensured that the input gate can open instantaneously when the electric perturbation potential factor Epe shows a spike, thus rapidly writing energy into the cell state. Then the Forgotten Gate takes over its long cooling process.
[0050] In one possible embodiment, certain tests, such as RF calibration, generate high-frequency current pulses lasting only 10ms with irregular intervals. Conventional temperature sensors, due to their typically 1Hz-10Hz sampling frequency and thermal conduction delay, often cannot react in time; the pulse ends before the reading rises. In this embodiment, the input gate is highly sensitively tuned to capture rapid abrupt changes in the electric field factor Epe. This allows the automated co-injury model to detect cell states even when the temperature sensor does not react. It will be injected with a high value because the input gate opens all at once, and will begin to follow... Slow decay. The automated collaborative damage model successfully established a virtual thermal stress model in the sensor blind zone. Although the current temperature reading is not high, the system has accumulated a very high risk of thermomechanical failure, thus avoiding ghost failure.
[0051] S35. Sample failure control strategies include: extracting test records from historical test data of automated testing equipment and setting sample labels. Where 1 represents a high-risk environmental sample, the labeling process is as follows: S351. Find all samples that meet the condition of failure in the initial test but pass in the retest and mark them as false failure positive samples. Mark the feature sequence recorded at the time of the initial test as y=1. S352. Mark the remaining failure samples except those in step S351 as ordinary failure negative samples; mark the feature sequence at this time as y=0; S353. Remove all samples that fail in both the initial test and the retest. These samples are usually considered to be due to hardware manufacturing defects in the chip itself, and have little to do with environmental fluctuations. To avoid interfering with the model’s learning of environmental features, these samples are removed from the training set.
[0052] S354. In the above dataset, the number of positive samples is far less than the number of negative samples, possibly reaching a ratio of 1:100. In the laboratory, 1% is an error; however, in mass production, 1% represents a huge financial loss, especially for the chip industry. Assuming a monthly production of 10 million chips, a 1% false failure rate means 100,000 chips are mistakenly rejected. Without retesting, discarding 100,000 good chips results in a yield loss; with full retesting, these 100,000 chips would have to be re-tested, wasting machine time. To address this issue, this embodiment introduces a class balance coefficient. : First, count the total number of common failure negative samples in the historical test data. Total number of positive samples with false failures For example, the system is configured to update statistical data every M batches or every 24 hours; and calculate the category balance coefficient. ; S355. During training, the network is trained using a weighted binary cross-entropy loss function as the objective function, and the class balance coefficients are... As the loss weight for the false failure positive samples, the loss value for each input sample i in the training step is... Represented as: in, For true labels (1 or 0); The predicted probability output by the model is the chip collaborative damage index; the loss weight assigned to the fake failure sample is set to be greater than that of the ordinary failure sample.
[0053] In this embodiment, when processing positive samples... The loss function becomes If the model predicts a very low chip-related synergistic damage index, such as 0.1, it means the risk has been missed. It is a large negative number, multiplied by a huge class balance coefficient. This can lead to a very large loss value. At this point, the loss function will generate a sharp gradient during backpropagation, strongly correcting the weight parameters inside the model. This mechanism forces the model to not miss any weak feature combinations that might cause false failures, even if it misreports a safe environment as a risky one, thereby greatly improving the model's recall rate for high-risk electrothermal environments.
[0054] Step S4: Compare the chip collaborative damage index with a preset risk threshold, and determine whether to trigger the elastic phase shift mechanism to suspend the test based on the comparison result; In existing multi-station testing solutions, there are two main extreme strategies: A. Fully Synchronous Mode: This mode uses a wait instruction, where all workstations must wait for the slowest one to complete before proceeding to the next. This results in significant wasted time and low output due to the bottleneck effect.
[0055] B. Fully Asynchronous Mode: Each workstation operates completely independently. However, this can lead to random collisions. For example, when Site1 is experiencing a high-current transient, Site2 might be entering a voltage-sensitive test, causing Site2 to falsely report a failure due to power fluctuations. This type of collision is random and unreproducible.
[0056] Therefore, this invention introduces a flexible phase shift. That is, when a high risk is detected, the testing process is not completely stopped, but rather, like a traffic light controlling an intersection, a microsecond-level, random time delay is applied to the current workstation. Specifically, this includes: S41. Dynamic Risk Threshold Determination: The system first obtains the sensitivity attribute of the current test item. For example, Efuse programming items have high sensitivity, while ordinary GPIO connectivity tests have low sensitivity, which can be set by those skilled in the art.
[0057] Based on this, the current dynamic risk threshold is determined. : At this point, the more sensitive the test item and the lower the threshold, the easier it is to trigger avoidance.
[0058] S42, Then, the chip collaborative damage index output in step S3 is compared with... A comparison is made. If the chip collaborative damage index is less than the dynamic risk threshold, the current environment is determined to be a safe window. The system immediately sends an execution command to the ATE's test engine, executing the current test item directly without introducing any additional delay.
[0059] When the chip collaborative damage index is greater than or equal to the dynamic risk threshold, it is determined that the current test station is in a high-risk interference window, the current test station is suspended, and a microsecond-level random backoff delay is inserted. : in, Basic avoidance time, basic avoidance time The value is based on the typical duration of high-load operations at neighboring workstations. The typical time for the target chip's internal phase-locked loop to start and stabilize is 4ms. To ensure complete avoidance of transient current surges from neighboring devices.
[0060] The random jitter time is, for example, a uniformly distributed random number between 0 and 10 ms. The upper limit of the value is set to By introducing This ensures that the wake-up time of each workstation is evenly distributed on the time axis, thus mathematically reducing the probability of multiple workstations re-occurring synchronous resonance.
[0061] In this embodiment, if all affected workstations wait a fixed 5ms, they will all wake up simultaneously after 5ms, pulling current again, leading to a new and more severe collision deadlock. Introducing random jitter discretizes the wake-up time of each workstation on the time axis, eliminating the risk of resonance. The time limit is very small (microseconds to milliseconds), and compared to the several seconds required to retest the entire chip, this tiny wait has a negligible impact on UPH.
[0062] S43. After the delay ends, steps S1 to S3 are re-executed to update the chip collaborative damage index until the chip collaborative damage index is less than the dynamic risk threshold, at which point the test engine is triggered to execute the test.
[0063] Step S5: Combine the test results with the chip collaborative damage index to generate an evaluation score, and update the parameters of the chip collaborative damage index based on the test results. Specific steps include: When the test result is passed, the confidence penalty logic is executed: a nonlinear decay function negatively correlated with the chip co-damage index is constructed, and the environmental confidence score is calculated. in The penalty coefficient is set to 0.8 in this embodiment. The non-linear factor is set to 2.0 in this embodiment, and csdi is the chip-related damage index. This formula results in fewer deductions under low risk, but the score decreases sharply and non-linearly as the risk increases.
[0064] If the environmental confidence score is greater than or equal to the first preset threshold, the chip is determined to be a high-quality product and is placed in a high-grade material warehouse; if the environmental confidence score is less than the first preset threshold but greater than or equal to the second preset threshold, the chip is determined to be an edge product and is placed in a downgraded material warehouse or a safety retest is triggered. When the test result is a failure, the environmental attribution logic is executed: a mapping function positively correlated with the chip's co-damage index is constructed, and the environmental confidence score is calculated. This characterizes the possibility of failure caused by environmental disturbances; E2 characterizes the probability that failure is caused by the environment.
[0065] If the environmental confidence score is greater than or equal to the third preset threshold, the current failure is determined to be a false failure, and the yield loss is not included, and the cleaning and retesting process is automatically triggered; if the environmental confidence score is less than the third preset threshold, the current failure is determined to be a hard failure and is sorted into the scrap silo. Finally, the feature sequences, test results, and retest results generated in this test are stored in the historical database and configured to periodically update the model parameters of the automated collaborative damage model.
[0066] To prevent infinite loops, for example, if a neighboring appliance breaks down and keeps overheating, a maximum number of retries is set. In this embodiment, three attempts are selected. If the environment remains harsh after three consecutive avoidance attempts, the test is forcibly executed, but the environmental confidence score of the chip is marked as extremely low in step S5, and it is then directly sent for retesting.
[0067] The threshold and weight settings involved in this embodiment can be set by default according to the present invention, or can be set by those skilled in the art.
[0068] Example 2 This embodiment provides an electronic device, including: a processor and a memory, wherein the memory stores a computer program that can be called by the processor; The processor executes the aforementioned chip multi-station parallel consistency test method by calling the computer program stored in memory.
[0069] This electronic device can vary considerably depending on its configuration or performance. It may include one or more processors (Central Processing Units, CPUs) and one or more memories, wherein the memory stores at least one computer program, which is loaded and executed by the processor to implement the chip multi-station parallel consistency testing method provided in the above-described embodiment. The electronic device may also include other components for implementing device functions; for example, it may have wired or wireless network interfaces and input / output interfaces for data input and output. Details will not be elaborated upon in this embodiment.
[0070] Example 3 This embodiment proposes a computer-readable storage medium on which an erasable and rewritable computer program is stored. When a computer program runs on a computer device, it causes the computer device to perform the aforementioned chip multi-station parallel consistency test method.
[0071] For example, computer-readable storage media can be read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices.
[0072] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0073] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0074] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0075] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0076] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for multi-station parallel consistency testing of chips, characterized in that: Includes the following steps: S1. Collect electrical and thermomechanical parameters of this workstation and adjacent workstations, perform time-domain alignment, and generate a time-synchronized original feature vector containing electrical interference characteristics and thermomechanical deformation characteristics. S2. Based on the original feature vector, calculate the electric domain perturbation potential energy factor and the thermo-mechanical coupling dispersion, respectively. S3. Construct an automated collaborative damage model, taking the electric domain perturbation potential energy factor and the thermomechanical coupling dispersion as inputs, and generating a chip collaborative damage index through multi-dimensional feature mapping; S4. Compare the chip collaborative damage index with a preset risk threshold, and determine whether to trigger the elastic phase shift mechanism based on the comparison result; S5. Combine the test results with the chip collaborative damage index to generate an evaluation score, and update the parameters of the chip collaborative damage index based on the test results.
2. The chip multi-station parallel consistency testing method according to claim 1, characterized in that: The process of constructing the original feature vector for time synchronization includes: establishing a dual-channel data buffer pool, wherein the first channel is configured to operate at a first sampling frequency. The system collects voltage data from the power supply pins of this workstation and current data from all neighboring workstations sharing the same power supply. The second channel is configured to use a second sampling frequency. Collect temperature and contact resistance data between workstations, and Subsequently, using a sliding time window algorithm, downsampling features are extracted from the data in the first channel, and the root mean square value and peak-to-peak value within the unit window are calculated to align the time granularity with the data in the second channel. The original feature vector is then generated, including the neighborhood aggregated load rate Nal based on the intensity of the electrical interference source, the power rail transient noise Ptn based on the degree of disturbance at this workstation, the spatial thermal gradient Stg based on the difference in thermal distribution, and the contact resistance dispersion Crd based on the mechanical contact stability.
3. The chip multi-station parallel consistency testing method according to claim 2, characterized in that: The electric domain perturbation potential energy factor is obtained by logarithmically smoothing the neighborhood aggregated load rate Nal; then, the percentage of the power rail transient noise Ptn relative to the chip's preset voltage safety margin is calculated to obtain the relative noise intensity; finally, the smoothed neighborhood aggregated load rate and the relative noise intensity are convolved and weighted to calculate the electric domain perturbation potential energy factor Epe. The thermomechanical coupling dispersion is determined based on the thermal expansion coefficients of the test socket and packaging material supporting the chip, and the thermal coupling weighting coefficient is determined accordingly. The spatial thermal gradient Stg is used as a gain factor and applied to the square term of the contact resistance dispersion Crd; based on the weighted calculation results, the thermomechanical coupling dispersion Tec is obtained.
4. The chip multi-station parallel consistency testing method according to claim 1, characterized in that: The construction process of the automated collaborative damage model includes: based on a dual-stream long short-term memory network, the time synchronization sequence of the electric domain perturbation potential energy factor and the thermo-mechanical coupling discreteness is used as input and fed into the automated collaborative damage model, including a first input stream configured to process the electric domain feature sequence and a second input stream configured to process the thermo-mechanical feature sequence; then, the forget gate in the model is adjusted through a heat dissipation mapping control strategy to simulate the rapid insertion process of transient electric shock on the system state; then, the sensitivity of the model to the recognition of the electrothermal environment combination is adjusted through a sample failure control strategy; and finally, the chip collaborative damage index is output.
5. The chip multi-station parallel consistency testing method according to claim 4, characterized in that: The heat dissipation mapping and control strategy includes: obtaining the physical and thermal characteristic parameters of the test socket and surrounding fixture materials supporting the chip, and calculating the thermal time constant of the system. Based on the thermal time constant, the forget gate parameters of the LSTM units in the network are initialized and adjusted, and the bias term of the forget gate is set so that its weight decay rate matches the thermal dissipation time constant of the test seat. At the same time, the input gate weights of the LSTM units in the network are adjusted by increasing the input weight coefficient for the electric perturbation potential factor Epe, thereby reducing the excitation threshold of the electric perturbation potential factor activating the input gate.
6. The chip multi-station parallel consistency testing method according to claim 4, characterized in that: The sample failure control strategy includes: extracting historical test data, marking samples that initially failed but passed the retest as false failure positive samples, and marking the remaining failure samples as ordinary failure negative samples; and counting the total number of ordinary failure negative samples in the historical test data. Total number of positive samples with false failures And calculate the category balance coefficient. The network is trained using a weighted binary cross-entropy loss function, and the class balance coefficients are... The loss weights assigned to the false failure positive samples are set to be greater than those assigned to ordinary failure samples.
7. The chip multi-station parallel consistency testing method according to claim 1, characterized in that: The triggering and determination process of the elastic phase shift mechanism includes: comparing the chip collaborative damage index with a preset dynamic risk threshold; when the chip collaborative damage index is greater than or equal to the dynamic risk threshold, determining that the current test station is in a high-risk interference window, suspending the current test station and inserting a microsecond-level random backoff delay; after the delay ends, triggering the re-execution of steps S1 to S3 to update the chip collaborative damage index, until the chip collaborative damage index is less than the dynamic risk threshold, and then triggering the test engine to execute the test.
8. The chip multi-station parallel consistency testing method according to claim 7, characterized in that: Step S5 includes: When the test result is passed, the confidence penalty logic is executed: a nonlinear decay function negatively correlated with the chip co-damage index is constructed to calculate the environmental confidence score; if the environmental confidence score is greater than or equal to the first preset threshold, the chip is determined to be a high-quality product and is assigned to a high-grade material warehouse; if the environmental confidence score is less than the first preset threshold but greater than or equal to the second preset threshold, the chip is determined to be an edge product and is assigned to a downgraded material warehouse or a safety retest is triggered. When the test result is a failure, the environmental attribution logic is executed: a mapping function positively correlated with the chip co-damage index is constructed, and the environmental confidence score is calculated; if the environmental confidence score is greater than or equal to a third preset threshold, the current failure is determined to be a false failure, and the yield loss is not included and the cleaning and retesting process is automatically triggered; if the environmental confidence score is less than the third preset threshold, the current failure is determined to be a hard failure and is sorted into the scrap silo. Finally, the feature sequences, test results, and retest results generated in this test are stored in the historical database and configured to periodically update the model parameters of the automated collaborative damage model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements any one of the chip multi-station parallel consistency testing methods according to claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements a chip multi-station parallel consistency testing method according to any one of claims 1-8.