ATE device health assessment and failure prediction method and system based on digital signal processing

CN122815293APending Publication Date: 2026-09-25BEIJING YUEXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610995785.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有技术通常未能将设备实时状态、校准基准、测试业务反馈和部件生命周期信息进行融合分析,导致健康度评估结果缺乏完整依据,故障预测准确性和维护建议针对性不足

Benefits of technology

本发明通过融合实时传感数据、校准基准数据、测试结果变化数据和周期性维护数据,构建了面向ATE设备的多源健康监测体系,并结合小波包变换、自适应阈值去噪、随机森林模型或LSTM模型,对设备运行状态进行多尺度特征提取、健康度评估和故障概率预测;同时,本发明可根据SoC ATE、Memory ATE和Power ATE等不同设备类型自适应配置传感器部署位置、特征提取参数和故障模式模板,提高跨设备适配能力;在满足预警条件时,系统能够输出包含设备标识、故障位置、关键监测参数、健康度评分、故障概率、推荐维护动作和预计剩余安全运行时间的结构化预警信息,并可结合实际故障事件对模型进行持续优化,从而提高早期故障识别准确性,降低误报和漏报风险,实现从设备状态监测到预测性维护决策的自动化闭环。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122815293A_ABST
    Figure CN122815293A_ABST
Patent Text Reader

Abstract

The application discloses an ATE device health degree evaluation and fault prediction method and system based on digital signal processing. The method comprises the following steps: collecting multi-source sensing signals, calibration data, test result change data and periodic maintenance data during the operation of the ATE device; performing time-frequency decomposition and adaptive denoising on the multi-source sensing signals by using wavelet packet transform, and extracting multi-dimensional fault sensitive features; inputting the features into a random forest model or an LSTM model after normalization fusion, and outputting device health degree scores and fault mode probability prediction values; and when the health degree score, the fault probability, the calibration deviation, the test result change or the maintenance period meet the early warning conditions, generating structured early warning information and outputting the same. The application can realize online monitoring, health degree evaluation and early fault warning of different types of automatic test equipment such as SoC ATE, Memory ATE and Power ATE, improve prediction accuracy and cross-device universality, and reduce equipment operation and maintenance costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault prediction technology, specifically to a method and system for health assessment and fault prediction of ATE equipment based on digital signal processing. Background Technology

[0002] In semiconductor manufacturing and integrated circuit testing, ATE (Automatic Test Equipment) devices, as core equipment for chip functional verification, performance testing, and quality screening, typically require continuous operation for extended periods. The operational stability of ATE devices directly impacts chip testing yield, testing efficiency, production line cycle time, and testing costs. For mass production testing scenarios, any malfunction in the ATE device can lead to the loss of batch test data, interruption of test tasks, equipment downtime for maintenance, and even disruption to subsequent production scheduling.

[0003] Existing ATE equipment operational status monitoring methods typically rely on a single signal source or limited data types, such as collecting only temperature, vibration, power supply, or test yield data. These methods lack comprehensive data mining capabilities and struggle to promptly detect early, weak characteristics of progressive faults, such as loose test head cage backplane connectors, aging digital channel drivers, degradation of power board devices, wear of cooling fan bearings, and deterioration of the reference clock source. Alarms are usually only triggered after obvious anomalies occur, resulting in significant warning delays. Different types of ATE equipment exhibit significant differences in hardware structure, workload, and fault manifestations. For example, anomalies in SoC ATE equipment may manifest as digital channel timing deviations, uneven load distribution in multi-site parallel testing, or abnormal output from the DPS power module; anomalies in Memory ATE equipment may manifest as reference clock drift, decreased timing control accuracy, or fluctuations in read / write test parameters; and anomalies in Power ATE equipment may manifest as overheating of the high-voltage power module, aging of power devices, abnormal power supply ripple, or decreased efficiency of the cooling system. Existing methods often require customized deployment for specific equipment models, lacking a unified health assessment framework that can be reused across equipment types. ATE testing environments typically contain complex sources of interference, including electromagnetic interference, mechanical vibration interference, test fixture insertion and removal shocks, and transient load changes. Traditional time-domain statistical methods or simple frequency-domain analysis methods are insufficient to effectively extract weak fault characteristics from multi-source signals, easily leading to false alarms or missed alarms.

[0004] On the other hand, ATE equipment-related data is typically scattered across different systems. For example, sensor data is stored in the equipment monitoring system, calibration data in the calibration management system, wafer yield and test parameter deviations are stored in the MES or test data platform, and maintenance cycles and historical repair records are stored in the operations and maintenance system. Existing technologies generally fail to integrate and analyze real-time equipment status, calibration benchmarks, test business feedback, and component lifecycle information, resulting in health assessment results lacking complete basis, and insufficient accuracy in fault prediction and targeted maintenance recommendations. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for health assessment and fault prediction of ATE equipment based on digital signal processing, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for health assessment and fault prediction of ATE equipment based on digital signal processing includes the following steps: Step S1: Multi-source data acquisition Based on the distribution of core components, equipment type, and failure mode characteristics of ATE equipment, sensors are deployed in key parts of the ATE equipment to collect multi-source sensor signals during equipment operation; at the same time, calibration data, test result change data, and periodic maintenance data are also collected.

[0007] The multi-source sensing signals include vibration signals, temperature signals, current signals, and digital channel timing signals. The calibration data includes sensor calibration data, ATE equipment resource board calibration data, timing calibration deviation data, power output accuracy calibration data, and reference clock source calibration data. The test result variation data includes wafer test yield fluctuations, chip test parameter deviations, test pass rate variations, batch test result consistency deviations, and the proportion of abnormal test data. The periodic maintenance data includes preset maintenance cycles, maintenance standards, historical maintenance times, maintenance content, and post-maintenance status parameters for components, boards, cooling systems, and pneumatic systems.

[0008] In one implementation, the sensor deployment scheme is adaptively configured according to the ATE device type: for SoCATE, sensors are mainly deployed in areas with dense digital channel boards, multi-site parallel test interfaces, test head cage backplane connectors, and DPS power modules; for Memory ATE, sensors are mainly deployed near timing control units, reference clock sources, and memory interface circuits; for Power ATE, sensors are mainly deployed near high-voltage power modules, power drive circuits, power device arrays, and heat dissipation systems.

[0009] Step S2: Data Preprocessing and Feature Extraction The analog signals acquired by the sensors are converted into digital signals using a high-speed data acquisition card and transmitted to the signal processing unit via a high-speed data transmission bus. Simultaneously, calibration data, test result change data, and periodic maintenance data are imported into the signal processing unit to perform time synchronization, anomaly removal, standardization processing, and feature extraction on various types of data.

[0010] In the signal processing unit, wavelet packet transform is used to perform multi-scale time-frequency decomposition on vibration signals, temperature signals, current signals and digital channel time-series signals, and an adaptive threshold denoising method is used to suppress transient noise, mechanical shock noise and electromagnetic interference noise.

[0011] Multidimensional fault-sensitive features are extracted from the preprocessed data, including: Vibration signal characteristics: frequency band energy ratio, spectral kurtosis index, and center frequency shift of key frequency bands; Temperature signal characteristics: rate of change of temperature gradient, uniformity of hot spot distribution, and temperature-current coupling characteristics; Current signal characteristics: transient impulse pulse interval, duty cycle parameter, and ripple amplitude fluctuation trend; Digital channel timing signal characteristics: edge deviation standard deviation, timing drift trend, timing margin degradation rate; Calibration data characteristics: calibration deviation trend, frequency of calibration errors exceeding the allowable range, and calibration consistency deviation between channels; Characteristics of test result changes: yield fluctuation range, mean and rate of change of test parameter deviation, decrease in pass rate, and consistency deviation of batch test results; Periodic maintenance characteristics: maintenance cycle deviation, average maintenance effect deviation, maintenance frequency anomaly coefficient, and duration of non-timely maintenance.

[0012] Step S3: Multi-source feature fusion and health modeling The various fault-sensitive features extracted in step S2 are normalized to form a unified multi-source feature vector. This multi-source feature vector is then input into a pre-trained health assessment and fault prediction model, which outputs the health score and fault mode probability prediction value of the ATE equipment.

[0013] The health assessment and fault prediction model employs either a random forest model or a Long Short-Term Memory (LSTM) network model. The random forest model is trained based on historical normal operation data, historical fault data, historical calibration data, historical test result data, and historical maintenance data. It outputs the current status level of the device through the ensemble analysis of multiple decision trees. The current status level of the device includes normal, attention, warning, and danger.

[0014] The LSTM model is trained and generated based on time series feature data. It is used to capture the long-term dependencies of equipment status changes over time, and outputs a health score of 0 to 100% and a probability prediction of the failure mode of key components within a preset future time period.

[0015] In one implementation, the health assessment and fault prediction model also incorporates partial least squares analysis to perform dimensionality reduction and screening of high-dimensional multi-source features, retaining features that are highly correlated with fault modes, thereby reducing model training complexity and improving prediction stability.

[0016] Step S4: Tiered early warning output The system presets dynamic thresholds for health scores, critical values ​​for failure mode probabilities, allowable ranges for calibration deviations, thresholds for abnormal test result changes, and maintenance warning thresholds. It then compares health scores with dynamic thresholds, predicted failure mode probabilities with critical values, calibration deviations with allowable ranges, test result changes with abnormal thresholds, and periodic maintenance data with maintenance warning thresholds in real time.

[0017] The system generates structured warning information when any of the following conditions are met: The device health score is below the preset dynamic threshold; The predicted probability value of any failure mode exceeds the preset probability threshold. The calibration deviation of resource boards, sensors, or critical components continues to exceed the allowable range; The test results showed changes exceeding the preset anomaly threshold, and the model analysis determined that these changes were related to abnormal equipment status. The actual operating time of the component is close to or exceeds the preset maintenance cycle.

[0018] The structured early warning information includes device identifier, early warning timestamp, current health score, probability values ​​of each fault mode, location of fault-related components, key monitoring parameters, recommended maintenance action type, and estimated remaining safe operating time. The structured early warning information is output in JSON or Protobuf format and pushed to the upper-level test task scheduling module, MES system, remote monitoring platform, or equipment operation and maintenance system.

[0019] On the other hand, the present invention also provides a health assessment and fault prediction system for ATE equipment based on digital signal processing, including a multi-source data acquisition module, a data acquisition and conversion module, a digital signal processing module, a health assessment and prediction module, and an early warning output interface.

[0020] The multi-source data acquisition module is deployed in the core part of the ATE equipment and includes vibration sensors, temperature sensors, current sensors, timing signal acquisition units, calibration data acquisition units, test result acquisition units, and periodic maintenance data acquisition units. It is used to acquire equipment operating status data, calibration reference data, test service feedback data, and maintenance life cycle data.

[0021] The data acquisition and conversion module is connected to the multi-source data acquisition module and includes a high-speed data acquisition card, a signal conditioning circuit, and a data import unit. It is used to convert analog sensing signals into digital signals and to transmit calibration data, test result change data, and periodic maintenance data to the digital signal processing module after standardization processing.

[0022] The digital signal processing module is connected to the data acquisition and conversion module, and has built-in wavelet packet transform unit, adaptive denoising unit, calibration data processing unit, test result processing unit, periodic maintenance data processing unit and feature extraction unit, which are used to preprocess, decompose, denoise and extract multidimensional fault-sensitive features from the received data.

[0023] The health assessment and prediction module is connected to the digital signal processing module and has built-in model training unit, health assessment unit, fault prediction unit and model optimization unit. It is used to input multi-source feature vectors into the prediction model and output the device health score, status level and fault mode probability prediction value.

[0024] The warning output interface is connected to the health assessment and prediction module, and has a built-in warning judgment unit, information generation unit and data push unit. It is used to generate standardized structured warning information when the warning conditions are met and push it to an external platform.

[0025] Compared with the prior art, the beneficial effects of the present invention are: This invention constructs a multi-source health monitoring system for ATE (Automatic Equipment) devices by integrating real-time sensor data, calibration benchmark data, test result change data, and periodic maintenance data. It combines wavelet packet transform, adaptive threshold denoising, random forest models, or LSTM models to perform multi-scale feature extraction, health assessment, and fault probability prediction of device operating status. Furthermore, this invention can adaptively configure sensor deployment locations, feature extraction parameters, and fault mode templates according to different device types such as SoC ATE, Memory ATE, and Power ATE, improving cross-device adaptability. When warning conditions are met, the system can output structured warning information including device identification, fault location, key monitoring parameters, health score, fault probability, recommended maintenance actions, and estimated remaining safe operating time. The system can continuously optimize the model based on actual fault events, thereby improving the accuracy of early fault identification, reducing the risk of false alarms and missed alarms, and achieving an automated closed loop from device status monitoring to predictive maintenance decision-making. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] In this invention, unless otherwise explicitly specified and limited, the terms installation, connection, linking, fixing, etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0030] Example 1: Health Assessment and Fault Prediction of SoC ATE Devices like Figure 1As shown, a semiconductor testing facility deploys multiple SoC ATE devices, primarily for wafer and finished product testing of SoC chips. Typical failures of this type of equipment include loose test head cage backplane connectors, timing degradation of digital channel boards, decreased consistency of multi-site parallel testing, and abnormal output of DPS power modules.

[0031] During the multi-source data acquisition phase, a triaxial accelerometer is installed at the connector on the backplane of the test head cage. A PT1000 platinum resistance temperature sensor and an open-loop Hall current sensor are deployed on the digital channel board and DPS power board, respectively. The timing signal acquisition unit connects to the test timing monitoring port of the digital channel board to acquire data on digital channel edge deviation, timing drift, and timing margin changes. The calibration data acquisition unit simultaneously acquires sensor calibration records, digital channel board timing calibration deviation, and DPS power board output accuracy calibration data. The test result acquisition unit simultaneously acquires chip functional test parameter deviations, multi-site test consistency deviations, test pass rate fluctuations, and batch test yield data. The periodic maintenance data acquisition unit imports the preset maintenance cycles for the digital channel board and DPS power module during equipment deployment and updates the maintenance time, maintenance content, and post-maintenance status parameters after each maintenance.

[0032] In the data preprocessing and feature extraction stage, a multi-channel synchronous sampling high-speed data acquisition card is used to convert vibration, temperature, and current analog signals into digital signals, which are then transmitted to the signal processing server via the PCIe bus. The signal processing server uses the Daubechies wavelet basis function to perform wavelet packet decomposition on the multi-source sensor signals. The number of decomposition levels can be configured according to the sampling frequency and the equipment vibration frequency band. In this embodiment, the energy ratio, spectral kurtosis, and center frequency offset in the 3kHz to 6kHz frequency band are extracted from the vibration signal of the test head cage backplate connector; the hot spot distribution uniformity and temperature rise gradient change rate are extracted from the temperature signal of the digital channel board; the transient impact pulse interval and ripple amplitude fluctuation trend are extracted from the current signal of the DPS power board; the edge deviation standard deviation and timing margin degradation rate are extracted from the timing signal of the digital channel; the calibration deviation trend is extracted from the calibration data; the yield fluctuation amplitude and pass rate decrease amplitude are extracted from the test result change data; and the maintenance cycle deviation and the duration of non-timely maintenance are extracted from the maintenance data.

[0033] During the health assessment and early warning phase, the normalized multi-source feature vectors are input into the trained LSTM health model. During the operation of a certain SoC ATE device, the system detected a continuous increase in the energy ratio of the 3kHz to 6kHz frequency band in the test head cage backplane connector area, a shift in the center frequency of the critical frequency band, uneven hotspot distribution on the digital channel board, irregular spikes in the DPS power supply current waveform, a significant increase in the timing calibration deviation of the digital channel board compared to the previous calibration cycle, a decrease in the test pass rate, and the digital channel board approaching its preset maintenance cycle. The LSTM model outputs a health score for the device below the dynamic threshold and outputs predicted failure modes for "loose backplane connector contact" and "digital channel board timing degradation." Based on this, the system generates a structured early warning message in JSON format and pushes it to the MES system and the equipment maintenance platform. Maintenance personnel tighten the test head cage backplane connector and clean the contact surfaces according to the early warning information, and recalibrate the digital channel board. After maintenance, the relevant monitoring characteristics return to normal ranges.

[0034] In the above implementation, the system does not rely solely on the abnormal value of a single sensor for alarm activation. Instead, it combines vibration frequency band characteristics, temperature distribution, current waveform, timing calibration deviation, test pass rate changes, and maintenance cycle data for joint judgment. Therefore, even if the change in a single signal is small, potential faults can be identified through the correlation between multiple source characteristics, improving the reliability of early fault warnings.

[0035] Example 2: Health Assessment and Fault Prediction of Memory ATE Devices A memory testing production line uses Memory ATE equipment to test DRAM, NAND Flash, or other memory chips. This type of equipment has high requirements for timing control accuracy, reference clock stability, and memory interface signal integrity. Typical faults include reference clock source degradation, timing control unit drift, poor memory interface contact, and DPS power supply loop abnormalities.

[0036] In this embodiment, a high-frequency vibration sensor is placed near the timing control board, a high-precision temperature sensor is placed near the reference clock oscillator, a high-frequency current probe is placed in the DPS power supply loop, and a timing signal acquisition unit is placed near the storage interface circuit. The calibration data acquisition unit focuses on acquiring calibration data from the timing control board and accuracy calibration data from the reference clock source. The test result acquisition unit focuses on acquiring data such as memory chip read / write speed deviation, timing parameter test deviation, wafer test yield fluctuation, and batch test result consistency deviation. The periodic maintenance data acquisition unit focuses on acquiring the maintenance cycle and historical maintenance records of the timing control board, reference clock source, and interface connection components.

[0037] The signal processing unit uses wavelet packet transform to decompose the reference clock-related signal and timing control-related signal, extracting sub-band energy distribution features related to clock jitter, edge drift, and timing margin degradation. For the temperature signal near the reference clock source, the system extracts the temperature rise gradient change rate and temperature stability index; for the timing signal, the system extracts the edge deviation standard deviation, timing drift trend, and timing margin degradation rate; for the test result variation data, the system extracts the mean read / write speed deviation, timing parameter deviation change rate, and wafer yield fluctuation amplitude.

[0038] During the health assessment phase, the LSTM model performs time series modeling of the timing drift characteristics of multiple consecutive test batches. When the system detects an abnormal increase in the reference clock source temperature gradient, a continuous expansion of the clock edge standard deviation, a timing control board calibration deviation continuously exceeding the allowable range, or abnormal fluctuations in wafer test yield, the model outputs a probability prediction value for clock synchronization failure within a preset future time period. If this probability prediction value exceeds a preset threshold, the warning output interface generates structured warning information and sends a task migration suggestion to the test task scheduling module, enabling the test task to be migrated to other ATE equipment with better health status before the fault escalates.

[0039] In the above implementation, the system can combine and analyze the signal characteristics of the timing control unit, the calibration data of the reference clock source, and the test result change data, thereby avoiding misjudging equipment failures based solely on yield fluctuations and avoiding ignoring the impact on the business layer based solely on clock parameter changes. This approach is particularly suitable for mass production testing scenarios of memory with high timing stability requirements.

[0040] Example 3: Power ATE Equipment Health Assessment and Fault Prediction A power device testing line uses Power ATE equipment to perform withstand voltage, withstand current, switching characteristics, and thermal characteristics tests on IGBT modules, power MOSFETs, diode modules, and other power devices. Typical failures of this type of equipment include overheating of high-voltage power supply modules, aging of power driver boards, abnormal power output ripple, decreased efficiency of the cooling system, and contact degradation at high-voltage circuit connection points.

[0041] In this embodiment, multiple temperature sensors are installed at the power MOSFET array, DC / DC converter output filter capacitor, power drive circuit, and heat dissipation channel of the high-voltage power supply module; voltage and current sensors are installed at the output of the high-voltage power supply to monitor output ripple, current overshoot, and transient characteristics of load discharge; vibration sensors are installed near the cooling fan or liquid cooling pump to monitor bearing wear or abnormal cooling cycle. The calibration data acquisition unit focuses on acquiring the output accuracy calibration data of the high-voltage power supply module and the calibration data of the power drive board. The test result acquisition unit focuses on acquiring the deviation of the withstand voltage test parameters, the deviation of the withstand current test parameters, the repeatability deviation of the test data, and the change in the test pass rate of the power devices. The periodic maintenance data acquisition unit focuses on acquiring the preset maintenance cycle and historical maintenance records of the high-voltage power supply module, power drive board, cooling system, and pneumatic system.

[0042] The signal processing unit uses wavelet packet transform to perform time-frequency decomposition on the current surge signal, power supply ripple signal, and heat dissipation system vibration signal, extracting the current overshoot peak amplitude, overshoot occurrence period, ripple amplitude fluctuation trend, heat dissipation system vibration frequency band energy ratio, and temperature rise gradient change rate. Based on these characteristics, the random forest model classifies fault modes such as high-voltage power module aging, heat dissipation system anomalies, and power drive board performance degradation.

[0043] When the system detects that the high-voltage power module's temperature rise rate exceeds the set range, the current waveform exhibits periodic overshoot spikes, the high-voltage power module's output accuracy calibration deviation increases, the test pass rate decreases, and the power drive board is approaching its preset maintenance cycle, the random forest model outputs the device's current status level as a warning or danger, and outputs the high-voltage power module's aging probability. The system generates structured early warning information, recommending that maintenance personnel perform power module calibration, power device array thermal status verification, and heat dissipation system checks.

[0044] For Power ATE equipment equipped with a liquid cooling system, the periodic maintenance data acquisition unit also collects data on coolant replacement cycle, coolant purity, level, and flow rate. When the coolant is overdue for replacement, its purity decreases, or its flow rate is insufficient, the system incorporates the cooling system maintenance characteristics into the health model and generates corresponding maintenance reminders. For test interfaces equipped with pneumatic systems, the system can also collect data on filter replacement cycle, air path dryness, and pipeline pressure to determine whether there are maintenance risks in the pneumatic clamping mechanism or test interface air path.

[0045] Example 4: Cold Start and Transfer Learning of Cross-Vendor ATE Equipment In another embodiment, the present invention is applied to a test facility that simultaneously deploys ATE equipment from multiple manufacturers and of various models. Because the data interfaces, number of channels, test head structures, and calibration formats of equipment from different manufacturers vary, the system first establishes a unified equipment profile based on the equipment type. This equipment profile includes the equipment type, number of channels, core board type, power module type, cooling method, sensor deployment location, supported data interfaces, historical fault modes, and maintenance cycle template.

[0046] When a new type of ATE device is first connected to the system, the system generates an initial health model based on historical training data from similar devices, and selects the corresponding fault mode template and feature extraction parameters according to the device profile of the new device. For example, digital channel timing degradation samples from existing SoC ATE devices are migrated to the newly connected SoC ATE device, and high-voltage power module aging samples from Power ATE devices are migrated to similar power testing equipment. For data with different dimensions or different sampling rates, the system generates a unified feature vector through feature normalization and feature mapping.

[0047] In the initial stage of new equipment operation, the model adopts a high safety redundancy threshold and continuously collects the equipment's own operational data, calibration data, test result change data, and maintenance records. Once the new equipment accumulates a certain number of normal and abnormal samples, the model optimization unit feeds the incremental new samples back into the training dataset to periodically update the model parameters. In this way, the present invention can shorten the model cold start cycle after the integration of new ATE equipment models, reduce reliance on a large number of historical fault samples, and improve cross-vendor deployment efficiency.

[0048] Example 5: Hierarchical Prediction Implementation of Edge Gateway and Cloud Collaboration In another embodiment, the present invention employs a collaborative deployment of edge gateways and a cloud platform. Each ATE device or group of ATE devices is configured with an edge signal processing gateway, which is deployed close to the device side to perform high-speed sensor signal acquisition, wavelet packet decomposition, adaptive denoising, and primary feature extraction. The cloud platform is used to perform cross-device data aggregation, model training, health trend analysis, and plant-wide operation and maintenance strategy generation.

[0049] In this implementation, the edge gateway performs local assessments of anomalies with high real-time requirements, such as high-voltage power supply overheating, current overshoot, abnormal test head vibration, or rapid reference clock drift. When the edge gateway determines that there is an immediate risk, it can directly output a local warning to the device control system or test task scheduling module. For faults requiring long-term trend analysis, such as gradual aging of digital channels, slow expansion of calibration deviation, or gradual decline in heat dissipation efficiency, the edge gateway uploads the normalized feature vector to the cloud platform, where a cloud-based LSTM model performs predictive analysis across batches, devices, and time windows.

[0050] By employing layered processing at the edge and cloud levels, this implementation reduces the amount of raw high-frequency signal data uploaded, lowers network bandwidth usage, and ensures rapid local response to high-risk events. For factories with high data security requirements, the system can also upload only the anonymized feature vectors and health results, without uploading the complete original test data.

[0051] Example 6: Self-evolution of a model based on maintenance closed-loop feedback In another implementation, the present invention incorporates the early warning results, actual maintenance actions, and post-maintenance status parameters into a self-evolving closed loop of the model. After each structured early warning message is generated, the system assigns a unique event number to the message and records the key characteristics that triggered the warning, the health score, the probability of the failure mode, and the recommended maintenance actions.

[0052] After maintenance personnel complete on-site inspections or repairs, the periodic maintenance data acquisition unit records the actual fault confirmation results, actual maintenance actions, information on replaced components, maintenance time, and post-maintenance equipment status parameters. The model optimization unit compares the actual fault confirmation results with the original prediction results: if the system predicts a fault and the on-site confirmation is a fault, the sample is marked as a valid early warning sample; if the system predicts a fault but no abnormality is found on-site, the sample is marked as a false alarm sample; if the equipment malfunctions but the system does not issue an early warning, the sample is marked as a missed alarm sample.

[0053] The model optimization unit periodically feeds back effective warning samples, false alarm samples, and missed alarm samples to the training dataset, and updates the tree structure weights of the random forest model or the parameters of the LSTM model. For frequently occurring false alarm feature combinations, the system reduces their warning weight; for newly emerging fault modes, the system establishes new fault labels and expands the fault mode template. In this way, the system can continuously optimize its prediction performance as equipment operation data and maintenance feedback data accumulate, improving its adaptability to specific plants, specific equipment groups, and specific process loads.

[0054] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for health assessment and fault prediction of ATE equipment based on digital signal processing, characterized in that, Includes the following steps: S1 collects multi-source sensor signals during the operation of the ATE equipment, and simultaneously collects calibration data, test result change data, and periodic maintenance data; S2, preprocess the multi-source sensor signals, calibration data, test result change data and periodic maintenance data, and use wavelet packet transform to perform time-frequency decomposition and noise reduction on the multi-source sensor signals to extract multi-dimensional fault-sensitive features; S3, normalize and fuse the multi-dimensional fault-sensitive features to form a multi-source feature vector, and input the multi-source feature vector into the health assessment and fault prediction model to output the health score and fault mode probability prediction value of the ATE equipment. S4. The health score, fault mode probability prediction value, calibration data, test result change data and periodic maintenance data are compared with the corresponding preset thresholds. When the warning conditions are met, structured warning information is generated and output.

2. The method for health assessment and fault prediction of ATE equipment based on digital signal processing according to claim 1, characterized in that: In step S1, the multi-source sensing signals include vibration signals, temperature signals, current signals, and digital channel timing signals; The calibration data includes sensor calibration data, ATE equipment resource board calibration data, timing calibration deviation data, power output accuracy calibration data, and reference clock source calibration data. The test result change data includes wafer test yield fluctuations, chip test parameter deviations, test pass rate changes, batch test result consistency deviations, and the proportion of abnormal test data; The periodic maintenance data includes the preset maintenance cycle, maintenance standards, historical maintenance time, maintenance content, and post-maintenance status parameters of components, circuit boards, cooling systems, or pneumatic systems.

3. The method for health assessment and fault prediction of ATE equipment based on digital signal processing according to claim 1, characterized in that: In step S1, the sensor deployment location and fault mode template are adaptively configured according to the ATE device type; For SoC ATE, sensors are deployed in the digital channel board area, multi-site parallel test interface, test head cage backplane connector, or DPS power module. For Memory ATE, the sensor is deployed near the timing control unit, reference clock source, or storage interface circuitry; For Power ATE, the sensor is deployed in the high-voltage power module, power drive circuit, power device array, or heat dissipation system.

4. The method for health assessment and fault prediction of ATE equipment based on digital signal processing according to claim 1, characterized in that: In step S2, the wavelet packet transform includes performing multi-level time-frequency decomposition on the vibration signal, temperature signal, current signal and digital channel time sequence signal using a preset wavelet basis function, and using an adaptive threshold denoising method to remove transient noise, mechanical shock noise and electromagnetic interference noise in the ATE test environment.

5. The method for health assessment and fault prediction of ATE equipment based on digital signal processing according to claim 1, characterized in that: In step S2, the multidimensional fault-sensitive features include the frequency band energy ratio, spectral kurtosis index, and key frequency band center frequency offset of vibration signals; the temperature rise gradient change rate, hot spot distribution uniformity, and temperature-current coupling characteristics of temperature signals; the transient impact pulse interval, duty cycle parameter, and ripple amplitude fluctuation trend of current signals; the edge deviation standard deviation, timing drift trend, and timing margin degradation rate of digital channel timing signals; the calibration deviation trend, frequency of calibration errors exceeding the allowable range, and inter-channel calibration consistency deviation of calibration data; the yield fluctuation amplitude, mean and rate of change of test parameter deviations, and the decrease in pass rate of test result change data; and the maintenance cycle deviation, mean maintenance effect deviation, maintenance frequency anomaly coefficient, and duration of untimely maintenance of periodic maintenance data.

6. The method for health assessment and fault prediction of ATE equipment based on digital signal processing according to claim 1, characterized in that: In step S3, the health assessment and fault prediction model is a random forest model or a long short-term memory network (LSTM) model. When a random forest model is used, the random forest model is trained and generated based on historical normal operation data, historical fault data, historical calibration data, historical test result data and historical maintenance data, and outputs the current status level of the device by integrating multiple decision trees; When using an LSTM model, the LSTM model is trained and generated based on time series feature data, and is used to output a health score of 0 to 100% and a predicted failure mode probability value of key components within a preset future time period.

7. The method for health assessment and fault prediction of ATE equipment based on digital signal processing according to claim 1, characterized in that: In step S4, the warning conditions include at least one of the following: The device health score is below the preset dynamic threshold; The predicted probability value of any failure mode exceeds the preset probability threshold. The calibration deviation of resource boards, sensors, or critical components continues to exceed the allowable range; The test results showed changes exceeding the preset anomaly threshold, and the model analysis determined that these changes were related to abnormal equipment status. The actual operating time of the component is close to or exceeds the preset maintenance cycle; The structured early warning information includes device identifier, early warning timestamp, current health score, probability value of each fault mode, location of fault-related components, key monitoring parameters, recommended maintenance action type and estimated remaining safe operating time, and is output in JSON or Protobuf format.

8. A system for health assessment and fault prediction of ATE equipment based on digital signal processing according to any one of claims 1-7, characterized in that, include: The multi-source data acquisition module is used to acquire multi-source sensor signals during the operation of ATE equipment, and simultaneously acquire calibration data, test result change data and periodic maintenance data; The data acquisition and conversion module, connected to the multi-source data acquisition module, is used to convert analog sensing signals into digital signals and to standardize the calibration data, test result change data, and periodic maintenance data. A digital signal processing module, connected to the data acquisition and conversion module, is used to perform time-frequency decomposition and denoising on the digital signal using wavelet packet transform, and to extract multi-dimensional fault-sensitive features. The health assessment and prediction module is connected to the digital signal processing module and is used to input the normalized and fused multi-source feature vector into the health assessment and fault prediction model, and output the health score and fault mode probability prediction value of the ATE equipment. The early warning output interface is connected to the health assessment and prediction module and is used to generate and output structured early warning information when the early warning conditions are met.

9. The ATE equipment health assessment and fault prediction system based on digital signal processing according to claim 8, characterized in that: The multi-source data acquisition module includes a vibration sensor, a temperature sensor, a current sensor, a timing signal acquisition unit, a calibration data acquisition unit, a test result acquisition unit, and a periodic maintenance data acquisition unit. The data acquisition and conversion module includes a high-speed data acquisition card, a signal conditioning circuit, and a data import unit; The digital signal processing module includes a wavelet packet transform unit, an adaptive denoising unit, a calibration data processing unit, a test result processing unit, a periodic maintenance data processing unit, and a feature extraction unit. The health assessment and prediction module includes a model training unit, a health assessment unit, a fault prediction unit, and a model optimization unit. The early warning output interface includes an early warning judgment unit, an information generation unit, and a data push unit.

10. The ATE equipment health assessment and fault prediction system based on digital signal processing according to claim 9, characterized in that: The model optimization unit is used to compare the actual fault events, on-site maintenance results, and post-maintenance status parameters with the health score and fault mode probability prediction value, to label the valid early warning samples, false alarm samples, and missed alarm samples, and to feed the labeled samples back into the training dataset to update the health assessment and fault prediction model.