Method and System for an Intelligent Fault Diagnosis Center for an Aging Device Under Test

By using a probabilistic machine learning system to automatically diagnose and transport fault devices, the problem of time-consuming and labor-consuming traditional aging tests is solved, and efficient fault detection and maintenance processes are realized, reducing resource consumption and time costs.

CN113779857BActive Publication Date: 2025-07-11DELL PROD LP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010517586.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-09
Publication Date
2025-07-11
Estimated Expiration
2040-06-09

AI Technical Summary

Technical Problem

Fault diagnosis during aging testing of traditional information processing systems is time-consuming and labor-intensive, requires a lot of manpower and resources, and technicians need a lot of training to diagnose various faults.

Method used

Probability machine learning systems, especially Naive Bayes classifiers, are used to automatically monitor and diagnose aging test failures, transport the faulty device to the repair station through automatic guide vehicles, and automatically order and replace spare parts to provide repair strategies.

Benefits of technology

It significantly reduces fault detection, diagnosis and maintenance time, reduces resource consumption, and improves the efficiency and cost-effectiveness of the aging testing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113779857B_ABST
    Figure CN113779857B_ABST
Patent Text Reader

Abstract

A mechanism is provided for automatically detecting, diagnosing, transporting, and repairing a device that fails during an aging test. Embodiments provide a system that monitors a device undergoing an aging test and detects when the device or internal components of the device fail the aging test. Then, embodiments can alert an aging test rack operator of the device failure. Embodiments can simultaneously determine the nature of the failure by applying a machine learning-based prediction model to a log file associated with the failed device. The diagnosis can be provided to a repair center along with a recommended repair strategy to help expedite the repair process. Additionally, the diagnosis can also be used to order spare parts from a spare parts warehouse for repair. In this way, embodiments can reduce the time required to detect, diagnose, and repair the failed device.
Need to check novelty before this filing date? Find Prior Art

Description

Background of the Invention Technical Field

[0002] The present invention relates to information processing systems. More specifically, embodiments of the present invention relate to automatically detecting, diagnosing, transporting, and repairing devices that fail during an aging test. Background Art

[0003] As the value and use of information continue to grow, individuals and businesses seek additional ways to process and store information. One option available to users is an information processing system. An information processing system typically processes, compiles, stores, and / or communicates information or data for commercial, personal, or other purposes, thereby allowing users to leverage the value of that information. Since technology and information processing needs and requirements vary among different users or applications, information processing systems can also vary in terms of what information is processed, how the information is processed, how much information is processed, stored, or communicated, and how quickly and efficiently the information can be processed, stored, or communicated. Variations in information processing systems allow the information processing system to be general-purpose or configured for a particular user or particular use (such as financial transaction processing, airline reservation, enterprise data storage, or global communication).

[0004] To provide flexibility in handling various information processing requirements, some information processing systems include a large number of hardware and software components (e.g., storage devices, communication devices, power supplies, processors, etc.) configured to process, store, and communicate information. To verify that these components work correctly individually and work properly with each other and to detect early failures in these components, a new assembled information processing system can be subjected to an aging test program. A typical aging test uses an expected operating cycle (lasting for a time period equivalent to several days) to provide an electrical test of the information processing system. Additionally, thermal stress and environmental stress screening can be performed. Aging tests detect failures that are typically caused by defects in the manufacturing and packaging processes. Such failures can affect one or more components in the information processing system.

[0005] When an information processing system fails during an aging test, the system is repaired and retested. Traditionally, diagnosing the cause of an aging failure and determining a solution to the failure are performed manually and can consume a large amount of human, time, and money resources. Summary of the Invention

[0006] A system, method, and computer-readable medium are disclosed for improving the diagnosis, alerting, and repair of devices that fail an aging test.

[0007] In one embodiment, a method for repairing an aging test failure of a device under test is provided. The method includes monitoring the status of one or more aging test devices, determining that a first type of device under test among the one or more aging test devices fails one or more aging tests, and diagnosing one or more reasons for the first type of device under test failing the one or more aging tests. A probabilistic machine learning system trained using a historical record set of aging test failure data of the first type of device is used to perform the diagnosis.

[0008] In one aspect of the above embodiment, the first type of device under test includes a first set of components, and the failure of the first aging test is associated with one or more of the components in the first set of components. In another aspect of the above embodiment, the probabilistic machine learning system includes a Naive Bayes classifier. In another aspect of the above embodiment, the method further includes instructing to restart the failed device under test in response to the diagnosis.

[0009] In another aspect of the above embodiment, the method further includes performing one or more of the following in response to the diagnosis: alerting an aging test center about the failed device under test; requesting one or more failed component replacements from a material handling system; and transmitting the diagnosis regarding the failed device under test to a device repair system. In another embodiment, the alerting, requesting, and transmitting are performed in parallel. In yet another embodiment, the alerting includes transmitting the identifier of the failed device under test to an aging test operator. In another embodiment, the alerting further includes transmitting an instruction to an automated guided vehicle to transport the failed device under test to a selected device repair station. In yet another aspect, the method further includes selecting the device repair station for the failed device under test. In yet another aspect, the alerting includes transmitting the identifier of the failed device under test to a mobile device associated with the aging test operator. In yet another aspect, requesting the one or more failed component replacements includes determining a recommended repair strategy in response to the diagnosis, determining recommended replacement components associated with the recommended repair strategy, and transmitting the identifiers of the recommended replacement components to the material handling system. In yet another aspect, the method further includes transmitting an instruction to an automated guided vehicle to transport the replacement components to the selected device repair station, where the failed device under test is also transported to the selected device repair station. In yet another aspect, the diagnosis includes using a log file associated with the aging test of the failed device under test as an input to the probabilistic machine learning system to identify one or more failed components of the failed device under test, and identifying one or more repair strategies for the failed device under test in response to the identification of the one or more failed components. In yet another aspect, the method further includes displaying the diagnosis at the selected repair station, where the failed device under test and the replacement components are transported to the selected repair station.

[0010] Another embodiment provides a system that includes: a processor; a data bus coupled to the processor; a network interface coupled to the data bus and a network; and a non-transitory computer-readable storage medium embodying computer program code, the non-transitory computer-readable storage medium being coupled to the data bus. The network interface is configured to communicate via the network with an aging test monitoring system, a material handling system, and a device repair system. The aging test monitoring system is coupled to one or more aging test devices. Each of the aging test devices includes one or more components. The computer program code interacts with a plurality of computer operations and includes instructions that are executable by the processor and configured to monitor the status of the one or more aging test devices, determine that a first type of device under test among the one or more devices under test fails one or more aging tests, and diagnose one or more reasons for the failure of the first type of device under test to pass the one or more aging tests, wherein a probabilistic machine learning system using a historical set of the first type of aging test failure data is trained to perform the diagnosis.

[0011] In another aspect of the above embodiment, the system further includes a machine learning accelerator processor coupled to the data bus and configured to execute instructions configured for the probabilistic machine learning system. In yet another aspect, the probabilistic machine learning system includes a Naive Bayes classifier.

[0012] In another aspect of the above embodiment, the computer program instructions further include instructions executable by the processor and configured to alert the aging test monitoring system of the failed device under test, request one or more replacement of failed components from the material handling system, and transmit the diagnosis regarding the failed device under test to the device repair system. In another aspect, the instructions are further configured to determine a recommended repair strategy in response to the diagnosis, determine recommended replacement components associated with the recommended repair strategy, and transmit identifiers of the recommended replacement components to the material handling system. In yet another aspect, the instructions are further configured to use a log file associated with the aging test of the failed device under test as an input to the probabilistic machine learning system to identify one or more failed components of the failed device under test, and identify one or more repair strategies for the failed device under test in response to the identification of the one or more failed components. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] By referring to the accompanying drawings, those skilled in the art can better understand the present invention, and many objects, features, and advantages of the present invention become apparent. The same reference numerals are used throughout the several drawings to refer to the same or similar elements.

[0014] Figure 1 A general illustration of the components of an information processing system implemented in the systems and methods of the present invention is shown.

[0015] Figure 2 Is a simplified block diagram showing an intelligent diagnostic system for burn-in testing according to an embodiment of the present invention.

[0016] Figure 3 Is a simplified flowchart showing a series of steps performed by an intelligent diagnostic system according to an embodiment of the present invention. Detailed Description

[0017] Disclosed is a system, method, and computer-readable medium for automatically detecting, diagnosing, transporting, and repairing a device that fails during burn-in testing. Embodiments provide a system for monitoring a device undergoing burn-in testing and detecting when the device or a component in the device fails the burn-in test. The embodiments can then alert the burn-in-rack monitor of the device failure while sending an automated guided vehicle (AGV) to the location of the failed device for transportation to a repair center. The embodiments can simultaneously determine the nature of the failure by applying a machine learning-based prediction model to the log file associated with the failed device. The diagnosis can be provided to the repair center along with a recommended repair strategy to help expedite the repair process. Additionally, the diagnosis can be used to order replacement parts from a parts depot for the repair while sending the AGV to the parts depot to transport the replacement parts to the repair center. In this way, the embodiments can reduce the time required to detect, diagnose, and repair a failed device.

[0018] For the purposes of the present disclosure, an information processing system can include any tool or collection of tools operable to compute, classify, process, transmit, receive, retrieve, originate, switch, store, display, manifest, detect, record, reproduce, process, or utilize any form of information, intelligence, or data for commercial, scientific, control, or other purposes. For example, an information processing system can be a personal computer, a network storage device, or any other suitable device, and can vary in size, shape, performance, functionality, and price. An information processing system can include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic, ROM, and / or other types of non-volatile memory. Additional components of an information processing system can include one or more disk drives, one or more network ports for communicating with external devices, and various input and output (I / O) devices such as a keyboard, a mouse, and a video display. An information processing system can also include one or more buses operable to transfer communications between the various hardware components. An information processing system can embody both embodiments of the present invention and the device under test managed by such embodiments.

[0019] Figure 1 FIG. 4 is a general illustration of an information processing system 100 that can be used to implement the systems and methods of the present invention. Information processing system 100 includes a processor (e.g., a central processing unit or “CPU”) 102, input / output (I / O) devices 104 (such as a display, a keyboard, a mouse, and associated controllers), a hard disk drive or disk storage device 106, and various other subsystems 108. In various embodiments, information processing system 100 also includes a network port 110 operable to connect to a network 140, which can also be accessed by an aging test monitoring system 142, a spare parts storage system 144, and a device repair system 146. In various embodiments, information processing system 100 also includes a wireless communication port 128 operable to communicate with remote devices via one or more wireless networking protocols, the remote devices including, for example, an automated guided vehicle 150. Information processing system 100 also includes a system memory 112 interconnected with the foregoing devices via one or more buses 114. System memory 112 also includes an operating system (OS) 116, and in various embodiments can also include an intelligent diagnostic system module 118.

[0020] The intelligent diagnostic system module 118 performs operations related to monitoring the device under test in the aging test monitoring system 142, diagnosing devices that fail the aging test, ordering replacement spare parts from the spare parts storage system 144 as indicated by the diagnosis, recommending a repair process to the device repair system 146, and in some embodiments, controlling the transportation of the failed device and the replacement spare parts to a device repair location associated with the device repair system 146. These operations will be discussed more fully below. The diagnostic operations are performed using the machine learning fault analysis module 120 in combination with the training module 122. The training module 122 uses a device under test (DUT) test fault training data set 124 stored, for example, in the hard disk drive / disk 106 to train the machine learning fault analysis module. In some embodiments, the machine learning accelerator processor 126 is coupled to the CPU 102 and the memory 112 via the bus 114. The machine learning accelerator is configured to execute instructions in the machine learning fault analysis module 120 more efficiently than the processor associated with the CPU 102, thus improving the performance of the diagnostic function of the intelligent diagnostic system 118, which is described more fully below. In other embodiments, the machine learning instructions are executed by the CPU 102 without using the machine learning accelerator. Using machine learning to automate device fault diagnosis and automatically order replacement spare parts in response to the diagnosis improves the overall efficiency of the aging fault and recovery cycle.

[0021] It should be understood that once the information processing system 100 is configured to perform the intelligent diagnostic operations described above, the information processing system 100 becomes a dedicated computing device specifically configured to perform the intelligent diagnostic operations, rather than a general-purpose computing device. Additionally, implementing the intelligent diagnostic operations on the information processing system 100 provides useful and specific results in terms of improving efficiency and reducing costs associated with aging faults and recovery cycles.

[0022] The aging test of electronic products, such as information processing systems, is a process of detecting early failures of product components, thereby improving the reliability of the components sold. The "early failure period" is the period during which early failures occur in the components and may be due to problems in the manufacturing process. During this early life cycle, the components may fail at a high rate, but the rate decreases over time. In some aging examples, the systems and system components are operated under extreme conditions (e.g., elevated temperature and voltage) or for long periods of time. This places stress on the device under test and eliminates the weak points in the product before customer delivery.

[0023] During traditional burn-in testing of information processing systems, a burn-in test monitoring system (e.g., a burn-in rack monitor) can monitor several information processing systems being tested simultaneously. The burn-in test monitoring system records log files, configuration files, and other data records for each information processing system. If one or more components of an information processing system fail during burn-in testing, the burn-in test monitoring system alerts the burn-in tester of the failure. Once the burn-in test staff sees the failure notification, they can go to the failed system, remove the failed system from the burn-in rack, and send the failed system to a device repair facility.

[0024] Once the failed device arrives at the device repair station, traditionally, the personnel at the device repair facility begin a manual process of diagnosing the cause of the failure. The logs and configuration files from the burn-in monitor are provided to the technicians at the device repair facility, who then use this information and their experience to diagnose the cause of the failure and determine how to fix the failure. Once the repair technician diagnoses the failure, the technician can request one or more replacement spare parts from a materials handling station to repair the failed device. When the replacement spare parts arrive, the technician can repair the device and then can return the device to burn-in testing for further testing, or can complete the testing at the repair facility.

[0025] The traditional processes of burn-in testing, transporting, diagnosing, ordering replacement spare parts, and repairing devices can take a significant amount of time. On average, for information processing system burn-in, the time from a burn-in failure to completion of system repair can exceed 120 minutes. In a facility where approximately 250,000 units fail burn-in testing each year, this amounts to 500,000 hours of work per year to detect and repair burn-in failures, consuming a significant amount of time, money, and manpower. Additionally, the requirement that repair technicians be able to diagnose all types of failures results in significant training of such personnel before they are fully qualified to work at the device repair station.

[0026] Embodiments of the present invention seek to reduce the resource consumption required to detect and repair burn-in failures by automatically detecting failed devices, automatically diagnosing the cause of the failure, and automatically ordering replacement spare parts to assist in repairing the failed device. As will be described more fully below, embodiments use a machine learning diagnostic system to determine the cause of the failure and how to solve the problem.

[0027] Figure 2is a simplified block diagram showing an intelligent diagnostic system 200 for aging testing according to an embodiment of the present invention. As described above, the aging test monitoring system 210 is used to manage the aging tests of a group of devices under test 220(1) to (N). Each device under test can be an information processing system, which includes many components, such as, for example, a processor, a storage device, a memory, a graphics card, a network communication card, etc. Alternatively, the device under test can be a specific-purpose grouping of specialized components (e.g., a network-attached storage system, edge computing resources, a media server, etc.). The aging test monitoring system 210 stores information about each device under test 220 in the aging record database 225. Such information can include log files, configuration files, and other information necessary for tasks and analysis results associated with each test performed on the device under test.

[0028] The aging test monitoring system 210 is configured to monitor the progress of each device under test 220 through a series of aging tests suitable for each type of device under test. The aging test monitoring system 210 stores the information associated with the aging test in the log files stored in the aging record database 225. When a device under test fails the aging test, the aging test monitoring system will be informed of such a failure, or the failure can be determined by other means.

[0029] Upon determining a failure of a device under test, the aging test monitoring system 210 can inform the intelligent diagnostic server 230 of the failure of the device under test via a communication link through the network 240. Once informed of a failure of the device under test 220, the intelligent diagnostic server 230 requests the log file and other information related to the aging test failure of the device under test, and can store the information in the database storage device 235. Additionally, the intelligent diagnostic server 230 can directly inform the personnel of the aging test rack of the existence and identification of the failed device under test. In one example, the intelligent diagnostic server 230 can wirelessly communicate the device identification information to the Andon watch 227 or other mobile messaging devices (e.g., a pager, a tablet, a phone), and notify the aging test personnel. This results in a reduction in the residence time of the failed device on the aging test rack before it is sent to the repair station.

[0030] The intelligent diagnostic server 230 uses probabilistic machine learning methods to perform a diagnosis on the cause of the aging test failure of the device under test using the information received from the aging test monitoring system. In some embodiments, historical fault logs, configuration file records, and historical repair data are used to train a Naive Bayes machine learning system to build a repair prediction model that can accurately determine the cause of new instances of aging failures among the devices under test. A set of tracked independent predictor variables associated with the components of the device under test (e.g., inputs, outputs, system models, fault information, system configuration (CPU model, DIMM size and type, hard drive size and type, PCI cards, etc.), system firmware version, repair codes, and other types of components, etc.), environmental factors, test time length, etc. are used to train the machine learning system. When information associated with the failure of the device under test is received, the machine learning system analyzes the information associated with these variables to determine the most likely cause of the failure. In some embodiments, the top three most likely causes of failure are used to generate a set of recommended repair actions as a guide for the subsequent steps to repair the failure of the device under test.

[0031] After performing a diagnosis on the cause of the aging failure, the intelligent diagnostic server 230 can provide the diagnostic information to the device repair system 250 via the network 240 for use by device repair technicians when they receive a failed device under test. Additionally, the intelligent diagnostic server 230 can communicate wirelessly with an automated guided vehicle (AGV) 260 to report to the aging test rack associated with the failed device under test to pick up the failed device and transport the failed device to a designated repair station 255. Further, the intelligent diagnostic server 230 can order spare parts indicated for repairing the failed device in response to the fault diagnosis. The spare parts can be ordered from the material handling system 270. Personnel associated with the material handling can select the indicated spare parts and provide these spare parts to another AGV 265 to transport these spare parts to the designated repair station 255.

[0032] Once the failed device and the indicated parts arrive at the repair station 255, the technician can use the diagnosis provided by the intelligent diagnostic server 232 to perform a repair on the failed device. In some embodiments, when multiple potential diagnoses are provided, the technician may need to determine which diagnoses are appropriate for a particular failure instance. However, by suggesting diagnoses and performing repairs considering these diagnoses with the spare parts, the repair speed of the device can be significantly accelerated. The parallel actions of detection, diagnosis, spare part request, and transportation of the failed device and the indicated spare parts can save a significant amount of time. In some known examples, the time from detecting a fault to repairing the device has been reduced to 45 minutes, or approximately 38% of a traditional system performing such repairs. This can significantly reduce the annual hours required to resolve aging failures and the total cycle time for manufacturing information processing systems that have undergone aging tests.

[0033] Figure 3 FIG. 300 is a simplified flow chart showing a series of steps performed by the intelligent diagnostic system 200 according to an embodiment of the present invention. During the aging test, the device under test (e.g., 220(1) to (N)) is monitored (305). As discussed above, such monitoring can be performed by the aging test monitoring system 210, or information can be provided directly from the aging test rack to the intelligent diagnostic server 230, which can directly perform device monitoring. The aging test monitoring system 210 or the intelligent diagnostic server 230 detects DUT test failures (310). Once the intelligent diagnostic server 230 detects a DUT test failure or is informed of the test failure, the intelligent diagnostic server uses machine learning tools (such as Naive Bayes) to diagnose the test failure as discussed above (315). If it is determined after diagnosis that a restart can resolve the test failure (320), the intelligent diagnostic server can instruct the aging test monitoring system to restart the failed device under test (325) and can restart the aging test (305).

[0034] If restarting the device under test does not resolve the failure, the intelligent diagnostic or 230 can perform several tasks in parallel. The aging test personnel can be alerted to the failed DUT (330). Such an alert can include information about the location of the failed DUT and, in some cases, the nature of the failure. Once the alert is received, the aging test personnel can remove the failed device from the aging test rack and place the failed device on the AGV for transportation to the repair station (335). The intelligent diagnostic server 230 can instruct the AGV to transport the failed DUT to the selected repair station (340).

[0035] While the process related to alerting the aging test personnel about the failed device is ongoing, the intelligent diagnostic server 230 can also request replacement spare parts for the failed device from the material handling (350). Such a request can be performed through inter-server communication between the intelligent diagnostic server 230 and the material handling system 270 and can take the form of any protocol utilized by the two systems. Then, the material handling personnel can locate the replacement spare parts and place the spare parts on another AGV for transportation to the selected repair station (355). The intelligent diagnostic server 230 can instruct the AGV to transport the spare parts to the selected device repair station (360).

[0036] During an additional parallel process, the intelligent diagnostic server can transmit diagnostics related to a failed device to a device repair system (e.g., 250)(370) associated with a selected device repair station. Additionally, the intelligent diagnostic server can provide a recommended repair strategy associated with the diagnostics. Alternatively, the device repair system can provide a recommended repair strategy in response to the received diagnostics, depending on the nature of the coupling system and the distribution of the associated database. The device repair technician can then use replacement spare parts, the diagnostics provided by the intelligent diagnostic server, and the recommended repair strategy provided by the intelligent diagnostic server to repair the failed device (375). Once repaired, the repaired device can return to the burn-in test, or further burn-in tests can be performed at the device repair station (380).

[0037] Embodiments of the present invention provide a mechanism for improving the efficiency of the burn-in test process of an information processing system. This is accomplished in part by using a machine learning program to automate the diagnosis of the causes of device failures during burn-in testing to associate information about the failures with previously known failures. Additionally, efficiency is achieved by alerting burn-in testers to failures and shipping requirements, requesting replacement spare parts from materials handling, and simultaneously providing diagnostics and recommended repair strategies to device repair technicians. Subsequently, the time to manufacture the information processing system is reduced, and the resource costs associated with the delays inherent in traditional methods of handling device burn-in failures are decreased. Additionally, by automating the diagnosis of device failures, the level of experience required by technicians to perform repairs on failed devices is reduced.

[0038] As will be appreciated by one skilled in the art, the present invention may be embodied as a method, system, or computer program product. These various embodiments may generally be collectively referred to herein as "circuits", "modules", or "systems". Additionally, the present invention may take the form of a computer program product on a computer-usable storage medium having computer-usable program code embodied in the medium.

[0039] Any suitable computer-usable or computer-readable medium can be utilized. A computer-usable or computer-readable medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable medium would include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, or a magnetic storage device. In the context of this document, a computer-usable or computer-readable storage medium can be any medium that can contain, store, communicate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0040] Computer program code for performing the operations of the present invention may be written in an object - oriented programming language such as Java, Smalltalk, C++. However, computer program code for performing the operations of the present invention may also be written in a conventional program programming language such as the "C" programming language or a similar programming language. The program code may execute entirely on the user's computer, partly on the user's computer, execute as a stand - alone software package, partly on the user's computer and partly on a remote computer, or execute entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0041] Embodiments of the present invention are described with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations or block diagrams, and combinations of blocks in the flowchart illustrations or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions / operations specified in one or more blocks of the flowchart and / or block diagram.

[0042] These computer program instructions may also be stored in a computer - readable memory, which can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer - readable memory produce an article of manufacture, the article of manufacture including instruction means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0043] The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable device, thereby producing a computer - implemented process, such that the instructions executed on the computer or other programmable device provide steps for implementing the functions / operations specified in one or more blocks of the flowchart and / or block diagram.

[0044] The present invention is highly suitable for obtaining the advantages mentioned and other advantages inherent therein. Although the present invention has been depicted, described, and defined by reference to specific embodiments of the present invention, such reference does not imply a limitation of the present invention, and no such limitation should be inferred. The present invention is capable of considerable modification, alteration, and equivalents in form and function, as will occur to those skilled in the relevant art. The embodiments depicted and described are merely exemplary and not exhaustive of the scope of the present invention.

[0045] Accordingly, the present invention is intended to be limited only by the spirit and scope of the appended claims, with full recognition of equivalents in all respects.

Claims

1. A computer - implementable method for repairing aging - test failures of a device under test, the method comprising: Monitoring the status of one or more devices being aging - tested; Determining that a first type of device under test (DUT) among the one or more devices being aging - tested fails one or more aging tests; And Diagnosing one or more reasons why the first - type DUT fails the one or more aging tests, wherein the diagnosis is performed using a probabilistic machine - learning system trained with a historical record set of aging - test failure data of the first - type devices, and the historical record set of aging - test failure data includes historical fault logs, configuration file records, and historical repair data; Transmitting instructions to an automated guided vehicle (AGV) to report to an aging - test rack associated with the faulty DUT and transporting the faulty DUT to a selected device repair station; And Transporting the faulty DUT to the selected device repair station by the AGV.

2. The method according to claim 1, wherein The first - type DUT includes a first set of components; and The failure of the first - type DUT to pass the one or more aging tests is associated with one or more components in the first set of components.

3. The method according to claim 1, wherein the probabilistic machine - learning system includes a Naive Bayes classifier.

4. The method according to claim 1, further comprising: Indicating to restart the faulty DUT in response to the diagnosis.

5. The method according to claim 1, further comprising performing one or more of the following in response to the diagnosis: Reminding an aging - test center of the faulty DUT; Requesting one or more faulty - component replacements from a material - handling system; and Transmitting the diagnosis of the faulty DUT to a device - repair system.

6. The method according to claim 5, wherein the reminder, request, and transmission are performed in parallel.

7. The method according to claim 5, wherein the reminder includes: Transmitting the identification of the faulty DUT to an aging - test operator.

8. The method according to claim 1, further comprising: Selecting the device repair station for the faulty DUT.

9. The method according to claim 7, wherein the reminder further includes: Transmitting the identification of the faulty DUT to a mobile device associated with the aging - test operator.

10. The method according to claim 7, wherein the requesting the one or more faulty - component replacements includes: Determining a recommended repair strategy in response to the diagnosis; Determining recommended replacement components associated with the recommended repair strategy; And Transmitting the identifiers of the recommended replacement components to a material - handling system.

11. The method according to claim 10, further comprising: Transmitting instructions to a second automated guided vehicle (AGV) to transport the replacement components to the selected device repair station; And Transporting the replacement components to the selected device repair station by the second AGV.

12. The method according to claim 5, wherein the diagnosis includes: Use a log file associated with the aging test of the faulty DUT as an input to the probabilistic machine learning system to identify one or more faulty components of the faulty DUT; and Identify one or more repair strategies for the faulty DUT in response to the identification of the one or more faulty components.

13. The method according to claim 12, further comprising: Display the diagnosis at the selected device repair station, where the faulty DUT and replacement components are transported to the selected device repair station.

14. A system, comprising: A processor; A data bus coupled to the processor; A network interface coupled to the data bus and a network and configured to communicate with an aging test monitoring system, a material handling system, and a device repair system via the network, where The aging test monitoring system is coupled to one or more devices being aging tested, and Each of the devices being aging tested includes one or more components; and A non-transitory computer-readable storage medium embodying computer program code, the non-transitory computer-readable storage medium coupled to the data bus, the computer program code interacting with a plurality of computer operations and including instructions that can be executed by the processor and are configured to Monitor the status of the one or more devices being aging tested, Determine that a first type of device under test (DUT) among one or more devices under test fails one or more aging tests, Diagnose one or more reasons for the first type of DUT failing the one or more aging tests, where the diagnosis is performed using a probabilistic machine learning system trained with a historical record set of aging test failure data for the first type of device, and the historical record set of aging test failure data includes historical fault logs, profile records, and historical repair data; Transmit instructions to an automated guided vehicle (AGV) to report to an aging test rack associated with the faulty DUT and transport the faulty DUT to a selected device repair station; and Transport the faulty DUT to the selected device repair station via the AGV.

15. The system according to claim 14, further comprising: A machine learning accelerator processor coupled to the data bus and configured to execute instructions configured for the probabilistic machine learning system.

16. The system according to claim 15, where the probabilistic machine learning system includes a Naive Bayes classifier.

17. The system according to claim 14, where the computer program code includes additional instructions that can be executed by the processor, and the additional instructions are further configured to: Remind the aging test monitoring system of the faulty DUT using the network interface; Request one or more faulty component replacements from the material handling system using the network interface; and Transmit the diagnosis regarding the faulty DUT to the device repair system using the network interface.

18. The system of claim 17, wherein the instructions configured to request one or more failed component replacements include additional instructions executable by the processor, the additional instructions being configured to determine a recommended repair strategy in response to the diagnosis; determine recommended replacement components associated with the recommended repair strategy; and transmit an identifier of the recommended replacement components to the material handling system using the network interface.

19. The system of claim 17, wherein the instructions configured to request one or more failed component replacements include additional instructions executable by a processor, the additional instructions being configured to use a log file associated with the burn-in test of the failed DUT as an input to the probabilistic machine learning system to identify one or more failed components of the failed DUT; and identify one or more repair strategies for the failed DUT in response to the identifying the one or more failed components.

Citation Information

Patent Citations

  • System and method of automated burn-in testing on integrated circuit devices

    CN111051902A