Test method and system of semiconductor memory test production line

Through dynamic load balancing and LSTM neural network optimization test parameters, combined with data binding and permission management, the problems of rigid resource configuration and data isolation in the semiconductor memory test production line are solved, and efficient and secure parallel multi-ticket testing and data transmission are achieved.

CN120336186APending Publication Date: 2025-07-18SHENZHEN JIAHE JINWEI ELECTRONICS TECH

Patent Information

Application Number
CN202510508843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing semiconductor memory test production lines have problems such as rigid resource configuration, isolated data management, extensive permission control and poor scalability in parallel testing, resulting in low testing efficiency, inaccurate data transmission and insufficient security.

Method used

The dynamic load balancing algorithm is used to divide the logical groups and allocate hardware resources. The fault prediction model is built based on the LSTM neural network, the test item priority and parameters are adjusted in real time, and the test results and SPD serial numbers are uploaded through asynchronous message queues, and data binding and closed-loop management are realized by combining the permission management module and the adaptive interface device.

Benefits of technology

It significantly improves the efficiency and accuracy of parallel testing of multiple tickets, enhances system security and fault prediction capabilities, realizes dynamic adjustment of hardware resources and reliable data transmission, and reduces the cost of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336186A_ABST
    Figure CN120336186A_ABST
Patent Text Reader

Abstract

The method comprises the steps of dividing a production line logic group through a dynamic load balancing algorithm, allocating hardware resources, and constructing a fault prediction model based on an LSTM neural network to optimize test parameters. Binding the test result, the SPD serial number and the work station ID number, and uploading the test result, the SPD serial number and the work station ID number to an MES system through an asynchronous message queue; the system integrates a modularized test configuration unit, a multi-dimensional data analysis platform, an authority management module and a self-adaptive interface device. According to the invention, high automation and intelligentization of the test process are realized, the test efficiency, the data accuracy and the system safety are obviously improved, the manual intervention cost is effectively reduced, the fault prediction capability is improved, and a comprehensively optimized technical solution is provided for a semiconductor memory test production line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semiconductor memory testing, and more particularly to a testing method and system for a semiconductor memory testing production line. Background Art

[0002] In the field of semiconductor memory testing, with the continuous improvement of integrated circuit manufacturing processes and the increasing complexity of chips, efficient and intelligent testing systems have become a crucial link in ensuring product quality and production efficiency. Modern testing technologies have evolved from single-functional devices to highly integrated platforms, playing an irreplaceable role in ensuring product performance consistency and reducing the defective product rate. Especially in the context of intelligent manufacturing, testing systems not only need to meet the basic requirements of rapid detection but also possess powerful data processing capabilities and broad applicability to cope with complex production process environments and diverse product requirements.

[0003] Currently, in addressing the issue of parallel testing of multiple work orders in a semiconductor production line, the industry generally adopts two mainstream approaches: one is to pre-set a fixed resource allocation plan applicable to various testing tasks; the other is to introduce a manual intervention mechanism, relying on technicians to manually adjust testing parameters and scheduling strategies. The former simplifies the initial deployment process of the testing system by meticulously planning the resource allocation plan in advance, but it appears inflexible in the face of dynamically changing workloads and is prone to phenomena such as idle or bottlenecked testing equipment resources. The latter emphasizes the value of human experience and can, to a certain extent, compensate for the deficiencies of static planning, but it significantly increases human resource costs and is difficult to ensure real-time performance and accuracy. Therefore, both of the aforementioned methods have certain limitations, especially in terms of the ability to achieve efficient parallel testing of multiple work orders in a volatile memory testing production line. The pre-set fixed resource allocation method is difficult to respond promptly to load fluctuations, often resulting in resource waste or underutilization; while excessive reliance on manual operations can, to some extent, alleviate the drawbacks of rigid planning, but it is limited by individual skill levels and has low efficiency. Therefore, how to achieve coordinated operation among multiple work orders in the physical environment of an existing testing production line while maintaining good dynamic adjustment capabilities has become a technical challenge that urgently needs to be overcome.

[0004] The published application number of the invention patent, CN108932196A, discloses a parallel automated testing method, system, device, and readable storage medium. Multiple testing units are configured, and each testing unit is provided with multiple test cases. One testing host is corresponded to one testing unit. The test cases in the testing unit are run to test the testing host. When the current test case is executed successfully, the next test case is executed until all the test cases in the testing unit have completed the testing of the testing host. A test report is output. The test cases are sorted out, and the sorted test cases are divided by module. After division, they are assigned to different machines to execute the test cases and generate the final test results. In the related prior art, the method can improve the execution efficiency of test cases and reduce resource waste. By dividing the test cases by module and assigning them to different machines for execution, it is ensured that the test cases run in parallel and are not executed repeatedly. The test cases have limitations on the execution time and number of times. If the time is exceeded, they will be forced to terminate to prevent the situation where the entire host resources are occupied due to the test cases getting stuck. However, in the environment of the same semiconductor memory testing production line, the parallel operation of the test cases will only be forced to terminate after the time extension consequence of resource occupation has occurred. It does not have the dynamic adjustment ability of coordinated operation between work cells, cannot intelligently optimize the established test process, the test cases can only be used and stopped, and there is a need for further improvement in the accuracy, traceability, and security of test data transmission under multiple work orders.

[0005] The published application number of the invention patent, CN1720586A, discloses an intelligent inspection of a multi-state memory, which programs the multi-state memory by using dynamic adjustment based on inspection results for a multi-state inspection range implemented based on sequential state inspection. This technology can improve the multi-state writing speed while maintaining reliable operation within the implementation of the sequential inspection of the multi-state memory. It achieves the above purpose by providing an "intelligent" method to minimize the number of sequential inspection operations in each programming / inspection / locking step of the writing sequence. In an exemplary embodiment of the writing sequence of the multi-state memory during the programming / inspection cycle sequence of the selected storage element, at the beginning of the process, in the detection stage, only the lowest state of the multi-state range into which the selected storage element is being programmed is detected. In the related prior art, it focuses on the internal state transition and its inspection mechanism of non-volatile memories (such as flash memory), and is committed to optimizing the writing performance of a single storage cell. There is no technical inspiration for how to comprehensively improve the test efficiency of volatile memories (such as DDR DRAM memory). The technical difficulty in testing volatile memories mainly lies in how to maintain the stability and consistency of data under high-speed operation. Since volatile memories need to continuously refresh data to maintain the stored information, how to optimize the refresh mechanism, reduce the refresh delay and power consumption becomes a key issue. Summary of the Invention

[0006] The main objective of the present invention is to provide a testing method and system for a semiconductor memory testing production line. The main progress lies in achieving coordinated operation among multiple work orders within the same semiconductor volatile memory testing production line, maintaining good dynamic adjustment capabilities, optimizing the refresh mechanism of the tested volatile memory, reducing the refresh latency of the tested volatile memory, reducing the power consumption of the tested volatile memory, and enhancing the system security of the tested volatile memory.

[0007] The first main objective of the present invention is achieved through the following technical solutions: A testing method for a semiconductor memory testing production line is proposed, including the following steps: S1. Divide a volatile memory testing production line into multiple logical groups, and allocate hardware resources through a dynamic load balancing algorithm to support parallel testing of multiple work orders; S3. Build a fault prediction model based on the LSTM neural network, and adjust the test item priorities and parameters in real time, including increasing the loop times of high-frequency test items and optimizing the BIOS parameter template of the hardware platform; S4. Automatically bind the test results, SPD serial numbers, and workstation ID numbers, and upload them to the MES system through an asynchronous message queue to achieve closed-loop management through a preset API interface.

[0008] By adopting the technical solution of the above method, the parallel testing of multiple work orders in a single production line in step S1 significantly improves the testing efficiency of volatile memories, and the effective allocation of hardware resources is carried out with a dynamic load balancing algorithm, which can reduce resource conflicts and waste of equipment investment in memory testing; in step S3, the priority and parameters of test items are adjusted in real time with a fault prediction model based on the LSTM neural network. The LSTM neural network can predict the DDR refresh failure rate and signal integrity faults such as bit flips and timing deviations. The real-time adjustment specifically includes increasing the number of cycles of high-frequency test items to more comprehensively cover possible fault situations, and optimizing the BIOS parameter template of the hardware platform to improve the testing accuracy and efficiency. The BIOS parameter optimization includes adjusting the CPU frequency and voltage of DDR testing to match the DRAM timing (such as the CL value and tRCD); the automatic binding in step S4 includes binding the test results with the SPD serial number (storing memory timing parameters) and the station ID number and uploading them to the MES system through an asynchronous message queue. Therefore, the data binding in the automatic binding includes binding the SPD serial number for tracing the DRAM performance, and completing the closed-loop management with the help of a preset API interface to ensure the reliability of data transmission during the testing process of volatile memories. Therefore, under high-frequency refresh testing, the dynamic load balancing optimizes the CPU and memory bandwidth allocation to support parallel refresh operations, realizes the seamless connection between the production line testing system and the volatile memory testing production line, reduces the testing cost of volatile memories and improves the overall management level of the semiconductor memory testing production line.

[0009] In a preferred method example of the present invention, it can be further configured that: in step S1, the dynamic load balancing algorithm considers the importance differences of each work order, and the hardware resource allocation result is obtained through comprehensive consideration by adding a weight coefficient.

[0010] By adopting the above preferred technical features, when allocating hardware resources through the dynamic load balancing algorithm, the importance differences of each work order are considered, and the specific manifestations are as follows: 1. The algorithm adds a weight coefficient to comprehensively consider the demand characteristics of each work order to ensure that key tasks obtain higher-priority resource allocation; 2. Avoid waste and conflicts of test resources, and significantly improve the testing efficiency and equipment utilization rate; 3. Optimize the flexibility and response speed of the overall testing process in the multi-work order parallel testing scenario.

[0011] In a preferred method example of the present invention, it can be further configured that: step S3 of constructing a fault prediction model based on the LSTM neural network further includes: multi-dimensional data analysis and data classification and sorting. According to the historical test data and current production status of different production lines, machine learning algorithms are used to predict the growth trend and abnormal patterns of data, which are used as the analysis basis for subsequent fault diagnosis and production optimization of the fault prediction model.

[0012] By adopting the above-mentioned preferred technical features, the effects that can be specifically achieved by constructing a fault prediction model based on the LSTM neural network include: 1. Perform multi-dimensional data analysis and classify and organize the data. According to the historical test data and current production status of different production lines, use machine learning algorithms to predict the growth trend and abnormal patterns of the data, providing a reliable analysis basis for subsequent fault diagnosis and production optimization; among them, multi-dimensional data analysis includes analyzing the DDR timing error distribution, such as DIMM slot hotspots; 2. Achieve a comprehensive mining of historical test data and current test status, ensuring that the fault prediction is more scientific and reasonable; 3. Use machine learning algorithms to predict the growth trend and abnormal patterns, effectively improving the early warning ability for unknown faults; 4. The analysis results directly serve the fault prediction model, support subsequent fault diagnosis and production optimization decisions, and further enhance the intelligent level of the test system.

[0013] In a preferred method example of the present invention, it can be further configured as: the multi-dimensional data analysis method in step S3 includes the following steps: S31. Receive test data from different production lines, classify and organize them according to time range, production line number, and hardware platform (specifically such as motherboard model, DDR generation), and generate an interactive chart. Step S31 includes: S31a. Draw a trend chart of the timing error rate to show the variation law of DRAM test items over time; S31b. Generate a heat map of the DIMM slot fault distribution to identify the abnormal hot spots of signal integrity; S31c. Provide a function to download structured files, including the original test data and SPD metadata (such as CL value, tRCD parameter); S32. Based on a spatio-temporal association algorithm (such as an association rule based on time stamp and physical location), screen the test data in multiple dimensions of time range, production line number, and hardware platform, and call the GPU acceleration rendering engine (such as Unity3D or WebGL) through the CUDA framework to make the interactive chart dynamically interactive.

[0014] By adopting the above-mentioned preferred technical features, multi-dimensional classification and organization of test data are realized, and an interactive chart is generated to facilitate users to intuitively understand the test situation; the specific effects are as follows: 1. Classify and organize the test data according to multiple dimensions such as time range and production line number, significantly improving the flexibility and intuitiveness of data processing, enabling users to more conveniently query and analyze the data, thereby accelerating the fault diagnosis process and reducing production costs; 2. Drawing a trend chart helps to observe the data change pattern, and a distribution heat map can visually present the fault concentration area. The provided structured file download function facilitates users to conduct in-depth analysis and archiving of the selected data; 3. Screening test data based on a spatio-temporal correlation algorithm and using a GPU-accelerated rendering engine to make the interactive chart dynamically interactive further improves the data analysis efficiency and enhances the user experience; 4. Support for filtering test data in multiple dimensions such as time range, production line number, and hardware platform enables users to more precisely locate the problem and improve the fault troubleshooting efficiency.

[0015] In a preferred method example of the present invention, it can be further configured that: step S31 of generating the interactive chart includes: S311. Adding an interaction function to the interactive chart to support clicking on a single data point (such as a DIMM slot number) to view detailed test parameters (number of timing errors, voltage fluctuation curve) and correlation analysis results (such as LSTM fault prediction confidence); S312. Adopting a data block loading technique to only render the test data within the current window range when dragging the time axis, ensuring that the dynamic update response time is less than 200 ms.

[0016] By adopting the above preferred technical features, adding an interaction function to the interactive chart enables users to view detailed test data and analysis results by clicking on a single data point, thus greatly enriching the user's information acquisition channels and facilitating quick positioning of the key data of concern; at the same time, supporting dragging the time axis to real-time update the data range displayed by the interactive chart enables users to flexibly adjust the observation window according to needs, significantly improving the convenience and intuitiveness of data query and analysis, and further enhancing the user experience.

[0017] In a preferred method example of the present invention, it can be further configured that: in step S3 of constructing a fault prediction model based on an LSTM neural network, the operation of optimizing the BIOS parameter template of the hardware platform includes: according to the DRAM hardware characteristics (such as DDR4 / DDR5 protocol) and test requirements, dynamically adjusting the following parameters by calling the BIOS programming interface through a UEFI Shell script: S3a. CPU frequency and memory controller clock synchronization parameters (such as the ratio of MEMCLK to UCLK); S3b. DRAM voltage (VDD / VPP) and timing configuration (tCL, tRCD, tRP); S3c. Cache prefetch policy (such as turning off redundant prefetch to reduce interference).

[0018] By adopting the above preferred technical features, based on the actual hardware characteristics and test requirements, the key parameters such as CPU frequency, voltage, and cache settings are adjusted using the BIOS programming interface to optimize the BIOS parameter template for a specific hardware platform. This can dynamically adapt to the requirements of different hardware configurations, ensure the best utilization state of hardware resources during the test process, and thus significantly improve the test accuracy and efficiency. The specific effects include but are not limited to: accurately adjusting hardware performance indicators to match the test standards, reducing unnecessary energy consumption and heat generation; enhancing system stability and avoiding potential failures caused by improper fixed parameter settings; promoting the standardization and automation of the test process, and reducing the uncertain factors brought by human intervention.

[0019] In a preferred method example of the present invention, it can be further configured as follows: The step S4 of automatically binding the test result, SPD serial number, and workstation ID number includes: listening for the test result generation event through a preset API interface (such as RESTful API), triggering the real-time capture of the SPD serial number and workstation ID number, using an atomic transaction lock mechanism to ensure that the three are bound in the same event, and uploading to the MES system after verifying the data integrity through SHA-256 hash. Therefore, the SPD serial number and workstation ID number can be automatically obtained at the same time as the test result is generated, and the above three are accurately matched and bound to ensure the accuracy and traceability of data transmission.

[0020] By adopting the above preferred technical features, the accurate matching and binding of the test result, SPD serial number, and workstation ID number are ensured; the specific effects are as follows: 1. Data transmission accuracy: With the help of a preset API interface, the SPD serial number and workstation ID number are automatically obtained at the same time as the test result is generated, and accurately matched and bound, effectively reducing the errors that may be caused by manual intervention and improving the reliability of data flow. 2. Enhanced traceability: By binding the test result with the corresponding SPD serial number and workstation ID number, it is ensured that each test result can be traced back to the specific memory unit and test site, providing a reliable basis for subsequent quality analysis and problem location. 3. Efficiency improvement: The automated binding process avoids the extra time consumed by manual operations, improves the efficiency of the overall test process, and reduces the risk of human errors at the same time.

[0021] In a preferred method example of the present invention, it can be further configured as follows: The test method has a hierarchical permission management function and further includes the following steps: S2. Use fingerprint recognition and dynamic tokens (such as the TOTP algorithm) to perform dual authentication of user identities, define role-production line-data permission groups based on the RBAC model, and record operation behaviors to the blockchain log (such as Hyperledger Fabric) to achieve tamper-proof traceability; S5. When a privilege violation operation is detected, send a Trap alarm message to the MES system through the SNMP protocol, freeze the account access token (such as JWT), and at the same time trigger hardware-level isolation (such as disabling the PCIe interface of the test motherboard).

[0022] By adopting the above preferred technical features, use fingerprint recognition and dynamic tokens to perform dual authentication of user identities to control access rights, define role-production line-data permission groups based on the RBAC model, and record operation behaviors to the blockchain log to achieve tamper-proof traceability. This measure effectively prevents illegal access, and at the same time ensures that each user can only access resources related to their responsibilities, enhancing data security and privacy protection; in addition, when a privilege violation operation is detected, the alarm mechanism is automatically triggered and the account is frozen, minimizing potential risks and strengthening the system's security protection ability; the specific effects are as follows: 1. Authentication accuracy: Dual authentication significantly improves the accuracy and reliability of authentication and effectively prevents illegal users from accessing the system; 2. Fine-grained privilege management: Define role-production line-data permission groups based on the RBAC model to ensure that each user can only access the production lines and data within the authorized scope, refining the privilege control granularity; 3. Enhanced security: Record operation behaviors to the blockchain log to achieve tamper-proof traceability, providing a reliable basis for system operation and maintenance, facilitating the timely discovery and solution of problems; 4. Risk prevention and control: When a privilege violation operation is detected, immediately trigger the alarm mechanism and freeze the account, minimizing security risks and ensuring the normal operation of the system.

[0023] The second main object of the present invention is achieved through the following technical solutions: A test system for a semiconductor memory test production line is proposed, which can be used to implement a test method for a semiconductor memory test production line as described above. The test system includes: A modular test configuration unit that supports single-line independent configuration and workshop-level expansion driven by a PXE server, performs hardware virtualization and dynamic load balancing, is used to parallel process multiple work orders for parallel testing on a volatile memory test production line, and dynamically adjusts test parameters according to the requirements of each work order; A multi-dimensional data analysis platform integrates spatio-temporal correlation algorithms and a GPU-accelerated rendering engine. It can receive test data from different production lines, classify and organize it according to multiple dimensions, and generate interactive charts to achieve real-time visualization of fault modes. The permission management module combines biometric recognition and the RBAC model to achieve triple binding of role-production line-data and two-factor authentication. It double-verifies the user identity through fingerprint recognition or facial recognition technology and dynamic tokens to control access permissions, and records operation behaviors in the blockchain log for tamper-proof traceability. The adaptive interface device is used for SPD serial number encapsulation and MES system communication. It uploads the SPD serial number asynchronously and ensures highly reliable transmission of massive data through an asynchronous message queue.

[0024] By adopting the above system technical solutions, the modular test configuration unit supports single-line independent configuration and workshop-level expansion, performs hardware virtualization and dynamic load balancing, significantly improves equipment utilization and test efficiency, ensures reasonable allocation of hardware resources in the multi-work order parallel test scenario, and avoids resource conflicts and waste. The multi-dimensional data analysis platform integrates spatio-temporal correlation algorithms and a GPU-accelerated rendering engine, can receive test data from different production lines, classify and organize it according to multiple dimensions, and generate interactive charts, significantly enhancing the real-time visualization ability of fault modes and greatly improving data analysis efficiency and accuracy. The permission management module combines biometric recognition and the RBAC model to achieve triple binding of role-production line-data and two-factor authentication, effectively preventing illegal access, and at the same time recording operation behaviors in the blockchain log for tamper-proof traceability, greatly enhancing the security and reliability of the system. The adaptive interface device is responsible for SPD serial number encapsulation and MES system communication, ensures highly reliable transmission of massive data through an asynchronous message queue, ensures automatic binding of test results and SPD serial numbers, realizes closed-loop data management, guarantees the accuracy and traceability of data transmission, and further strengthens production collaboration capabilities.

[0025] In a preferred system example, the present invention can be further configured as follows: The modular test configuration unit can use a Kubernetes cluster to manage hardware resource virtualization, deploy the DRAM test firmware image through a PXE server, and support dynamic allocation of CPU cores and memory channels according to work order requirements. The adaptive interface device can achieve two-way communication with the MES system based on the gRPC protocol, and append the SPD serial number hash value (such as SHA-3) and the workstation ID digital signature (such as RSA-2048) when encapsulating test results. The modular test configuration unit further includes a sensor array for collecting temperature and current and feeding back to the central processor to update the test variable values including voltage and frequency.

[0026] By adopting the above-mentioned preferred system technical features, key data such as temperature and current can be collected through the sensor array and fed back to the central processor to dynamically update the values of test variables such as voltage and frequency. This measure significantly improves the intelligence level and response speed of the test system, and can adjust parameters in a timely manner according to the actual environmental changes during the test process, avoiding test errors or hardware damage risks caused by fixed settings, thereby improving the test accuracy and the operation stability of the equipment. The specific effects are as follows: 1. Dynamic parameter adjustment: The key data collected in real time by the sensor array is fed back to the central processor, and the system can dynamically adjust the values of test variables such as voltage and frequency according to the actual situation to ensure that the test process is always in the best state; 2. Improve test accuracy: Based on the real-time feedback mechanism, parameter fine-tuning is carried out to effectively reduce the deviation caused by external environmental factors and improve the accuracy of the overall test results; 3. Enhance equipment stability: By continuously monitoring key indicators such as temperature and current, potential overload or other abnormal conditions can be prevented, the service life of the equipment can be extended, and the maintenance cost can be reduced.

[0027] In a preferred system example of the present invention, it can be further configured that: the multi-dimensional data analysis platform further includes an anomaly detection unit, which deeply analyzes the test data based on advanced algorithms, accurately locates potential fault sources, and combines the particle-level fault tracing technology to quickly lock the corresponding DIMM slot and related number information.

[0028] By adopting the above-mentioned preferred system technical features, deeply analyzing the test data based on advanced algorithms can accurately locate potential fault sources; combining the particle-level fault tracing technology can quickly lock the corresponding DIMM slot and related number information, thereby significantly improving the efficiency and accuracy of fault diagnosis, shortening the repair cycle and reducing production losses; The specific effects are as follows: 1. Based on advanced algorithms, deeply analyzing the test data can timely discover hidden problems and improve the sensitivity and accuracy of fault prediction; 2. Combining the particle-level fault tracing technology to refine the positioning range and quickly lock the specific DIMM slot and related number information, greatly reducing the time cost of fault troubleshooting.

[0029] In a preferred system example, the present invention can be further configured as follows: The permission management module includes a two-factor authentication interface under the LSTM neural network, integrating one or more biometric identification mechanisms of a fingerprint recognition subsystem (FIDO2 protocol) and a face recognition subsystem, and a dynamic token generator based on the TOTP algorithm. The two work together to verify the user's identity. The user needs to pass both biometric and one-time password verification. The administrator can intuitively define permission rules through a visual Web interface to ensure that operators can only access the production lines and data within the authorized scope.

[0030] By adopting the above preferred system technical features, one or more biometric identification mechanisms of a fingerprint recognition subsystem and a face recognition subsystem, as well as a dynamic token generator based on the TOTP algorithm, can be integrated through the two-factor authentication interface in the permission management module. The two work together to verify the user's identity. The administrator can intuitively define permission rules through a visual Web interface to ensure that operators can only access the production lines and data within the authorized scope, thereby enhancing the security and reliability of the system. The specific effects include: 1) improving the accuracy and reliability of identity verification; 2) refining the permission control granularity to ensure that each user can only access the resources related to their responsibilities; 3) recording operation behaviors in the blockchain log for tamper-proof traceability, providing a reliable basis for system operation and maintenance; 4) automatically triggering an alarm mechanism and freezing the account when unauthorized operations are detected, minimizing potential risks and strengthening the security protection ability.

[0031] In a preferred system example, the present invention can be further configured as follows: The adaptive interface device realizes the automatic binding of test results, SPD serial numbers, and workstation ID numbers through a preset API interface and completes closed-loop management.

[0032] By adopting the above preferred system technical features, the adaptive interface device realizes the automatic binding of test results, SPD serial numbers, and workstation ID numbers through a preset API interface and completes closed-loop management. The specific effects are as follows: 1. Ensure the accuracy of data transmission: Automatically obtain SPD serial numbers and workstation ID numbers through a preset API interface, and perform precise matching and binding, reducing errors that may be caused by manual intervention. 2. Improve data traceability: Bind test results with SPD serial numbers and workstation ID numbers to ensure that each test result can be traced back to specific memory units and test sites, facilitating later quality control and problem location. 3. Strengthen the system coordination ability: With the help of the closed-loop management mechanism, efficient cooperation between the test system and the MES system is realized, promoting the smooth flow of test data to the production management link. 4. Enhance system reliability: By automating the binding process, the uncertainty caused by human factors is reduced, and the stability and credibility of the entire test system are improved.

[0033] In a preferred system example of the present invention, it can be further configured that: the test system further includes a blockchain log audit component and a hardware-level isolation unit. The blockchain log audit component records operation behaviors to private chain nodes (based on Hyperledger Fabric) to ensure that the logs cannot be tampered with; the hardware-level isolation unit is linked with the MES system. Once an unauthorized operation occurs, the PCIe interface of the test mainboard is remotely disabled through the IPMI protocol, and an SNMP Trap alarm is sent to the MES system.

[0034] By adopting the above preferred system technical features, all operation behaviors can be comprehensively recorded through the log audit component, and a linkage mechanism is formed between the hardware-level isolation unit and the MES system; once an unauthorized operation is detected, the alarm mechanism is immediately triggered and the relevant accounts are frozen, effectively preventing data leakage or system damage caused by illegal operations, and significantly improving the security and reliability of the system; the specific effects include: first, ensuring the complete traceability of the operation chain; second, quickly responding to abnormal events and reducing potential risks; third, strengthening the permission management system and ensuring data security.

[0035] In summary, the present invention includes at least one of the following technical effects that contribute to the prior art: 1. The dynamic load balancing algorithm supports parallel testing of multiple work orders through a weight allocation strategy (such as work order priority × equipment utilization rate). For example, in the scenario of four-work order parallel testing, the equipment utilization rate is increased by ≥40%, and the test cycle is shortened by 30%; 2. A fault prediction model is constructed based on the LSTM neural network to adjust the test item priority and parameters in real time, optimize the BIOS settings of the hardware platform, and the early warning accuracy rate for DRAM timing errors (such as tRCD exceeding the limit) is ≥95%, and the fault undetected rate is reduced to less than 2%; 3. The test results, SPD serial numbers, and work station ID numbers are automatically bound through a preset API interface and uploaded to the MES system through an asynchronous message queue to ensure the accuracy and traceability of data transmission, strengthen the production collaboration ability, and the data binding and asynchronous upload mechanism supports processing 100,000 test records per second, and the transmission error rate <0.001%. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Draw a schematic diagram of the main step process of a test method for a semiconductor memory test production line according to an embodiment of the present invention; Figure 2Illustrate the overall test architecture diagram of a semiconductor memory test production line according to an embodiment of the present invention, showing the hierarchical relationship among the CCE (Central Control End), LCE (Production Line Control End), test main board, and PXE server; Figure 3 Illustrate the execution of Figure 1 The system architecture diagram of step S1; Figure 4 Illustrate the execution of Figure 1 The system architecture diagram of step S2; Figure 5 Illustrate the execution of Figure 1 The system architecture diagram of step S3; Figure 6 Illustrate the execution of Figure 1 The system architecture diagram of step S4; Figure 7 Illustrate the execution of Figure 1 The system architecture diagram of step S5; Figure 8 Illustrate the dynamic load balancing algorithm flow chart according to an embodiment of the present invention, explaining the resource allocation logic; Figure 9 Illustrate the schematic diagram of the operation mode of the data visualization Web interface, showing multi-dimensional screening conditions and interactive charts; Figure 10 Illustrate the schematic diagram of the operation mode of the permission management interface, showing the role-production line-data binding relationship; Figure 11 Illustrate the MES docking data flow diagram, describing the SPD serial number binding and asynchronous message queue transmission mechanism.

[0037] Reference numerals: 100, volatile memory test production line; 101, test station; 102, test main board; 103, PXE server; 104, station ID number; 105, DIMM slot; 120, MES system; 10, modular test configuration unit; 11, logic group management module; 12, sensor array; 13, central processor; 14, PXE server interface; 20, multi-dimensional data analysis platform; 21, GPU accelerated rendering engine; 22, interactive chart; 23, anomaly detection unit; 24, test item; 25, BIOS programming interface; 30, permission management module; 31, fingerprint recognition subsystem; 32, dynamic token generator; 33, RBAC model; 34, blockchain log audit component; 40, adaptive interface device; 41, data interaction layer; 42, preset API interface; 43, hardware-level isolation unit; 44, asynchronous message queue; 45, test result; 50, expansion module; 51, Web interface; 52, data analysis interface; 60, LSTM neural network; 61, fault prediction model; 62, two-factor authentication interface; 70, central control layer; 71, dynamic load balancing algorithm module; 91, work order; 92, volatile memory; 93, SPD serial number. Detailed implementation manners

[0038] The technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments are only part of the embodiments for understanding the inventive concept of the present invention and cannot represent all embodiments, nor are they interpreted as the only embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention under the premise of understanding the inventive concept of the present invention fall within the scope of protection of the present invention. The embodiments are used to explain the technical solutions of the present invention and do not constitute a limitation on the scope of protection of the present invention. Those skilled in the art can reasonably adjust the embodiments based on the actual application scenarios on the premise of understanding the core idea of the present invention.

[0039] In the field of semiconductor volatile memory testing, the following problems exist in simple and well-known testing systems: Rigid configuration: It is impossible to independently adjust test items for a single production line, and the parallel testing efficiency of multiple work orders is low; Data isolation: The test data lacks the ability of multi-dimensional analysis, and it is difficult to achieve particle-level fault tracing; Coarse-grained permissions: The operation permissions are not strongly bound to the production line and data, and there is a risk of data leakage; Poor scalability: The existing architecture is difficult to be compatible with the future expansion requirements of workshop-level equipment.

[0040] Although it is possible to attempt to integrate simple data collection and preliminary analysis functions into the test system, it is still in its infancy and lacks refined permission management and security measures. Usually, a test system with a fixed configuration is adopted, relying on manual adjustment of parameters to adapt to different test scenarios, integrating limited data collection and preliminary analysis functions, and establishing a basic authentication process to distinguish the permissions of different levels of employees. However, these methods have problems such as long configuration adjustment time, low efficiency in parallel testing of multiple work orders, lack of effective linkage between data management and permission control, low troubleshooting efficiency, and insufficient system security.

[0041] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly. For a more convenient understanding of the technical solution of the present invention, the test method and system of a semiconductor memory test production line of the present invention will be further described and explained in detail below, but it does not constitute the protected scope defined by the present invention.

[0042] What is shown in the drawings is only the part common to multiple embodiments. The parts with differences or distinctions are described in words or presented by comparing with the drawings. Therefore, based on the industrial characteristics and technical essence, those skilled in the art should correctly and reasonably understand and judge whether the following described individual technical features or any combination of them can be characterized in the same embodiment, or whether multiple technically mutually exclusive technical features can only be characterized in different variant embodiments respectively.

[0043] Refer to Figure 1 , a test method of a semiconductor memory test production line disclosed in an embodiment of the present invention mainly includes steps S1 to S5. Among them, step S1 is dynamic load balancing and logic group division, step S2 is hierarchical permission management and security control, step S3 is LSTM fault prediction and BIOS parameter optimization, step S4 is data binding and closed-loop management, and step S5 is to send an alarm to the MES system and freeze the token to trigger hardware-level isolation when an unauthorized operation occurs; steps S1, S3, and S4 are necessary steps, and steps S2 and S5 are optional steps. And the corresponding system architecture can be seen in Figure 2 .

[0044] In Figure 1 's method flow architecture and Figure 2 's system architecture, refer to Figure 3, Step S1 involves dynamic load balancing and logical group division, and dynamically defines the hardware configuration through logical group division for parallel testing of multiple work orders. The role of the dynamic load balancing algorithm in dividing logical groups and allocating hardware resources is as follows: dividing a production line into multiple logical groups, and allocating hardware resources through the dynamic load balancing algorithm to support parallel testing of multiple work orders, avoiding resource conflicts and waste, and improving testing efficiency. Figure 2 It is the system architecture diagram used for this testing method. A volatile memory test production line 100 has multiple test workstations 101. The volatile memory test production line 100, as the hardware resource layer, can perform binding and transmission of test data. In a specific embodiment, the inputs of Step S1 for dynamic load balancing and logical group division are: the hardware resources (CPU cores, memory bandwidth) of the volatile memory test production line 100 and the priority weights of work order 91; the outputs of Step S1 are: logical group division (GroupA / B / C / D) and resource allocation results; in the DRAM adaptation, for the parallel refresh requirements of DDR, the memory bandwidth and refresh cycle resources are dynamically allocated.

[0045] Refer to Figure 3 , In this embodiment, when specifically implementing the dynamic load balancing algorithm and the parallel testing process of multiple work orders, the modular test configuration unit 10, as the production line execution layer (LCE), can drive memory testing and upload test data. The modular test configuration unit 10 includes a logical group management module 11 and a sensor array 12. The test system also includes a central control layer 70 (CCE). The central control layer 70 is used to issue test instructions and receive feedback on test results. The central control layer 70 includes a dynamic load balancing algorithm module 71. The logical group management module 11 can divide one or more volatile memory test production lines 100 into multiple logical groups based on each production line, and the dynamic load balancing algorithm module 71 allocates hardware resources through the dynamic load balancing algorithm to support parallel testing of multiple work orders. In this embodiment, the logical group management module 11 performs logical group division, dividing a volatile memory test production line 100 into four logical groups (Group A / B / C / D). Each group is equipped with an independent test main board 102 and a sensor array 12. Each test main board 102 is set with a work station ID number 104. Physical isolation is achieved between logical groups through solid-state switches to avoid signal interference. Then, the dynamic load balancing algorithm module 71 is used for dynamic resource allocation. Specifically, the Kubernetes cluster is used to manage hardware resources (such as CPU cores, memory channels), and resources are allocated through the dynamic load balancing algorithm module 71.

[0046] The role of the dynamic load balancing algorithm considering the importance difference of work orders is as follows: By adding a weight coefficient to comprehensively consider the importance of each work order, reasonably allocate hardware resources, and ensure that critical tasks are executed first. The algorithm specifically introduces a weight coefficient (such as work order priority × device utilization rate). The specific factors affecting the weight coefficient of work order priority include: 1. Allocate more CPU resources (the allocation weight can reach 70%) to critical work orders (such as DDR5 modules with high yield requirements); 2. Allocate the remaining resources to ordinary work orders (such as a 30% allocation weight); 3. The sensor array 12 collects temperature and current data in real time and feeds it back to the central processor 13 of the modular test configuration unit 10 to adjust the voltage (such as ±0.05V) and frequency (such as ±50MHz), and adjust the preset weight to the actual operation weight according to the feedback message to ensure test stability. Therefore, in a specific example, the dynamic load balancing algorithm described in step S1 can consider the importance difference of each work order and obtain the hardware resource allocation result through comprehensive consideration by adding a weight coefficient.

[0047] The test process of step S1 also includes: S11. Work order configuration: Based on the information of work order 91, the operator configures test items (such as "Hammer Test") for logic group Group A through the Web interface 51, sets the number of cycles to be specifically 1000 times, and binds the SPD serial number 93 in the template; S12. Parallel execution: The system deploys the DRAM test firmware to each logic group through the PXE server 103 of the volatile memory test production line 100 and starts the tests of four work orders 91 synchronously; The volatile memory 92 to be tested has been installed in the DIMM slot 105; The test data is stored in the MES system 120 of the independent database partition in real time to ensure data isolation (such as Figure 2 shown).

[0048] In Figure 1 the method process architecture and Figure 2 the system architecture of Figure 4, Step S2 involves hierarchical permission management and security control. The role of the hierarchical permission management method is as follows: double-verify the user identity using fingerprint recognition (or / and face recognition) and dynamic tokens, define role-production line-data permission groups based on the RBAC model, record operation behaviors to private chain nodes, and prevent unauthorized operations. Double-verify the user identity using fingerprint recognition and dynamic tokens to control access permissions, define role-production line-data permission groups based on the RBAC model, and record operation behaviors to the blockchain log to achieve tamper-proof traceability. The permission management module 30 implements the aforementioned two-factor authentication. The permission management module 30 includes a fingerprint recognition subsystem 31 (or a face recognition subsystem), a dynamic token generator 32, and an RBAC model 33. First, perform two-factor authentication with the fingerprint recognition subsystem 31 and the dynamic token generator 32. During the authentication process through the two-factor authentication interface 62, the operator needs to pass both the fingerprint recognition subsystem 31 and the dynamic token issued by the dynamic token generator 32 to complete the double authentication. Refer to Figure 10 , after the double authentication is passed, perform permission binding, and as Figure 4 shown, define role-production line-data permission groups based on the RBAC model 33. For example: Engineers can access the test data of logic groups Group A / B; Administrators can configure the parameters of the entire production line. The input of step S2 for hierarchical permission management and security control is: double-authentication identity verification information. After the permission management module 30 confirms the permissions, the permission distribution is synchronized to the central control layer 70 and the modular test configuration unit 10 serving as the production line execution layer.

[0049] In Figure 1 the method process architecture and Figure 2 the system architecture of Figure 5, Step S3 involves LSTM fault prediction and BIOS parameter optimization. A fault prediction model 61 is constructed based on the LSTM neural network 60 to adjust the priority and parameters of test items in real time, including increasing the loop count of high-frequency test items and optimizing the BIOS parameter template of the hardware platform. The role of constructing a fault prediction model based on the LSTM neural network is to perform multi-dimensional data analysis and data classification and sorting, use machine learning algorithms to predict data growth trends and abnormal patterns, provide an analysis basis for the fault prediction model, adjust the priority and parameters of test items in real time, increase the loop count of high-frequency test items, optimize the BIOS parameter template of the hardware platform, and improve the test accuracy and fault prediction ability. The operation of optimizing the BIOS parameter template of the hardware platform includes: adjusting key parameters including the CPU frequency, voltage, and cache settings of the test station 101 through the BIOS programming interface 25 according to the actual hardware characteristics and test requirements to improve the test accuracy and efficiency. The inputs of Step S3 of LSTM fault prediction and BIOS parameter optimization are historical test data (timing errors, refresh failure rate) and real-time sensor data (temperature, voltage); the outputs of Step S3 are the adjustment of test item priority (such as increasing the loop count of the "Row Hammer Test") and BIOS parameter optimization (CL value, tRCD). In the DRAM adaptation, the focus of model training is on the unique timing deviation and signal integrity faults of DRAM.

[0050] The method for constructing a fault prediction model in Step S3 includes: S3a. Data collection and training: Construct a fault prediction model 61 of the LSTM neural network 60 based on historical test data (such as timing error rate, voltage fluctuation); the input features include: tCL and tRCD parameters of the DIMM slot 105; environmental data collected by the temperature sensor in the sensor array 12; and the fault prediction model 61 outputs the fault prediction confidence level (such as triggering an alarm when the DDR refresh failure rate ≥ 95%). S3b. Real-time parameter adjustment: Based on the priority of the test item 24 and the BIOS programming interface 25, when the fault prediction model 61 detects an increase in the failure rate of the "Bit Fade Test" for a certain batch of memory (such as Figure 5 shown), automatically increase the loop count of this test item to 1500 times; call the BIOS programming interface 25 through the UEFI Shell script to adjust the VDDQ voltage of the Z790 motherboard to 1.2V and optimize the CL value (tCL = 18); the test results show that the fault undetected rate is reduced from 5% to 1.8%, and the test cycle is shortened by 30% (such as Figure 8 shown). Among them, Figure 8 shows a dynamic load balancing process.

[0051] In a specific embodiment of step S3, it includes multi-dimensional data analysis and visualization. According to the historical test data and current production status of different production lines, machine learning algorithms are used to predict the growth trend and abnormal patterns of data, providing an analysis basis for subsequent fault diagnosis and production optimization of the fault prediction model 61. Therefore, the role of the multi-dimensional data analysis method is to receive test data from different test production lines, classify and organize them according to multiple dimensions such as time range, production line number, etc., generate interactive charts, and accelerate the fault diagnosis process.

[0052] The multi-dimensional data analysis method of step S3 includes: S31. Generating an interactive chart based on multi-dimensional screening conditions. The multi-dimensional data analysis platform 20 receives test data from different production lines, classifies and organizes them according to the time range, production line number, and hardware platform (specifically, motherboard model, DDR generation), and then generates an interactive chart 22; refer to Figure 9 , the classification mode of the test data is more specifically executed according to the timing parameter CL value, DIMM slot position (corresponding to the production line number), and DDR generation (such as DDR4 / DDR5). Among them, Figure 9 It shows an operating mode of a data visualization web interface.

[0053] Step S31 includes: S311. Adding an interaction function to the interactive chart 22 to view detailed test data and analysis results by clicking on a single data point; S312. Adopting a data chunk loading technology, when dragging the time axis, only the test data within the current window range is rendered, ensuring that the dynamic update response time is less than 200 ms.

[0054] The visualization method of step S3 includes: S32. GPU-accelerated rendering: Based on a spatio-temporal association algorithm (such as an association rule based on timestamp and physical location), test data is screened according to multiple dimensions of time range, production line number, and hardware platform, and the CUDA framework is used to call the GPU-accelerated rendering engine 21 (such as Unity3D or WebGL) to generate a dynamic and interactive interactive chart 22. The interactive chart 22 includes a trend chart (such as a timing error rate curve) and a heat map (such as a DIMM slot fault distribution).

[0055] Therefore, the role of generating a dynamic and interactive interactive chart is to add an interaction function to the chart. Users can view detailed test data by clicking on a single data point and drag the time axis to real-time update the data range displayed on the chart. Therefore, the interactive chart 22 has an interaction function. When a user clicks on a specific DIMM slot 105 in the heat map as shown in Figure 9 , the detailed voltage fluctuation curve and the LSTM prediction result can be viewed. When dragging the time axis of the interactive chart 22, the test system adopts a data chunk loading technology, and the response time of the interactive chart 22 is <200 ms (as shown in Figure 9 ).

[0056] In Figure 1 's method process architecture and Figure 2 's system architecture, referring to Figure 6 , step S4 involves data binding and closed-loop management, including binding the test result with the SPD serial number. Its function is: automatically obtain the SPD serial number through a preset API interface and precisely match it with the test result to ensure the accuracy and traceability of data transmission. In Figure 6 , the test result 45, the SPD serial number 93, and the station ID number 104 of the station are automatically bound, and uploaded to the MES system 120 through the asynchronous message queue 44, and closed-loop management is achieved through the preset API interface 42. The function of binding the test result with the SPD serial number and uploading it to the MES system through the asynchronous message queue is: to ensure the accuracy and traceability of data transmission, achieve closed-loop management of production line testing and test production, and reduce the cost of manual intervention. In a specific example, through the preset API interface 42, the SPD serial number 93 and the station ID number 104 are automatically obtained simultaneously when the test result 45 is generated, and the above three are precisely matched and bound to ensure the accuracy and traceability of data transmission. The inputs of step S4 of data binding and closed-loop management are: the test result 45, the SPD serial number 93 (including DDR timing parameters), and the station ID number 104 of the test station 101; the output of step S4 is: the JSON data packet is uploaded to the MES system 120 through Kafka, triggering a production rework instruction. In the DRAM adaptation, SPD binding ensures the traceability of DRAM performance parameters (such as frequency and latency).

[0057] The method for binding the SPD serial number 93 in step S4 specifically includes: S4a. Automated binding process: Refer to Figure 6 and Figure 11 , after the memory test is completed, the system grabs the SPD serial number 93 and the station ID number 104 through the preset API interface 42 of RESTful, and binds the above three into a JSON data packet using the atomic transaction lock mechanism plus the test result 45; S4b. Asynchronous upload: After the JSON data packet is verified by SHA-256 hashing, it is uploaded to the MES system 120 through the Kafka asynchronous message queue 44, and the transmission rate of this upload can reach 100,000 pieces / second, and the error rate < 0.001%.

[0058] The closed-loop management method of step S4 includes: S4c. MES feedback closed-loop, specifically including: The MES system 120 generates a rework work order according to the test result 45 and issues it to the volatile memory test production line 100 through the reverse of the preset API interface 42, realizing seamless docking between the production line test system and the test production line (such as Figure 6as shown).

[0059] In Figure 1 the method process architecture of and Figure 2 under the system architecture of, referring to Figure 7 , step S5 involves unauthorized operation alarm and account freezing. When an unauthorized operation is detected, an alarm message is sent to the MES system 120, and the account access token is frozen. At the same time, hardware-level isolation is triggered. Specifically, the blockchain log audit component 34 and the hardware-level isolation unit 43 are used to record blockchain logs and perform hardware isolation respectively. The operation audit of the blockchain log audit component 34 is to record all operations to the Hyperledger Fabric private chain node to ensure that the logs cannot be tampered with. When unauthorized access is detected, the system performs unauthorized processing. The hardware-level isolation unit 43 disables the PCIe interface of the test motherboard 102 through the IPMI protocol and sends an SNMP alarm to the MES system 120 (as Figure 7 shown). Therefore, step S5 includes log recording, unauthorized detection, and alarm freezing.

[0060] In summary, the synergistic effect of the dynamic load balancing algorithm (seen in S1) and the fault prediction model based on the LSTM neural network (seen in S3) is as follows: The dynamic load balancing algorithm ensures the reasonable allocation of hardware resources, while the LSTM neural network adjusts the test item priorities and parameters in real time. The combination of the two significantly improves the test efficiency and fault prediction ability. The synergistic effect of the hierarchical permission management method in step S2 and the log audit component is as follows: The hierarchical permission management method controls access permissions through double verification and role binding, and the log audit component records operation behaviors to the private chain node. The combination of the two enhances the security and traceability of the system. The synergistic effect of the multi-dimensional data analysis method in step S3 and the generation of interactive charts is as follows: The multi-dimensional data analysis method classifies and organizes test data, and the generation of interactive charts provides an intuitive data display method. The combination of the two accelerates the fault diagnosis process. The synergistic effect of binding the test results with the SPD serial number and uploading them to the MES system through the asynchronous message queue in step S4 is as follows: Binding the test results with the SPD serial number ensures the accuracy and traceability of the data, and uploading to the MES system through the asynchronous message queue realizes the closed-loop management of production line testing and test production, reducing the manual intervention cost.

[0061] The test method of a semiconductor memory test production line in the above example of the present invention has the following system synergistic effects, including: 1. Collaborating with the modular test configuration unit 10 and the multi-dimensional data analysis platform 20, the dynamic load balancing ensures efficient resource utilization, the GPU-accelerated chart improves the fault location speed, and the repair cycle is shortened by 50%; 2. The fingerprint recognition subsystem 31 of the linkage permission management module 30, the two-factor authentication with the dynamic token generator 32, and the blockchain audit of the blockchain log audit component 34 achieve an interception rate of 99.9% for unauthorized operations, with the risk of data leakage approaching zero, enhancing the system security; 3. The data interaction layer 41 cooperating with the adaptive interface device 40 and the sensor array 12 of the modular test configuration unit 10 optimize the test parameters with real-time environmental data feedback, the test yield can be increased by 15%, and the data accuracy is improved; 4. Realize the dynamic adjustment of test resources for the test production line of volatile memories such as DRAM memories, obtain more suitable test resources when testing different production lines (such as DDR4 and DDR5), ensuring the maximization of test efficiency; in terms of specific industrial applicability, the present invention can be applied to the DDR5 test production line of semiconductor manufacturers, achieving parallel testing of four work orders in a single line, with the equipment utilization rate increased by 42% and the fault prediction accuracy rate reaching 96.5%.

[0062] Refer to again Figure 2 , the embodiment of the present invention also discloses a test system for a semiconductor memory test production line, which can be applied and matched with a volatile memory test production line 100. The volatile memory test production line 100 includes a plurality of test workstations 101, each test workstation 101 has a work station ID number 104 corresponding to its identity, and a test main board 102 is arranged on the test workstation 101 and is driven by connecting to a PXE server 103. A plurality of DIMM slots 105 are provided in the test main board 102 for pluggable engagement with the volatile memory 92 to be tested. The test system mainly includes: a modular test configuration unit 10, a multi-dimensional data analysis platform 20, a permission management module 30, and an adaptive interface device 40.

[0063] The modular test configuration unit 10 includes a logic group management module 11, a sensor array 12, and a central processor 13, which can support single-line independent configuration relative to the volatile memory test production line 100, and support workshop-level expansion driven by the PXE server 103 through the PXE server interface 14, perform hardware virtualization and dynamic load balancing, and are used for parallel processing of multiple work orders 91 in parallel testing on a volatile memory test production line 100, and dynamically adjust test parameters according to the requirements of each work order 91. That is, the modular test configuration unit 10 further includes a reserved PXE server interface 14 for increasing the expandability of the modular test configuration unit 10 and supporting workshop-level expansion (such as from a single line to 2048 devices). Therefore, the role of the modular test configuration unit 10 is to support single-line independent configuration and workshop-level expansion driven by the PXE server, perform hardware virtualization and dynamic load balancing, and dynamically adjust test parameters.

[0064] The multi-dimensional data analysis platform 20 integrates a spatio-temporal correlation algorithm and a GPU accelerated rendering engine 21, can receive test data from different production lines, sort and organize it according to multiple dimensions, and generate interactive charts 22 to achieve real-time visual presentation of fault modes; the multi-dimensional screening conditions can include: DIMM slot position, DDR generation (DDR4 / DDR5), timing parameters (CL value). Therefore, the functions of the multi-dimensional data analysis platform 20 are: integrating a spatio-temporal correlation algorithm and a GPU accelerated rendering engine, receiving test data from different production lines, generating interactive charts, and achieving real-time visual presentation of fault modes. The combination of the multi-dimensional data analysis platform 20 and the modular test configuration unit 10 produces a collaborative effect of visually dynamically adjusting test parameters under multi-work order parallel testing, significantly improving test efficiency and data analysis capabilities.

[0065] The permission management module 30 combines biometric identification with the RBAC model 33 to achieve triple binding of role-production line-data (where administrators can access all production lines), and its biometric mechanism specifically includes a fingerprint recognition subsystem 31 and a dynamic token generator 32 for two-factor authentication. Through the fingerprint recognition technology (or facial recognition technology) of the fingerprint recognition subsystem 31 and the dynamic token provided by the dynamic token generator 32 based on the TOTP algorithm, the user identity is double-verified to control access permissions. The permission management module 30 also includes a blockchain log auditing component 34, which is linked with the MES system 120. The blockchain log auditing component 34 records operation behaviors to the blockchain log to achieve tamper-proof traceability, automatically freezes operation permissions and issues an alarm when an unauthorized operation is triggered, enhancing system security. Therefore, the functions of the permission management module 30 are: combining biometric identification with the RBAC model, achieving triple binding of role-production line-data and two-factor authentication, controlling access permissions and recording operation behaviors. In addition, based on the combination of the permission management unit of the permission management module 30 and the blockchain log auditing component 34, access permissions are first controlled by two-factor verification and role binding, and then all operation behaviors are recorded to the private chain node and linked with the MES system. The combination of the two enhances the security and traceability of the system.

[0066] The adaptive interface device 40 includes a data interaction layer 41, which is used to encapsulate the SPD serial number 93 data and communicate with the MES system 120, asynchronously upload the SPD serial number 93, and ensure the highly reliable transmission of massive data through the asynchronous message queue 44. This data encapsulation automatically binds the test result 45, the SPD serial number 93, and the workstation ID number 104, and then generates a JSON data packet. Therefore, the role of the adaptive interface device 40 is to encapsulate the SPD serial number and communicate with the MES system, and ensure the highly reliable transmission of massive data through the message queue. This asynchronous upload is uploaded to the MES system 120 through the asynchronous message queue 44 (such as Kafka) to ensure the transmission integrity of a large amount of data. Therefore, in combination with the sensor array 12, the adaptive interface device 40 ensures the binding of the test result and the SPD serial number and the reliability of data transmission, collects key data and feeds it back to the central processor, and the combination of the two improves the collaborative effect of test accuracy and data transmission efficiency.

[0067] In a specific embodiment, the modular test configuration unit 10 uses a Kubernetes cluster to manage the virtualization of hardware resources, deploys the DRAM test firmware image through the PXE server 103, and supports the dynamic allocation of CPU cores and memory channels according to the work order requirements; the adaptive interface device 40 realizes bidirectional communication with the MES system 120 based on the gRPC protocol, and attaches the hash value of the SPD serial number 93 and the digital signature of the workstation ID number 104 when encapsulating the test result 45; the modular test configuration unit 10 further includes a sensor array 12 that can be used to collect temperature, current, or other key data, and feedback to the central processor 13 for updating the test variable values (including voltage, frequency) to improve the test accuracy. Refer to Figure 11 And Figure 6 , in the docking data stream of the MES system 120, the test result 45, the SPD serial number 93, and the workstation ID number 104 data are encapsulated into a JSON data packet and sent to the MES system 120 through the Kafka asynchronous message queue 44 to ensure data integrity under high concurrency. Through the adaptive interface device 40, the MES system 120 can perform closed-loop feedback, generate a rework instruction according to the timing parameters in the SPD serial number 93 (such as adjusting the tRFC value), and the generation of the rework work order can optimize the test production process.

[0068] In a variant embodiment, the modular test configuration unit 10 is a distributed cluster composed of a series of microcomputer nodes. A small Linux operating system is installed on each node to execute various task scheduling and service request commands. These nodes are connected by gigabit optical fibers to form a local area network and share a storage pool as a temporary buffer to save intermediate calculation results. To enhance disaster tolerance, a mirror copy is also backed up in a remote location to prevent accidental loss. In addition, a sensor array 12 is added to sense environmental changes and give immediate modification suggestions for certain sensitive parameters. For example, when it is found that the temperature is too high in a certain place, a warning is immediately issued and the working intensity of the corresponding area is appropriately reduced until it returns to normal.

[0069] In a specific embodiment, the multi-dimensional data analysis platform 20 can rely on a GPU cluster to perform large-scale matrix calculation operations. Users can select an interested time period or the product pipeline number of a specific type on the front-end interface, and the background program will quickly retrieve the corresponding indicators that meet the conditions and make them into corresponding statistical graphs for display. In addition to the basic trend curve, there are also pie charts showing the proportion and bar charts comparing the performance gaps in different time periods for reference.

[0070] In a more specific embodiment, the multi-dimensional data analysis platform 20 further includes an anomaly detection unit 23, which deeply analyzes the test data based on advanced algorithms, accurately locates potential fault sources, and combines with the particle-level fault tracing technology to quickly lock the corresponding DIMM slot 105 and related numbers (specifically, it can include the workstation ID number 104 and the SPD serial number 93). The anomaly detection unit 23 can use advanced mathematical formulas to deeply analyze the existing data to find clues of small probability events that are unusual but valuable and give key attention, reminding the staff to follow up and verify in time to avoid greater losses. Therefore, the role of the anomaly detection unit 23 is to: locate potential fault sources (such as locking the DIMM slot number), accelerate the fault diagnosis process, and improve the intelligence level of the test system.

[0071] Refer to Figure 5 And Figure 9 , the interactive charts 22 provided by the multi-dimensional data analysis platform 20 specifically include trend charts and heat maps. The trend charts show the change of the timing error rate over time (such as tRCD overrun events), and the heat maps identify the areas where DRAM failures are concentrated (such as the refresh failure hotspots of specific DIMM slots and the Bank conflict areas). The above information can be exported to an Excel file by the LSTM neural network 60, including the original test data and SPD metadata.

[0072] In a specific embodiment, refer to Figure 4, the framework of the permission management module 30 can adopt the RBAC role-based access control policy. The permission management module 30 includes a two-factor authentication interface 62 under the LSTM neural network 60. The permission management module 30 integrates one or more biometric identification mechanisms of the fingerprint recognition subsystem 31 and the face recognition subsystem, and a dynamic token generator 32. The fingerprint recognition subsystem 31 performs biometric identification, and the dynamic token generator 32 provides dynamic tokens. The two work together to verify the user's identity. The administrator can intuitively define permission rules through the visual Web interface 51 of the extension module 50 to ensure that the operator can only access the production lines and data within the authorized scope.

[0073] In a specific embodiment, the adaptive interface device 40 plays a bridging role, connecting the aforementioned software and hardware facilities in series to operate normally as an organic whole. On the one hand, it can receive the original information transmitted from the volatile memory test production line 100; on the other hand, it can reorganize and arrange it according to the established rules and then send it to the production line test system to ensure that there are no errors in the subsequent operations. The whole process is completely automated without human intervention, greatly improving work efficiency, reducing cost consumption, and effectively preventing problems such as possible human negligence and omission during the process.

[0074] In a more specific embodiment, the test system may include a central control layer (CCE) 70, which is used to issue test instructions to the volatile memory test production line 100 and receive information feedback from the volatile memory test production line 100. The central control layer 70 specifically includes a dynamic load balancing algorithm module 71, which is used to allocate hardware resources to support the parallel testing of multiple work orders 91. The allocation mode of hardware resources includes dynamic load balancing optimization of CPU and memory bandwidth allocation under high-frequency refresh testing to support parallel refresh operations. Refer to Figure 8 , the algorithm logic of the dynamic load balancing algorithm module 71 includes: inputting the work order weight, resource allocation strategy, and the output result of logical group resource allocation. The input work order weight allocates resources according to the urgency of the work order (for example, the DDR5 verification requirement is greater than the DDR4 mass production); the resource allocation strategy preferentially allocates memory bandwidth to high-frequency test items (such as DDR5-6400 verification), and real-time feedback of sensor data (temperature, current) is used to trigger dynamic resource adjustment (specifically, such as frequency reduction to prevent overheating); the logical group resource allocation includes the occupancy ratio of CPU cores and memory channels.

[0075] Specifically, the dynamic load balancing algorithm module 71 includes a multi-component separation unit, such as a separation unit composed of a relay array or a solid-state switch. These separation units are responsible for physically isolating the resource sharing between logical groups to avoid interference. The separation unit can select a relay made of high-performance metal contact materials or use a solid-state switch device to replace the mechanical relay to improve the switching speed and lifespan. Each logical group is equipped with an independent resource management unit, which consists of a microcontroller and an embedded operating system, and controls resource allocation through a dedicated communication protocol. Assuming that there are a total of N available cores in the entire production line, then approximately n = N / m cores are evenly allocated to each active logical group for its use according to the number of work orders m currently being executed. In addition, a weight coefficient w can be introduced to reflect the importance difference of each work order, and a more reasonable allocation result can be obtained through comprehensive consideration.

[0076] In a more specific embodiment, the LSTM neural network 60 is constructed with a fault prediction model 61. The core of the construction is a group of sensor arrays 12. Common types include PT100 thermistor temperature measurement probes and Hall effect current transformers, which are used to collect key data such as temperature and current of the corresponding logical group in the test station 101 and feedback to the central processor 13 to update test variable values such as voltage or frequency. After receiving the sensor data from the sensor array 12, the central processor 13 inputs it into the fault prediction model 61 of the LSTM neural network 60 for processing. Through learning and analysis of historical test data, the fault prediction model 61 can adjust the priority and parameters of test items in real time, such as increasing the loop times of high-frequency test items and optimizing the BIOS parameter template of the hardware platform. For the Z790 motherboard, key parameters such as CPU frequency, voltage, and cache settings can be adjusted through the BIOS programming interface to improve test accuracy and efficiency.

[0077] In addition, the priority and parameters of the test item 24 are adjusted in real time through the BIOS programming interface 25, including increasing the loop times of high-frequency test items and optimizing the BIOS parameter template of the hardware platform. As Figure 4 shown, the expansion module 50 includes a Web interface 51 for independently configuring the test items (such as test loop times, voltage parameters) of a single production line. The expansion module 50 has a network topology function, supports the unified management of the chip workshop and the module workshop, and is compatible with DDR4 / DDR5 test platforms. Refer to Figure 4 And Figure 5 , the expansion module 50 may also include a data analysis interface 52 for displaying an interactive chart 22 on the human-machine operation interface, allowing human intervention in addition to automatic machine adjustment. In terms of permission hierarchy, the administrator can access the data of the entire workshop, and the production line engineer is only limited to the authorized production line.

[0078] In a specific embodiment, the adaptive interface device 40 implements the automatic binding of the test result 45, the SPD serial number 93, and the station ID number 104 through a preset API interface 42, and completes the closed-loop management. The preset API interface 42 is used to implement the automatic binding function in the data binding and uploading part. After the memory test is completed, the test system will automatically obtain the SPD serial number 93 on the corresponding memory chip, and package it together with the station ID number 104 into a JSON format data packet. Then it is uploaded to the MES system 120 by using an asynchronous message queue 44 (such as Kafka) to ensure data integrity in a high-concurrency scenario. At the same time, the MES system 120 side will also feedback production instructions (such as the generation of rework work orders) through the reverse of the preset API interface 42, so as to realize the closed-loop management of testing and production.

[0079] In a specific embodiment, the adaptive interface device 40 of the test system further includes a hardware-level isolation unit 43. The blockchain log audit component 34 records operation behaviors to the private chain node. The hardware-level isolation unit 43 is linked with the MES system 120. Once an unauthorized operation occurs, the hardware-level isolation unit 43 remotely disables the PCIe interface of the test main board 102 and sends an alarm signal to the MES system 120.

[0080] The implementation principle of this embodiment is as follows: reasonably allocate hardware resources through a dynamic load balancing algorithm to avoid resource conflicts and waste; based on the LSTM neural network, adjust the test item priorities and parameters in real time to improve the test accuracy and intelligence level; data binding and uploading ensure the accuracy and traceability of data transmission, and realize the closed-loop management of testing and production. The combined effect of the above three measures significantly improves the overall efficiency and management level of the semiconductor memory test production line.

[0081] Therefore, a test system for a semiconductor memory test production line disclosed in an embodiment of the present invention enables components to work together. With dynamic load balancing, intelligent fault prediction, and data binding as the core, it realizes an efficient and accurate semiconductor memory test process, significantly improving the test efficiency, data accuracy, system security, and fault prediction ability of the semiconductor memory test production line. The test system of the embodiment of the present invention forms a secure and scalable closed-loop management system through a modular architecture and a multi-dimensional data analysis platform, combined with permission management and adaptive interfaces. The two work together to solve the technical bottlenecks of traditional test systems in aspects such as rigid resource allocation, data isolation, and extensive permissions, providing a comprehensively optimized solution for the semiconductor memory test production line. Specifically, it realizes parallel testing of four or more work orders in one or more DDR DRAM test production lines (or other known semiconductor memory test production lines), can dynamically adjust test parameters to improve test efficiency; enhances data security through refined permission management to ensure that engineers can only access authorized production lines and data; integrates test data and production data for pattern recognition and performance prediction to improve the accuracy of fault prediction and the technical solution for formulating maintenance plans.

[0082] The embodiments of this specific implementation manner are all preferred embodiments for conveniently understanding or implementing the technical solution of the present invention, and do not limit the protection scope of the present invention accordingly. Any equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the scope of the claims of the present invention.

Claims

1. A test method for a semiconductor memory test production line, characterized in that, Including the following steps: S1. Divide a volatile memory test production line into multiple logical groups, and allocate hardware resources through a dynamic load balancing algorithm to support parallel testing of multiple work orders; S3. Build a fault prediction model based on the LSTM neural network, and adjust the test item priorities and parameters in real time, including increasing the loop times of high-frequency test items and optimizing the BIOS parameter template of the hardware platform; S4. Automatically bind the test results, SPD serial numbers, and workstation ID numbers, and upload them to the MES system through an asynchronous message queue to achieve closed-loop management through a preset API interface.

2. The test method according to claim 1, characterized in that, In step S1, the dynamic load balancing algorithm considers the importance differences of each work order, and obtains the hardware resource allocation result through comprehensive consideration by adding weight coefficients.

3. The test method according to claim 1, characterized in that, Step S3 of building a fault prediction model based on the LSTM neural network further includes: multi-dimensional data analysis and data classification and sorting. According to the historical test data and current production status of different production lines, use machine learning algorithms to predict the growth trend and abnormal patterns of data, which are used as the analysis basis for subsequent fault diagnosis and production optimization of the fault prediction model.

4. The test method according to claim 3, wherein The multi-dimensional data analysis method in step S3 includes the following steps: S31. Receive test data from different production lines, classify and sort them according to multiple dimensions including at least the time range and production line number, and generate an interactive chart, including drawing a trend chart, a distribution heat map, and providing a function to download a structured file containing the original data and metadata; S32. Based on the spatio-temporal correlation algorithm, support filtering test data according to multiple dimensions of time range, production line number, and hardware platform, and use a GPU accelerated rendering engine to make the interactive chart dynamically interactive.

5. The test method according to claim 4, wherein Step S31 of generating the interactive chart includes: S311. Add an interactive function to the interactive chart to view detailed test data and analysis results by clicking on a single data point; S312. Adopt a data chunk loading technology, and only render the test data within the current window range when dragging the time axis to ensure that the dynamic update response time is less than 200 ms.

6. The test method according to claim 1, wherein In step S3 of building a fault prediction model based on the LSTM neural network, the operation of optimizing the BIOS parameter template of the hardware platform includes: adjusting key parameters including CPU frequency, voltage, and cache settings through the BIOS programming interface according to the actual hardware characteristics and test requirements to improve the test accuracy and efficiency.

7. The test method according to claim 1, wherein Step S4 of automatically binding the test results, SPD serial numbers, and workstation ID numbers includes: through a preset API interface, automatically obtain the SPD serial number and workstation ID number at the same time when the test results are generated, and accurately match and bind the above three to ensure the accuracy and traceability of data transmission.

8. The test method according to any one of claims 1-7, characterized in that, The test method has a hierarchical permission management function, and further includes the following steps: S2. Use fingerprint recognition and dynamic tokens to double-verify the user identity to control access rights, define role-production line-data permission groups based on the RBAC model, and record operation behaviors in the blockchain log to achieve tamper-proof traceability; S5. When unauthorized operations are detected, send an alarm message to the MES system, freeze the account access token, and trigger hardware-level isolation.

9. A test system for a semiconductor memory test production line, characterized in that, Including: A modular test configuration unit that supports single-line independent configuration and workshop-level expansion driven by a PXE server, performs hardware virtualization and dynamic load balancing, is used to parallel process multiple work orders on a volatile memory test production line for parallel testing, and dynamically adjusts test parameters according to the requirements of each work order; A multi-dimensional data analysis platform that integrates spatio-temporal correlation algorithms and a GPU accelerated rendering engine, can receive test data from different production lines, classifies and organizes them according to multiple dimensions, generates interactive charts, and realizes real-time visualization of fault modes; A permission management module that combines biometrics and the RBAC model to achieve triple binding of role-production line-data and two-factor authentication, double-verifies the user identity through fingerprint recognition or facial recognition technology and a dynamic token to control access permissions, and records operation behaviors to the blockchain log to achieve tamper-proof traceability; An adaptive interface device for SPD serial number encapsulation and MES system communication, asynchronously uploading the SPD serial number, and ensuring highly reliable transmission of massive data through an asynchronous message queue.

10. The test system according to claim 9, wherein The modular test configuration unit uses a Kubernetes cluster to manage hardware resource virtualization, deploys DRAM test firmware images through a PXE server, and supports dynamically allocating CPU cores and memory channels according to work order requirements; the adaptive interface device realizes two-way communication with the MES system based on the gRPC protocol, and attaches the SPD serial number hash value and workstation ID digital signature when encapsulating test results; the modular test configuration unit further includes a sensor array for collecting temperature and current, and feeding back to the central processor to update the test variable values including voltage and frequency; The multi-dimensional data analysis platform further includes an anomaly detection unit that deeply analyzes test data based on advanced algorithms, accurately locates potential fault sources, and combines granular-level fault tracing technology to quickly lock the corresponding DIMM slot and related number information; The permission management module includes a two-factor authentication interface under the LSTM neural network, integrating one or more biometric identification mechanisms of a fingerprint recognition subsystem and a facial recognition subsystem and a dynamic token generator, and the two work together to verify the user identity. The administrator intuitively defines permission rules through a visual Web interface to ensure that operators can only access the production lines and data within the authorized scope; The adaptive interface device realizes automatic binding of test results, SPD serial numbers, and workstation ID numbers through a preset API interface and completes closed-loop management; The test system further includes a blockchain log audit component and a hardware-level isolation unit. The blockchain log audit component records operation behaviors to private chain nodes; the hardware-level isolation unit is linked with the MES system. Once unauthorized operations occur, it remotely disables the PCIe interface of the test motherboard and alarms to the MES system.

Citation Information

Patent Citations

  • Method, system, and device for parallel automated test, and readable storage medium

    CN108932196A

  • Smart verify for multi-state memories

    CN1720586A

Cited By

  • DDR3 memory chip integrated circuit test method and system based on particle swarm optimization

    CN120544646A

  • Test method of multi-site semiconductor device test equipment, equipment and medium

    CN121299399A

  • Multifunctional test point distribution method and system based on adaptive switch matrix

    CN121389830A