Chip package test production line performance control method based on reinforcement learning

By adopting reinforcement learning strategy models and causal rule graphs in the chip packaging and testing production line to generate frequency band reservation, forced disabling and quantum control instructions, the problem of response lag in existing technologies is solved, real-time control and efficient intervention of chip packaging and testing are achieved, and the stability and intelligence level of the production line are improved.

CN120631674BActive Publication Date: 2025-10-21弘润半导体(苏州)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511121948.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-10-21
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing reinforcement learning models lack effective modeling of time series dependencies and causal logic in chip packaging testing, resulting in the inability to achieve rapid intervention and closed-loop control, especially delayed response when facing sudden anomalies or high-risk conditions.

Method used

A chip packaging and testing production line performance control method based on reinforcement learning is adopted. By scanning the RF band occupancy data, a joint state vector is generated. The pre-trained reinforcement learning strategy model is used to trigger causal intervention, generate frequency band reservation, forced disablement and quantum control instructions, and combine the causal rule graph to perform real-time decision-making and incremental training to optimize the operational decision logic.

Benefits of technology

It realizes intelligent perception and causal reasoning of the chip packaging and testing production line, can identify high-risk conditions in real time, generate accurate operation instructions, improve the stability and efficiency of the testing process, and enhance the ability to respond quickly to abnormal conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631674B_ABST
    Figure CN120631674B_ABST
Patent Text Reader

Abstract

The application discloses a chip packaging test production line performance control method based on reinforcement learning, and relates to the technical field of quantum regulation and control, and comprises the following steps: a joint state vector is generated; the joint state vector is input into a pre-trained reinforcement learning strategy model; when a high-risk state is detected, a causal intervention flag is triggered, a frequency band reservation instruction, a forced disable instruction set and a quantum regulation instruction set are generated; the frequency band reservation instruction, the forced disable instruction set and the quantum regulation instruction set are transmitted to a radio frequency test machine and a chip packaging device through a communication interface, a low-load test channel is distributed, a high-risk channel is disabled, and temperature and pressure are adjusted; test data of chip function testing are collected, compliance of the test data is verified, and a qualified report is generated. Through the reinforcement learning strategy model and the causal rule graph, intelligent perception and causal reasoning of the running state of the chip packaging test production line are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of quantum control technology, and in particular to a chip packaging and testing production line performance control method based on reinforcement learning. Background Art

[0002] With the continuous evolution of semiconductor manufacturing processes, chip packaging and testing are becoming increasingly critical in the overall production process. Especially with the growing demand for high-density, high-frequency applications, radio frequency (RF) testers are playing an increasingly prominent role in chip functional verification. Currently, the industry generally uses methods based on fixed rules or statistical analysis to monitor physical parameters during the packaging process, using preset thresholds to determine whether to trigger intervention mechanisms. In recent years, artificial intelligence technology, particularly reinforcement learning algorithms, has demonstrated excellent adaptability and optimization capabilities in the field of complex system control and is gradually being introduced into the scheduling and resource management of chip test equipment. For example, research has applied deep reinforcement learning to the selection and allocation of RF test channels to improve test efficiency and reduce energy consumption.

[0003] However, existing methods still have significant limitations in multi-dimensional state perception and causal reasoning. Most reinforcement learning models rely on static feature inputs and lack effective modeling of time series dependencies and causal logic. This leads to delayed responses to sudden anomalies or high-risk conditions, making rapid intervention and closed-loop control impossible. For example, some improved reinforcement learning methods attempt to introduce temporal difference mechanisms or long-short-term memory networks to enhance the model's perception of time series and improve the overall response speed to dynamic environmental changes. However, this approach primarily relies on experience replay and offline training, lacks explicit modeling of causal relationships, and struggles to achieve efficient, real-time risk identification and closed-loop control in complex and changing packaging and testing environments. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a chip packaging and testing production line performance control method based on reinforcement learning to solve the problem of being unable to achieve rapid intervention and closed-loop regulation.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a chip packaging and testing production line performance control method based on reinforcement learning, which includes scanning the occupancy data of each radio frequency band to generate a joint state vector; the joint state vector includes an electrostatic level value, a frequency band number sequence, a frequency band occupancy sequence, and packaging parameters;

[0008] The joint state vector is input into the pre-trained reinforcement learning policy model. When a high-risk state is detected, the causal intervention flag is triggered to generate frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets.

[0009] Transmit frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets to RF testers and chip packaging equipment through the communication interface, allocate low-load test channels, disable high-risk channels, and adjust temperature and pressure. After the adjustments, start chip functional testing;

[0010] Collect test data from chip functional tests, verify the compliance of the test data, generate a qualified report, optimize the reinforcement learning strategy model through abnormal features, and complete the performance control of the chip packaging and testing production line.

[0011] As a preferred solution of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, the joint state vector is input into a pre-trained reinforcement learning strategy model to generate an operation decision vector to control the performance of the chip packaging and testing production line, and the node association strength of the causal rule graph is updated through incremental training of the packaging conflict log;

[0012] The operation decision vector includes a frequency band reservation instruction, a forced disable instruction set and a quantum control instruction set.

[0013] As a preferred solution of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, wherein: the node association strength of the causal rule graph is updated through incremental training of the packaging conflict log, which means that when the causal intervention flag and the packaging intervention flag are activated, the joint state vector, operation decision vector and quantum control instruction set in the packaging conflict log are sorted as input data of the reinforcement learning strategy model, and the association strength between the current electrostatic level value node, the current frequency band occupancy node, the current temperature node and the current pressure node of the causal rule graph is updated.

[0014] As a preferred embodiment of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, the packaging intervention flag is activated after detecting the status of the current temperature node and the current pressure node through a causal rule graph and determining that either the temperature or pressure node reaches a high-risk packaging state;

[0015] The encapsulation conflict log is a data set recorded by the reinforcement learning strategy model after the encapsulation intervention flag is activated.

[0016] As a preferred solution of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, the triggering causal intervention flag is a flag that is activated after detecting the status of the current electrostatic level value node and the current frequency band occupancy node through the causal rule graph, confirming that the current electrostatic level value node and the current frequency band occupancy node reach the predefined electrostatic level threshold and frequency band exceeding standard criterion.

[0017] As a preferred solution of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, wherein: the low-load test channel detects the frequency band occupancy node status through a causal rule graph and identifies the radio frequency test channel whose frequency band occupancy is lower than the load threshold;

[0018] The high-risk channel detects the status of the current frequency band occupancy rate node and the current static electricity level value node through the causal rule graph, and identifies the radio frequency test channel where either the frequency band occupancy rate or the static electricity level value exceeds the standard.

[0019] As a preferred solution of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, the temperature and pressure adjustment refers to transmitting the quantum control instruction set to the chip packaging equipment through the communication interface, using the gallium nitride quantum dot sensor array to detect the temperature and pressure status of the packaging equipment in real time, and automatically adjusting the bonding temperature and patch pressure of the chip packaging equipment to the standard range according to the temperature adjustment instructions and pressure adjustment instructions in the quantum control instruction set.

[0020] As a preferred solution of the chip packaging and testing production line performance control method based on reinforcement learning described in the present invention, the method of collecting test data of chip functional testing and verifying the compliance of the test data refers to collecting the event count and functional results of the chip functional testing through a radio frequency testing machine, analyzing whether the event count meets the predetermined standard and verifying whether the functional result meets the JEDEC standard to determine the qualification of the chip test.

[0021] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the chip packaging and testing production line performance control method based on reinforcement learning as described in the first aspect of the present invention is implemented.

[0022] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the chip packaging and testing production line performance control method based on reinforcement learning as described in the first aspect of the present invention.

[0023] The beneficial effects of the present invention are: through the reinforcement learning strategy model and causal rule map, intelligent perception and causal reasoning of the operating status of the chip packaging and testing production line are realized. The reinforcement learning strategy model can identify high-risk states in real time in complex and changing test environments, and generate frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets based on causal intervention logic, thereby realizing dynamic scheduling and closed-loop control of RF testers and chip packaging equipment. In addition, the reinforcement learning strategy model also supports an incremental training mechanism based on packaging conflict logs, continuously optimizing the node association strength in the decision logic and causal rule map, improving adaptability in actual scenarios such as process fluctuations and equipment aging, achieving rapid response and precise intervention to abnormal states, and enhancing the stability, efficiency, and intelligence level of the testing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 Flowchart of the chip packaging and testing production line performance control method based on reinforcement learning.

[0026] Figure 2 Generate flowcharts for reinforcement learning decisions and instructions.

[0027] Figure 3 Execute flow charts for equipment control and testing.

[0028] Figure 4 Flowchart for test validation and model optimization. DETAILED DESCRIPTION

[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0031] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0032] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides a chip packaging test production line performance control method based on reinforcement learning, comprising the following steps:

[0033] S1. Scan the occupancy data of each radio frequency band and generate a joint state vector.

[0034] Deploy a quantum entangled state frequency band probe array, place a GaN quantum dot sensor array on the surface of the production line equipment, and initialize the GaN quantum dot sensor array to the quantum entangled state;

[0035] A sweeping detection signal is emitted to the quantum entangled state frequency band probe array, and quantum correlation collapse events are monitored by single-photon detectors. The collapse timing and amplitude are recorded, and a frequency band response matrix containing the frequency band number and quantum correlation amplitude is output. The frequency band occupancy rate is generated based on the quantum correlation amplitude, indicating the occupancy level (percentage) of each RF band (referring to the RF band used by the RF tester during the chip testing process) and recorded in the frequency band response matrix. The correlation collapse event includes the change in the state of the photon pair, the collapse timing and amplitude, and the collapse data corresponding to the frequency band number is recorded.

[0036] The frequency band response matrix is ​​loaded into the three-dimensional coordinate space: the frequency band number axis is loaded with the frequency band number of the frequency band response matrix, the time attenuation axis is loaded with the second-level attenuation factor, and the quantum correlation strength axis is loaded with the quantum correlation amplitude of the frequency band response matrix; the three-dimensional coordinate space is constructed by assigning values ​​to the coordinate axes; the second-level attenuation factor is a property of the frequency band response matrix. In the quantum sensing scenario, the signal of the correlation collapse event decays with time, and the attenuation characteristics are recorded on a time scale of seconds to form a second-level attenuation factor; the purpose of generating the second-level attenuation factor is to provide a time attenuation characteristic for the frequency band topological field, and to integrate it into the frequency band occupancy sequence of the joint state vector to ensure that the operation decision vector can accurately control the frequency band allocation and risk channel disabling based on dynamic time information, and avoid the judgment ambiguity caused by the missing time dimension in subsequent steps.

[0037] Electric field sensors measure the surface charge density of production line equipment, outputting voltage signals that are converted to ESD levels through analog-to-digital conversion. Predefined weights are determined using the JEDEC standard (a semiconductor device specification that defines test methods and classifications for device sensitivity to electrostatic discharge). These weights are then superimposed on the quantum correlation intensity axis of the frequency band response matrix to generate a frequency band topological field. The predefined weights are: if the ESD level is <500V (low risk), the weight is 0.1; if the ESD level is between 500V and 2000V (medium risk), the weight is 0.5; and if the ESD level is >2000V (high risk), the weight is 1.0.

[0038] A GaN quantum dot sensor array is placed on the surface of the chip packaging equipment to monitor the bonding temperature and pressure during the packaging process and generate packaging parameters containing temperature and pressure. The packaging parameters are loaded into a three-dimensional coordinate space, with the temperature axis loading the temperature, the pressure axis loading the pressure, and the quantum correlation strength axis loading the quantum correlation amplitude of the packaging parameters. The packaging topology field is then output. The packaging topology field is then combined with the frequency band topology field to generate a space-time topological tensor.

[0039] A one-dimensional convolution kernel (size 3, stride 1, covering 3 band numbers) is applied to the band topology field along the band number. The distribution characteristics of the band number are extracted by sliding the one-dimensional convolution kernel to generate a band convolution feature containing the distribution characteristics. The band convolution feature is processed by mean pooling (selecting 2 adjacent windows along the band number, averaging the band convolution feature values ​​within the window, and compressing the band number dimension) to generate a band scalar set. A two-dimensional convolution kernel (size 3×3, stride 1, covering 3×3 temperature and pressure regions) is applied to the spatiotemporal topology tensor along the temperature axis and pressure axis. The distribution characteristics of temperature and pressure are extracted by sliding the two-dimensional convolution kernel to generate a temperature-pressure convolution feature containing the distribution characteristics. The temperature-pressure convolution feature is processed by mean pooling (selecting 2×2 adjacent windows along the temperature axis and pressure axis, averaging the convolution feature values ​​within the window, and compressing the temperature axis and pressure axis dimensions) to generate the packaging parameters.

[0040] Arrange the frequency band number sequence in ascending order according to the frequency band number, associate the frequency band scalar set to the corresponding frequency band number sequence, append the electrostatic level value to the end of the frequency band number sequence, the frequency band occupancy sequence and the packaging parameter, and output a joint state vector consisting of the frequency band number sequence, the frequency band occupancy sequence, the packaging parameter and the electrostatic level value.

[0041] The frequency band occupancy rate sequence is generated based on the frequency band occupancy rate and is arranged in ascending order according to the frequency band number.

[0042] S2. Input the joint state vector into the pre-trained reinforcement learning strategy model. When a high-risk state is detected, the causal intervention flag is triggered to generate frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets.

[0043] Construct a causal rule map, set static electricity level thresholds and frequency band exceeding standards criteria, and analyze the causal relationship between historical nodes in the causal rule map based on historical test data (including static electricity level value records, frequency band occupancy records, temperature records, pressure records, and test results) to generate a causal rule map with a time dimension. The historical test data is obtained based on production line operation records through electric field sensors, quantum entangled state frequency band probe arrays, and gallium nitride quantum dot sensor arrays.

[0044] Construct a causal rule graph, set the electrostatic level threshold and frequency band exceeding standard criterion, specifically: first, obtain the electrostatic level value, frequency band occupancy sequence and packaging parameters from the joint state vector as the input variables of the causal rule graph; then, define the electrostatic level value, frequency band occupancy sequence, temperature and pressure as the current electrostatic level value node, current frequency band occupancy node, current temperature node and current pressure node in the causal rule graph respectively, and each current node represents a key variable; according to the predefined weight, set the electrostatic level threshold to electrostatic level value>2000V, for example: if the electrostatic level value is 2500V, it is judged to be a high-risk state; according to the frequency band occupancy sequence, set the frequency band exceeding standard criterion to frequency band occupancy>80%, for example: if the occupancy of a certain frequency band is 85%, it is judged to be an exceeding standard state; the electrostatic level threshold and frequency band exceeding standard criterion are stored in the current electrostatic level value node and the current frequency band occupancy node respectively;

[0045] The causal relationship between historical nodes in the causal rule map is analyzed based on historical test data. Specifically, first, the electrostatic level value records, frequency band occupancy records, temperature records, pressure records and test results (success or failure) are extracted from the historical test data, and are defined as historical electrostatic level value nodes, historical frequency band occupancy nodes, historical temperature nodes, historical pressure nodes and historical test result nodes in the causal rule map respectively. Each historical node represents a key variable; correlation analysis is performed on the historical test data: the Pearson correlation coefficient is used to compare the correlation between the electrostatic level value records, frequency band occupancy records, temperature records, pressure records and test results to determine the historical electrostatic level. The correlation strength between the value node, the historical frequency band occupancy rate node, the historical temperature node and the historical pressure node and the historical test result node is analyzed; then, a time series analysis is performed on the historical test data: by observing the changing trend of the electrostatic level value record, the frequency band occupancy rate record, the temperature record and the pressure record over time, the time-dependent relationship of each record is identified, and the temporal correlation between the historical electrostatic level value node, the historical frequency band occupancy rate node, the historical temperature node and the historical pressure node and the historical test result node is determined; based on the correlation strength and temporal correlation, the directed edge of the causal rule graph is constructed, from the historical electrostatic level value node, the historical frequency band occupancy rate node, the historical temperature node or the historical pressure node to the historical The test result node establishes a directed edge to reflect the causal effect of each historical node variable on the test result; a time label is added to each directed edge of the causal rule map to record the causal effect delay time from the historical electrostatic level value node, the historical frequency band occupancy node, the historical temperature node or the historical pressure node to the historical test result node; finally, the directed edges and time labels of the historical electrostatic level value node, the historical frequency band occupancy node, the historical temperature node, the historical pressure node, the historical test result node, and the causal rule map are integrated to generate a causal rule map containing a time dimension, which is used to guide the real-time operation of the current electrostatic level value node, the current frequency band occupancy node, the current temperature node, and the current pressure node. Decision-making, triggering the operation decision vector; among them, the triggering operation decision vector is based on the electrostatic level threshold, frequency band exceeding standard criterion, temperature threshold and pressure threshold, and detects the state of the current node (current electrostatic level value node, current frequency band occupancy node, current temperature node, current pressure node) through the causal rule map to determine the activation intervention flag and decide whether to generate an operation decision vector and what kind of instructions to generate; generating the operation decision vector is a reinforcement learning model based on the trigger result and the joint state vector, and the output includes the frequency band reservation instruction, forced disable instruction set and quantum control instruction set. The trigger is the cause and the generation is the result, and the two are closely related through the decision logic of the causal rule map.

[0046] The joint state vector is input into the pre-trained reinforcement learning policy model to generate the action decision vector, as follows:

[0047] If the current static level value node reaches the static level threshold and the current frequency band occupancy rate node exceeds the frequency band exceeding the criterion, the causal intervention flag is activated;

[0048] If the current temperature node or the current pressure node reaches a high-risk packaging state (temperature > 150°C or pressure > 1MPa, based on historical test data statistics), the packaging intervention flag is activated;

[0049] If the causal intervention flag is not activated, the frequency band reservation instruction (including frequency band allocation, bandwidth configuration, priority marking and time scheduling) in the operation decision vector is output;

[0050] When both the causal intervention flag and the package intervention flag are activated, based on the causal relationship of the causal rule graph, the operation instructions with excessive frequency band occupancy are blocked, and a mandatory disable instruction set is generated to prohibit the RF tester from using high-risk frequency bands;

[0051] When the packaging intervention flag is activated, the temperature and pressure status of the packaging equipment are detected, compared with the preset standard values, and a determination is made as to whether they exceed the standard range. A quantum control instruction set (including temperature increase instructions, temperature reduction instructions, pressure increase instructions, pressure reduction instructions, status monitoring instructions, and calibration instructions) is generated to optimize the temperature or pressure. The temperature standard value range is: 20°C to 30°C, and the pressure standard value range is: 0.8 to 1.2 standard atmospheres (atm), which are set based on the accuracy requirements of the packaging equipment operation.

[0052] When the package intervention flag is activated, the electrostatic level value, frequency band occupancy sequence, temperature and pressure in the joint state vector, the decision type and adjustment parameters in the operation decision vector, and the temperature or pressure adjustment instructions in the quantum control instruction set are recorded to generate a package conflict log; wherein, the decision type refers to the specific optimization action selected when the package intervention flag is activated, including: temperature optimization, pressure optimization, joint optimization (simultaneous adjustment of temperature and pressure) and maintaining the status quo; the adjustment parameter refers to the specific adjustment amount or target value corresponding to each decision type, including: temperature optimization parameters: target temperature (e.g., 25°C), adjustment amplitude (e.g., temperature increase by 2°C) and adjustment duration (e.g., 10 seconds); pressure optimization parameters: target pressure (e.g., 1.0atm), adjustment amplitude (e.g., pressure increase by 0.1atm) and adjustment duration (e.g., 5 seconds); joint optimization parameters: specifying the target value and adjustment amplitude of temperature and pressure at the same time (e.g., temperature increase by 1°C and pressure increase by 0.05atm);

[0053] When the encapsulation intervention flag is activated, the joint state vector, operation decision vector, and quantum control instruction set are recorded to generate an encapsulation conflict log;

[0054] Based on the encapsulation conflict log, incremental training is started to improve the accuracy and real-time performance of the causal rule graph in guiding the operation decision vector and quantum control instruction set. Specifically, when the accumulation of encapsulation conflict logs reaches a specified number (for example, 100), the joint state vector, operation decision vector, and quantum control instruction set in the encapsulation conflict log are sorted as input data for incremental training. The current electrostatic level value node, current frequency band occupancy node, current temperature node, current pressure node, and the association strength of the directed edges are read from the causal rule graph, and the directed edge association strength from the current electrostatic level value node to the current frequency band occupancy node in the causal rule graph is updated, as well as the directed edge association strength from the current temperature node and current pressure node to the operation decision vector.

[0055] The causal rule graph guides the real-time decision-making of chip packaging equipment through the current electrostatic level value node, current frequency band occupancy node, current temperature node, current pressure node and the correlation strength of directed edges, and generates an operation decision vector for controlling the operation of chip packaging equipment;

[0056] Retrieve the frequency band that triggers the causal intervention flag in the causal rule graph, send a forced disable instruction set to the RF tester, and generate an extended operation decision vector (including the frequency band disable instruction, decision type, and adjustment parameters);

[0057] Based on the packaging conflict log, the reinforcement learning strategy model is optimized, and the expanded operation decision vector is adjusted to improve the control accuracy of the temperature, pressure and frequency band occupancy of the chip packaging equipment; specifically, the temperature and pressure are obtained from the gallium nitride quantum dot sensor array, the frequency band occupancy sequence is obtained from the quantum entangled state frequency band probe array, the correlation strength of the causal rule map is updated, and the final operation decision vector (including frequency band reservation instructions, forced disable instruction set and quantum control instruction set) is generated.

[0058] S3. Transmit the frequency band reservation instruction, forced disable instruction set, and quantum control instruction set to the RF tester and chip packaging equipment through the communication interface, allocate low-load test channels, prohibit high-risk channels, and adjust temperature and pressure. After adjustment, start the chip functional test.

[0059] The frequency band reservation instruction and the forced disabling instruction set are transmitted to the RF tester through the communication interface; the frequency band reservation instruction allocates low-load test channels based on the channels in the frequency band occupancy sequence whose frequency band occupancy is lower than the load threshold (such as 0-80%), ensuring that the RF tester operates in a stable test environment; the forced disabling instruction set configures the RF tester through the communication interface to prohibit the use of high-risk test channels based on the high-risk status of excessive frequency band occupancy or electrostatic level value greater than 2000V identified by the causal rule map, to prevent electrical interference or equipment failure during the test; among which, the load threshold is set based on the current frequency band occupancy node status based on the causal rule map, and is specifically determined by the frequency band occupancy sequence analysis to identify low-load test channels and ensure that the RF tester operates in a stable test environment;

[0060] The quantum control instruction set is transmitted to the chip packaging equipment through the communication interface. The quantum control instruction set includes temperature adjustment instructions and pressure adjustment instructions, which drive the gallium nitride quantum dot sensor array to adjust the bonding temperature and patch pressure of the chip packaging equipment in real time.

[0061] The causal rule graph identifies the low-risk state of the RF tester through the current frequency band occupancy node (frequency band occupancy <80%) and the current static electricity level value node (<2000V). It then selects the corresponding RF tester as the target RF tester, activates the communication port of the target RF tester through the communication interface, establishes a connection with the low-load test channel, and restricts non-target devices from accessing the high-load test channel to avoid resource competition and scheduling conflicts.

[0062] A test start signal is sent to the target RF tester through the communication interface to load the preset chip functional test, including logic function test, timing test and power consumption test, to start the chip test; wherein, the preset chip functional test is set based on the chip design specifications and test requirements; the logic function test verifies whether the logic function of the chip is correct and checks whether the chip output meets the design specifications. For example, the logic gates (such as AND and OR) of the digital circuit or the instruction execution of the chip's built-in microcontroller are tested to ensure that there are no logical errors; the timing test checks whether the chip's test signal timing meets the requirements at the specified clock frequency, and measures the test signal propagation delay, setup time and hold time, for example, to verify whether the chip can synchronize correctly during high-speed operation to avoid timing violations (such as data loss); the power consumption test measures the power consumption of the chip in different working modes (such as idle and full load) to verify whether it meets the design power consumption range. For example, static power consumption (when in standby) and dynamic power consumption (when running) are tested to ensure that the chip is energy-saving and not overheating;

[0063] During the test, the gallium nitride quantum dot sensor array collects the packaging parameters of the chip packaging equipment every second. If the temperature or pressure deviates from the target range of the quantum control instruction set, for example, the temperature rises to 32°C or the pressure drops to 0.7 standard atmospheres, the chip packaging equipment's bonding temperature and placement pressure are automatically adjusted according to the temperature adjustment instructions and pressure adjustment instructions of the quantum control instruction set. For example, the temperature adjustment instruction is used to increase the cooling power of the chip packaging equipment, or the pressure adjustment instruction is used to increase the pressure of the chip packaging equipment to the target value of the quantum control instruction set. This ensures process stability and records the adjustment log, including temperature and pressure adjustment records, in the data recording memory.

[0064] S4. Collect test data from chip functional tests, verify the compliance of the test data, generate a qualified report, optimize the reinforcement learning strategy model through abnormal features, and complete the performance control of the chip packaging and testing production line.

[0065] After completing the chip functional test, the RF tester collects test data, including event counts (number of errors during the test, such as the number of abnormal triggers of the test signal) and functional results (whether the test passed, such as the logic function verification result);

[0066] Analyze event count compliance using an RF tester (for example, the number of errors is less than the maximum allowable event count of 5 times as recorded in the test results of historical test data) and verify that the functional results comply with JEDEC standards.

[0067] If the event count is compliant and the functional results pass, a qualified report is generated and records the batch yield (e.g., 98%), test efficiency (e.g., 1,000 chips tested per hour), and the average packaging parameters of the chip packaging equipment (e.g., temperature 25°C and pressure 1.0 standard atmospheres), and stores them in the data recording memory for subsequent performance analysis;

[0068] If an anomaly is detected (event count exceeds the limit or functional result fails), a defect report is generated, abnormal features are extracted (for example, temperature exceeds the limit to 35°C and pressure fluctuates to 1.5 standard atmospheres, or the electrostatic level value is abnormal), and the current temperature node and current pressure node of the causal rule map are associated. The causal relationship between the abnormal features and the event count exceeds the limit or the functional result fails is analyzed through the causal rule map, and the impact of process parameters (temperature, pressure and electrostatic level) on chip test failure is determined. The analysis is specifically as follows: the abnormal features in the defect report are mapped to the current temperature node, current pressure node and current electrostatic level value node of the causal rule map, the Pearson correlation coefficient is used to calculate the correlation strength between the abnormal features and the test failure, the time series correlation between the occurrence time of the abnormal features and the test failure is checked, and the impact of the packaging parameters and electrostatic level values ​​of the chip packaging equipment on the chip test failure is inferred by combining the directed edges of the causal rule map. Quantified results (such as the probability of failure due to temperature exceeding the limit) are generated and stored in the defect report to provide a basis for optimizing the reinforcement learning strategy model.

[0069] The formula for the Pearson correlation coefficient is:

[0070] ;

[0071] in, is the Pearson correlation coefficient, ranging from [-1, 1], where 1 represents a strong positive correlation, -1 represents a strong negative correlation, and 0 represents no correlation. is the number of test data, that is, the total number of tests in the analysis, For the Abnormal characteristic values ​​of the test (such as temperature 35°C), is the mean of the abnormal eigenvalues, For the Test failure metrics for each test (such as event counts or 0 / 1 function results), is the mean of the test failure indicators, is the index of the test, indicating the The test number, used to traverse the data points of each test;

[0072] It should be noted that this formula is a standard method in statistics, proposed by Karl Pearson in the late 19th century. It is widely used in data analysis and is a well-known mathematical formula;

[0073] Abnormal features in defect reports are input through the communication interface. Combined with the current ESD level value node, current temperature node, and current pressure node of the causal rule graph, the system uses the test results recorded in historical test data as a reward function to adjust the parameters of the reinforcement learning strategy model. This optimizes the ESD level threshold (for example, from 2000V to 1800V) and the temperature and pressure risk ranges (for example, from >30°C to >28°C) set in the causal rule graph. This improves the reinforcement learning strategy model's adaptability to environmental changes on the chip packaging and test production line, generates more accurate operation decision vectors and quantum control instruction sets, and ensures the accuracy of production line performance control.

[0074] The reinforcement learning strategy model is updated through incremental training. The model's parameters are adjusted using the chip packaging equipment's packaging parameters and test data, improving its adaptability to environmental changes on the chip packaging and test production line (such as equipment aging or batch differences). During the update period, the electrostatic level threshold set by the causal rule map is maintained to ensure the continuity of the chip packaging and test production line. Once the reinforcement learning strategy model is verified through analysis of the causal rule map (based on causal relationship evaluation recorded in historical test data), a seamless switch to the updated reinforcement learning strategy model is performed to improve control accuracy.

[0075] The updated reinforcement learning strategy model is further calibrated according to the current electrostatic level value node, current frequency band occupancy node, current temperature node and current pressure node of the causal rule graph to generate optimized operation decision vectors and quantum control instruction sets to ensure control accuracy.

[0076] This embodiment also provides a computer device, which is suitable for the chip packaging and testing production line performance control method based on reinforcement learning, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the chip packaging and testing production line performance control method based on reinforcement learning proposed in the above embodiment.

[0077] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.

[0078] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for implementing chip packaging and testing production line performance control based on reinforcement learning as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0079] In summary, the present invention realizes intelligent perception and causal reasoning of the operating status of the chip packaging and testing production line through the reinforcement learning strategy model and causal rule map. The reinforcement learning strategy model can identify high-risk states in real time in a complex and changeable test environment, and generate frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets based on causal intervention logic, thereby realizing dynamic scheduling and closed-loop control of RF testers and chip packaging equipment; in addition, the reinforcement learning strategy model also supports an incremental training mechanism based on packaging conflict logs, continuously optimizing the node association strength in the decision logic and causal rule map, improving adaptability in actual scenarios such as process fluctuations and equipment aging, achieving rapid response and precise intervention to abnormal states, and enhancing the stability, efficiency, and intelligence level of the testing process.

[0080] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A chip packaging and testing production line performance control method based on reinforcement learning, characterized by: include, Scanning occupancy data of each radio frequency band to generate a joint state vector; the joint state vector includes an electrostatic level value, a frequency band number sequence, a frequency band occupancy sequence, and packaging parameters; The joint state vector is input into the pre-trained reinforcement learning policy model. When a high-risk state is detected, the causal intervention flag is triggered to generate frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets. Transmit frequency band reservation instructions, forced disable instruction sets, and quantum control instruction sets to RF testers and chip packaging equipment through the communication interface, allocate low-load test channels, disable high-risk channels, and adjust temperature and pressure. After the adjustments, start chip functional testing; Collect test data from chip functional tests, verify the compliance of the test data, generate a qualified report, optimize the reinforcement learning strategy model through abnormal features, and complete the performance control of the chip packaging and testing production line; The joint state vector is input into the pre-trained reinforcement learning policy model to generate an operation decision vector, control the performance of the chip packaging test production line, and update the node association strength of the causal rule graph through incremental training of the packaging conflict log; The operation decision vector includes a frequency band reservation instruction, a forced disable instruction set, and a quantum control instruction set; The updating of the node association strength of the causal rule graph through incremental training of the encapsulated conflict log means that when the causal intervention flag and the encapsulated intervention flag are activated, the joint state vector, operation decision vector and quantum control instruction set in the encapsulated conflict log are sorted as input data of the reinforcement learning strategy model, and the association strength between the current electrostatic level value node, the current frequency band occupancy node, the current temperature node and the current pressure node of the causal rule graph is updated.

2. The chip packaging and testing production line performance control method based on reinforcement learning according to claim 1, characterized in that: The packaging intervention flag is activated after detecting the status of the current temperature node and the current pressure node through the causal rule graph and determining that either the temperature or pressure node reaches a high-risk packaging state; The encapsulation conflict log is a data set recorded by the reinforcement learning strategy model after the encapsulation intervention flag is activated.

3. The chip packaging and testing production line performance control method based on reinforcement learning according to claim 1, characterized in that: The trigger causal intervention flag is a flag that is activated after detecting the status of the current electrostatic level value node and the current frequency band occupancy rate node through the causal rule graph, confirming that the current electrostatic level value node and the current frequency band occupancy rate node reach the predefined electrostatic level threshold and frequency band exceeding standard criterion.

4. The chip packaging and testing production line performance control method based on reinforcement learning according to claim 1, characterized in that: The low-load test channel detects the frequency band occupancy node status through a causal rule graph and identifies a radio frequency test channel whose frequency band occupancy is lower than a load threshold; The high-risk channel detects the status of the current frequency band occupancy rate node and the current static electricity level value node through the causal rule graph, and identifies the radio frequency test channel where either the frequency band occupancy rate or the static electricity level value exceeds the standard.

5. The chip packaging and testing production line performance control method based on reinforcement learning according to claim 1, characterized in that: Adjusting the temperature and pressure refers to transmitting the quantum control instruction set to the chip packaging equipment through the communication interface, using the gallium nitride quantum dot sensor array to detect the temperature and pressure status of the packaging equipment in real time, and automatically adjusting the bonding temperature and patch pressure of the chip packaging equipment to the standard range according to the temperature adjustment instructions and pressure adjustment instructions in the quantum control instruction set.

6. The chip packaging and testing production line performance control method based on reinforcement learning according to claim 1, characterized in that: The collecting of test data of chip functional test and verifying compliance of test data refers to collecting event counts and functional results of chip functional test through RF test machine, analyzing whether event counts meet predetermined standards and verifying whether functional results meet JEDEC standards, and determining the qualification of chip test.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the chip packaging and testing production line performance control method based on reinforcement learning are implemented in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the chip packaging and testing production line performance control method based on reinforcement learning are implemented.

Citation Information

Patent Citations

  • Man-machine cooperation intelligent control system based on AIGC

    CN119940425A

  • Test screening method and system for enhancing electrical interference resistance of chip

    CN120370141A