Chip design and manufacturing optimization system and method based on edge reinforcement learning

Through edge reinforcement learning, chip design and manufacturing optimization systems have solved the problems of insufficient adaptability and real-time control in the existing technology, efficient device status recognition and strategy optimization are achieved, and the adaptability and coordination of chip manufacturing systems are improved, and suitable for large-scale complex chip manufacturing environments.

CN120337836AActive Publication Date: 2025-07-18GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510807314.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing chip manufacturing systems lack adaptive learning capabilities and are difficult to automatically adjust decision strategies according to environmental changes. Traditional centralized control systems cannot meet the real-time control needs of industrial sites, and there is a risk of design errors and performance overload.

Method used

The chip design and manufacturing optimization system based on edge reinforcement learning is adopted, and through the scene identification module, policy selection module, decision correction module and cloud update module, the precise collection and processing of the operating status of the equipment is achieved, low-latency policy loading and cross-node decision correction, and has self-stabilization and energy consumption management capabilities.

Benefits of technology

It improves the adaptability and coordination of chip manufacturing systems, reduces the probability of local optimality and strategy conflicts, and realizes millisecond-level response and self-learning capabilities, which are suitable for large-scale complex chip manufacturing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337836A_ABST
    Figure CN120337836A_ABST
Patent Text Reader

Abstract

The invention discloses a chip design and manufacturing optimization system and method based on edge reinforcement learning, and relates to the field of data analysis, and the system comprises a scene recognition module, a strategy selection module, a decision correction module and a cloud updating module. A fuzzy recognition and alarm mechanism is introduced to enhance the adaptability; inter-process linkage optimization is realized by means of a double-index strategy library and an influence gradient transmission mechanism, and the process collaboration is improved; through dynamic uploading of the state tetrad and cloud oscillation identification, the system can automatically enter cold-quiet period locking under strategy abnormity and energy consumption fluctuation, and abnormal diffusion is effectively inhibited. Meanwhile, a lightweight edge deployment and reinforcement learning mechanism enables the system to have high response, low communication and self-learning capabilities, and the system is suitable for a large-scale complex chip manufacturing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis, and specifically to a chip design and manufacturing optimization system and method based on edge reinforcement learning. Background Art

[0002] Semiconductor chip manufacturing is a highly complex process involving hundreds of precision steps, mainly divided into stages such as wafer preparation, pattern transfer, doping, thin film deposition, etching, and packaging and testing. Modern chip manufacturing may take months to complete. These processes require the use of extremely specialized equipment, including lithography machines, etching systems, chemical vapor deposition (CVD) equipment, and various metrology and measurement instruments. Traditional industrial control systems mostly adopt fixed rules or models and lack the ability of adaptive learning, making it difficult to automatically adjust decision-making strategies according to environmental changes. Existing reinforcement learning applications mainly focus on single processes or equipment and lack the collaborative learning ability among multiple agents.

[0003] The prior art, such as the invention patent with the publication number: CN118194790B, is a chip design method and a chip design system. The chip design method includes the following steps: reading register transfer level code data; identifying multiple register transfer level codes in the register transfer level code data to classify multiple first registers and multiple second registers corresponding to the multiple register transfer level codes, where the multiple first registers are not electrically connected to the interfaces, and the multiple second registers are electrically connected to the multiple interfaces; performing multi-bit register merging on the multiple first registers to generate at least one first multi-bit register; and performing multi-bit register merging on the multiple second registers according to the initial physical position information of the multiple second registers to generate at least one second multi-bit register.

[0004] The prior art, such as the invention patent with the publication number: CN114818553B, is a chip integrated design method, which includes: defining the hierarchical structure of each module of the chip in a template; setting the RTL file path of the module in the template, and at the same time setting the module name and whether it belongs to the module for which RTL code is to be generated; analyzing the RTL files of the modules for which RTL code has been generated to extract the port connection information and parameter information of the modules; receiving the connection information of the unconnected ports between the modules added by the user in the template and the instantiation parameter values of the modules; using a script tool to analyze the added template to generate the RTL code of the module for which RTL code is to be generated; generating the corresponding port connections for the modules for which port connection information has been defined, and instantiating the modules for which parameter values have been defined using the defined parameter values.

[0005] Based on the above solutions, it can be seen that the existing solutions mainly rely highly on manual maintenance, with risks of design errors or performance overload. In addition, traditional centralized control systems need to transmit all data to a central server for processing, which cannot meet the requirements of real-time control in industrial fields. Moreover, they mostly adopt fixed rules or models. Therefore, a chip design system that can be flexibly controlled and respond quickly is needed. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention provides a chip design and manufacturing optimization system and method based on edge reinforcement learning. To achieve the above objectives, the present invention is realized through the following technical solutions: A chip design and manufacturing optimization system based on edge reinforcement learning, including: A scenario recognition module, which is used to collect the original device operation data on the process equipment during chip manufacturing, and perform state feature extraction and scenario classification to obtain the scenario labels of the process equipment.

[0007] A policy selection module, which is used to, after the scenario recognition is completed, select the corresponding policy model from the local policy library as the adjustment policy model through the edge agent of the process equipment according to the scenario label.

[0008] A decision correction module, which is used to calculate the influence gradient of the adjustment policy model on the downstream process equipment through the edge agent of the process equipment, and at the same time, the edge agent of the downstream process equipment introduces this influence gradient as a penalty factor during the decision-making process for decision correction.

[0009] A cloud update module, which is used to continuously update the local state quadruple to the cloud through the edge agent of the process equipment. The cloud identifies whether there is policy oscillation and energy consumption anomaly. If so, it controls the process equipment to enter the cooling period locking mode.

[0010] A chip design and manufacturing optimization method based on edge reinforcement learning, including: Collect the original device operation data on the process equipment during chip manufacturing, and perform state feature extraction and scenario classification to obtain the scenario labels of the process equipment.

[0011] After the scenario recognition is completed, select the corresponding policy model from the local policy library as the adjustment policy model through the edge agent of the process equipment according to the scenario label.

[0012] Calculate the influence gradient of the adjustment policy model on the downstream process equipment through the edge agent of the process equipment, and at the same time, the edge agent of the downstream process equipment introduces this influence gradient as a penalty factor during the decision-making process for decision correction.

[0013] The edge agent of the process equipment continuously updates the local state quadruple to the cloud. The cloud identifies whether there is policy oscillation and abnormal energy consumption. If so, it controls the process equipment to enter the cooling period locking mode.

[0014] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects: (1) The present invention provides a chip design and manufacturing optimization system based on edge reinforcement learning. By deploying a state perception feature channel on the process equipment side, it realizes the accurate acquisition and structured processing of the original equipment operation data. By combining with the device model feature channel mapping library, it can extract state features with high discrimination, and uses a confidence scoring mechanism to quantitatively identify multiple candidate scenarios, effectively avoiding misidentification and weak label interference. When the confidence score does not meet the standard or there are multiple optimal labels, a fuzzy recognition and alarm mechanism is introduced to provide more robustness and fault tolerance for downstream decision-making, ensuring the adaptability and accuracy of the system to complex and changing working conditions.

[0015] (2) According to the identified scenario labels, the edge agent of the present invention establishes double-index anchor points in the local policy library to achieve the precise loading of the low-latency policy model; under fuzzy recognition, a multi-model fusion mechanism is introduced to improve generalization. At the same time, the system designs an "influence gradient" evaluation and compressed transmission mechanism for downstream process equipment, enabling cross-node decision correction based on feedback factors among process equipment, forming an edge linkage mode that takes into account local independence and system synergy. This mechanism effectively improves the process coordination degree between equipment and reduces the occurrence probability of local optimum and policy conflict problems.

[0016] (3) Through the continuous update of the state quadruple and the identification of policy oscillation in the cloud, the present invention realizes the dynamic tracking of the long-term operation state and policy fluctuation of the equipment. When detecting abnormal changes in the policy network or abnormal energy consumption, it gives an alarm and conducts remote intervention, automatically calculates the length of the cooling period and freezes the edge adjustment policy, thereby avoiding the spread of problems. It designs an emergency upload mechanism with channel optimization to ensure the timely and reliable reporting of abnormal states. This makes the system have excellent self-stabilization ability, energy consumption management ability and operation and maintenance controllability, and is especially suitable for chip manufacturing scenarios with large scale, multiple nodes and complex process chains.

[0017] (4) By deploying lightweight agents, the decision-making process can be directly completed near the data source, greatly reducing communication latency. This enables the system to meet the millisecond-level response requirements in chip manufacturing. Through reinforcement learning, the present invention continuously learns from experience and optimizes decision-making strategies, and can cope with the appearance of various defects and interferences, improving the adaptability of the system to process changes.

[0018] Of course, it is not necessary for any product implementing the present invention to achieve all the above advantages simultaneously. Description of the Drawings

[0019] Figure 1 It is a schematic diagram of the system modules of the present invention.

[0020] Figure 2 It is a schematic diagram of the method flow of the present invention.

[0021] Figure 3 It is a schematic diagram of the logic flow of the present invention.

[0022] Figure 4 It is a page diagram of the process list of the chip manufacturing operation and maintenance system involved in the embodiment of the present invention.

[0023] Figure 5 It is a page diagram of the real-time monitoring of the chip manufacturing operation and maintenance system involved in the embodiment of the present invention.

[0024] Figure 6 It is a page diagram of the scene recognition of the chip manufacturing operation and maintenance system involved in the embodiment of the present invention.

[0025] Figure 7 It is a page diagram of the policy management of the chip manufacturing operation and maintenance system involved in the embodiment of the present invention.

[0026] Figure 8 It is a page diagram of the cool-down period control of the chip manufacturing operation and maintenance system involved in the embodiment of the present invention. Detailed Embodiments

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inner", "periphery", etc. indicating orientation or positional relationships are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention.

[0029] Please refer to Figure 1 As shown, the embodiment of the present invention provides a chip design and manufacturing optimization system based on edge reinforcement learning. At the same time, refer to Figure 3The following is a schematic diagram of the logic flow involved in the embodiments of the present invention. Its overall process starts with device data collection and finally decides whether to enter the cooling-off period state based on whether there are significant anomalies, thereby ensuring the stability of device operation and the reliability of the policy model. It includes: In the embodiments of the present invention, a display page of a chip manufacturing operation and maintenance system is provided, such as Figure 4 The following is a flowchart page diagram of the chip manufacturing operation and maintenance system involved in the embodiments of the present invention. It is used for managers to overview the entire chip manufacturing process and provide corresponding interactions for operations.

[0030] Such as Figure 5 The following is a real-time monitoring page diagram of the chip manufacturing operation and maintenance system involved in the embodiments of the present invention. It is used to monitor relevant information of the current manufacturing process, including raw device operation data and scenario confidence scores, etc., to facilitate managers to monitor the current manufacturing process in real time.

[0031] Such as Figure 6 The following is a scenario recognition page of the chip manufacturing operation and maintenance system involved in the embodiments of the present invention. It is used to view the scenario recognition process of the edge intelligent agent, including information such as feature extraction of data.

[0032] A scenario recognition module is used to collect raw device operation data on the process equipment of chip manufacturing, and perform state feature extraction and scenario classification to obtain the scenario labels of the process equipment.

[0033] Collect raw device operation data on the process equipment of chip manufacturing, and perform state feature extraction and scenario classification to obtain the scenario labels of the process equipment. The specific process is as follows: The raw device operation data includes process parameter data, electrical signal data, and environmental parameter data.

[0034] It should be noted that in the embodiments of the present invention, since this system is applied to the full-link process in the semiconductor chip manufacturing field, including wafer preparation, lithography, etching, thin film deposition, doping, metallization, and packaging and testing, there are various types of process equipment. Therefore, when the present invention embodiments involve obtaining and processing parameters related to the equipment, the parameters of the lithography equipment will be used as an example for elaboration.

[0035] The raw device operation data includes process parameter data, electrical signal data, and environmental parameter data. Among them, the process parameter data is obtained by reading through the built-in control system of the device in the embodiments of the present invention, including lithography light intensity, lithography time, and light source wavelength; the electrical signal data is monitored by the built-in electrical monitoring module, including supply voltage, supply current, and average power; the environmental parameter data is measured by the environmental monitoring sensor, including the internal temperature of the device, average vibration frequency, and average vibration intensity.

[0036] Extract the model of the extraction process equipment, map and match it with the mapping set of the equipment model - state perception feature channels pre - stored in the database to obtain the state perception feature channels of the process equipment, and input the original equipment operation data into the state perception feature channels of the process equipment to extract state features to obtain the respective preliminary scenario labels of the process equipment, specifically including: It should be noted that the state perception feature channels involved in the embodiments of the present invention are a set of predefined algorithms and processing flows for different equipment models, and are used to extract key state indicators from the original equipment operation data.

[0037] The process of state feature extraction includes pre - processing the original equipment operation data and then extracting the indicators defined in the state perception feature channels, including process parameter indicators, electrical signal indicators, and environmental parameter indicators. Among them, in the embodiments of the present invention, the process parameter indicators include the lithography light intensity stability coefficient, lithography time deviation, and light source wavelength volatility; the electrical signal indicators include the current instantaneous mutation rate, average power stability, and voltage overshoot / undershoot rate; the environmental parameter indicators include the internal temperature stability coefficient, average vibration frequency principal component, and average vibration intensity anomaly index. These indicators are used to identify the equipment operation state and are stored in the database.

[0038] Calculate the statistical features, time - domain features, and frequency - domain features of the original equipment operation data. The statistical features include the mean, variance, maximum value, and minimum value; the time - domain features include the instantaneous change rate, change slope, and peak - valley difference; the frequency - domain features include the main frequency components and power spectral density. Combine them with the indicators defined in the state perception feature channels to obtain the data features of process parameters, electrical parameters, and environmental parameters. For example, if the mean and peak - valley values of the lithography light intensity conform to the mean and peak - valley values defined in the lithography light intensity stability coefficient, then there is a data feature of stable lithography light intensity in the process parameters.

[0039] Based on the auto - encoder, fuse the features of process parameters, electrical signals, and environmental parameters to form a comprehensive state feature vector. Use the multi - label neural network classifier, input the comprehensive state feature vector, and output the respective preliminary scenario labels of the process equipment. In the embodiments of the present invention, taking the lithography equipment as an example, the preliminary scenario labels include but are not limited to the stable operation scenario, light intensity unstable scenario, voltage overshoot interference scenario, and vibration interference scenario, etc.

[0040] Extract the equipment operation reference data of each preliminary scenario label, including process parameter reference data, electrical signal reference data, and environmental parameter reference data. Compare them with the original equipment operation data to obtain a set of equipment operation data deviation values, and after weighted coupling, obtain the confidence score values of each preliminary scenario label. In the embodiments of the present invention, taking the original lithography equipment operation data as a calculation example, it specifically includes: ; ; ; ; Among them, is the process parameter confidence score value of the th preliminary scenario label, is the electrical signal parameter confidence score value of the th preliminary scenario label, is the environmental parameter confidence score value of the th preliminary scenario label, is the confidence score value of the th preliminary scenario label, is the lithography light intensity, is the lithography time, is the light source wavelength, is the supply voltage, is the supply current, is the average power, is the internal temperature of the device, is the average vibration frequency, is the average vibration intensity, is the th lithography light intensity of the preliminary scenario label, is the th lithography time of the preliminary scenario label, is the th light source wavelength of the preliminary scenario label, is the th supply voltage of the preliminary scenario label, is the th supply current of the preliminary scenario label, is the th average power of the preliminary scenario label, is the th internal temperature of the device of the preliminary scenario label, is the th average vibration frequency of the preliminary scenario label, is the th average vibration intensity of the preliminary scenario label, is the lithography light intensity weighting factor, is the lithography time weighting factor, is the light source wavelength weighting factor, is the supply voltage weighting factor, is the supply current weighting factor, is the average power weighting factor, is the internal temperature of the device weighting factor, is the weighted factor of the average vibration frequency, is the weighted factor of the average vibration intensity, is the weighted factor of the confidence score value of the process parameter, is the weighted factor of the confidence score value of the electrical signal parameter, is the weighted factor of the confidence score value of the environmental parameter, is the label number of the preliminary scenario, , is the number of labels of the preliminary scenario.

[0041] It should be noted that in the embodiments of the present invention, a dynamic weight adjuster is configured to dynamically adjust each target weight according to the current chip manufacturing scenario and management requirements.

[0042] The weighted factor of the lithography light intensity is used to adjust the contribution degree of the lithography light intensity parameter in the comprehensive feature, reflecting the importance of this parameter for equipment status recognition; the weighted factor of the lithography time represents the weight of the lithography time parameter in the overall feature fusion, reflecting the relative strength of its influence on the process quality; the weighted factor of the light source wavelength is used to measure the importance degree of the light source wavelength parameter in the status analysis. The weighted factor of the supply voltage reflects the influence weight of the supply voltage parameter on the equipment operation stability, the weighted factor of the supply current measures the contribution of the current parameter to the electrical state feature, and the weighted factor of the average power represents the weight of the average power parameter in the electrical signal category. The weighted factor of the internal temperature of the equipment is used to adjust the influence of the temperature parameter on the environmental state recognition, the weighted factor of the average vibration frequency reflects the sensitivity of the vibration frequency to the mechanical state of the equipment, and the weighted factor of the average vibration intensity is used to measure the importance of the vibration intensity parameter in the environmental monitoring. The weighted factor of the confidence score value of the process parameter represents the weighting of the credibility of the entire process parameter category data, reflecting its data quality and stability; the weighted factor of the confidence score value of the electrical signal parameter is used to adjust the influence of the overall confidence of the electrical signal type parameters; the weighted factor of the confidence score value of the environmental parameter reflects the adjustment effect of the confidence level of the environmental parameter category data on the comprehensive feature fusion.

[0043] It should also be noted that in the embodiments of the present invention, taking a lithography device as an example, there is a close correlation among process parameter data, electrical signal data, and environmental parameter data, specifically including: Process parameter data such as lithography light intensity, lithography time, and light source wavelength directly determine the quality and accuracy of the lithography process, and the stability of these parameters is the key to achieving high-precision manufacturing. The supply voltage, supply current, and average power in the electrical signal data reflect the operating conditions of the device power supply system. Fluctuations in electrical parameters may directly affect the performance of the process parameters of the lithography device. For example, abnormal voltage or current may cause instability in lithography light intensity and light source wavelength. Environmental parameter data includes the internal temperature of the device, average vibration frequency, and average vibration intensity, which reflect the physical environment and mechanical state of the device. Changes in the environment will affect the electrical system and process through thermal effects or mechanical vibrations. For example, an increase in temperature may cause a drift in the light source wavelength, and abnormal vibrations may cause fluctuations in lithography light intensity. Therefore, these three types of parameters influence and restrict each other.

[0044] Compare the confidence score values of each preliminary scenario label with the pre-stored confidence score threshold in the database, and count the preliminary scenario labels whose confidence score values are greater than or equal to the confidence score threshold as confidence scenario labels, and select the confidence scenario label with the largest confidence score value from them as the scenario label of the process equipment.

[0045] If there are two or more confidence scenario labels with the largest confidence score value at the same time, it is determined that the demand execution fuzzy recognition processing mode is to be performed, and these confidence scenario labels are recorded as each fuzzy scenario label at the same time.

[0046] If the confidence score values of each preliminary scenario label are all less than the confidence score threshold, an alarm reminder is given.

[0047] As Figure 7 shown is the strategy management page diagram of the chip manufacturing operation and maintenance system involved in the embodiments of the present invention, which is used for administrators to view and manage strategies, including the current device strategy and strategy change records.

[0048] The strategy selection module is used to, after the scenario recognition is completed, select the corresponding strategy model from the local strategy library as the adjustment strategy model through the edge agent of the process equipment according to the scenario label.

[0049] Select the corresponding strategy model from the local strategy library as the adjustment strategy model. The specific process is as follows: After receiving the scenario label, the edge agent on the process equipment establishes a two-level index anchor point of the process equipment model - scenario label, and inputs it into the local strategy library for indexing, obtains the strategy model corresponding to the two-level index anchor point and loads it as the adjustment strategy model of the process equipment. The strategy model includes strategy network parameters and control adjustment parameters.

[0050] In the embodiment of the present invention, the secondary index anchor of the equipment model - scenario label is an index mechanism for quickly retrieving and loading local policy models. Its definition method is to splice the model identifier of the process equipment and the identified scenario label in a standardized manner to form a unique index identifier. For example, if the model of the lithography equipment is Litho - M3400 - Canon, and its current scenario label is abnormal light source power fluctuation with the scenario label identifier being S002, the formed secondary index anchor is Litho - M3400 - Canon#S002. This anchor is used to quickly locate the corresponding adjustment policy model in the local policy library and achieve dynamic loading.

[0051] In the embodiment of the present invention, taking the lithography equipment as an example, the system loads the corresponding adjustment policy model from the policy library through the index anchor Litho - M3400 - Canon#S002. The adjustment policy model includes a set of policy network parameters and a set of control adjustment parameters. The policy network consists of a three - layer feed - forward neural network. The input is the comprehensive state feature vector, and the output is the recommended ratio of adjustable parameters, including the exposure time adjustment ratio, the light source power correction ratio, and the light intensity stability filtering coefficient, etc. Inside the network, reasoning processing is carried out through the weight matrix and the activation function, and finally, the decision output is formed through the Softmax function. At the same time, the control adjustment parameters provide directly executable operation instructions, such as recommending to increase the exposure time by 0.12 seconds, increase the light source power by 8%, and set the filtering coefficient of the light intensity change to 0.92. Through the above mechanism, the edge agent can quickly and accurately adapt and adjust the state of the lithography equipment.

[0052] If it is determined to execute the fuzzy recognition processing mode for the demand, then the fuzzy recognition processing mode is executed for the process equipment. The specific process includes: establishing each fuzzy index anchor of the process equipment model - fuzzy scenario label, indexing it in the local policy library, obtaining each policy model corresponding to the fuzzy index anchor, loading each policy model and averaging each policy model to obtain the adjustment policy model of the process equipment.

[0053] In an embodiment of the present invention, when the process equipment is in the fuzzy recognition processing mode, the fuzzy index anchor of the equipment model - fuzzy scenario label constructed is a multi-label coupling index structure, and its definition method is to combine the equipment model with multiple fuzzy scenario labels with confidence levels in a certain order and uniformly encode them to form a unique index identifier. For example, for a lithography equipment with the model Litho-M3400-Canon, the current fuzzy scenario labels include abnormal light source power fluctuation, initial stage of light source aging, and frequent voltage fluctuation. The scenario label for the initial stage of light source aging is identified as S004, and the scenario label for frequent voltage fluctuation is identified as S007. Then, Litho-M3400-Canon#S002+S004+S007 is constructed as the fuzzy index anchor for retrieving multiple policy models that match this fuzzy state in the local policy library.

[0054] After loading multiple policy models, model averaging processing will be performed to form a unified adjustment policy model. This processing process includes two levels: at the policy network level, weighted averaging is performed on the neural network weight parameters of each policy model. For example, if the confidence levels corresponding to the three labels are set to 0.5, 0.3, and 0.2 respectively, then the weight matrix and bias term at the same structural position in its policy network are weighted and fused to form a new network parameter set; at the control adjustment parameter level, control parameters in different policy models, such as exposure time adjustment, light source correction ratio, temperature control adjustment value, etc., are also weighted and averaged. The finally constructed unified adjustment policy model has the ability to cooperate and respond under fuzzy states, can effectively take into account the adjustment requirements of each potential failure scenario, and improve the adaptive regulation accuracy and robustness of the system.

[0055] The decision correction module is used to calculate the influence gradient of the adjustment policy model on the downstream process equipment through the edge agent of the process equipment. At the same time, the edge agent of the downstream process equipment introduces this influence gradient as a penalty factor in the decision-making process for decision correction.

[0056] Calculating the influence gradient of the adjustment policy model on the downstream process equipment through the edge agent of the process equipment specifically includes: The edge agent of the process equipment calls the local proxy model, inputs the control adjustment parameters of the adjustment policy model, and the proxy model outputs the predicted influence value of the production quality parameters of the downstream process equipment after the application of the control adjustment parameters. After performing differential calculation on the predicted influence value, the influence gradient on the downstream process equipment is obtained.

[0057] It should be noted that the edge agent calls the locally deployed lightweight proxy model, which is a deep network model trained through historical production data and can simulate the parameter linkage relationship between this equipment and its downstream equipment. The input of the proxy model is the control adjustment parameters in the adjustment policy model.

[0058] In the specific calculation process, the edge agent inputs the above control adjustment parameters into the proxy model in vector form. After the proxy model outputs the control parameters and applies them to the current process equipment, the predicted values of several key production quality parameters caused in the downstream process equipment are obtained. In the embodiment of the present invention, the process equipment is a lithography equipment, and the downstream process equipment is an etching equipment. Then, the predicted values of the key production quality parameters include, but are not limited to, pattern transfer accuracy, etch depth uniformity, or film thickness deviation, etc.

[0059] To further quantify the sensitivity of these predicted values to the performance of the downstream equipment, at the output end, the rate of change of each predicted value with respect to the input control parameters is differentiated to obtain the corresponding influence gradient. The differential calculation is implemented using an automatic differentiation framework. For example, a small perturbation is applied to each control adjustment parameter separately using the finite difference method, and the change amplitude of the predicted output value is calculated and normalized to obtain the influence gradient. This influence gradient is used to measure the quality fluctuation degree brought by the current adjustment strategy to the subsequent process flow after execution. This mechanism helps to avoid the risk of overall process degradation caused by local optimization.

[0060] The edge agent of the process equipment compresses the influence gradient and sends it to the edge agent of the downstream process equipment through the edge communication channel.

[0061] The edge agent of the downstream process equipment introduces this influence gradient as a penalty factor in the decision-making process for decision correction, specifically including: After receiving the influence gradient, the edge agent of the downstream process equipment inputs the influence gradient into the mapping set of influence gradient - penalty factor pre-stored in the database for mapping and matching to obtain the penalty factor of the downstream process equipment, and introduces this penalty factor into the loss function of the policy network parameters of the adjustment strategy model of the downstream process equipment for decision correction.

[0062] In the embodiments of the present invention, after the edge agent of the downstream process equipment receives the influence gradient from the upstream process equipment, in order to achieve the collaborative regulation of the process chain and quality risk control, the following specific implementation steps are executed to complete the introduction of the penalty factor and the correction process of the policy model: The edge agent takes the received influence gradient as input and searches for a preset influence gradient - penalty factor mapping set in the local or collaborative database. This mapping set is formed by training a large number of historical regulation experiments and simulation data, and describes the influence degree of different influence gradient levels on the product quality stability of the downstream equipment. Through this mapping, a quantified penalty factor is obtained, which is used to punish the potential quality deviation brought by the upstream parameter change to the downstream. This penalty factor is introduced into the adjustment policy model of the current downstream equipment, specifically by embedding it into the loss function in the policy model. Taking the policy network as an example, the loss function is usually the expected return loss, action deviation loss in reinforcement learning, or the mean square error in supervised learning, etc. At this time, the edge agent modifies the loss function into a combined form with a penalty term, for example , where is the original loss term, is the penalty factor, is the penalty function for the influence gradient , and is the corrected loss term.

[0063] This mechanism of decision correction makes the policy network automatically tend to select regulation actions with less impact on the downstream equipment during the training or inference process, realizing flexible cooperation across equipment. It ensures the data-driven linkage regulation ability between the upstream and downstream equipment, and significantly improves the overall quality coordination and dynamic adaptability in the complex chip manufacturing process.

[0064] The cloud update module is used to continuously update the local state quadruple to the cloud through the edge agent of the process equipment. The cloud identifies whether there is policy oscillation and abnormal energy consumption. If so, it controls the process equipment to enter the cooling period lock mode.

[0065] The edge agent of the process equipment continuously updates the local state quadruple to the cloud, which specifically includes: The state quadruple includes the original equipment operation data, the adjustment policy model, the adjustment feedback data, and the influence gradient. The parameters in the original equipment operation data, the adjustment policy model, and the adjustment feedback data are numbered respectively to obtain the original equipment operation data parameter set, the adjustment policy model parameter set, and the adjustment feedback data parameter set.

[0066] In an embodiment of the present invention, taking a lithography device as an example, the adjusted feedback data represents the control effect parameters actually fed back by the process device after executing the policy model. Specifically, it includes the process index response delay and the feedback offset. The process index response delay includes the exposure time response delay, the light source power adjustment response delay, and the light intensity filtering adjustment response delay. The feedback offset includes the exposure time offset, the light source power offset, and the filtering stability offset. These parameters are measured and obtained through the local closed-loop control logic or the feedback interface. The exposure time response delay refers to the time required for the actual exposure duration adjustment to take effect after the policy is issued to the device exposure unit. The light source power adjustment response delay refers to the time required for the device light source output to reach the target power after the light source power adjustment instruction is issued. The light intensity filtering adjustment response delay refers to the time required for the light intensity signal collected by the sensor to stabilize at the new filtering state after adjusting the filtering coefficient. The exposure time offset is the difference between the target exposure time and the actual exposure time. The light source power offset is the difference between the target light source power and the actual output power. The filtering stability offset is the deviation between the light intensity fluctuation degree of the actual filtering result and the expected filtering result.

[0067] Compare the state quadruple with the preset parameter threshold set in the database. The parameter threshold set includes the original device operation data threshold set, the adjustment policy model threshold set, the adjusted feedback data threshold set, and the influence gradient threshold. The original device operation data threshold set includes the process parameter data threshold set, the electrical signal data threshold set, and the environmental parameter data threshold set.

[0068] After obtaining the deviation value, perform weighted coupling to obtain the state quadruple deviation evaluation value, specifically including: ; Among them, is the state quadruple deviation evaluation value, is the value of the th parameter in the original device operation data parameter set, is the threshold of the th parameter in the original device operation data parameter set, is the weighting factor of the th parameter in the original device operation data parameter set, satisfying , is the value of the th parameter in the adjustment policy model parameter set, is the threshold of the th parameter in the adjustment policy model parameter set, is the weighting factor of the th parameter in the adjustment policy model parameter set, satisfying , is the value of the The value of a parameter, is the threshold value of the th parameter in the adjusted feedback data parameter set, is the th parameter's weighting factor in the adjusted feedback data parameter set, satisfying , is the influence gradient, is the influence gradient threshold value, is the weighting factor of the original device operation data, is the weighting factor of the adjusted strategy model, is the weighting factor of the adjusted feedback data, is the weighting factor of the influence gradient, is the number of the parameter in the original device operation data parameter set, , is the total number of parameters in the original device operation data parameter set, fl is the number of the parameter in the adjusted strategy model parameter set, f , is the total number of parameters in the adjusted strategy model parameter set, is the number of the parameter in the adjusted feedback data parameter set, , is the total number of parameters in the adjusted feedback data parameter set.

[0069] It should be noted that the weighting factor of the original device operation data, the weighting factor of the adjusted strategy model, the weighting factor of the adjusted feedback data, and the weighting factor of the influence gradient satisfy in the embodiments of the present invention. When calculating the deviation evaluation value of the state quadruple, the weighting factor of the original device operation data is used to reflect the relative importance of the original operation data such as process parameters, electrical signals, and environmental parameters in the overall deviation evaluation, ensuring the accurate characterization of the actual operation state of the device; the weighting factor of the adjusted strategy model is used to measure the contribution degree of the strategy model parameters to the deviation evaluation, and by assigning higher weights to the key strategy parameters, the sensitivity to the strategy adjustment effect is improved; the weighting factor of the adjusted feedback data reflects the weight of the feedback information in the deviation evaluation, emphasizing the attention to the actual adjustment effect and feedback response, and is helpful to timely detect the anomalies or deviations in the strategy execution; the weighting factor of the influence gradient is used to adjust the weight of the influence gradient on the downstream device, reflecting the contribution of different influencing factors to the deviation of the overall system state, so as to support more comprehensive state evaluation and anomaly detection.

[0070] It should also be noted that in the embodiments of the present invention, taking a lithography device as an example, the original device operation data includes process parameter data, electrical signal data, and environmental parameter data, and the internal parameter correlation has been described above.

[0071] The adjustment strategy model includes a set of policy network parameters and a set of control adjustment parameters. The policy network parameters include the exposure time adjustment ratio, the light source power correction ratio, and the light intensity stability filtering coefficient. The control adjustment parameters include the exposure time adjustment value, the light source power adjustment value, and the filtering coefficient adjustment value. There is a certain correlation between these parameters. Specifically, in the embodiments of the present invention, the policy network parameters are adjustment direction indicators obtained through the learning and reasoning of edge agents, specifically including the exposure time adjustment ratio, the light source power correction ratio, and the light intensity stability filtering coefficient, etc. They are used to guide the device in which direction to adjust under the current working conditions to optimize the operating state or maintain system stability. The control adjustment parameters are the specific numerical execution results of these policy direction indicators, including the exposure time adjustment value, the light source power adjustment value, and the filtering coefficient adjustment value. They are calculated based on the actual operating data of the current device (such as the current exposure time, the current light source power, etc.) and the ratio parameters given by the policy network, and are used to drive the actual hardware to execute the adjustment instruction. Specifically, the exposure time adjustment value is obtained by multiplying the current exposure time by the exposure time adjustment ratio. For example, if the current exposure time is 5 seconds and the adjustment ratio is +0.12, the adjustment value is +0.6 seconds; similarly, the light source power adjustment value is calculated based on the current light source power and the correction ratio; the filtering coefficient adjustment value directly participates in the signal processing link and is used to control the response sensitivity of the device to light intensity changes. This conversion process from policy ratio to execution value reflects a closed-loop structure of policy guidance - control execution.

[0072] In the embodiments of the present invention, taking a lithography device as an example, the adjustment feedback data includes the process index response delay and the feedback offset. There is a close correlation between the process index response delay and the feedback offset, manifested as causality, synergy, and dynamic coupling. Generally speaking, the greater the response delay, the higher the possibility and degree of feedback offset. This is because the device may enter the next round of policy adjustment before the response is completed, resulting in over-adjustment or lag compensation. In addition, if the response delay shows fluctuations in different cycles, this instability will directly amplify the fluctuations of the feedback offset, making the policy model frequently corrected but difficult to converge stably, forming a policy oscillation. Therefore, the response delay and the feedback offset are not only two independent feedback data indicators, but also important factors that interact with each other and jointly determine the execution effect of the adjustment strategy and the energy consumption stability of the device in a dynamic system.

[0073] There is also a certain correlation among several major categories of data, namely, the original device operation data, the adjustment strategy model, the adjustment feedback data, and the influence gradient. In the embodiments of the present invention, the original device operation data serves as the perception basis of the entire system and mainly includes process parameters, electrical signals, and environmental data. These data reflect the current operation state of the device and the process execution environment, and are the input basis for the adjustment strategy model. Based on the input of the original data, the adjustment strategy model generates a set of control adjustment parameters and policy network parameters, specifically determining adjustment actions such as exposure time, light source power, and filtering coefficient. This model seeks the optimal control path during continuous updates to adapt to environmental fluctuations and changes in process objectives. After the parameters output by the policy model act on the device, the system will collect adjustment feedback data in real time, mainly including the response delay and offset of process indicators. These feedback data reflect the actual execution effect of the current policy model and provide direct evidence for the correction and optimization of the policy model. Through the paired analysis between the feedback data and the original data, the system can further evaluate whether the policy adjustment has achieved the expected goal. Based on the feedback, the system further calculates the influence gradient, that is, the sensitivity of the adjustment of a certain policy parameter to the performance of the downstream process equipment. This gradient information is not only used to evaluate the influence of the current policy but also to control the policy convergence speed and prevent policy oscillation. Finally, these four types of data continuously interact and optimize in a closed-loop manner. The original device operation data drives the policy adjustment, the policy model acts on the device to form feedback data, the feedback data is analyzed to generate the influence gradient, and the influence gradient acts on the policy model again to form a new adjustment strategy.

[0074] If the deviation of the state quadruple from the evaluation value is greater than or equal to the preset deviation evaluation threshold in the database, an emergency upload is triggered. The state quadruple is encapsulated into a structured data packet and directly sent to the cloud after selecting the edge upload channel.

[0075] If the deviation of the state quadruple from the evaluation value is less than the deviation evaluation threshold, the edge agent of the process equipment encapsulates the state quadruple into a structured data packet and performs timed batch upload based on the preset upload interval period in the database.

[0076] The cloud identifies whether there is policy oscillation and energy consumption anomaly, specifically including: After receiving the structured data packet of the state quadruple of the process equipment, the cloud preferentially parses the structured data packet of the emergency upload, subtracts the deviation evaluation value of the state quadruple of the emergency upload from the deviation evaluation threshold to obtain the state quadruple deviation difference, and compares it with the pre-stored deviation difference threshold in the database. If the state quadruple deviation difference is greater than or equal to the deviation difference threshold, it is determined that the process equipment has a fluctuation anomaly and an alarm reminder is given.

[0077] If the deviation of the status quadruple is less than the fluctuation deviation threshold, it is determined that there is no abnormal fluctuation in the process equipment, and the urgently uploaded structured data packet is marked and sent back to the edge intelligent agent for local proxy model training.

[0078] After the cloud parses the structured data packet of the status quadruple uploaded regularly in batches, it performs time series alignment. When the cloud parses each structured data packet, it reads the uploaded timestamp and the unique device identifier carried in it, sorts all the uploaded data packets of the same device in ascending order of the timestamp, and constructs an ordered status sequence.

[0079] Obtain the change values of each data in the policy network parameters in each upload interval period, including: ; Among them, is the change value of the u-th data in the policy network parameters in the t-th upload interval period, is the value of the u-th data in the policy network parameters at the t-th upload interval period, is the value of the u-th data in the policy network parameters at the (t - 1)-th upload interval period, is the data number of the policy network parameters, , is the total number of data in the policy network parameters, is the upload interval period number, , is the total number of upload interval periods.

[0080] After being corrected with the corresponding unit weighting factor and coupled, the change amplitude of the policy network parameters in each upload interval period is obtained, specifically including: ; Among them, is the change amplitude of the policy network parameters in the t-th upload interval period, is the change value of the u-th data in the policy network parameters in the t-th upload interval period, is the unit weighting factor corresponding to the u-th data in the policy network parameters, satisfying , u is the data number of the policy network parameters, , is the total number of data in the policy network parameters, is the upload interval period number, , is the total number of upload interval periods.

[0081] Compare the change amplitude of the policy network parameters in each upload interval period with the threshold of the change amplitude of the policy network parameters pre-stored in the cloud database, and record the policy network parameters with the change amplitude of the policy network parameters greater than or equal to the threshold of the change amplitude of the policy network parameters as significantly changed parameters.

[0082] Count the upload interval periods with significantly changed parameters. If there are also significantly changed parameters in adjacent upload interval periods, they are recorded as consecutive periods. Obtain the number of consecutive periods. If the number of consecutive periods exceeds the consecutive period threshold, it is determined that there is a policy oscillation in the edge agent of the process equipment. Policy oscillation means that within a period of time, the parameters in the policy model of the process equipment are adjusted frequently and significantly, manifested as significant fluctuations in the policy network parameters in multiple consecutive upload periods. Its essence reflects that the edge agent fails to converge stably to a certain optimized policy, which may be caused by external interference, data anomalies, or the failure of the feedback mechanism, resulting in frequent policy updates without obvious effects. Policy oscillation usually means a decrease in system stability.

[0083] When there is a policy oscillation in the edge agent of the process equipment, retrieve the original equipment operation data in the state quadruple within the consecutive period, and calculate the energy consumption data of the process equipment at each upload interval period within the consecutive period based on the original equipment operation data in the Python + engineering library. In the embodiment of the present invention, the power consumption value is used as the energy consumption data.

[0084] Subtract the energy consumption data of adjacent upload interval periods in the consecutive period to obtain the energy consumption change value of each upload interval period in the consecutive period, and perform weighted coupling averaging to obtain the energy consumption data change amplitude value in the consecutive period, specifically including: ; Among them, is the energy consumption data change amplitude value in the consecutive period, is the energy consumption change value of the th upload interval period in the consecutive period, is the weight factor of the th upload interval period in the consecutive period, satisfying , is the upload interval period number in the consecutive period, , is the total number of upload interval periods in the consecutive period.

[0085] If the energy consumption data change amplitude value exceeds the normal change amplitude value preset by the equipment, it is determined that there is an energy consumption anomaly in the process equipment.

[0086] Control the process equipment to enter the cooling period locking mode. The specific processing conditions are: The difference between the number of consecutive periods and the consecutive period threshold is obtained as the consecutive period difference, and the difference between the energy consumption data change amplitude value and the normal change amplitude value is processed to obtain the energy consumption change difference.

[0087] The consecutive period difference is multiplied by the period number unit factor to eliminate the unit. After multiplying the energy consumption change difference by the energy consumption unit factor to eliminate the unit, they are added to obtain the cooling period length pointing factor. The cooling period length pointing factor is input into the mapping set of cooling period length pointing factor - cooling period length pre-stored in the cloud database for mapping and matching to obtain the cooling period length of the process equipment.

[0088] Based on the cooling period length, the process equipment is controlled to enter the cooling period locking mode. The cooling period locking mode is specifically that the adjustment strategy model of the edge agent remains frozen, the process equipment operates in the existing state and caches all data in the edge agent, and the cloud sends the process equipment number to the operation and maintenance management terminal for reminder.

[0089] As Figure 8 shown is the cooling period control page diagram of the chip manufacturing operation and maintenance system involved in the embodiment of the present invention, which is used for the administrator to view and manage the cooling period, including information such as the display of equipment in the cooling period and the operation control interaction panel.

[0090] It further includes an emergency upload module, which is used to select the edge upload channel when an emergency upload is triggered, specifically including: Obtain the channel real-time performance data of each edge upload channel, including general index data and extended index data. Extract the hard thresholds of the general index data from the database, and after comparison, screen out the edge upload channels that meet all the hard thresholds, which are recorded as general upload channels.

[0091] It should be noted that the general index data includes bandwidth occupancy rate, latency, packet loss rate, and throughput; the extended index data includes the median data transmission delay, the number of retransmissions, and the bandwidth fluctuation value, and these parameters are automatically collected and calculated by the network performance monitoring module built into the edge agent.

[0092] The extended index data of each general upload channel is compared with the set of extended index data standard values extracted from the database, including the median data transmission delay standard value, the number of retransmissions standard value, and the bandwidth fluctuation standard value, and then weighted and coupled to obtain the extended index value of each general upload channel, specifically including: ; Among them, is the extended index value of the th general upload channel, is the median data transmission delay of the th general upload channel, is the The number of retransmissions of a general upload channel, is the bandwidth fluctuation value of the th general upload channel, is the standard value of the data transmission delay median, is the standard value of the number of retransmissions, is the standard value of the bandwidth fluctuation, is the weighted factor of the data transmission delay median, is the weighted factor of the number of retransmissions, is the weighted factor of the bandwidth fluctuation, , is the total number of general upload channels.

[0093] It should be noted that the weighted factor of the data transmission delay median, the weighted factor of the number of retransmissions, and the weighted factor of the bandwidth fluctuation all have a value range between 0 and 1 and satisfy . The weighted factor of the data transmission delay median is a weight coefficient that reflects the degree of attention of the system to the stability of the channel transmission delay. The weighted factor of the number of retransmissions reflects the degree of importance of the system to the reliability of data transmission. This factor is used to measure the impact weight of retransmission behavior caused by network instability, interference, or packet loss. If the system has extremely high requirements for data integrity, this factor should be set higher. The weighted factor of the bandwidth fluctuation is used to characterize the sensitivity of the system to the stability of the bandwidth. When the task requires continuous large-bandwidth transmission, this weight should be increased to ensure that the evaluation process pays more attention to the impact of bandwidth fluctuation on the channel availability. These three weighted factors are generally preset according to the characteristics of the system operation tasks, and the sum of the three is normalized to 1, which is used as the basis for weighted fusion to calculate the channel expansion performance index. The final channel expansion index score can be obtained by summing the products of the three indicators and their corresponding weighted factors, reflecting the comprehensive performance level of the channel.

[0094] It should also be noted that there is a certain correlation among several parameters such as the median data transmission delay, the number of retransmissions, and the bandwidth fluctuation. The median data transmission delay reflects the typical delay level during the data packet transmission process and is an important parameter for evaluating the network response speed. When data packets are lost or in error during network transmission, the system needs to perform retransmission operations. The increase in the number of retransmissions will significantly increase the overall transmission delay, thereby increasing the median delay. Therefore, there is a direct positive correlation between the number of retransmissions and the median delay. Bandwidth fluctuation measures the stability of the available bandwidth of the channel. A large fluctuation indicates that the network may be congested or jittery during certain periods, resulting in unstable transmission of burst data. When the bandwidth is unstable, there may be a situation of sudden bandwidth drop, causing some data packets to be unable to be transmitted in time or lost, triggering the retransmission mechanism. Therefore, the expansion of bandwidth fluctuation often leads to an increase in the number of retransmissions. At the same time, bandwidth fluctuation will also cause inconsistent transmission rates, exacerbate the delay of data packets during queuing and waiting, and lead to an increase in the median delay. These three indicators form a dynamic feedback chain: the expansion of bandwidth fluctuation may lead to an increase in the number of retransmissions, which in turn increases the median data transmission delay. The three together reflect the volatility and uncertainty of the upload channel. Therefore, when evaluating the performance of the edge upload channel, these three indicators should be used as a coupled evaluation system and weighted coupling processing should be carried out to achieve a comprehensive determination of the reliability and stability of the channel.

[0095] Select the general upload channel with the maximum expansion index value as the emergency upload channel.

[0096] Such as Figure 2 As shown, in this embodiment, the present invention provides a method for optimizing chip design and manufacturing based on edge reinforcement learning, including: Collect the original device operation data on the process equipment of chip manufacturing, and perform state feature extraction and scenario classification to obtain the scenario labels of the process equipment.

[0097] After the scenario recognition is completed, the edge agent of the process equipment selects the corresponding policy model from the local policy library as the adjustment policy model according to the scenario label.

[0098] The edge agent of the process equipment calculates the influence gradient of the adjustment policy model on the downstream process equipment. At the same time, the edge agent of the downstream process equipment introduces this influence gradient as a penalty factor during the decision-making process for decision correction.

[0099] The edge agent of the process equipment continuously updates the local state quadruple to the cloud. The cloud identifies whether there is policy oscillation and abnormal energy consumption. If so, it controls the process equipment to enter the cooling period locking mode.

[0100] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0101] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. As long as it does not deviate from the structure of the present invention or exceed the scope defined by the present invention, it should fall within the protection scope of the present invention.

Claims

1. A chip design and manufacturing optimization system based on edge reinforcement learning, characterized in that, Including: A scene recognition module, which is used to collect the original device operation data on the process equipment in chip manufacturing, and perform state feature extraction and scene classification to obtain the scene label of the process equipment; A policy selection module, which is used to, after the scene recognition is completed, select the corresponding policy model from the local policy library as the adjustment policy model through the edge agent of the process equipment according to the scene label; A decision correction module, which is used to calculate the influence gradient of the adjustment policy model on the downstream process equipment through the edge agent of the process equipment, and at the same time, the edge agent of the downstream process equipment introduces this influence gradient as a penalty factor during the decision-making process for decision correction; A cloud update module, which is used to continuously update the local state quadruple to the cloud through the edge agent of the process equipment, and the cloud identifies whether there is policy oscillation and energy consumption anomaly. If so, it controls the process equipment to enter the cooling period locking mode.

2. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, wherein: The specific process of collecting the original device operation data on the process equipment in chip manufacturing, and performing state feature extraction and scene classification to obtain the scene label of the process equipment is as follows: The original device operation data includes process parameter data, electrical signal data, and environmental parameter data; Extract the model number of the process equipment, perform mapping matching with the mapping set of device model number - state perception feature channels pre-stored in the database to obtain the state perception feature channels of the process equipment, and input the original device operation data into the state perception feature channels of the process equipment for state feature extraction to obtain the preliminary scene labels of the process equipment; Extract the device operation reference data of each preliminary scene label, including process parameter reference data, electrical signal reference data, and environmental parameter reference data, compare it with the original device operation data to obtain a set of device operation data deviation values, and perform weighted coupling to obtain the confidence score value of each preliminary scene label; Compare the confidence score value of each preliminary scene label with the confidence score threshold pre-stored in the database, count the preliminary scene labels whose confidence score value is greater than or equal to the confidence score threshold as the confidence scene labels, and select the confidence scene label with the largest confidence score value as the scene label of the process equipment; If there are two or more confidence scene labels with the largest confidence score value at the same time, execute the fuzzy recognition processing mode, and record these confidence scene labels as each fuzzy scene label; If the confidence score value of each preliminary scene label is less than the confidence score threshold, an alarm reminder is given.

3. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, wherein: The specific process of selecting the corresponding policy model from the local policy library as the adjustment policy model is as follows: After receiving the scene label, the edge agent on the process equipment establishes a secondary index anchor point of device model number - scene label, and inputs it into the local policy library for indexing, obtains the policy model corresponding to the secondary index anchor point and loads it as the adjustment policy model of the process equipment. The policy model includes policy network parameters and control adjustment parameters; If the fuzzy recognition processing mode is executed, each fuzzy index anchor of the process equipment model - fuzzy scenario label is established and indexed in the local policy library to obtain each policy model corresponding to the fuzzy index anchor. Each policy model is loaded and averaged to obtain the adjustment policy model of the process equipment.

4. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, wherein: Calculating the influence gradient of the adjustment policy model on the downstream process equipment through the edge agent of the process equipment specifically includes: The edge agent of the process equipment calls the local proxy model, inputs the control adjustment parameters of the adjustment policy model, and the proxy model outputs the predicted influence value of the production quality parameters of the downstream process equipment after the application of the control adjustment parameters. After differential calculation of the predicted influence value, the influence gradient on the downstream process equipment is obtained; The edge agent of the process equipment compresses the influence gradient and sends it to the edge agent of the downstream process equipment through the edge communication channel.

5. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, characterized in that: When making a decision, the edge agent of the downstream process equipment introduces the influence gradient as a penalty factor to correct the decision, specifically including: After receiving the influence gradient, the edge agent of the downstream process equipment inputs the influence gradient into the mapping set of influence gradient - penalty factor pre - stored in the database for mapping and matching to obtain the penalty factor of the downstream process equipment, and introduces the penalty factor into the loss function of the policy network parameters of the adjustment policy model of the downstream process equipment for decision correction.

6. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, characterized in that: The edge agent of the process equipment continuously updates the local state quadruple to the cloud, specifically including: The state quadruple includes the original equipment operation data, the adjustment policy model, the adjustment feedback data, and the influence gradient; Compare the state quadruple with the preset parameter threshold set in the database. The parameter threshold set includes the original equipment operation data threshold set, the adjustment policy model threshold set, the adjustment feedback data threshold set, and the influence gradient threshold. After obtaining the deviation value, weighted coupling is performed to obtain the state quadruple deviation evaluation value. If the state quadruple deviation evaluation value is greater than or equal to the deviation evaluation threshold, an emergency upload is triggered; If the state quadruple deviation evaluation value is less than the deviation evaluation threshold, the edge agent of the process equipment encapsulates the state quadruple into a structured data packet and performs timed batch upload based on the preset upload interval period.

7. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, wherein: The cloud identifies whether there is policy oscillation and energy consumption anomaly, specifically including: After receiving the structured data packet of the state quadruple of the process equipment, the cloud preferentially parses the structured data packet of the emergency upload. Subtract the state quadruple deviation evaluation value of the emergency upload from the deviation evaluation threshold to obtain the state quadruple deviation difference, and compare it with the deviation difference threshold pre - stored in the database. If the state quadruple deviation difference is greater than or equal to the deviation difference threshold, it is determined that the process equipment has a fluctuation anomaly, and an alarm reminder is given; If the state quadruple deviation difference is less than the fluctuation deviation difference threshold, it is determined that the process equipment has no fluctuation anomaly. The structured data packet of the emergency upload is marked and sent back to the edge agent for local proxy model training; After structuring the data packet of the status quadruple for timed batch upload in the cloud, perform time series alignment, obtain the change values of each data in the policy network parameters in each upload interval period, and after correcting with the corresponding unit weighting factors, couple them to obtain the change amplitude of the policy network parameters in each upload interval period. Compare it with the threshold of the change amplitude of the policy network parameters stored in the cloud database, and record the policy network parameters with the change amplitude of the policy network parameters greater than or equal to the threshold of the change amplitude of the policy network parameters as significantly changed parameters; Count the upload interval periods with significantly changed parameters. If there are also significantly changed parameters in adjacent upload interval periods, record them as continuous periods, obtain the number of continuous periods. If the number of continuous periods exceeds the continuous period threshold, it is determined that there is a policy oscillation in the edge agent of the process equipment; When there is a policy oscillation in the edge agent of the process equipment, retrieve the original equipment operation data in the status quadruple within the continuous period. Based on the original equipment operation data, calculate the energy consumption data of the process equipment at each upload interval period within the continuous period, and calculate the difference from the energy consumption data of the adjacent upload interval period in the continuous period to obtain the change value of the energy consumption at each upload interval period in the continuous period. After weighted coupling averaging, obtain the change amplitude value of the energy consumption data in the continuous period. If the change amplitude value of the energy consumption data exceeds the normal change amplitude value preset by the equipment, it is determined that there is an energy consumption anomaly in the process equipment.

8. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 1, wherein: Controlling the process equipment to enter the cooling period locking mode, the specific processing conditions are: Subtract the number of continuous periods from the continuous period threshold to obtain the continuous period difference, and perform a difference operation on the change amplitude value of the energy consumption data and the normal change amplitude value to obtain the energy consumption change difference; Combine the continuous period difference and the energy consumption change difference with the period number unit factor and the energy consumption unit factor respectively and then add them to obtain the cooling period length pointing factor. Map the cooling period length pointing factor into the mapping set of the cooling period length pointing factor - cooling period length stored in the cloud database for mapping and matching to obtain the cooling period length of the process equipment; Control the process equipment to enter the cooling period locking mode based on the cooling period length. The cooling period locking mode is specifically that the adjustment strategy model of the edge agent remains frozen, the process equipment operates in the existing state and caches all data in the edge agent, and the cloud sends the process equipment number to the operation and maintenance management terminal for reminder.

9. The chip design and manufacturing optimization system based on edge reinforcement learning according to claim 6, wherein: It also includes an emergency upload module for selecting an edge upload channel when an emergency upload is triggered, specifically including: Obtain the real-time channel performance data of each edge upload channel, including general index data and extended index data. Extract the hard threshold of the general index data from the database, and after comparison, screen out the edge upload channels that meet all the hard thresholds and record them as general upload channels; Compare the extended index data of each general upload channel with the set of standard values of the extended index data extracted from the database, and after weighted coupling, obtain the extended index value of each general upload channel. Select the general upload channel with the largest extended index value as the emergency upload channel.

10. A method applied to the chip design and manufacturing optimization system based on edge reinforcement learning described in claims 1-9, characterized in that: Collect the original equipment operation data on the process equipment of chip manufacturing, and perform state feature extraction and scenario classification to obtain the scenario labels of the process equipment; After the scenario recognition is completed, the edge agent of the process equipment selects the corresponding policy model from the local policy library as the adjustment policy model according to the scenario label; The edge agent of the process equipment calculates the influence gradient of the adjustment policy model on the downstream process equipment, and at the same time, the edge agent of the downstream process equipment introduces this influence gradient as a penalty factor in the decision-making process for decision correction; The edge agent of the process equipment continuously updates the local state quadruple to the cloud, and the cloud identifies whether there is policy oscillation and energy consumption anomaly. If so, it controls the process equipment to enter the calm period locking mode.

Citation Information

Patent Citations

  • A chip integration design method

    CN114818553B

  • Chip design method and chip design system

    CN118194790B

  • Wafer yield prediction method based on deep learning model

    CN109636026A

  • Layout generation system and method for multi-chip assembly assembling line

    CN116484795A

  • Methods for risk-informed chip layout generation

    US7590968B1

Cited By

  • Semiconductor anomaly detection method and system based on physical causal relationship modeling

    CN120781269A

  • Semiconductor anomaly detection method and system based on physical causal relationship modeling

    CN120781269B

  • Preparation process optimization method and system and yoga clothing antibacterial breathable fabric

    CN120975299A

  • An industrial internet-based production data edge collaborative processing method

    CN122507527A

  • An industrial internet-based production data edge collaborative processing method

    CN122507527B