Industrial robot semiconductor doping system based on multi-dimensional intelligent decision
Through the multi-dimensional intelligent decision-making method, combined with multi-sensor fusion, data acquisition, reinforcement learning and virtual simulation technology, the problems of accuracy, uniformity and efficiency in industrial robot doping are solved, and high-precision, uniformity and efficient doping effects are achieved.
Patent Information
- Application Number
- CN202510374843.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
AI Technical Summary
The existing industrial robot doping technology is difficult to meet the requirements of high precision, uniformity and efficiency, resulting in unstable performance, high waste rate and low production efficiency of semiconductor devices.
The multi-dimensional intelligent decision-making method is adopted, and the accuracy control, uniformity optimization and efficiency improvement of the doping process is achieved through multi-sensor fusion and real-time correction, multi-modal data acquisition and adaptive regulation, multi-objective optimization and reinforcement learning, digital twins and virtual simulation, and multi-robot collaboration and intelligent scheduling.
Significantly improve doping accuracy and uniformity, reduce waste rate, improve production efficiency, meet large-scale production needs, and optimize resource utilization.
Smart Images

Figure CN120233746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semiconductor manufacturing, and in particular to an industrial robot semiconductor doping system based on multi-dimensional intelligent decision-making. Background Art
[0002] Semiconductor doping is a key process that introduces specific impurity atoms into semiconductor materials to change their electrical properties and achieve specific functions. With the continuous development of semiconductor technology, the integration and performance requirements of chips continue to increase, and higher standards are put forward for the accuracy, uniformity and efficiency of semiconductor doping processes. Industrial robots have been widely used in the operation and control of semiconductor doping equipment due to their automation, high precision and programmability. However, there are still some problems that need to be solved in the existing industrial robot doping technology, which seriously restricts the further development of semiconductor doping technology.
[0003] Doping accuracy is difficult to meet high-precision requirements: In advanced semiconductor manufacturing processes, the accuracy requirements for doping concentration and position are extremely high. Traditional industrial robot doping systems are mainly operated based on preset parameters and simple positioning methods, and cannot accurately adapt to the differences in semiconductor material properties, microscopic morphology changes on the wafer surface, and environmental factors during the doping process. During the doping process, factors such as the flow fluctuation of the impurity source, the deviation of the robot positioning, and the drift of process parameters will cause a large deviation between the actual doping concentration and position and the design requirements, affecting the performance and reliability of semiconductor devices, increasing the scrap rate, and reducing production efficiency.
[0004] Doping uniformity is difficult to guarantee: Doping uniformity refers to the consistency of doping concentration on the entire surface of semiconductor materials or in a specific area. For large-sized semiconductor wafers or semiconductor devices with complex structures, it is crucial to ensure doping uniformity, which is directly related to the consistency and yield of device performance. The existing industrial robot doping system uses fixed doping process parameters and a single doping path, and cannot be dynamically adjusted according to the local characteristics and real-time changes of the wafer surface. Due to the structural and process limitations of the doping equipment, there are often inconsistent doping concentrations between the edge and center of the wafer and between different areas, resulting in large differences in semiconductor device performance, which seriously affects product quality and market competitiveness.
[0005] Low doping efficiency: Doping efficiency is an important factor affecting the production cost and production cycle of semiconductor manufacturing. At present, industrial robots mainly rely on manual experience to set doping parameters and operating procedures in the semiconductor doping process, which makes it difficult to optimize the doping process. At the same time, due to the limited operating speed and automation of doping equipment, as well as insufficient processing capabilities for complex processes, the doping process takes a long time, the production efficiency is low, and it cannot meet the needs of large-scale production. Summary of the invention
[0006] The present invention provides an industrial robot semiconductor doping system based on multi-dimensional intelligent decision-making, including:
[0007] An intelligent control algorithm module for doping accuracy based on multi-sensor fusion and real-time correction: Install various sensors such as an impurity source flow sensor, a position sensor, and a concentration sensor on the doping equipment of the industrial robot. Let the flow rate collected by the impurity source flow sensor be Q, the position information obtained by the position sensor be (x, y, z), and the doping concentration detected by the concentration sensor be C; Use multi-sensor fusion to construct a real-time state model S of the doping process through the formula S = f(Q, x, y, z, C), where f is a multi-sensor fusion function based on deep learning. Combine the real-time correction algorithm. According to the state model S, calculate the doping parameters of the industrial robot through the formula ΔP = g(S), where g is the real-time correction function, including the impurity source flow adjustment amount ΔQ, the doping position adjustment amount Δ(x, y, z), and the doping time adjustment amount Δt, and the adjustment amount ΔP, to achieve intelligent control of doping accuracy;
[0008] An optimization algorithm module for doping uniformity based on multi-modal data acquisition and adaptive regulation: Integrate various data acquisition devices such as a laser scanning microscope, an atomic force microscope, and a spectroscopic analyzer in the doping equipment. Let the surface topography information of the wafer obtained by the laser scanning microscope be M, the material property information collected by the atomic force microscope be N, and the doping concentration distribution information obtained by the spectroscopic analyzer be O; Use multi-modal data acquisition to construct a comprehensive information model I of the wafer surface through the formula I = h(M, N, O), where h is a multi-modal data acquisition fusion function based on deep learning; Combine the adaptive regulation algorithm. According to the comprehensive information model I and the real-time changes in the doping process, calculate the doping process parameters of the industrial robot through the formula ΔT = k(I), where k is the adaptive regulation function, such as the impurity source distribution mode adjustment amount ΔD, the doping path adjustment amount ΔR, and the doping speed adjustment amount ΔV, and the adjustment amount ΔT, to achieve the optimization of doping uniformity;
[0009] An algorithm module for improving doping efficiency based on multi-objective optimization and reinforcement learning: Take doping efficiency E, doping accuracy A, and doping uniformity U as the indicators of multi-objective optimization. Define the state space Sstate, including various parameters and state information such as impurity source flow rate, doping position, doping time, temperature, and pressure during the doping process; Define the action space Aaction, including various adjustment actions for doping parameters; Design a reward function R = q(E, A, U), where q is the reward calculation function, and comprehensively consider factors such as doping efficiency, doping accuracy, and doping uniformity to reward or punish the actions of the robot; Use the iterative formula of the deep Q network (DQN) of reinforcement learning
[0010] Q(Sstate, Aaction) ← Q(Sstate, Aaction) + α[R + γAaction′ max Q(Sstate′, Aaction′) - Q(Sstate, Aaction)], where α is the learning rate and γ is the discount factor or the relevant update formula of the Proximal Policy Optimization algorithm (PPO). Let the industrial robot learn the optimal doping parameter adjustment strategy through continuous interaction with the doping environment, improve the doping efficiency, and ensure the doping accuracy and uniformity at the same time;
[0011] Doping process pre-planning and optimization algorithm module based on digital twin and virtual simulation: Establish a digital twin model of the semiconductor doping system through modeling software and sensor data to achieve real-time interaction and synchronization between the physical entity and the virtual model. Use virtual simulation technology to pre-plan and simulate the semiconductor doping process. Let the simulation result obtained according to the process parameters Pparam and the process flow Fprocess during the simulation be Rsim. Optimize and adjust the parameters and process flow of the semiconductor doping process through the formulas ΔPparam = r(Rsim), where r is the optimization function, and ΔFprocess = s(Rsim), where s is the optimization function. At the same time, collect the historical operation data and fault data of the equipment, establish an equipment status evaluation model, and use the digital twin model to monitor and analyze the operation status of the doping equipment in real time;
[0012] Large-scale doping production optimization algorithm module based on multi-robot cooperation and intelligent scheduling: Regard multiple industrial robots in the semiconductor doping workshop as a multi-agent system, define the state space Ss, including the positions xr, yr of the robots, the task progress ptask, the working status swork, and the environmental information E of the workshop, that is
[0013] Ss = (xr, yr, ptask, swork, E); Define the action space Aa, including various motion actions mmove and task assignment decisions dassign of the robots, that is Aa = (mmove, dassign), and design the reward function
[0014] Rew = t(eefficiency, qquality, uutilization), where t is the reward calculation function, eefficiency is the doping production efficiency, qquality is the task completion quality, and uutilization is the equipment utilization rate); Use the reinforcement learning algorithm Deep Q-Network (DQN) or Proximal Policy Optimization algorithm (PPO) to let the robot learn the optimal cooperative doping production scheduling strategy through continuous interaction with the environment, achieve efficient cooperative work among multiple robots, and improve the overall efficiency and resource utilization rate of large-scale doping production.
[0015] Furthermore, in the intelligent control algorithm module for doping accuracy based on multi-sensor fusion and real-time correction, the multi-sensor data fusion adopts an algorithm based on deep learning to fuse the data collected by different sensors and construct a real-time state model of the doping process.
[0016] Furthermore, in the doping uniformity optimization algorithm module based on multi-modal data acquisition and adaptive regulation, the multi-modal data acquisition technology adopts an algorithm based on deep learning to fuse the data collected by different devices and construct a comprehensive information model of the wafer surface.
[0017] Furthermore, in the doping efficiency improvement algorithm module based on multi-objective optimization and reinforcement learning, the state space includes various parameters and state information in the doping process, the action space includes various adjustment actions for doping parameters, and the reward function comprehensively considers factors such as doping efficiency, doping accuracy, and doping uniformity.
[0018] Furthermore, in the doping process pre-planning and optimization algorithm module based on digital twin and virtual simulation, the semiconductor doping process is pre-planned and simulated through virtual simulation technology, the process parameters and flow are optimized according to the simulation results, and the digital twin model is used to monitor and analyze the operating state of the equipment in real time.
[0019] Furthermore, in the large-scale doping production optimization algorithm module based on multi-robot cooperation and intelligent scheduling, the state space includes the positions, task progress, working states of the robots, and the environmental information of the workshop, the action space includes various motion actions of the robots and task allocation decisions, and the reward function takes doping production efficiency, equipment utilization rate, and task completion quality as reward indicators.
[0020] Furthermore, in the intelligent control algorithm module for doping accuracy based on multi-sensor fusion and real-time correction, the sensors collect data in real time at a set frequency and transmit it to the data processing unit, and the data processing unit uses an algorithm based on deep learning to fuse and analyze the data.
[0021] Furthermore, in the doping uniformity optimization algorithm module based on multi-modal data acquisition and adaptive regulation, the data acquisition devices collect data in real time at a set frequency and transmit it to the data processing center, and the data processing center uses an algorithm based on deep learning to fuse and analyze the data.
[0022] Furthermore, in the doping efficiency improvement algorithm module based on multi-objective optimization and reinforcement learning, a large amount of doping experiment data is collected, the reinforcement learning algorithm is used to train and optimize the reinforcement learning model, and the industrial robot selects actions and adjusts strategies according to the real-time doping state.
[0023] Furthermore, the large-scale doping production optimization algorithm module based on multi-robot collaboration and intelligent scheduling defines multiple industrial robots in the semiconductor doping workshop as multiple agents, establishes a communication network among the multi-agents, collects historical operation data, trains and optimizes the reinforcement learning model using the reinforcement learning algorithm, and the agents select actions and adjust strategies according to real-time status information.
[0024] Beneficial effects:
[0025] Significantly improve doping accuracy: The intelligent control algorithm for doping accuracy based on multi-sensor fusion and real-time correction can comprehensively utilize the information of multiple sensors, adjust the doping parameters of industrial robots in real time, compensate and correct the errors in the doping process in real time, realize the intelligent control of doping accuracy, meet the requirements of high-precision doping, improve the performance and reliability of semiconductor devices, and reduce the rejection rate.
[0026] Effectively optimize doping uniformity: The optimization algorithm for doping uniformity based on multi-modal data acquisition and adaptive regulation can comprehensively utilize the information of multiple data acquisition devices, adjust the doping process parameters of industrial robots in real time, adapt to the local characteristics and real-time changes of the wafer surface, realize the optimization of doping uniformity, and improve the consistency and yield of semiconductor device performance.
[0027] Greatly improve doping efficiency: The algorithm for improving doping efficiency based on multi-objective optimization and reinforcement learning can realize the improvement of doping efficiency through reinforcement learning and multi-objective optimization, while ensuring doping accuracy and uniformity, improve the production speed, reduce production costs, and improve production efficiency on the premise of ensuring product quality.
[0028] Plan and optimize the doping process in advance: The pre-planning and optimization algorithm for doping process based on digital twin and virtual simulation can predict and analyze potential problems in the doping process in advance through virtual simulation technology, optimize and adjust the parameters and processes of the doping process, and at the same time use the equipment status evaluation model to detect potential problems in time and optimize them, improve the reliability and stability of the equipment, and reduce production interruptions.
[0029] Optimize large-scale doping production scheduling: The large-scale doping production optimization algorithm based on multi-robot collaboration and intelligent scheduling can realize the efficient collaborative work among multiple robots, dynamically adjust task allocation and production plans according to real-time production conditions and equipment status, improve the overall efficiency and resource utilization rate of large-scale doping production, and meet the needs of large-scale production. Description of the drawings
[0030] Figure 1 Schematic diagram of the module process. Detailed implementation manners
[0031] Example 1:
[0032] An industrial robot semiconductor doping system based on multi-dimensional intelligent decision-making, comprising:
[0033] An intelligent control algorithm module for doping accuracy based on multi-sensor fusion and real-time correction: Install various sensors such as an impurity source flow sensor, a position sensor, and a concentration sensor on the doping equipment of the industrial robot. Let the flow rate collected by the impurity source flow sensor be Q, the position information obtained by the position sensor be (x, y, z), and the doping concentration detected by the concentration sensor be C; Using multi-sensor fusion, through the formula S = f(Q, x, y, z, C), where f is a multi-sensor fusion function based on deep learning, construct a real-time state model S of the doping process, and combine the real-time correction algorithm. According to the state model S, through the formula ΔP = g(S), where g is a real-time correction function, calculate the doping parameters of the industrial robot, the impurity source flow adjustment amount ΔQ, the doping position adjustment amount Δ(x, y, z), the doping time adjustment amount Δt, and the adjustment amount ΔP, to achieve intelligent control of the doping accuracy;
[0034] An optimization algorithm module for doping uniformity based on multi-modal data acquisition and adaptive regulation: Integrate various data acquisition devices such as a laser scanning microscope, an atomic force microscope, and a spectral analyzer in the doping equipment. Let the wafer surface topography information obtained by the laser scanning microscope be M, the material property information collected by the atomic force microscope be N, and the doping concentration distribution information obtained by the spectral analyzer be O; Using multi-modal data acquisition, through the formula I = h(M, N, O), where h is a multi-modal data acquisition fusion function based on deep learning, construct a comprehensive information model I of the wafer surface; Combine the adaptive regulation algorithm. According to the comprehensive information model I and the real-time changes in the doping process, through the formula ΔT = k(I), where k is an adaptive regulation function, calculate the doping process parameters of the industrial robot, such as the impurity source distribution mode adjustment amount ΔD, the doping path adjustment amount ΔR, the doping speed adjustment amount ΔV, and the adjustment amount ΔT, to achieve optimization of the doping uniformity;
[0035] An algorithm module for improving doping efficiency based on multi-objective optimization and reinforcement learning: Take the doping efficiency E, the doping accuracy A, and the doping uniformity U as the indicators of multi-objective optimization. Define the state space Sstate, including various parameters and state information such as the impurity source flow rate, the doping position, the doping time, the temperature, and the pressure during the doping process; Define the action space Aaction, including various adjustment actions for the doping parameters; Design the reward function R = q(E, A, U), where q is a reward calculation function, comprehensively consider the factors of doping efficiency, doping accuracy, and doping uniformity, and reward or punish the actions of the robot; Use the iterative formula of the deep Q network (DQN) of reinforcement learning
[0036] Q(Sstate,Aaction ) ← Q(Sstate,Aaction ) + α[R + γAaction′ max Q(Sstate′,Aaction′ ) - Q(Sstate,Aaction)], where α is the learning rate and γ is the discount factor or the relevant update formula of the Proximal Policy Optimization algorithm (PPO). Let the industrial robot learn the optimal doping parameter adjustment strategy through continuous interaction with the doped environment, improve the doping efficiency, and at the same time ensure the doping accuracy and uniformity;
[0037] Doping process pre-planning and optimization algorithm module based on digital twin and virtual simulation: Establish a digital twin model of the semiconductor doping system through modeling software and sensor data to achieve real-time interaction and synchronization between the physical entity and the virtual model. Using virtual simulation technology, pre-plan and simulate the semiconductor doping process. Let the simulation result obtained according to the process parameters Pparam and the process flow Fprocess during the simulation be Rsim. Optimize and adjust the parameters and process flow of the semiconductor doping process through the formulas ΔPparam = r(Rsim), where r is the optimization function, and ΔFprocess = s(Rsim), where s is the optimization function. At the same time, collect the historical operation data and fault data of the equipment, establish an equipment status evaluation model, and use the digital twin model to monitor and analyze the operation status of the doping equipment in real time;
[0038] Large-scale doping production optimization algorithm module based on multi-robot cooperation and intelligent scheduling: Regard multiple industrial robots in the semiconductor doping workshop as a multi-agent system, define the state space Ss, including the positions xr, yr of the robots, the task progress ptask, the working status swork, and the environmental information E of the workshop, that is
[0039] Ss = (xr, yr, ptask, swork, E); Define the action space Aa, including various movement actions mmove and task assignment decisions dassign of the robots, that is Aa = (mmove, dassign), and design the reward function
[0040] Rew = t(eefficiency, qquality, uutilization), where t is the reward calculation function, eefficiency is the doping production efficiency, qquality is the task completion quality, and uutilization is the equipment utilization rate); Use the reinforcement learning algorithm Deep Q-Network (DQN) or Proximal Policy Optimization algorithm (PPO) to let the robots learn the optimal cooperative doping production scheduling strategy through continuous interaction with the environment, achieve efficient cooperative work among multiple robots, and improve the overall efficiency and resource utilization rate of large-scale doping production.
[0041] Implementation of an Intelligent Control Algorithm for Doping Precision Based on Multi-Sensor Fusion and Real-Time Calibration
[0042] Sensor Deployment and Data Acquisition: Reasonably deploy impurity source flow sensors, position sensors, concentration sensors, etc. on the doping equipment of industrial robots to ensure accurate acquisition of multi-dimensional parameter information during the doping process. The sensors collect data in real time according to the set frequency and transmit it to the data processing unit through the network.
[0043] Multi-Sensor Data Fusion and State Model Construction: The data processing unit uses a multi-sensor data fusion algorithm based on deep learning, such as a fusion network based on the attention mechanism, to fuse the data collected by different sensors and construct a real-time state model of the doping process. Use deep learning algorithms, such as convolutional neural networks (CNNs), to analyze and process the fused data and extract key feature information.
[0044] Real-Time Calibration Algorithm Implementation and Doping Execution: According to the constructed real-time state model, the real-time calibration algorithm calculates the adjustment amount of the doping parameters of the industrial robot. The robot control system adjusts the doping parameters of the robot in real time according to the adjustment amount to complete the doping task. During the doping process, continuously monitor the sensor data in real time and dynamically adjust the doping parameters to ensure doping precision. Implementation of an Optimization Algorithm for Doping Uniformity Based on Multi-Modal Data Acquisition and Adaptive Regulation
[0045] Data Acquisition Equipment Deployment and Data Acquisition: Reasonably deploy laser scanning microscopes, atomic force microscopes, spectrometers, etc. in the doping equipment to ensure comprehensive acquisition of multi-modal data on the wafer surface. The data acquisition equipment collects data in real time according to the set frequency and transmits it to the data processing center through the network.
[0046] Multi-Modal Data Acquisition and Comprehensive Information Model Construction: The data processing center uses multi-modal data acquisition technology based on deep learning to fuse the data collected by different devices and construct a comprehensive information model of the wafer surface. Use deep learning algorithms, such as convolutional neural networks (CNNs), to analyze and process the fused data and extract key feature information.
[0047] Adaptive Regulation Algorithm Implementation and Doping Execution: According to the constructed comprehensive information model and real-time changes during the doping process, the adaptive regulation algorithm calculates the adjustment amount of the doping process parameters of the industrial robot. The robot control system adjusts the doping process parameters of the robot in real time according to the adjustment amount to complete the doping task. During the doping process, continuously monitor the data of the data acquisition equipment in real time and dynamically adjust the doping process parameters to ensure doping uniformity.
[0048] Implementation of an Algorithm for Improving Doping Efficiency Based on Multi-Objective Optimization and Reinforcement Learning
[0049] Definition of State Space, Action Space, and Reward Function: Clearly define the state space, action space, and reward function of reinforcement learning. The state space includes various parameters and state information during the doping process, such as impurity source flow rate, doping position, doping time, temperature, pressure, etc.; the action space includes various adjustment actions for doping parameters, such as increasing or decreasing the impurity source flow rate, adjusting the doping position and time, etc.; the reward function comprehensively considers factors such as doping efficiency, doping accuracy, and doping uniformity, and rewards or punishes the actions of the robot.
[0050] Reinforcement Learning Model Training and Application: Collect a large amount of doping experiment data, and use reinforcement learning algorithms, such as Deep Q-Network (DQN) or Proximal Policy Optimization (PPO), to train and optimize the reinforcement learning model. During the actual doping process, the industrial robot selects actions based on the real-time doping state, and adjusts the strategy according to the reward feedback after executing the actions, gradually optimizing the doping parameter adjustment strategy, improving the doping efficiency, and ensuring doping accuracy and uniformity at the same time.
[0051] Implementation of Doping Process Pre-Planning and Optimization Algorithm Based on Digital Twin and Virtual Simulation
[0052] Digital Twin Model Construction: Use modeling software and sensor data to establish a digital twin model of the semiconductor doping system, map physical entities such as industrial robots, doping equipment, and semiconductor materials in the physical world to the virtual space, and achieve real-time interaction and synchronization between physical entities and virtual models.
[0053] Virtual Simulation and Process Pre-Planning: Use virtual simulation technology to pre-plan and simulate the semiconductor doping process, and predict and analyze possible problems during the doping process in advance in the virtual environment. According to the results of virtual simulation, optimize and adjust the parameters and processes of the semiconductor doping process, and formulate the optimal doping process plan.
[0054] Equipment State Monitoring and Optimization: Use the digital twin model to monitor and analyze the operating state of the doping equipment in real time, collect the historical operating data and fault data of the equipment, and establish an equipment state evaluation model. During the actual operation process, input the real-time monitored data into the equipment state evaluation model, discover potential problems in time and optimize them to improve the reliability and stability of the equipment.
[0055] Implementation of Large-Scale Doping Production Optimization Algorithm Based on Multi-Robot Collaboration and Intelligent Scheduling
[0056] Definition of Agents and Establishment of Communication Mechanism: Define multiple industrial robots in the semiconductor doping workshop as multiple agents, and each agent has independent decision-making capabilities and goals. Establish a communication network between multiple agents to achieve real-time information interaction between agents.
[0057] Reinforcement Learning Model Training and Application: Define the state space, action space, and reward function of reinforcement learning. Collect historical operation data of the semiconductor doping workshop and use reinforcement learning algorithms such as Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) to train and optimize the reinforcement learning model. In actual production, the agent selects actions based on real-time state information and adjusts the policy according to the reward feedback after executing the actions, gradually optimizing the collaborative doping production scheduling strategy, achieving efficient collaborative work among multiple robots, and improving the overall efficiency and resource utilization rate of large-scale doping production.
Claims
1. An industrial robot semiconductor doping system based on multi-dimensional intelligent decision-making, characterized in that: include: Intelligent control algorithm module for doping accuracy based on multi-sensor fusion and real-time correction: install multiple sensors such as impurity source flow sensor, position sensor, concentration sensor, etc. on the doping equipment of the industrial robot. Assume that the flow collected by the impurity source flow sensor is Q, the position information obtained by the position sensor is (x, y, z), and the doping concentration detected by the concentration sensor is C; use multi-sensor fusion, through the formula S = f (Q, x, y, z, C), where f is the multi-sensor fusion function based on deep learning, to build a real-time state model S of the doping process, combined with the real-time correction algorithm, according to the state model S, through the formula ΔP = g (S), g is the real-time correction function to calculate the doping parameters of the industrial robot, impurity source flow adjustment ΔQ, doping position adjustment Δ (x, y, z), doping time adjustment Δt, adjustment ΔP, to achieve intelligent control of doping accuracy; Doping uniformity optimization algorithm module based on multimodal data acquisition and adaptive control: Integrate multiple data acquisition devices such as laser scanning microscope, atomic force microscope, and spectrometer in the doping equipment. Suppose the wafer surface morphology information acquired by the laser scanning microscope is M, the material property information acquired by the atomic force microscope is N, and the doping concentration distribution information obtained by the spectrometer is O; Using multimodal data acquisition, the comprehensive information model I of the wafer surface is constructed through the formula I = h(M, N, O), where h is the multimodal data acquisition fusion function based on deep learning; Combined with the adaptive control algorithm, according to the comprehensive information model I and the real-time changes of the doping process, the doping process parameters of the industrial robot are calculated through the formula ΔT=k(I), where k is the adaptive control function, such as the adjustment amount ΔD of the impurity source distribution mode, the adjustment amount ΔR of the doping path, the adjustment amount ΔV of the doping speed, and the adjustment amount ΔT, to achieve the optimization of doping uniformity; Doping efficiency improvement algorithm module based on multi-objective optimization and reinforcement learning: take doping efficiency E, doping accuracy A and doping uniformity U as indicators of multi-objective optimization, define the state space Sstate, including various parameters and state information in the doping process of impurity source flow, doping position, doping time, temperature and pressure; define the action space Aaction, including various adjustment actions for doping parameters; design the reward function R = q (E, A, U), q is the reward calculation function, comprehensively consider the factors of doping efficiency, doping accuracy and doping uniformity, and reward or punish the robot's actions; use the iterative formula Q (Sstate, Aaction) ← Q (Sstate, Aaction) + α [R + γAaction′max Q (Sstate′, Aaction′) - Q (Sstate, Aaction)] of the reinforcement learning deep Q network (DQN), α is the learning rate, γ is the discount factor or the relevant update formula of the proximal policy optimization algorithm (PPO), so that the industrial robot can learn the optimal doping parameter adjustment strategy in the continuous interaction with the doping environment, so as to improve the doping efficiency while ensuring the doping accuracy and uniformity; Doping process pre-planning and optimization algorithm module based on digital twin and virtual simulation: A digital twin model of the semiconductor doping system is established through modeling software and sensor data to achieve real-time interaction and synchronization between the physical entity and the virtual model. Using virtual simulation technology, the semiconductor doping process is pre-planned and simulated. The simulation result obtained according to the process parameters Pparam and the process Fprocess during the simulation is assumed to be Rsim. The parameters and processes of the semiconductor doping process are optimized and adjusted through the formula ΔPparam=r(Rsim), r is the optimization function, and ΔFprocess=s(Rsim), s is the optimization function. At the same time, the historical operation data and fault data of the equipment are collected, an equipment status evaluation model is established, and the operation status of the doping equipment is monitored and analyzed in real time using the digital twin model; Large-scale doping production optimization algorithm module based on multi-robot collaboration and intelligent scheduling: multiple industrial robots in the semiconductor doping workshop are regarded as a multi-agent system, and the state space Ss is defined, including the robot's position xr, yr, task progress ptask, work status swork and workshop environment information E, that is, Ss = (xr, yr, ptask, swork, E); the action space Aa is defined, including the robot's various motion actions mmove and task allocation decisions dassign, that is, Aa = (mmove, dassign), and the reward function Rew = t(eefficiency, qquality, uutilization) is designed, t is the reward calculation function, eefficiency is the doping production efficiency, qquality is the task completion quality, and uutilization is the equipment utilization rate); using the reinforcement learning algorithm deep Q network (DQN) or the proximal policy optimization algorithm (PPO), the robot learns the optimal collaborative doping production scheduling strategy in the continuous interaction with the environment, realizes efficient collaborative work between multiple robots, and improves the overall efficiency and resource utilization of large-scale doping production.
2. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the doping precision intelligent control algorithm module based on multi-sensor fusion and real-time correction, multi-sensor data fusion adopts an algorithm based on deep learning to fuse the data collected by different sensors and build a real-time state model of the doping process.
3. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the doping uniformity optimization algorithm module based on multimodal data acquisition and adaptive control, the multimodal data acquisition technology adopts an algorithm based on deep learning to fuse the data collected by different devices and construct a comprehensive information model of the wafer surface.
4. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the doping efficiency improvement algorithm module based on multi-objective optimization and reinforcement learning, the state space includes various parameters and state information in the doping process, the action space includes various adjustment actions on the doping parameters, and the reward function comprehensively considers factors such as doping efficiency, doping accuracy and doping uniformity.
5. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the doping process pre-planning and optimization algorithm module based on digital twin and virtual simulation, the semiconductor doping process is pre-planned and simulated through virtual simulation technology, the process parameters and processes are optimized according to the simulation results, and the digital twin model is used to monitor and analyze the equipment operation status in real time.
6. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the large-scale doping production optimization algorithm module based on multi-robot collaboration and intelligent scheduling, the state space includes the robot's position, task progress, working status and workshop environment information, the action space includes the robot's various motion actions and task allocation decisions, and the reward function uses doping production efficiency, equipment utilization, and task completion quality as reward indicators.
7. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the doping precision intelligent control algorithm module based on multi-sensor fusion and real-time correction, the sensor collects data in real time at a set frequency and transmits it to the data processing unit, and the data processing unit uses an algorithm based on deep learning to fuse and analyze the data.
8. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: In the doping uniformity optimization algorithm module based on multimodal data acquisition and adaptive control, the data acquisition device collects data in real time at a set frequency and transmits it to the data processing center, and the data processing center uses an algorithm based on deep learning to fuse and analyze the data.
9. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1 is characterized in that: The doping efficiency improvement algorithm module based on multi-objective optimization and reinforcement learning collects a large amount of doping experimental data, uses the reinforcement learning algorithm to train and optimize the reinforcement learning model, and the industrial robot selects actions and adjusts strategies according to the real-time doping status.
10. The industrial robot semiconductor doping system based on multi-dimensional perception and intelligent decision-making according to claim 1, characterized in that: The large-scale doping production optimization algorithm module based on multi-robot collaboration and intelligent scheduling defines multiple industrial robots in a semiconductor doping workshop as multiple intelligent agents, establishes a communication network among the multiple intelligent agents, collects historical operation data, and uses a reinforcement learning algorithm to train and optimize the reinforcement learning model. The intelligent agent selects actions and adjusts strategies based on real-time status information.