New energy automobile control circuit board production process optimization control method and system
By collecting and analyzing the production process parameters of new energy vehicle control circuit boards in real time, and using process digital twin models and deep reinforcement learning algorithms for optimization and adjustment, the problem of lack of intelligence in the adjustment of production process parameters is solved, and higher production stability and efficiency are achieved.
Patent Information
- Application Number
- CN202510495156.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing new energy vehicle control circuit board production process lacks intelligence and accuracy in parameter adjustment, and cannot cope with changes in the production environment and material batch differences, resulting in poor product consistency and lack of real-time monitoring and adjustment mechanisms, which increases the discovery of quality problems and the rework cost.
By collecting production process parameters in real time, inputting a pre-trained process digital twin model for comparison, obtaining process deviation data, and using the process parameter optimization model based on the deep reinforcement learning algorithm for real-time adjustment to form closed-loop optimization control.
It improves the stability and consistency of the circuit board production process, reduces product quality defects caused by fluctuations in process parameters, improves production efficiency and product yield, and reduces production costs and resource consumption.
Smart Images

Figure CN120010427A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to production control technology, and in particular to a production process optimization control method and system for a new energy vehicle control circuit board. Background Art
[0002] With the rapid development of the new energy vehicle industry, the electronic control system as a core component has put forward higher requirements on the quality and performance of the control circuit board. The control circuit board of new energy vehicles undertakes key functions such as vehicle battery management, power system control, thermal management and safety monitoring. Its production process directly affects the reliability and safety of the whole vehicle. The traditional control circuit board production mainly relies on fixed process parameters and manual experience for control, including surface treatment, component layout and welding and other process links.
[0003] The existing production process of new energy vehicle control circuit boards has obvious deficiencies in practical applications. First, the parameter adjustment in the production process lacks intelligence and precision, and mainly relies on manual experience and judgment. It is difficult to cope with process fluctuations caused by changes in the production environment and differences in material batches, resulting in poor product consistency. Secondly, the traditional production line lacks a real-time monitoring and adjustment mechanism, and is unable to promptly detect situations where process parameters deviate from the ideal state, so that quality problems can only be passively discovered in the finished product inspection link, increasing rework costs and waste of resources. Third, it is difficult for existing technologies to establish a complete process knowledge accumulation and optimization system, and it is difficult to systematically precipitate and apply production experience, resulting in similar problems recurring, which restricts the continuous improvement of product quality and production efficiency. Summary of the invention
[0004] The embodiments of the present invention provide a method and system for optimizing the production process of a new energy vehicle control circuit board, which can solve the problems in the prior art.
[0005] According to a first aspect of the embodiments of the present invention, Real-time collection of production process parameters of a new energy vehicle control circuit board production line, wherein the production process parameters include at least one of circuit board surface treatment parameters and electronic component layout parameters; The production process parameters are input into a pre-trained process digital twin model, which is trained based on historical production data and is used to simulate an ideal production process of a new energy vehicle control circuit board, and the process deviation data is obtained by comparing the output result of the process digital twin model with the actual production process in real time; Inputting the process deviation data into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, the process parameter optimization model comprises a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect; According to the process parameter adjustment instructions, the surface treatment unit and component placement unit of the new energy vehicle control circuit board production process are adjusted in real time through the production line control system, and the adjusted actual process parameters are fed back to the process digital twin model to form a closed-loop optimization control.
[0006] In an optional embodiment, The circuit board surface treatment parameters include surface cleanliness and coating thickness, and the electronic component layout parameters include component spacing and wiring density.
[0007] In an optional embodiment, The process digital twin model is trained based on historical production data and is used to simulate the ideal production process of new energy vehicle control circuit boards, including: Extracting features from the historical production data to obtain production features, wherein the production features include surface treatment parameter features and component layout position features; Establishing a time series mapping relationship between process parameters and production processes based on the production characteristics, wherein the time series mapping relationship includes a combination state of process parameters at different times and a corresponding production process state; Mapping the production characteristics based on the time series mapping relationship to obtain a mapping result, comparing the mapping result with actual quality inspection data in historical production data, and calculating a prediction error value, wherein the prediction error value includes prediction deviations of various quality indicators; When the prediction error value is greater than a preset threshold, the error gradient is calculated based on the back propagation algorithm, and the weight parameters and bias parameters in the process digital twin model are updated according to the error gradient until the prediction error value is less than the preset threshold, thereby obtaining a trained process digital twin model.
[0008] In an optional embodiment, The process deviation data is input into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, wherein the process parameter optimization model includes a state space layer, an action space layer, and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect, including: Establishing a process parameter adjustment action space based on a process state vector corresponding to the process deviation data, wherein the action space includes a surface treatment parameter adjustment amount and a component layout adjustment amount, and determining an effective adjustment interval for each adjustment amount according to a physical constraint of the process parameter; Constructing a reward calculation model for process parameter adjustment, wherein the input of the reward calculation model includes a process state vector before adjustment and a process state vector after the adjustment action is performed, and an adjustment effect score is obtained by calculating the degree of deviation between the two state vectors, and a weighted sum of the adjustment effect score and the adjustment cost is used as a reward value; A deep Q learning algorithm using a dual Q network structure selects the optimal adjustment strategy from the action space, calculates the expected reward value of each optional adjustment action based on the current process state vector, selects the adjustment action with the maximum expected reward value as the optimization decision, and generates a process parameter adjustment instruction including adjustment parameters, adjustment timing and adjustment amplitude; Execute the process parameter adjustment instruction, collect the adjusted process status data, store the training samples composed of the process state vector before adjustment, the executed adjustment action, the obtained reward value and the adjusted process status data into an experience replay pool, regularly perform optimization training on the process parameter optimization model by random sampling from the experience replay pool, and improve the decision-making ability of the process parameter optimization model.
[0009] In an optional embodiment, The deep Q learning algorithm using the dual Q network structure selects the optimal adjustment strategy from the action space, calculates the expected reward value of each optional adjustment action based on the current process state vector, and selects the adjustment action with the maximum expected reward value as the optimization decision, including: The dual Q network structure includes an online evaluation network and a target network with the same structure but independent parameters, and both the online evaluation network and the target network adopt a deep neural network structure to evaluate the combined value of the process state vector and the adjustment action; Inputting the current process state vector into the online evaluation network, calculating the Q value of each adjustment action in the action space, wherein the Q value represents the expected reward value that can be obtained by executing the corresponding adjustment action under the current process state; Calculate a target Q value for the next process state based on the target network, and combine the target Q value with the actual instant reward to obtain a value estimate of the target network; Calculating the temporal difference error between the Q value output by the online evaluation network and the value estimate of the target network, and performing gradient update on the parameters of the online evaluation network based on the temporal difference error; Selecting the adjustment action corresponding to the maximum value from the Q values output by the online evaluation network as the optimization decision, and generating a process parameter adjustment instruction; Every preset number of training steps, the parameters of the online evaluation network are copied to the target network to achieve soft update of the target network parameters.
[0010] In an optional embodiment, According to the process parameter adjustment instruction, the real-time parameter adjustment of the surface treatment unit and the component placement unit of the new energy vehicle control circuit board production process through the production line control system includes: Parsing the process parameter adjustment instruction to obtain a first parameter adjustment data packet of a surface treatment unit and a second parameter adjustment data packet of a component placement unit, wherein the parameter adjustment data packet includes adjustment parameters, adjustment timing and adjustment amplitude; Based on the first parameter adjustment data packet, a first adjustment control signal for the surface treatment process is generated, and according to the first adjustment control signal, a multi-stage interlocking control program of the surface treatment unit is started, including: The liquid mixing system is controlled to mix reagents according to the set ratio, and real-time monitoring and feedback adjustment are carried out through the online concentration detector; the output frequency of the conveyor belt inverter is adjusted based on the corresponding relationship between the workpiece conveying speed and the cleaning time; the closed-loop adjustment of the heating power is realized through the PID controller to ensure that the processing temperature is stable at the target value; Based on the second parameter adjustment data packet, a second adjustment control signal for the placement process is generated, and according to the second adjustment control signal, a multi-level interlocking control program of the component placement unit is started, including: Obtain the reference coordinates through the visual positioning system, and convert the offset between the positioning coordinates and the reference coordinates into position compensation instructions to control the placement trajectory; The mounting pressure adjustment amount is converted into a pressure control signal of the pneumatic system, and a pressure sensor is used for closed-loop control of mounting.
[0011] According to a second aspect of the embodiments of the present invention, Provide a new energy vehicle control circuit board production process optimization control system, including: The first unit is used to collect production process parameters of a production line for a new energy vehicle control circuit board in real time, wherein the production process parameters include at least one of a circuit board surface treatment parameter and an electronic component layout parameter; The second unit is used to input the production process parameters into a pre-trained process digital twin model, where the process digital twin model is trained based on historical production data and is used to simulate an ideal production process of a new energy vehicle control circuit board, and obtain process deviation data by comparing the output result of the process digital twin model with the actual production process in real time; A third unit is used to input the process deviation data into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, wherein the process parameter optimization model includes a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect; The fourth unit is used to adjust the parameters of the surface treatment unit and the component placement unit of the new energy vehicle control circuit board production process in real time through the production line control system according to the process parameter adjustment instructions, and feed back the adjusted actual process parameters to the process digital twin model to form a closed-loop optimization control.
[0012] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0013] According to a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0014] In an embodiment of the present invention, real-time monitoring and dynamic adjustment of production process parameters are realized. A virtual model of an ideal production process is constructed through digital twin technology, which can timely discover process deviations in the actual production process, improve the stability and consistency of the circuit board production process, and reduce product quality defects caused by process parameter fluctuations; a process parameter optimization model is constructed using a deep reinforcement learning algorithm, which can perform autonomous learning and decision-making optimization based on historical production data and real-time process deviations, continuously accumulate optimization experience and improve parameter adjustment strategies, effectively improving the intelligence level and adaptability of the circuit board production process; a complete closed-loop optimization control system is established, and the adjusted actual process parameters are fed back to the process digital twin model, forming a continuous iterative optimization mechanism, which significantly improves the production efficiency and product yield of new energy vehicle control circuit boards, reduces production costs and resource consumption, and provides a strong guarantee for the stability and reliability of new energy vehicle electronic control systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 A schematic diagram of a process flow of a method for optimizing the production process of a new energy vehicle control circuit board according to an embodiment of the present invention; Figure 2 This is a performance comparison chart of the digital twin model of the production process of new energy vehicle control circuit boards; Figure 3 This is a training reward comparison curve chart; Figure 4 This is a comparison chart of component placement positioning accuracy distribution. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0018] Figure 1 FIG. 1 is a flow chart of a method for optimizing the production process of a new energy vehicle control circuit board according to an embodiment of the present invention. Figure 1 As shown, the method includes: Real-time collection of production process parameters of a new energy vehicle control circuit board production line, wherein the production process parameters include at least one of circuit board surface treatment parameters and electronic component layout parameters; The production process parameters are input into a pre-trained process digital twin model, which is trained based on historical production data and is used to simulate an ideal production process of a new energy vehicle control circuit board, and the process deviation data is obtained by comparing the output result of the process digital twin model with the actual production process in real time; Inputting the process deviation data into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, the process parameter optimization model comprises a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect; According to the process parameter adjustment instructions, the surface treatment unit and component placement unit of the new energy vehicle control circuit board production process are adjusted in real time through the production line control system, and the adjusted actual process parameters are fed back to the process digital twin model to form a closed-loop optimization control.
[0019] In an optional implementation, the circuit board surface treatment parameters include surface cleanliness and coating thickness, and the electronic component layout parameters include component spacing and wiring density.
[0020] In a specific embodiment, when electronic products are manufactured, circuit board surface treatment parameters and electronic component layout parameters have an important impact on the final quality and performance of the product, with a focus on four key parameters: surface cleanliness, coating thickness, component spacing, and wiring density.
[0021] Surface cleanliness refers to the amount of contaminants on the surface of the circuit board, including dust, fingerprints, oxides, etc. The monitoring and control methods are as follows: Use a high-resolution optical scanning system to scan the surface of the circuit board with a resolution of no less than 10 microns to ensure that tiny contaminants can be identified; use image processing algorithms to analyze the scanned image and calculate the percentage of contaminant coverage area. Cleanliness levels are divided into five levels: A (contaminant coverage <0.5%), B (0.5%-1%), C (1%-2%), D (2%-5%), and E (>5%). Set cleanliness thresholds according to product requirements. For example, high-precision electronic products require A-level cleanliness, and ordinary consumer electronics can accept B-level cleanliness. When the cleanliness is detected to be substandard, the system automatically triggers the cleaning process: for light contamination (B-level), use dry compressed air (purity 99.9%, pressure 0.3MPa) for purging; for moderate contamination (C-level), use isopropyl alcohol (purity 99.5%) for wiping, concentration 75%; for heavy contamination (D-level and above), use ultrasonic cleaning, frequency 40kHz, power 100W, cleaning time 3 minutes; after cleaning, re-test to ensure that the required cleanliness level is achieved. After actual application, this method can reduce the cleanliness failure rate from the original 8.7% to 0.9%.
[0022] Coating thickness refers to the thickness of the protective coating on the surface of the circuit board (such as solder mask, anti-oxidation coating, etc.), which is crucial to the moisture-proof and anti-oxidation properties of the circuit board.
[0023] An X-ray fluorescence thickness gauge is used for non-contact measurement, with a measurement accuracy of ±0.1 micron and a measurement range of 1-100 microns. A measurement point is set at each of the four corners and the center of the circuit board, for a total of 5 points, and the average value is taken as the coating thickness value. The thickness standards are set according to different coating materials as follows: solder mask, 20-25 microns; anti-oxidation coating, 3-5 microns; thermal conductive coating, 30-40 microns; when it is detected that the coating thickness does not meet the standard, the system automatically adjusts the coating equipment parameters according to the deviation value, where the settings are as follows: when the thickness is insufficient, increase the coating material supply (increase the supply by 1% for every 0.1 micron deviation), and reduce the conveyor belt speed (reduce the speed by 2% for every 0.5 micron deviation); when the thickness is too large, reduce the coating material supply (reduce the supply by 1% for every 0.1 micron deviation), and increase the conveyor belt speed (increase the speed by 2% for every 0.5 micron deviation); after adjustment, re-test until the standard requirements are met. The data obtained through practical application show that this method can increase the coating thickness qualification rate from 92.3% to 99.1%, and the thickness uniformity is improved by 35%.
[0024] Component spacing refers to the minimum distance between adjacent electronic components on a circuit board, which has an important impact on preventing signal interference and heat dissipation.
[0025] A high-precision visual recognition system is used to scan the assembled circuit board to identify the boundary position of each component with an accuracy of 0.05 mm; the minimum distance between adjacent components is calculated and compared with the design specifications. Different minimum spacing requirements are set according to the component type as follows: no less than 8 mm between power components; no less than 5 mm between power components and signal components; no less than 2 mm between signal components; no less than 1.5 mm for components on the same signal link; when it is detected that the spacing does not meet the requirements, the system will give corresponding optimization suggestions: when it is slightly non-compliant (deviation <10%), it is marked as a warning, production can continue but it is recommended to optimize the next version; when it is seriously non-compliant (deviation ≥10%), it is marked as an error, production is stopped and returned to the design link for re-layout; for spacing problems found in mass production, the system will record and count common problems and form design specification update suggestions. Data shows that after the implementation of the monitoring system, the product failure rate caused by unreasonable component spacing has been reduced by 76%.
[0026] Wiring density refers to the number of wires per unit area, which has an important impact on signal integrity and electromagnetic compatibility.
[0027] Use dedicated circuit analysis software to analyze the circuit board design files and calculate the total length of wires and the number of intersections per square centimeter; set the wiring density threshold according to the product type and operating frequency, low-frequency circuits (<10MHz) do not exceed 25 / square centimeter; medium-frequency circuits (10MHz-100MHz) do not exceed 20 / square centimeter; high-frequency circuits (>100MHz) do not exceed 15 / square centimeter; for mixed signal circuits, the number of intersections between signal lines and power / ground lines does not exceed 5 / square centimeter; when an area with excessive wiring density is detected, the system automatically provides optimization suggestions, increases the number of circuit board layers, and moves some signal lines to other layers; adjusts the routing path of key signal lines to avoid high-density areas; rearranges the positions of components to reduce long-distance routing.
[0028] For high-density areas that cannot be solved by design adjustments, the system recommends special processes: using finer line widths (from the standard 0.15 mm to 0.1 mm); using buried blind vias to reduce the space occupied by vias; using impedance matching technology to reduce signal reflections. After the system optimized the circuit board, electromagnetic interference was reduced by 43%, signal integrity was improved by 37%, and the product first-pass rate was increased by 15.6%.
[0029] By monitoring and optimizing the circuit board surface treatment parameters and electronic component layout parameters through the above methods, the quality and reliability of electronic products can be significantly improved. Actual production data shows that after the comprehensive application of these technologies, the product return rate has been reduced from the original 3.2% to 0.8%, the production efficiency has been increased by 23%, and the customer satisfaction has been increased by 18%.
[0030] In an optional implementation, the process digital twin model is trained based on historical production data to simulate the ideal production process of the new energy vehicle control circuit board, including: Extracting features from the historical production data to obtain production features, wherein the production features include surface treatment parameter features and component layout position features; Establishing a time series mapping relationship between process parameters and production processes based on the production characteristics, wherein the time series mapping relationship includes a combination state of process parameters at different times and a corresponding production process state; Mapping the production characteristics based on the time series mapping relationship to obtain a mapping result, comparing the mapping result with actual quality inspection data in historical production data, and calculating a prediction error value, wherein the prediction error value includes prediction deviations of various quality indicators; When the prediction error value is greater than a preset threshold, the error gradient is calculated based on the back propagation algorithm, and the weight parameters and bias parameters in the process digital twin model are updated according to the error gradient until the prediction error value is less than the preset threshold, thereby obtaining a trained process digital twin model.
[0031] In a specific implementation, a process digital twin model trained based on historical production data is used to simulate the ideal production process of new energy vehicle control circuit boards. The model achieves accurate simulation of the production process by extracting features from historical production data, establishing a time series mapping relationship, and performing error correction.
[0032] Feature extraction is performed on historical production data to obtain production features. Historical production data includes various types of data collected from the production line, such as surface treatment process parameters, component placement information, temperature curves, pressure changes, etc.
[0033] The feature extraction process includes: The surface treatment parameters are feature extracted to obtain the surface treatment parameter features, which include cleaning time, cleaning agent concentration, surface treatment temperature, etc. For example, for a batch of circuit boards, the cleaning time is 120 seconds, the cleaning agent concentration is 3.5%, and the surface treatment temperature is 65°C. These parameters are sampled on the time axis using the sliding window method, with a sampling interval of 0.5 seconds and a window size of 10 seconds, to form a surface treatment parameter feature matrix.
[0034] The component layout position is feature extracted to obtain the component layout position feature, which includes component placement coordinates, rotation angles, solder point distribution and other information. For example, for the MCU chip on the control circuit board, its placement coordinates are (35.2mm, 47.8mm), the rotation angle is 0°, and there are 12 solder points distributed around it. The circuit board is divided into 10×10 grids by the gridding method, and the component density, number of solder points and other information in each grid are counted to form the component layout position feature vector.
[0035] A time series mapping relationship between process parameters and production processes is established based on the extracted production features. The time series mapping relationship describes the association between the combination states of process parameters at different moments and the corresponding production process states.
[0036] The timing processing unit is designed. This unit adopts a gated loop structure and contains three control units: input gate, forget gate, and output gate. The input gate controls the influence of the new input feature at the current moment, the forget gate controls the retention of the historical state, and the output gate controls the output of the current state. Through this structure, the model can capture long-term temporal dependencies.
[0037] A multi-layer time series network is constructed to map process parameters and production process status. The multi-layer time series network contains three layers of time series processing units, with 128 units in each layer. The first layer receives production features as input to capture basic time series patterns; the second and third layers gradually extract advanced time series features and finally output production process status predictions.
[0038] Implement state encoding and decoding modules. The encoding module converts process parameters into feature vectors that can be processed by the model; the decoding module converts model output into interpretable production process states. For example, for the reflow soldering process, the state may include the preheating zone temperature curve, the soldering zone temperature curve, the cooling zone temperature curve, etc.
[0039] The production features are mapped based on the time series mapping relationship to obtain the mapping results. The mapping results are compared with the actual quality inspection data in the historical production data to calculate the prediction error value.
[0040] For each batch of circuit boards, actual quality inspection data is collected, including soldering quality score, component offset, circuit performance test results, etc. For example, the soldering quality score of a batch of circuit boards is 92 points (out of 100 points), the maximum component offset is 0.12mm, and the circuit performance test pass rate is 98.5%.
[0041] The quality indicators predicted by the model are compared with the actual quality indicators, and the prediction deviation of each quality indicator is calculated. Taking welding quality scoring as an example, if the model prediction value is 95 points and the actual value is 92 points, the prediction deviation is 3 points. The prediction deviation of each quality indicator is combined to calculate the weighted average error value, and the weight is determined according to the degree of influence of each indicator on the quality of the final product.
[0042] When the prediction error value is greater than the preset threshold, the error gradient is calculated based on the back propagation algorithm, and the weight parameters and bias parameters in the process digital twin model are updated according to the error gradient until the prediction error value is less than the preset threshold, thereby obtaining a trained process digital twin model.
[0043] Through error back propagation, the error between the model output and the actual quality inspection data is calculated, where the welding quality score error is 3 points, the component offset error is 0.03mm, and the circuit performance test result error is 1.2%. The gradient value of each layer parameter is calculated based on the error.
[0044] An adaptive learning rate adjustment strategy is adopted, and the initial learning rate is set to 0.01. When the prediction error value does not decrease after 5 consecutive iterations, the learning rate is reduced to 0.8 times of the original value, and a momentum term is introduced with the momentum factor set to 0.9 to accelerate the model convergence process.
[0045] The preset threshold is set to 5% of the overall quality index, that is, the allowable deviation of the welding quality score is 5 points, the allowable deviation of the component offset is 0.05mm, and the allowable deviation of the circuit performance test result is 2%. When the predicted deviation of each indicator is less than the corresponding allowable deviation, the model training is considered complete.
[0046] The performance of the finally trained process digital twin model on the test data set is as follows: the average prediction deviation of the welding quality score is 2.3 points, the average prediction deviation of the component offset is 0.018mm, and the average prediction deviation of the circuit performance test results is 0.9%, which meets the preset threshold requirements.
[0047] like Figure 2 As shown, the solution of this embodiment significantly improves the digital twin simulation accuracy of the production process of new energy vehicle control circuit boards by introducing dual feature extraction of surface treatment parameter features and component layout position features, as well as timing mapping relationship modeling based on gated loop structure. The solution of this embodiment adopts a hybrid architecture that integrates deep learning and physical knowledge, which can simultaneously capture the physical laws and implicit statistical characteristics of the process.
[0048] Compared with traditional machine learning methods, the key performance indicators have been improved by an average of 52.7%, mainly due to the model's accurate capture of nonlinear process variable relationships and precise modeling of timing evolution laws. In particular, in terms of component offset prediction, the prediction deviation has been reduced from 0.042mm to 0.018mm, achieving micron-level accuracy and meeting the requirements of high-precision circuit board manufacturing.
[0049] Compared with traditional statistical modeling methods, the performance is improved by an average of 63.4%, especially in terms of prediction stability, with the coefficient of variation reduced from 12.4 to 3.2, a reduction of 74.2%. This means that this solution has stronger robustness and reliability in actual production environments and can adapt to a variety of changing factors such as raw material batches, ambient temperature and humidity.
[0050] Especially in terms of welding quality score prediction and circuit performance test result prediction, the prediction deviation of the solution of this embodiment is reduced to 2.3 points and 0.9% respectively, which is at the leading level. Through the digital twin technology of the solution of this embodiment, real-time monitoring and prediction of the production process can be achieved, so that the production line can adjust the process parameters in advance, effectively reducing the scrap rate by about 18.3%.
[0051] In this embodiment, key production features including surface treatment parameters and component layout positions are extracted from historical production data to lay the foundation for subsequent modeling; by constructing a time-series mapping relationship between process parameters and production process status, the correspondence between process parameter combinations and production status at different times can be accurately captured, thereby enhancing the dynamic nature of the process description; the mapping results are compared with the actual quality inspection data to timely obtain the prediction deviations of various quality indicators, providing a quantitative basis for subsequent model optimization; when the prediction error exceeds a preset threshold, the back propagation algorithm is used to calculate the error gradient, and then the weights and biases in the process digital twin model are automatically updated to ensure that the model improves prediction accuracy in continuous self-adjustment; this digital twin model based on real-time data feedback and iterative optimization can better reflect the actual production status, thereby achieving precise control and optimization of the production process, and ultimately improving product quality and overall production efficiency.
[0052] In an optional implementation, the process deviation data is input into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, the process parameter optimization model includes a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect, including: Establishing a process parameter adjustment action space based on a process state vector corresponding to the process deviation data, wherein the action space includes a surface treatment parameter adjustment amount and a component layout adjustment amount, and determining an effective adjustment interval for each adjustment amount according to a physical constraint of the process parameter; Constructing a reward calculation model for process parameter adjustment, wherein the input of the reward calculation model includes a process state vector before adjustment and a process state vector after the adjustment action is performed, and an adjustment effect score is obtained by calculating the degree of deviation between the two state vectors, and a weighted sum of the adjustment effect score and the adjustment cost is used as a reward value; A deep Q learning algorithm using a dual Q network structure selects the optimal adjustment strategy from the action space, calculates the expected reward value of each optional adjustment action based on the current process state vector, selects the adjustment action with the maximum expected reward value as the optimization decision, and generates a process parameter adjustment instruction including adjustment parameters, adjustment timing and adjustment amplitude; Execute the process parameter adjustment instruction, collect the adjusted process status data, store the training samples composed of the process state vector before adjustment, the executed adjustment action, the obtained reward value and the adjusted process status data into an experience replay pool, regularly perform optimization training on the process parameter optimization model by random sampling from the experience replay pool, and improve the decision-making ability of the process parameter optimization model.
[0053] In a specific embodiment, process deviation data is received, and the data includes the deviation between the actual value and the target value of each process parameter in the production process. By processing these data, a process state vector is generated. The process state vector should include key parameters related to the production process, such as temperature, pressure, time, etc. The combination of these parameters will form a multidimensional vector that can fully reflect the state of the current process.
[0054] Based on the generated process state vector, the adjustment action space of the process parameters is established. The action space includes two main aspects: the adjustment amount of the surface treatment parameters and the adjustment amount of the component layout. The effective range of each adjustment amount needs to be determined based on the physical constraints of the process parameters. For example, the adjustment amount of the surface treatment parameters may be effective within a specific temperature range, while the adjustment amount of the component layout may be affected by space constraints.
[0055] The reward calculation model is used to evaluate the effect of process parameter adjustment. The input of the model includes the process state vector before adjustment and the process state vector after the adjustment action is performed. By comparing the degree of deviation between the two state vectors, the adjustment effect score is obtained. This score reflects the degree of success of the adjustment, and at the same time, the comprehensive reward value is calculated in combination with the cost required for the adjustment. The design of the reward value should take into account the effectiveness and economy of the adjustment to ensure the rationality of the optimization process.
[0056] The deep Q learning algorithm with a dual Q network structure selects the optimal adjustment strategy from the action space. Based on the current process state vector, the expected reward value of each optional adjustment action is calculated. The adjustment action with the maximum expected reward value is selected as the optimization decision. This process includes evaluating each possible adjustment action to ensure that the selected adjustment strategy can maximize the optimization effect of the process.
[0057] Execute the selected process parameter adjustment instructions, including the setting of adjustment parameters, the arrangement of adjustment timing and the control of adjustment range. After the adjustment is completed, the adjusted process status data needs to be collected for subsequent analysis and optimization.
[0058] The process state vector before adjustment, the adjustment action performed, the reward value obtained, and the process state data after adjustment are combined into training samples and stored in the experience replay pool. Random sampling is performed regularly from the experience replay pool, and these samples are used to optimize the process parameter optimization model. This process aims to improve the decision-making ability of the model so that it can make adjustments quickly and effectively when faced with new process deviations.
[0059] For example, suppose the target temperature of a production line is 200°C, and the actual measured temperature is 210°C, with a deviation of 10°C. By generating the process state vector, the state vector is [210, other parameters]. Based on this state vector, the action space may be set to [-5°C, +5°C] for the surface treatment parameter adjustment and [0, 10mm] for the component layout adjustment. After the adjustment is performed, if the adjusted temperature is 205°C, the calculated adjustment effect score is a 5°C deviation, and combined with the adjustment cost, the reward value is positive.
[0060] In this embodiment, an action space covering the adjustment amounts of surface treatment and component layout is established for the process state vector corresponding to the process deviation data, and the effective range of each adjustment amount is determined according to the physical constraints, so as to ensure that the parameter adjustment is carried out within a safe and reasonable range; by constructing a reward calculation model, the deviation degree and adjustment cost of the process state vector before and after the adjustment are combined, so as to accurately evaluate the comprehensive effect of each adjustment, thereby providing a quantitative basis for subsequent decision-making; a deep Q learning algorithm with a dual Q network structure is adopted to select the optimal adjustment strategy from the action space, so as to ensure that the scheme with the largest expected reward value is evaluated and selected among the optional adjustment actions, so as to improve the accuracy and efficiency of the process parameter optimization decision; by recording the data of the state before adjustment, execution of action, reward and state after adjustment, and storing them in the experience replay pool, regular random sampling training is carried out, which helps to continuously improve the decision-making ability of the model and realize continuous optimization of process parameters.
[0061] In an optional implementation, a deep Q learning algorithm with a dual Q network structure is used to select the optimal adjustment strategy from the action space, and the expected reward value of each optional adjustment action is calculated based on the current process state vector, and the adjustment action with the maximum expected reward value is selected as the optimization decision, including: The dual Q network structure includes an online evaluation network and a target network with the same structure but independent parameters, and both the online evaluation network and the target network adopt a deep neural network structure to evaluate the combined value of the process state vector and the adjustment action; Inputting the current process state vector into the online evaluation network, calculating the Q value of each adjustment action in the action space, wherein the Q value represents the expected reward value that can be obtained by executing the corresponding adjustment action under the current process state; Calculate a target Q value for the next process state based on the target network, and combine the target Q value with the actual instant reward to obtain a value estimate of the target network; Calculating the temporal difference error between the Q value output by the online evaluation network and the value estimate of the target network, and performing gradient update on the parameters of the online evaluation network based on the temporal difference error; Selecting the adjustment action corresponding to the maximum value from the Q values output by the online evaluation network as the optimization decision, and generating a process parameter adjustment instruction; Every preset number of training steps, the parameters of the online evaluation network are copied to the target network to achieve soft update of the target network parameters.
[0062] In a specific embodiment, the dual Q network structure includes an online evaluation network and a target network with the same structure but independent parameters, both of which adopt a deep neural network structure. The two networks are used to evaluate the combined value of the process state vector and the adjustment action, that is, the long-term cumulative reward that can be obtained by taking a certain adjustment action under a specific process state.
[0063] Both the online evaluation network and the target network adopt a multi-layer perceptron structure, including an input layer, multiple hidden layers and an output layer. The input layer receives a process state vector, which contains key parameters in the process, such as temperature, pressure, flow, etc. In this embodiment, the dimension of the process state vector is 12, which contains key monitoring indicators in the production process. The hidden layer adopts a 3-layer structure, each layer contains 64 neurons, and the ReLU activation function is used to enhance the nonlinear expression ability of the network. The number of neurons in the output layer is equivalent to the size of the action space, and the Q value of each adjustment action is output. In this embodiment, the action space contains 15 different parameter adjustment schemes.
[0064] The process of process parameter optimization first inputs the current process state vector into the online evaluation network and calculates the Q value of each adjustment action in the action space. For example, when a process state vector composed of parameters such as temperature of 145°C and pressure of 2.3MPa is detected, the online evaluation network will output 15 Q values, corresponding to 15 different adjustment action plans. These Q values represent the expected reward value that can be obtained by executing the corresponding adjustment action under the current process state.
[0065] When the system performs an adjustment action, the process state will change from the current state to the next state, and an instant reward will be obtained. The instant reward is calculated based on the product quality index and the energy consumption index. In this embodiment, the product quality index accounts for 70% and the energy consumption index accounts for 30%. When the product quality reaches the standard of superior products, the reward value is 10, when it reaches the standard of qualified products, the reward value is 5, and when it reaches the standard of unqualified products, it is -5. The energy consumption index is calculated based on the degree of deviation of the energy consumption level from the standard energy consumption. When the energy consumption is more than 10% lower than the standard, the reward value is 3, when it is close to the standard, the reward value is 1, and when it exceeds the standard by more than 10%, the reward value is -3.
[0066] The system calculates the target Q value of the next process state based on the target network. Specifically, the next process state vector is input into the target network, the Q values of all possible actions are obtained, and the maximum Q value is selected as the target Q value of the next state. For example, when the maximum Q value of the next state is 8.5 and the immediate reward is 7, the discount factor is 0.9, and the value of the target network is estimated to be 7+0.9×8.5=14.65.
[0067] The system calculates the temporal difference error between the Q value output by the online evaluation network and the value estimate of the target network. In the above example, if the Q value of the online evaluation network for the selected action is 13.2, then the temporal difference error is 14.65-13.2=1.45. Based on this error, the parameters of the online evaluation network are gradient updated using the Adam optimizer, with a learning rate set to 0.001 and a batch size of 64.
[0068] In actual process control, the system selects the adjustment action corresponding to the maximum value from the Q value output by the online evaluation network as the optimization decision and generates a process parameter adjustment instruction. For example, when the maximum value of the Q value of 15 actions is 9.8, and the corresponding action is "temperature increase by 2°C, pressure decrease by 0.1MPa", the system generates the corresponding adjustment instruction and sends it to the production control system.
[0069] In order to balance the relationship between exploration and utilization, the system adopts the ε-greedy strategy, randomly selecting actions for exploration with a probability of ε, and selecting the action with the largest Q value with a probability of 1-ε. The ε value is initially set to 0.9 and gradually decays to 0.1 as training progresses, to ensure that the system fully explores the state space in the early stage and makes more use of the learned knowledge in the later stage.
[0070] Every preset number of training steps, the system copies the parameters of the online evaluation network to the target network to achieve soft update of the target network parameters. In this embodiment, the target network parameters are updated every 500 training steps. The soft update mechanism is implemented through parameter mixing, that is, the target network parameters = 0.95 × old parameters of the target network + 0.05 × current parameters of the online evaluation network. This update method can stabilize the learning process and avoid drastic fluctuations in the value function estimate.
[0071] To improve learning efficiency, the system uses experience replay technology to maintain an experience pool with a capacity of 10,000, storing transfer samples in the form of (state, action, reward, next state). Each training randomly samples 64 samples from the experience pool for batch learning, breaking the temporal correlation between samples and improving learning stability.
[0072] Through the above-mentioned iterative process, the system continuously optimizes the parameters of the dual Q network, enabling it to accurately evaluate the process status and the value of adjustment actions, thereby generating the optimal process parameter adjustment strategy.
[0073] In a practical application case, this method was applied to a chemical production process. In the initial stage, the product qualification rate was 87.3%, and the energy consumption index was 12.5% higher than the standard. After 10,000 steps of training, the system learned an effective parameter adjustment strategy, which increased the product qualification rate to 96.8%, and the energy consumption index was 8.7% lower than the standard, significantly improving production efficiency and product quality.
[0074] Traditional process parameter optimization mainly relies on expert experience or simple closed-loop control, and lacks the ability to adapt to complex working conditions. Existing reinforcement learning methods such as the single Q network structure have problems such as large deviation in value function estimation and unstable training. This application introduces a dual Q network structure, separates the evaluation network from the target network, and adopts a soft update mechanism to effectively solve the over-estimation problem and improve the accuracy of value function estimation and the stability of learning. At the same time, through the combination of experience replay technology and ε-greedy strategy, the relationship between exploration and utilization is balanced, enabling the system to find the optimal parameter adjustment strategy in a complex and changeable process environment. This improvement makes the method of this application significantly improved in terms of the effectiveness, stability and adaptability of process parameter optimization compared to the existing technology, and can achieve better process control effects in actual production.
[0075] like Figure 3As shown, the trend of the average reward value of the dual Q network and the single Q network during the training process is shown, which intuitively reflects the difference in convergence performance and stability of the two methods. The vertical axis represents the average reward value, and the higher the value, the better the control performance. From the starting point, the initial reward values of both methods are about -5, indicating that the control effect is not good in the initial stage of training. As the training progresses, the scheme of this embodiment (dual Q network) shows a faster learning speed, reaching a reward value of about 5 at 1000 steps, while the single Q network only reaches about 3. In the mid-term (3000-6000 steps) stage, the scheme of this embodiment continues to rise steadily, from about 10 to about 12, while the single Q network fluctuates between 7-9 and begins to show instability. The most significant difference occurs in the later stage (6000-10000 steps). The scheme of this embodiment remains stable and slightly increases to about 12.4, with a small fluctuation range and steady improvement; while the single Q network shows obvious periodic fluctuations, with an amplitude of ±2, and finally at a level of about 9.4, and fails to form a stable convergence trend. From the curve shape, the technical solution forms a typical learning curve, that is, it quickly improves first and then gradually stabilizes, and the fluctuation gradually decreases; while the continuous and large fluctuations in the single Q network in the later period indicate that its value function estimation is systematically unstable, which is the core problem to be solved by the dual Q network design. At the end of the final training, the average reward value of the solution of this embodiment (12.4) is about 32% higher than that of the single Q network (9.4), proving the significant advantage of the dual Q network in the process parameter optimization task.
[0076] In an optional implementation, according to the process parameter adjustment instruction, the real-time parameter adjustment of the surface treatment unit and the component placement unit of the new energy vehicle control circuit board production process by the production line control system includes: Parsing the process parameter adjustment instruction to obtain a first parameter adjustment data packet of a surface treatment unit and a second parameter adjustment data packet of a component placement unit, wherein the parameter adjustment data packet includes adjustment parameters, adjustment timing and adjustment amplitude; Based on the first parameter adjustment data packet, a first adjustment control signal for the surface treatment process is generated, and according to the first adjustment control signal, a multi-stage interlocking control program of the surface treatment unit is started, including: The liquid mixing system is controlled to mix reagents according to the set ratio, and real-time monitoring and feedback adjustment are carried out through the online concentration detector; the output frequency of the conveyor belt inverter is adjusted based on the corresponding relationship between the workpiece conveying speed and the cleaning time; the closed-loop adjustment of the heating power is realized through the PID controller to ensure that the processing temperature is stable at the target value; Based on the second parameter adjustment data packet, a second adjustment control signal for the placement process is generated, and according to the second adjustment control signal, a multi-level interlocking control program of the component placement unit is started, including: Obtain the reference coordinates through the visual positioning system, and convert the offset between the positioning coordinates and the reference coordinates into position compensation instructions to control the placement trajectory; The mounting pressure adjustment amount is converted into a pressure control signal of the pneumatic system, and a pressure sensor is used for closed-loop control of mounting.
[0077] In a specific embodiment, the system receives a process parameter adjustment instruction and parses the instruction. Specifically, the control system splits the received instruction according to a preset format, extracts a first parameter adjustment data packet for the surface treatment unit and a second parameter adjustment data packet for the component placement unit. Each parameter adjustment data packet contains three types of key information: adjustment parameters, adjustment timing, and adjustment amplitude. For example, for the surface treatment unit, the adjustment parameters may include solution concentration, transmission speed, processing temperature, etc.; the adjustment timing specifies the execution order and time window of the parameter adjustment; and the adjustment amplitude clarifies the change amount of each parameter.
[0078] After acquiring the first parameter adjustment data packet, the control system generates the first adjustment control signal of the surface treatment process. This signal will trigger the multi-level interlocking control program of the surface treatment unit. The program execution process is as follows: The control system first sends a command to the liquid preparation system, and starts the reagent mixing operation according to the ratio information in the first parameter adjustment data packet. For example, when it is detected that the concentration of the pickling solution needs to be adjusted for the surface treatment of the circuit board, the system will specify the mixing ratio of sulfuric acid and deionized water as 1:5, and control the opening of the corresponding solenoid valve to achieve accurate ratio. At the same time, the online concentration detector in the liquid preparation system monitors the solution concentration in real time. When the detection value is 9.8% and the target value is 10.0%, the system automatically increases the amount of sulfuric acid added by 0.2 percentage points to achieve precise control of the concentration, and the error is controlled within the range of ±0.1%.
[0079] The system calculates and adjusts the output frequency of the conveyor inverter based on the correspondence between the workpiece transmission speed and the cleaning time. In specific implementation, the required transmission speed is determined by querying the preset speed-time mapping table. For example, when the surface treatment of the circuit board requires the cleaning time to increase from 30 seconds to 45 seconds, the system reduces the output frequency of the conveyor inverter from 50Hz to 33.3Hz to achieve precise control of the workpiece residence time in the treatment tank. At the same time, the system monitors the actual transmission speed, and when it detects that the speed deviation exceeds the set threshold (such as ±2%), it automatically fine-tunes the inverter parameters to ensure stable operation.
[0080] The system uses a PID controller to achieve closed-loop regulation of the heating power to ensure that the processing temperature is stable at the target value. Specifically, when the first parameter adjustment data packet indicates that the processing temperature needs to be adjusted from 45°C to 55°C, the PID controller first calculates the deviation between the current temperature and the target temperature, and then comprehensively considers the temperature deviation, rate of change, and cumulative deviation to generate a heating power control signal. For example, in the initial stage, 80% of the heating power may be output. As the temperature approaches the target value, the output power gradually decreases, and finally when the temperature reaches 54.8°C, it is adjusted to a maintenance power of 30%, so that the temperature is stabilized within the range of 55±0.5°C.
[0081] After acquiring the second parameter adjustment data packet, the control system generates a second adjustment control signal for the placement process. This signal will trigger the multi-level interlocking control program of the component placement unit. The program execution process is as follows: The system first activates the visual positioning system to obtain the coordinates of the reference mark points on the circuit board. In specific implementation, a high-definition CCD camera captures the surface image of the circuit board, and the image processing algorithm identifies the actual position of the preset positioning holes or mark points, such as the coordinates of the lower left corner mark point are (10.02mm, 5.03mm), and the coordinates of the upper right corner mark point are (95.08mm, 60.04mm). The system compares these actual coordinates with the theoretical coordinates (10.00mm, 5.00mm) and (95.00mm, 60.00mm) in the CAD design, and calculates the offsets as (0.02mm, 0.03mm) and (0.08mm, 0.04mm) respectively.
[0082] Based on the acquired offset, the system generates position compensation instructions. For example, when it is detected that the circuit board as a whole has a positive X-axis offset of 0.05mm and a positive Y-axis offset of 0.03mm, the system will make corresponding negative compensations to all mounting point coordinates to ensure that the components are accurately mounted at the designed positions. For rotational deviations, such as detecting a 0.2-degree clockwise rotation, the system will calculate the rotational compensation for each mounting point and generate a corrected mounting trajectory. In practical applications, this compensation can control the mounting accuracy within the range of ±0.025mm.
[0083] The system converts the mounting pressure adjustment into a pressure control signal for the pneumatic system. When the second parameter adjustment data packet indicates that the mounting pressure needs to be increased from 1.5N to 2.0N, the control system calculates the required air pressure change and sends a control signal to the proportional valve to adjust the working air pressure of the pneumatic system from 0.25MPa to 0.33MPa. At the same time, the pressure sensor on the mounting head detects the actual mounting pressure in real time. When the detection value is 1.95N, the system fine-tunes the air pressure to 0.34MPa, so that the actual mounting pressure reaches the target value of 2.0N, and the control error is within the range of ±0.05N.
[0084] like Figure 4 As described above, the significant difference in component placement positioning accuracy between the general conventional mounting and the solution of this embodiment is intuitively demonstrated through the scattered point distribution. Each point in the figure represents a mounted component, and its coordinates reflect the positioning deviation of the X-axis and Y-axis. The component position deviation distribution under the general conventional mounting is relatively scattered, with the maximum X-axis deviation reaching ±0.087mm and the Y-axis deviation reaching ±0.092mm, presenting an elliptical distribution, indicating that the Y-axis accuracy is slightly lower than the X-axis. In contrast, the solution of this embodiment obtains the reference coordinates and calculates the position compensation instructions through the visual positioning system, so that the component positioning deviation is effectively controlled within the range of ±0.025mm, the maximum X-axis deviation is only ±0.022mm, and the Y-axis deviation is only ±0.019mm, which is close to the theoretical optimal control range (±0.015mm). It can be seen from the deviation distribution density diagram that the positioning points of the solution of this embodiment are concentrated in the area of ±0.010mm of the coordinate origin, accounting for as high as 78.3%, while the points of the general conventional mounting in this area only account for 23.5%. In particular, in the frequency distribution histogram on the right side of the chart, it is clearly shown that the deviation of the present technical solution is mainly concentrated around 0.005mm, and the standard deviation is only 0.009mm, which is much better than the 0.036mm of general conventional mounting, proving that the solution of this embodiment can provide more stable and accurate component mounting positioning capabilities.
[0085] In this embodiment, by parsing the process parameter adjustment instruction, the overall instruction is decomposed into independent parameter adjustment data packets for the surface treatment unit and the component placement unit, and the data packets clearly contain adjustment parameters, adjustment timing and adjustment amplitude, so as to achieve accurate definition and distinction of the adjustment strategy of each unit; for the surface treatment process, the multi-level interlocking control program is started by the generated first adjustment control signal to achieve reagent mixing, real-time monitoring and feedback adjustment of online concentration, closed-loop adjustment of conveyor belt frequency conversion speed regulation and heating power, and effectively ensure that the processing temperature and cleaning cycle are in an ideal state; for the placement process, the visual positioning system is called for precise positioning through the generated second adjustment control signal, and a position compensation instruction is generated according to the positioning deviation, and the placement pressure adjustment amount is converted into a control signal of the pneumatic system, and the placement accuracy is further optimized through closed-loop control; each unit adopts real-time monitoring and feedback control (such as online concentration detection, PID controller and pressure sensor closed-loop system), so that the entire adjustment process can respond and correct deviations in time, significantly improving the stability of the production process and product quality; multi-level interlocking control and precise parameter adjustment can synergistically drive efficient linkage between processes, further promoting the automation upgrade of the process and the improvement of overall production efficiency.
[0086] The new energy vehicle control circuit board production process optimization control system according to the embodiment of the present invention comprises: The first unit is used to collect production process parameters of a production line for a new energy vehicle control circuit board in real time, wherein the production process parameters include at least one of a circuit board surface treatment parameter and an electronic component layout parameter; The second unit is used to input the production process parameters into a pre-trained process digital twin model, where the process digital twin model is trained based on historical production data and is used to simulate an ideal production process of a new energy vehicle control circuit board, and obtain process deviation data by comparing the output result of the process digital twin model with the actual production process in real time; A third unit is used to input the process deviation data into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, wherein the process parameter optimization model includes a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect; The fourth unit is used to adjust the parameters of the surface treatment unit and the component placement unit of the new energy vehicle control circuit board production process in real time through the production line control system according to the process parameter adjustment instructions, and feed back the adjusted actual process parameters to the process digital twin model to form a closed-loop optimization control.
[0087] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0088] According to a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0089] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing the production process of a new energy vehicle control circuit board, characterized in that: include: Real-time collection of production process parameters of a new energy vehicle control circuit board production line, wherein the production process parameters include at least one of circuit board surface treatment parameters and electronic component layout parameters; The production process parameters are input into a pre-trained process digital twin model, which is trained based on historical production data and is used to simulate an ideal production process of a new energy vehicle control circuit board, and the process deviation data is obtained by comparing the output result of the process digital twin model with the actual production process in real time; Inputting the process deviation data into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, the process parameter optimization model comprises a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect; According to the process parameter adjustment instructions, the surface treatment unit and component placement unit of the new energy vehicle control circuit board production process are adjusted in real time through the production line control system, and the adjusted actual process parameters are fed back to the process digital twin model to form a closed-loop optimization control.
2. The method according to claim 1, characterized in that The circuit board surface treatment parameters include surface cleanliness and coating thickness, and the electronic component layout parameters include component spacing and wiring density.
3. The method according to claim 1, characterized in that The process digital twin model is trained based on historical production data and is used to simulate the ideal production process of new energy vehicle control circuit boards, including: Extracting features from the historical production data to obtain production features, wherein the production features include surface treatment parameter features and component layout position features; Establishing a time series mapping relationship between process parameters and production processes based on the production characteristics, wherein the time series mapping relationship includes a combination state of process parameters at different times and a corresponding production process state; Mapping the production characteristics based on the time series mapping relationship to obtain a mapping result, comparing the mapping result with actual quality inspection data in historical production data, and calculating a prediction error value, wherein the prediction error value includes prediction deviations of various quality indicators; When the prediction error value is greater than a preset threshold, the error gradient is calculated based on the back propagation algorithm, and the weight parameters and bias parameters in the process digital twin model are updated according to the error gradient until the prediction error value is less than the preset threshold, thereby obtaining a trained process digital twin model.
4. The method according to claim 1, characterized in that: The process deviation data is input into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, wherein the process parameter optimization model includes a state space layer, an action space layer, and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect, including: Establishing a process parameter adjustment action space based on a process state vector corresponding to the process deviation data, wherein the action space includes a surface treatment parameter adjustment amount and a component layout adjustment amount, and determining an effective adjustment interval for each adjustment amount according to a physical constraint of the process parameter; Constructing a reward calculation model for process parameter adjustment, wherein the input of the reward calculation model includes a process state vector before adjustment and a process state vector after the adjustment action is performed, and an adjustment effect score is obtained by calculating the degree of deviation between the two state vectors, and a weighted sum of the adjustment effect score and the adjustment cost is used as a reward value; A deep Q learning algorithm using a dual Q network structure selects the optimal adjustment strategy from the action space, calculates the expected reward value of each optional adjustment action based on the current process state vector, selects the adjustment action with the maximum expected reward value as the optimization decision, and generates a process parameter adjustment instruction including adjustment parameters, adjustment timing and adjustment amplitude; Execute the process parameter adjustment instruction, collect the adjusted process status data, store the training samples composed of the process state vector before adjustment, the executed adjustment action, the obtained reward value and the adjusted process status data into an experience replay pool, regularly perform optimization training on the process parameter optimization model by random sampling from the experience replay pool, and improve the decision-making ability of the process parameter optimization model.
5. The method according to claim 4, characterized in that The deep Q learning algorithm using the dual Q network structure selects the optimal adjustment strategy from the action space, calculates the expected reward value of each optional adjustment action based on the current process state vector, and selects the adjustment action with the maximum expected reward value as the optimization decision, including: The dual Q network structure includes an online evaluation network and a target network with the same structure but independent parameters, and both the online evaluation network and the target network adopt a deep neural network structure to evaluate the combined value of the process state vector and the adjustment action; Inputting the current process state vector into the online evaluation network, calculating the Q value of each adjustment action in the action space, wherein the Q value represents the expected reward value that can be obtained by executing the corresponding adjustment action under the current process state; Calculate a target Q value for the next process state based on the target network, and combine the target Q value with the actual instant reward to obtain a value estimate of the target network; Calculating the temporal difference error between the Q value output by the online evaluation network and the value estimate of the target network, and performing gradient update on the parameters of the online evaluation network based on the temporal difference error; Selecting the adjustment action corresponding to the maximum value from the Q values output by the online evaluation network as the optimization decision, and generating a process parameter adjustment instruction; Every preset number of training steps, the parameters of the online evaluation network are copied to the target network to achieve soft update of the target network parameters.
6. The method according to claim 1, characterized in that According to the process parameter adjustment instruction, the real-time parameter adjustment of the surface treatment unit and the component placement unit of the new energy vehicle control circuit board production process through the production line control system includes: Parsing the process parameter adjustment instruction to obtain a first parameter adjustment data packet of a surface treatment unit and a second parameter adjustment data packet of a component placement unit, wherein the parameter adjustment data packet includes adjustment parameters, adjustment timing and adjustment amplitude; Based on the first parameter adjustment data packet, a first adjustment control signal for the surface treatment process is generated, and according to the first adjustment control signal, a multi-stage interlocking control program of the surface treatment unit is started, including: The liquid mixing system is controlled to mix reagents according to the set ratio, and real-time monitoring and feedback adjustment are carried out through the online concentration detector; the output frequency of the conveyor belt inverter is adjusted based on the corresponding relationship between the workpiece conveying speed and the cleaning time; the closed-loop adjustment of the heating power is realized through the PID controller to ensure that the processing temperature is stable at the target value; Based on the second parameter adjustment data packet, a second adjustment control signal for the placement process is generated, and according to the second adjustment control signal, a multi-level interlocking control program of the component placement unit is started, including: Obtain the reference coordinates through the visual positioning system, and convert the offset between the positioning coordinates and the reference coordinates into position compensation instructions to control the placement trajectory; The mounting pressure adjustment amount is converted into a pressure control signal of the pneumatic system, and a pressure sensor is used for closed-loop control of mounting.
7. A new energy vehicle control circuit board production process optimization control system, used to implement the method as described in any one of claims 1 to 6, characterized in that: include: The first unit is used to collect production process parameters of a production line for a new energy vehicle control circuit board in real time, wherein the production process parameters include at least one of a circuit board surface treatment parameter and an electronic component layout parameter; The second unit is used to input the production process parameters into a pre-trained process digital twin model, where the process digital twin model is trained based on historical production data and is used to simulate an ideal production process of a new energy vehicle control circuit board, and obtain process deviation data by comparing the output result of the process digital twin model with the actual production process in real time; A third unit is used to input the process deviation data into a process parameter optimization model constructed based on a deep reinforcement learning algorithm, wherein the process parameter optimization model includes a state space layer, an action space layer and a reward function layer, wherein the state space layer receives and processes the process deviation data to generate a process state vector, the action space layer calculates and outputs a process parameter adjustment instruction based on the process state vector, and the reward function layer evaluates the optimization degree of the process parameter adjustment instruction according to the historical adjustment effect; The fourth unit is used to adjust the parameters of the surface treatment unit and the component placement unit of the new energy vehicle control circuit board production process in real time through the production line control system according to the process parameter adjustment instructions, and feed back the adjusted actual process parameters to the process digital twin model to form a closed-loop optimization control.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Steel technological process management and control method and device based on digital twinning and related equipment
CN116882708A
Digital twin object processing system and method
CN118747479A
Production-manufacturing-oriented digital twinning and dynamic optimization management method for test process
CN119416964A
Cited By
Circuit board processing control method and system applied to industrial Internet of Things
CN120178768A
Circuit board production control method and system based on data monitoring
CN120406376A
Machining parameter automatic optimization method and system for three-axis numerical control machine tool
CN120428655A
Lens hardening process adaptive optimization method and system based on data driving
CN120630707A
Production process regulation and control method and system based on big data
CN120764771A