Radio frequency plasma spheroidization powder preparation method and system based on multi-objective optimization

CN122500206APending Publication Date: 2026-08-04SHAANXI WEINA ZHIYAN NEW MATERIAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI WEINA ZHIYAN NEW MATERIAL TECHNOLOGY CO LTD
Filing Date
2026-05-09
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

首先参数优化方法局限,现有方法本质上是基于局部搜索的贪心策略,缺乏对参数-质量映射关系的全局建模能力,无法有效处理高维、多峰、多目标的优化问题

Benefits of technology

本发明首先利用量子点标记与高速光谱采集技术,结合时空图神经网络,实现了对缺陷演化的超前预警,将传统的事后检测转变为事前预测;其次,通过贝叶斯多目标优化,在极少量实验下便高效寻得帕累托最优解集,成功平衡了球化率、氧含量、粒度分布与流动性等相互制约的关键指标;再次,通过空间映射矩阵将全局参数解耦为六分区差异化控制指令,精准顺应了等离子体炬内部非均匀温度场,实现了精细化的热历史管理;最后,引入基于近端策略优化的强化学习智能体,对残余缺陷进行在线闭环修复,赋予系统自适应能力。本发明有效解决了传统工艺中多目标难兼顾、热场控制粗放、缺陷响应滞后等痛点,大幅提升了粉末产品的综合质量、一致性与生产过程的智能化水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122500206A_ABST
    Figure CN122500206A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of powder metallurgy and industrial artificial intelligence, and in particular to a radio frequency plasma spheroidization powder preparation method and system based on multi-objective optimization, the method comprising: collecting particle and process data and preprocessing; constructing a process space-time heterogeneous graph, predicting a defect probability field using ST-GCN and GAT; building a Bayesian multi-objective optimization model based on early warning and historical data, and solving a Pareto solution set; decoupling into six partition instructions using a mapping matrix to fine-tune the thermal history; modeling residual defects as POMDP, training an Actor-Critic network using PPO, and repairing defects in a closed loop online. The present application effectively solves the problems of multi-objective difficulty in traditional processes, extensive heat field control, and defect response lag, and greatly improves the comprehensive quality, consistency, and intelligent level of the production process of the powder product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of powder metallurgy and industrial artificial intelligence, specifically to a method and system for preparing radio frequency plasma spheroidized powder based on multi-objective optimization for dynamic prediction of subsidence areas. Background Technology

[0002] Radio frequency plasma spheroidization technology utilizes a high-frequency electromagnetic field (typically 3-4 MHz) to ionize a working gas (primarily argon, with the addition of hydrogen) to form a high-temperature plasma torch (reaching temperatures up to 10°C). 4 The plasma (above K) causes irregularly shaped metal powder particles to rapidly absorb heat and melt as they pass through the high-temperature plasma zone, sphericalizing under surface tension. These particles then rapidly condense in the cooling chamber to form spherical powder. This technology boasts advantages such as high energy density, large plasma torch volume, rapid heat transfer and cooling, no electrode contamination, and controllable reaction atmosphere. It has become a core technology for preparing high-performance spherical metal powders and is widely used in additive manufacturing, thermal spraying, powder metallurgy, and other fields.

[0003] However, the current equipment relies solely on the operator's experience to repeatedly adjust the process parameters. The existing technical methods mainly include the following: ① Linear parameter optimization method: Orthogonal experimental design: Using statistical methods to design a limited number of experimental combinations, and using linear analysis to study the influence of each parameter on the quality index; Offline sampling and testing: After spheroidization, samples are taken and their morphology is observed by scanning electron microscopy (SEM), particle size distribution is analyzed by laser particle size analyzer, and oxygen content is detected by oxygen and nitrogen analyzer.

[0004] ② Quality inspection methods: Limited online monitoring: Some devices are equipped with simple temperature or pressure sensors, but lack the ability to monitor the state of powder particles in real time in situ; Offline sampling and testing: After spheroidization, samples are taken and their morphology is observed by scanning electron microscopy (SEM), particle size distribution is analyzed by laser particle size analyzer, and oxygen content is detected by oxygen and nitrogen analyzer.

[0005] ③ Process control methods: Global unified control: The power and gas flow rate settings are uniformly applied to the entire plasma torch area, without taking into account the differences in thermodynamic properties between different regions inside the plasma torch; Open-loop control or simple feedback: manual parameter adjustment based on offline detection results, lacking real-time closed-loop control capability.

[0006] ④ Defect handling methods: Offline screening and grading: The spheroidized powder is sieved and its morphology is screened, and unqualified products are downgraded for use or reworked and remelted. Lack of online repair methods: There is no effective online intervention and repair capability for defects generated during the spheroidization process.

[0007] Each of the above technologies has its own drawbacks: Firstly, parameter optimization methods have limitations. Existing methods are essentially greedy strategies based on local search, lacking the ability to globally model the parameter-quality mapping relationship and thus failing to effectively handle high-dimensional, multi-peak, and multi-objective optimization problems. Specifically, the RF plasma spheroidization process involves 6-8 key process parameters (powder feed rate, plasma power, carrier gas flow rate, sheath gas flow rate and gas ratio, powder particle size distribution, feed location, system pressure, etc.), and these parameters have complex nonlinear coupling relationships. Empirical trial-and-error methods require hundreds or even thousands of experiments to find optimal parameters, with optimization cycles lasting from weeks to months, resulting in high experimental costs (the cost of consumables for a single experiment is several thousand yuan). In addition, although orthogonal experimental design reduces the number of experiments, it can only analyze single-factor or two-factor interactions, failing to capture high-order coupling effects of multiple parameters and making it difficult to locate the global optimum in a complex multi-parameter space. Furthermore, adjusting a single parameter often has a chain reaction effect on multiple quality indicators (e.g., increasing power can improve spheroidization rate but increases oxidation risk), and traditional methods lack the ability to perform multi-objective collaborative optimization.

[0008] Furthermore, the detection of quality defects is lagging, making real-time monitoring and feedback control of the manufacturing process impossible. Existing monitoring technologies lack highly sensitive labeling methods capable of stable operation in high-temperature plasma environments, and the temporal and spectral resolutions of signal acquisition systems are insufficient to capture millisecond-level particle dynamics. Offline sampling inspections suffer from time delays ranging from hours to days. By the time quality problems are detected, a large number of defective products have already been generated, making early warning and preventative control impossible before defects form. Secondly, multi-objective trade-offs in process control are difficult, and the mutual constraints among quality indicators make coordinated optimization challenging. Existing methods lack quantitative proxy models between process parameters and multiple quality indicators, and lack systematic multi-objective optimization algorithm support. Specifically, improving spheroidization rate often requires higher plasma power and longer residence time, but this leads to increased oxygen content, and oxidation is exacerbated at high temperatures, resulting in coarsening of particle size. Common quality evaluation of spheroidized powder involves 4-6 mutually constraining indicators: spheroidization rate, particle size distribution concentration, oxygen content increment, powder flowability, and loose packing density. These indicators are mutually constraining, but algorithms merely treat them in isolation or simply as linear. Traditional process optimization models typically employ a single-objective successive optimization strategy (e.g., first optimize spheroidization rate, then control oxygen content), leading to local optima rather than global optima, and therefore are unsuitable for our process.

[0009] Furthermore, the defect repair capabilities are insufficient. Existing control systems are open-loop or simple closed-loop, lacking feedforward control based on predictive models and adaptive repair strategies based on learning. Specifically, even after optimized control, 5%-15% residual defects (incompletely spherical satellite spheres, hollow spheres, surface oxidation, etc.) will still be generated during the spheroidization process. Existing technologies can only screen for residual defects offline and cannot repair them through online fine-tuning of process parameters; they also lack cross-timescale collaborative control capabilities, such as the lack of effective connection between millisecond-level particle flight processes and second-level process parameter adjustments; on the other hand, existing technologies lack adaptive compensation mechanisms for sudden process disturbances (such as unstable powder feeding and gas flow fluctuations).

[0010] Finally, some technologies utilize zoned control, but due to a lack of precise modeling of the spatiotemporal evolution characteristics within the plasma torch, a mapping relationship between spatial zones and process parameters has not been established. For example, most controls do not consider the spatial non-uniformity of the plasma torch. Generally, the plasma torch exhibits a significant temperature gradient along the axial direction (center > 10000 K, edge < 3000 K) and also shows differences in temperature distribution along the radial direction. Existing technologies use uniform power and gas settings for the entire plasma torch, failing to differentiate control based on the thermodynamic characteristics of different regions. The powder feeding position and carrier gas injection angle are fixed, making it impossible to adjust heat input conditions according to the real-time position of particles within the torch. The temperature gradient control in the cooling chamber is coarse, failing to achieve graded cooling rate control for powders of different particle sizes.

[0011] Therefore, it is crucial to address the technical challenges in radio frequency plasma spheroidization technology, such as reliance on trial and error based on human experience, lack of in-situ monitoring of particle state, difficulty in coordinating multiple target parameters, extensive spatial control, and inability to repair defects online, which lead to long process optimization cycles, high costs, and unstable product quality. Summary of the Invention

[0012] Based on the technical problems existing in the prior art, the present invention provides a method and system for preparing radio frequency plasma spheroidized powder based on multi-objective optimization.

[0013] As a first aspect of the present invention, a method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization is provided, comprising the following steps: S1. Obtain particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and perform feature extraction and noise reduction preprocessing using signal processing algorithms; S2. Based on the particle features extracted in step S1, construct a process spatiotemporal heterogeneous graph with process parameters, powder properties, monitoring features and quality indicators as nodes, and spatiotemporal evolution relationship and parameter coupling as connection edge rules. A spatiotemporal information aggregation between nodes is realized through spatiotemporal graph convolutional network and graph attention mechanism to generate a spatial distribution field of defect probability that represents the defect evolution law. S3. Based on the spatial distribution field of the defect probability and historical data in step S2, construct a Bayesian multi-objective optimization model with process parameters as decision variables and the objective functions of maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability as the set of objective functions. Use a multi-output Gaussian process proxy model and the expected hypervolume improvement acquisition strategy to iteratively optimize the multi-objective optimization model and obtain the Pareto front solution set. S4. Based on the Pareto optimal parameter solution set output in step S3, the global variables are decoupled into differentiated control commands for six axial and radial zones through a preset spatial mapping matrix, so as to achieve fine thermal history control in accordance with the non-uniform temperature field inside the plasma torch. S5. Based on the residual defect state after adjustment in step S4 and the real-time monitoring data, a partially observable Markov decision process is modeled. An Actor-Critic network is trained using a proximal policy optimization algorithm to adaptively output power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects.

[0014] Based on the above scheme, step S2, which involves constructing a spatiotemporal heterogeneous graph of a process based on the particle features extracted in step S1, with process parameters, powder properties, monitoring features, and quality indicators as nodes, and spatiotemporal evolution relationships and parameter coupling as connection rules, and using a spatiotemporal graph convolutional network and graph attention mechanism to aggregate spatiotemporal information between nodes to generate a spatial distribution field of defect probability representing the defect evolution law, specifically includes: S201. Heterogeneous graph construction: Construct a heterogeneous graph that includes process parameters, powder properties, monitoring features, and quality nodes; (4) in, For process parameters, The normalized value; These are powder properties, namely, particle size, initial morphology, velocity, and position. For monitoring characteristics, including temperature field, melt index, stability index, etc.; Quality indicators include sphericity, oxygen content, particle size distribution, and flowability. Edge type construction, setting For fully connected parametric coupling edges, The edges represent spatiotemporal evolution, with weights based on kinematic predictions. The edges representing causal relationships are determined using the Granger causality test. S202, Spatiotemporal Graph Convolutional Network Model Training and Feature Aggregation: Learning complex relationships between nodes using spatiotemporal graph convolutional networks; The spatiotemporal graph convolutional network model is represented as follows: (5) in, Let K be the spatial adjacency matrix of the k-th order Chebyshev polynomial approximation (K=3). The features of the l-th layer node, These are learnable weights; the time dimension is processed through convolution with a kernel size of 3 and an expansion rate that increases with the number of layers. The graph attention mechanism is represented as: (6) In the formula, For nodes With nodes Attention coefficient between them Attention weight vector transpose, The weight matrix is ​​a linear transformation matrix. For nodes The input feature vector, For nodes The input feature vector; (7) In the formula, For nodes With nodes The initial weights of the edges between them; For the first A quality objective function, Belonging to the set {1,2,3,4}, it corresponds to the sphericity, oxygen content increment, particle size distribution standard deviation, and Hall flow rate, respectively; For nodes The corresponding process parameters, For nodes Corresponding process parameters; (8) In the formula, For nodes The updated feature vector, For nodes The neighborhood set, The normalized attention coefficient. For neighboring nodes Features after linear transformation; S203, Output Decoding and Defect Prediction Output decoding: (9) In the formula, Output the tensor for the defect probability distribution field. This is the node feature matrix of the last (third) layer of the graph neural network. Dimension labeling for real tensors, representing It is a 128*128*4 three-dimensional real tensor; It is a multilayer perceptron (fully connected neural network), which contains 2-3 layers of linear transformations and activation functions, mapping the node features output by the graph neural network to the defect probability space; The four channels correspond to the probabilities of incomplete spheroidization, satellite spheres, hollow spheres, and oxidative contamination, respectively. The loss function is: (10) In the formula, This is the total loss function value. For real labels, To predict probabilities for the model, For positive sample loss terms, is the L2 regularization coefficient, which controls the strength of the penalty for model complexity; The squared L2 norm of all learnable parameters of the model.

[0015] Based on the above scheme, step S3 involves constructing a Bayesian multi-objective optimization model with process parameters as decision variables and maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability as objective functions. The step of iteratively optimizing the multi-objective optimization model using a multi-output Gaussian process surrogate model and an expected hypervolume improvement acquisition strategy to obtain the Pareto front solution set specifically includes: S301, The multi-objective function is expressed as: (26) in, For sphericity, For the increase in oxygen content, The standard deviation of particle size distribution. To improve powder flowability; The process parameter vector is as follows: (31) in, The power output of the plasma equipment is expressed in kW. This refers to the powder delivery rate, expressed in g / min. This refers to the carrier gas flow rate, expressed in L / min. The sheath gas H2 / Ar ratio is expressed in % (%). Powder particle size, in μm; This indicates the feed position, in mm. S302. Constructing a multi-output Gaussian process proxy model For each target Establish independent but related Gaussian process models: (11) For the process parameter vector (decision variable vector), see formula (31); The Matern 5 / 2 kernel function is used to capture the nonlinear relationships between parameters: (12) in, , where l is the length dimension. The signal variance; Construct the covariance structure of the multi-output Gaussian process surrogate model using the correlation matrix B between tasks: (13) In the formula, For the extended covariance matrix of a multi-output Gaussian process, Given the input sample point set, For another set of new test input sample points, It is the spatial kernel function matrix; S303, Optimization of the Desired Supervolume Improved Acquisition Function Define the current Pareto frontier The hypervolume index relative to the reference point r is: (15) in, This represents the current Pareto frontier, with r as the reference point. Lebesgue measure; Calculate new points The expected increase in hypervolume that can be brought to the current Pareto frontier: (17) In the formula, Let S be the desired supervolume improvement value, and S be the number of Monte Carlo samplings. For the current Pareto frontier, hypervolume, This represents the extended Pareto front after adding new sampling points. Let s be the target vector sampled from the posterior of GP in the s-th iteration. Use as a reference point; Using the L-BFGS-B algorithm Perform numerical optimization to find the next candidate parameter point that maximizes EHI. ; S304, Bayesian Iterative Update and Convergence Judgment.

[0016] Based on the above scheme, step S4, which uses the Pareto optimal parameter solution set output in step S3 to decouple global variables into differentiated control commands for six axial and radial zones through a preset spatial mapping matrix, is a step to achieve refined thermal history control in accordance with the non-uniform temperature field inside the plasma torch. Specifically, this includes: S401. Divide the plasma torch space into six regions. The area of ​​0-300mm at the top of the torch along the axial direction is divided into a central high-temperature zone for initial rapid melting of powder, a transitional melting zone for complete melting and spheroidization of particles, and an edge cooling zone for rapid solidification of particles. The region with a radius of 0-50mm is divided radially into a core area where energy is most concentrated, a middle area, and an outer area near the cooling wall; S402, Construct the partition control parameter mapping matrix The global parameters output by Bayesian optimization Mapped to partition mapping instructions : (18) Where M is a 6*6 mapping matrix, This is the region bias vector. This refers to the operating power of the plasma equipment. To ensure powder delivery rate, Carrier gas flow rate, The ratio of sheath gas H2 / Ar is [missing information]. Powder particle size, This is the feed location.

[0017] Based on the above scheme, step S5, which involves modeling the residual defect state after adjustment in step S4 and real-time monitoring data as a partially observable Markov decision process, training an Actor-Critic network using a proximal policy optimization algorithm, and adaptively outputting power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects, specifically includes: S501. The defect repair problem is formalized as a POMDP quintuple (S,A,T,R,O). S502. Construct an Actor-Critic network architecture, which includes: a shared feature extraction layer, a policy network, and a value network; S503. An Actor-Critic network is trained using a proximal strategy optimization algorithm to achieve online closed-loop defect repair.

[0018] Based on the above scheme, step S501, which formalizes the defect repair problem into a POMDP quintuple (S, A, T, R, O), specifically includes: (1) State space : (19) Among them, the defect type codes are: [1,0,0,0] = incomplete spheroidization, [0,1,0,0] = satellite sphere, [0,0,1,0] = hollow sphere, [0,0,0,1] = oxidation contamination; (2) Action space A A={ },in, For power fine-tuning, For carrier gas pulse, As compensation for the second batch of powder delivery, For inert purging, No operation is performed; State transition T: describes the rules that describe how the environment evolves from the current state to the next state after the agent performs an action; ; Approximation of the prediction model using graph neural networks: (20) (4) Reward function The comprehensive system includes rewards for defect elimination, incentives for quality improvement, penalties for operational costs, and penalties for deterioration. : (twenty one) In the formula, To map to the set of real numbers, the output is a scalar reward value, with positive numbers encouraging and negative numbers penalizing. The instantaneous reward at time t; This represents the change in the severity of the defect. (5) Observation O: Based on the real-time monitoring data and defect prediction results of step S4.

[0019] Based on the above scheme, step S502, constructing an Actor-Critic network architecture, includes the steps of sharing a feature extraction layer, a policy network, and a value network, specifically including: First, a two-layer fully connected network with 128 neurons each is used as a shared feature layer to extract features from the input state, employing the ReLU activation function. Then, a policy network, based on these shared features, is constructed through a hierarchical structure from 128 neurons to 64 neurons and then to 5 neurons, using the Softmax activation function to output the probability distribution of each action, guiding action selection. The value network, also using the shared features as input, is constructed through a hierarchical structure from 128 neurons to 64 neurons and then to 1 neuron, without employing an activation function, to evaluate the value of the current state, providing a value reference for policy optimization.

[0020] Based on the above scheme, step S1, which involves acquiring particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and performing feature extraction and noise reduction preprocessing using signal processing algorithms, specifically includes: The quantum dot fluorescence intensity model is expressed as: (1) in, To excite light intensity, For temperature-dependent quantum yield, For the emission spectral profile, Thermally activated energy, Local ambient temperature; The particle temperature inversion algorithm based on dual-wavelength colorimetry is expressed as follows: (2) Through experiments, obtain =600nm, =650nm, solve for the particle surface temperature. ; Melt State Index The calculation is as follows: (3) Among them, the weighting coefficient of the temperature term =0.6, spectral term weighting coefficient β=0, The full width at half maximum (FWHM) of the quantum dot fluorescence spectrum. For reference, half-height and full width.

[0021] As a second aspect of the present invention, a system is provided according to the radio frequency plasma spheroidizing powder preparation method based on multi-objective optimization.

[0022] The system includes: The sensing layer is used to acquire particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and to perform feature extraction and noise reduction preprocessing using signal processing algorithms. The prediction layer is used to construct a spatiotemporal heterogeneous graph of the process based on the extracted particle features. The graph uses process parameters, powder properties, monitoring features and quality indicators as nodes and spatiotemporal evolution relationship and parameter coupling as connection edge rules. The spatiotemporal information between nodes is aggregated through spatiotemporal graph convolutional network and graph attention mechanism to generate a spatial distribution field of defect probability that represents the defect evolution law. The optimization layer is used to construct a Bayesian multi-objective optimization model based on the spatial distribution field of defect probability and historical data. The model has process parameters as decision variables and the objective functions are maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability. The multi-output Gaussian process proxy model and the expected hypervolume improvement acquisition strategy are used to iteratively optimize the multi-objective optimization model to obtain the Pareto front solution set. The control layer is used to decouple global variables into differentiated control commands for six axial and radial zones based on the Pareto optimal parameter solution set based on the output. This is achieved by using a preset spatial mapping matrix to adapt to the non-uniform temperature field inside the plasma torch and realize fine thermal history control. The repair layer is used to model the residual defect state after regulation and real-time monitoring data as a partially observable Markov decision process. The Actor-Critic network is trained using a proximal policy optimization algorithm to adaptively output power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects.

[0023] As a third aspect of the invention, a system is provided for quality control of radio frequency plasma spheroidized powder.

[0024] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention first utilizes quantum dot labeling and high-speed spectral acquisition technology, combined with a spatiotemporal graph neural network, to achieve advanced early warning of defect evolution, transforming traditional post-event detection into pre-event prediction. Second, through Bayesian multi-objective optimization, it efficiently finds the Pareto optimal solution set with minimal experimental data, successfully balancing key indicators such as sphericity, oxygen content, particle size distribution, and flowability. Third, by decoupling global parameters into six-zone differentiated control commands through a spatial mapping matrix, it precisely adapts to the non-uniform temperature field inside the plasma torch, achieving refined thermal history management. Finally, it introduces a reinforcement learning agent based on near-end policy optimization to perform online closed-loop repair of residual defects, endowing the system with adaptive capabilities. This invention effectively solves the pain points of traditional processes, such as difficulty in simultaneously achieving multiple objectives, coarse thermal field control, and delayed defect response, significantly improving the overall quality, consistency, and intelligent level of powder products and the production process.

[0025] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0027] Figure 1 This is a graph neural network training and prediction performance curve shown according to an exemplary embodiment (showing the loss convergence curve). Figure 2 This is a graph neural network training and prediction performance curve shown according to an exemplary embodiment (showing ROC-AUC curves for four types of defects). Figure 3 This is a graph neural network training and prediction performance curve illustrated according to an exemplary embodiment (showing the Precision-Recall curve: for the defect category of imbalanced samples). Figure 4 This is a Bayesian multi-objective optimization convergence curve shown according to an exemplary embodiment (showing the independent convergence trajectory of each objective function). Figure 5 This is a GP agent model prediction accuracy verification curve / scatter plot shown according to an exemplary embodiment; Figure 6 This is a PPO reinforcement learning training curve (showing cumulative reward) illustrated according to an exemplary embodiment. Figure 7 This is a PPO reinforcement learning training curve shown according to an exemplary embodiment (showing Actor loss / Critic loss / entropy three-axis curves). Figure 8This is a PPO reinforcement learning training curve shown according to an exemplary embodiment (showing a KL divergence monitoring curve). Figure 9 This is an online closed-loop repair dynamic effect curve shown according to an exemplary embodiment. Detailed Implementation

[0028] The following description and accompanying drawings fully illustrate specific embodiments of this application to enable those skilled in the art to practice them. Some parts and features of some embodiments may be included in or replace parts and features of other embodiments.

[0029] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0030] This application provides a method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization, comprising the following steps: S1. Data Collection and Preprocessing Particle state and process parameter data are obtained through quantum dot labeling and high-speed spectral acquisition technology, and feature extraction and noise reduction preprocessing are performed using signal processing algorithms. Specifically, the quantum dot fluorescence intensity model is expressed as: (1) in, To excite light intensity, For temperature-dependent quantum yield, For the emission spectral profile, Thermally activated energy, Local ambient temperature; The particle temperature inversion algorithm (based on dual-wavelength colorimetry) is expressed as follows: (2) Through experiments, obtain =600nm, =650nm, solve for T; The melt state index is calculated as follows: (3) Where α=0.6, β=0, To broaden the spectral lines.

[0031] S2, Graph Neural Network Spatiotemporal Prediction like Figures 1-3As shown, based on the particle features extracted in step S1, a process spatiotemporal heterogeneous graph is constructed with process parameters, powder properties, monitoring features and quality indicators as nodes, and spatiotemporal evolution relationship and parameter coupling as connection edge rules. The spatiotemporal information aggregation between nodes is realized through spatiotemporal graph convolutional network (ST-GCN) and graph attention mechanism (GAT) to generate a spatial distribution field of defect probability that characterizes the defect evolution law. Specifically, S201, Heterogeneous graph construction: Constructing a heterogeneous graph that includes process parameters, powder properties, monitoring features, and quality nodes; (4) in, For process parameters, The normalized value; These are powder properties, namely, particle size, initial morphology, velocity, and position. For monitoring characteristics, including temperature field, melt index, stability index, etc.; Quality indicators include sphericity, oxygen content, particle size distribution, and flowability. Edge type construction, setting For fully connected parametric coupling edges, The edges represent spatiotemporal evolution, with weights based on kinematic predictions. The edges representing causal relationships are determined using the Granger causality test. S202 and ST-GCN Model Training and Feature Aggregation: Learning Complex Relationships Between Nodes Using Spatiotemporal Graph Convolutional Networks; Mathematical Construction of Spatiotemporal Graph Convolutional Network (ST-GCN) (5) in, Let K be the spatial adjacency matrix of the k-th order Chebyshev polynomial approximation (K=3). The features of the l-th layer node, These are learnable weights; the time dimension is processed through convolution with a kernel size of 3 and an expansion rate that increases with the number of layers. Graph Attention (GAT) Mechanism Construction In processing graph-structured data, to more accurately aggregate neighbor information and generate better node representations, a graph attention mechanism (GAT) is constructed: (6) In the formula, For nodes With nodes Attention coefficient between them Attention weight vector transpose, The weight matrix is ​​a linear transformation matrix. For nodes The input feature vector, For nodes The input feature vector; (7) In the formula, For nodes With nodes The initial weights of the edges between them; For the first A quality objective function, Belonging to the set {1,2,3,4}, it corresponds to the sphericity, oxygen content increment, particle size distribution standard deviation, and Hall flow rate, respectively; For nodes The corresponding process parameters, such as power P, powder feeding rate F, etc.; Such as carrier gas flow rate Q, sheath gas ratio R, etc.; (8) In the formula, For nodes The updated feature vector; For nodes The neighborhood set in the heterogeneous graph includes parameter-coupled neighbors, spatiotemporal evolution neighbors, and causal relationship neighbors; The normalized attention coefficient. For neighboring nodes Features after linear transformation; S203, Output Decoding and Defect Prediction Output decoding: (9) In the formula, Output a tensor for the defect probability distribution field, representing a three-dimensional array where each element is a probability value in the range [0, 1]. This is the node feature matrix of the last (third) layer of the graph neural network; Dimension labeling for real tensors, representing It is a 128*128*4 three-dimensional real tensor; 128 is the spatial resolution, the cross-section of the plasma torch is discretized into a 128*128 pixel grid; 4 is the number of defect categories, the four channels correspond to incomplete spheroidization, satellite sphere, hollow sphere, and oxidation contamination, respectively; The four channels correspond to the probabilities of incomplete spheroidization, satellite spheres, hollow spheres, and oxidative contamination, respectively. The loss function is: (10) In the formula, The total loss function value is the objective to be minimized during training. The smaller the value, the more accurate the model prediction and the stronger the generalization ability. The label is the actual label, with a value of 0 or 1, representing the spatial location. , Does a type c defect exist at position (1 indicates presence, 0 indicates absence)? The predicted probability for the model is a continuous value between [0,1]. For positive sample loss terms, when When =1, it is required that If the value approaches 1, otherwise the penalty term approaches infinity; is the L2 regularization coefficient (hyperparameter); The squared L2 norm of all learnable parameters of the model.

[0032] S3, Bayesian Multi-Objective Optimization Based on the spatial distribution field and historical data of the defect probability in step S2, a Bayesian multi-objective optimization model is constructed with process parameters as decision variables and the objective functions of maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability as the set of objective functions. The model utilizes a multi-output Gaussian process proxy and expected hypervolume improvement acquisition strategy, and iteratively optimizes the model using the L-BFGS-B algorithm to obtain the Pareto front solution set. Specifically, such as Figure 4 As shown, S301, multi-objective function The objective function vector is optimized as follows: (26) in, For sphericity, For the increase in oxygen content, The standard deviation of particle size distribution. To improve powder flowability; The algorithm aims to maximize sphericity, minimize oxygen content, minimize particle size distribution standard deviation, and maximize powder flowability. Specifically: (27) This represents the ratio of the total number of spheroidized products to the total number of products under sampling testing. (28) The difference between the final oxygen quantity and the initial oxygen quantity; (29) (30) The process parameter vector is as follows: (31) in, The power output of the plasma equipment is expressed in kW. This refers to the powder delivery rate, expressed in g / min. This refers to the carrier gas flow rate, expressed in L / min. The sheath gas H2 / Ar ratio is expressed in % (%). Powder particle size, in μm; This indicates the feed position, in mm. The constraints are as follows: (1) Display boundary constraints: (32) Based on experimental experience and equipment parameters, the constraints are as follows: 150 2 8 0 50 (2) Implicit process constraints, as follows: To prevent powder evaporation: Ensure the raw materials are fully melted: Cooling capacity matching: ; The values ​​in the above formulas are provided by PLC data.

[0033] S302, Construction of Multi-Output Gaussian Process Proxy Model like Figure 5 As shown, for each target Establish independent but related GP models: (11) The Matern 5 / 2 kernel function is used to capture the nonlinear relationships between parameters: (12) in, l is the length dimension (naturally optimized by maximizing the log-likelihood). The signal variance; Construct the covariance structure of the multi-output GP using the correlation matrix B between tasks: (13) In the formula, For the extended covariance matrix of a multi-output Gaussian process, Given the input sample point set, For another set of new test input sample points, It is the spatial kernel function matrix; This is a task relevance matrix, describing the strength of the correlation between the four quality objectives (sphericity, oxygen content, particle size distribution, and flowability). S303, Optimization of the Desired Supervolume Improved Acquisition Function Define the current Pareto frontier The hypervolume index relative to the reference point r is: (15) in, This represents the current Pareto frontier, with r as the reference point. Lebesgue measure; Calculate new points The expected increase in hypervolume that can be brought to the current Pareto frontier: (17) In the formula, Let be the scalar of the expected hypervolume improvement, representing the expected hypervolume increment to the current Pareto frontier after experimental evaluation at candidate point x; S is the number of Monte Carlo samplings. For the current Pareto frontier, hypervolume, This represents the extended Pareto front after adding new sampling points. Let s be the target vector sampled from the posterior of GP in the s-th iteration. Use as a reference point; Using the L-BFGS-B algorithm Perform numerical optimization to find the next candidate parameter point that maximizes EHI. .

[0034] S304, Bayesian Iterative Update and Convergence Judgment Initial sampling: Latin hypercube sampling was used to generate initial experimental points.

[0035] Loop execution: running on physical devices Obtain the actual product quality data by using the corresponding process parameters. ; New data ( , Add it to dataset D; Retrain the multi-output GP model; Update the current Pareto frontier .

[0036] Termination condition: Stop when the change in the Pareto front is less than a set threshold or when the maximum number of iterations is reached.

[0037] S4, Zoned Control Execution Based on the Pareto optimal parameter solution set output in step S3, the global variables are decoupled into differentiated control commands for six axial and radial zones through a preset spatial mapping matrix, so as to achieve fine thermal history control in accordance with the non-uniform temperature field inside the plasma torch. Specifically, S401 divides the plasma torch space into six regions. Along the axial direction, the torch top 0-300mm is divided into a central high-temperature zone for initial rapid melting of powder (0-50mm, 8000-12000K), a transitional melting zone for complete melting and spheroidization of particles (50-150mm, 4000-8000K), and an edge cooling zone for rapid solidification of particles (150-300mm, below 4000K). The radial area is divided into a core zone (0-15mm, about 10000K) with the highest energy concentration, an intermediate zone (15-30mm, 6000-10000K), and an outer zone near the cooling wall (30-50mm, 3000-6000K).

[0038] S402, Construct the partition control parameter mapping matrix The global parameters output by Bayesian optimization Mapped to partition mapping instructions : (18) Where M is a 6*6 mapping matrix, This is the region bias vector. This refers to the operating power of the plasma equipment. To ensure powder delivery rate, Carrier gas flow rate, The ratio of sheath gas H2 / Ar is [missing information]. Powder particle size, This is the feed location.

[0039] S403, Differentiated Control Command Execution and Thermal History Control The mapped control commands for each zone are then precisely applied to the corresponding physical areas via actuators (such as RF power supplies, servo motors, flow valves, and linear modules). This zoned differential control allows powder particles of different sizes and initial positions to undergo customized thermal history trajectories as they pass through the plasma torch.

[0040] S5, Reinforcement Learning Defect Repair like Figures 6-8As shown, based on the residual defect state after step S4 and real-time monitoring data, it is modeled as a partially observable Markov decision process (POMDP). The Actor-Critic network is trained using the Proximal Policy Optimization (PPO) algorithm to adaptively output power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects.

[0041] S501, Modeling of Partially Observable Markov Decision Processes (POMDP) The defect repair problem is formalized as a POMDP quintuple (S, A, T, R, O), defined as: (1) State space : (19) Among them, the defect type codes are: [1,0,0,0] = incomplete spheroidization, [0,1,0,0] = satellite sphere, [0,0,1,0] = hollow sphere, [0,0,0,1] = oxidation contamination; (2) Action space A A={ },in, For power fine-tuning, For carrier gas pulse, As compensation for the second batch of powder delivery, For inert purging, No operation is performed; (3) State transition T The above formula is the standard mathematical representation of the state transition function in reinforcement learning.

[0042] in: For state transition functions, which describe how the environment transitions from one state to the next; The state space in the patent includes a 16-dimensional vector containing defect type (4-dimensional), defect location (3-dimensional), defect severity (1-dimensional), current process parameters (6-dimensional), and historical features (2-dimensional). The discrete action space, as described in the patent, refers to five actions: power fine-tuning, carrier gas pulse, secondary powder feeding, inert purging, and no operation. It is a Cartesian product, that is, a combination of states and actions; The probability distribution in the state space indicates that the transfer result has a certain degree of randomness. In the patent, due to the complexity and uncertainty of the plasma spheroidization process, the change in the defect state after the repair action is not deterministic, but follows a certain probability distribution.

[0043] Approximation of the prediction model using graph neural networks: (20) Reward function R: Combining defect elimination reward, quality improvement incentive, operating cost penalty, and deterioration penalty; : (twenty one) In the formula, To map to the set of real numbers, the output is a scalar reward value, with positive numbers encouraging and negative numbers penalizing. The instantaneous reward at time t; The change in defect severity is the difference in defect severity after the repair action is performed, ranging from [-1, 1]. 10 is the indicator function (significant elimination), and 10 is the weighting coefficient for the significant elimination reward. 5 is the indicator function (partial elimination), and 5 is the weighting coefficient for the partial elimination reward. 2 represents the change in sphericity, and 2 represents the weighting coefficient for the quality improvement reward. The repair action selected for time t; As an indicator function (deterioration), when the severity of the defect increases ( When a penalty is triggered, -5 is the weighting coefficient for the worsening penalty; Formula (21) is the standard mathematical representation of the reward function in reinforcement learning. This is the reward function, used to evaluate the quality of performing a certain action in a given state; This is the state space, used to contain the set of all possible states; Action space, which contains the set of all possible actions; It is a Cartesian product, that is, a combination of states and actions.

[0044] (5) Observation O: Based on the real-time monitoring data (such as particle temperature and location) and defect prediction results from step S4.

[0045] S502, Constructing the Actor-Critic Network Architecture The Actor-Critic network architecture includes a neural network that shares a feature extraction layer, a policy network (Actor), and a value network (Critic). First, a two-layer fully connected network (128→128 neurons, ReLU activation) is used as a shared feature layer to extract features from the input state. Then, the policy network, based on the shared features, outputs the probability distribution of each action through a (128→64→5 neurons, Softmax activation) structure to guide action selection. The value network, also with the shared features as input, evaluates the value of the current state through a (128→64→1 neuron, no activation) structure, providing a value reference for policy optimization.

[0046] S503 and PPO algorithm training process A Proximal Policy Optimization (PPO) algorithm is used to train an Actor-Critic network to achieve online defect closure repair. The training process is as follows: First, set hyperparameters such as pruning coefficient Ɛ=0.2, discount factor γ=0.95, GAE parameter λ=0.95, learning rate 3×10⁻⁴, and batch size 64. Then, execute the current policy in the environment, collecting state, action, reward, value, and action probability to complete trajectory collection. Next, calculate the advantage using Generalized Advantage Estimation (GAE) to measure the superiority of actions relative to the average level. Finally, update the network through multiple epochs, using the pruning objective function min( , ), MSE value loss and entropy regularization (encourage exploration) to optimize the total loss L= +0.5 -0.01 .

[0047] S504, Online Defect Repair Execution Logic like Figure 9 As shown, the trained policy network is applied to a real-time scenario to complete the defect closed-loop repair.

[0048] Based on the above-mentioned radio frequency plasma spheroidizing powder preparation method based on multi-objective optimization, this application provides a radio frequency plasma spheroidizing powder preparation system based on multi-objective optimization, comprising: The sensing layer is used to acquire particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and to perform feature extraction and noise reduction preprocessing using signal processing algorithms. The prediction layer is used to construct a spatiotemporal heterogeneous graph of the process based on the extracted particle features. The graph uses process parameters, powder properties, monitoring features and quality indicators as nodes and spatiotemporal evolution relationship and parameter coupling as connection edge rules. The spatiotemporal information between nodes is aggregated through spatiotemporal graph convolutional network and graph attention mechanism to generate a spatial distribution field of defect probability that represents the defect evolution law. The optimization layer is used to construct a Bayesian multi-objective optimization model based on the spatial distribution field of defect probability and historical data. The model uses process parameters as decision variables and the objective functions are maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability. The model uses multi-output Gaussian process proxy and expected hypervolume to improve the acquisition strategy and performs iterative optimization through the L-BFGS-B algorithm to obtain the Pareto front solution set. The control layer is used to decouple global variables into differentiated control commands for six axial and radial zones based on the Pareto optimal parameter solution set based on the output. This is achieved by using a preset spatial mapping matrix to adapt to the non-uniform temperature field inside the plasma torch and realize fine thermal history control. The repair layer is used to model the residual defect state after regulation and real-time monitoring data as a partially observable Markov decision process. The Actor-Critic network is trained using the Proximal Policy Optimization (PPO) algorithm to adaptively output power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects.

[0049] Based on the above-mentioned multi-objective optimization-based radio frequency plasma spheroidizing powder preparation method, this embodiment is a titanium alloy powder spheroidizing experiment conducted on a 200kW radio frequency plasma spheroidizing furnace. Through the closed-loop synergy of five modules—quantum dot labeling monitoring, graph neural network defect prediction, Bayesian multi-objective optimization, partitioned control execution, and reinforcement learning defect repair—the initial hydrogenated dehydrogenated raw material powder with a spheroidization rate of 85%, an oxygen content of 1500ppm, and a Hall flow rate of 32s / 50g was improved to a high-quality spherical powder for additive manufacturing with a spheroidization rate of 98.5%, an oxygen content of 1850ppm, and a Hall flow rate of 22s / 50g after 50 rounds of iterative optimization. The monitoring data and quality indicators of the entire process are gathered in real time on a large data screen, realizing millisecond-level process visualization and adaptive process control from raw material feeding to finished product collection.

[0050] S1. Data Collection and Preprocessing 1.1 Data Source (1) Quantum dot labeling monitoring data: The raw materials were labeled and doped with Eu to enable them to withstand high temperatures. 3+ The oxides were collected using a Hamamatsu C11440-22CU high-speed spectrophotometer with a frame rate of 2000fps, combined with a multi-channel photodetector array and a fiber-coupled plasma spectroscopy diagnostic unit.

[0051] The experimental signal parameters are as follows: Quantum dot fluorescence signal: wavelength range 550-750nm, spectral resolution 2nm, temporal resolution 0.5ms (2000fps). Plasma emission spectroscopy: wavelength range 200-900 nm, used for electronic temperature diagnostics (based on the intensity ratio of Ar I 696.5 nm and Ar II 480.6 nm spectral lines). Background radiation signal: used for subtraction and correction; (2) Process parameter data RF power supply output power P (kW): 150-250kW, sampling frequency 10Hz; Powder feeding rate F (g / min): 20-100g / min, recorded in real time by a loss-in-weight powder feeder; Carrier gas flow rate Q (L / min): 2-8 L / min, collected by mass flow meter; Sheath gas H2 / Ar ratio R: 0-10% by volume, recorded by a gas ratio controller; Powder particle size D (μm): measured offline by laser particle size analyzer, D10, D50, D90; Feed position L (mm): Axial position of powder feeding probe, encoder feedback; (3) Quality index data of product targets Sphericity: SEM image analysis shows the percentage of particles with a sphericity > 0.9; Oxygen content increment: measured by an oxygen and nitrogen analyzer (LECO TC-600); Particle size distribution standard deviation: laser particle size analyzer (Malvern Mastersizer 3000); Hall effect flow rate: Hall effect flow meter (time for a 50g sample to pass through a standard funnel); 1.2 Data Preprocessing def preprocess_data(raw_data): # 1. Wavelet denoising (for fluorescence signals) denoised_signal = wavelet_denoising( signal=raw_data.fluorescence, wavelet='db4', # Daubechies 4th order wavelet level=5, threshold_mode='soft', threshold=universal_threshold ) # 2. Baseline correction (for plasma background radiation) baseline = airPLS( spectrum=raw_data.plasma_emission, lambda_=100, # Smoothing parameter porder=1, itermax=20 ) corrected_spectrum = raw_data.plasma_emission - baseline # 3. Spectral unmixing (separating quantum dot fluorescence from plasma radiation) # Construct the mixing matrix A = [quantum_dot_spectrum, plasma_basis_vectors] A = construct_mixing_matrix() # Nonnegative matrix factorization sources = NMF_unmixing(corrected_spectrum, A, n_components=3) qd_signal = sources[0] # Quantum dot composition # 4. Feature Extraction features = { 'particle_temperature': estimate_temperature( planck_fit(qd_signal, wavelength_range=[550,750]) ), 'melting_index': calculate_liquid_solid_ratio( viscosity_indicator=qd_signal.peak_width, temperature=features['particle_temperature'] ), 'flight_velocity': track_particle( consecutive_frames=5, pixel_resolution=0.05mm, time_interval=0.5ms ), 'plasma_stability': calculate_variance( temperature_field=corrected_spectrum[696.5nm], spatial_window=[128,128] ) } return features 1.3 Summary of Data Preprocessing Table 1 Data Preprocessing II. Model Establishment 2.1 Definition of Multi-Objective Function Optimize the target vector: (26) in, For sphericity, For the increase in oxygen content, The standard deviation of particle size distribution. To achieve powder flowability, the algorithm aims to maximize sphericity, minimize oxygen content, minimize particle size distribution standard deviation, and maximize powder flowability.

[0052] Under sampling testing, the ratio of the total number of spheroidized products to the total number of products: (27) The difference between the final oxygen level and the initial oxygen level: (28) (29) (30) 2.2 Definition of Decision Variables Process parameter vector: (31) in, The power output of the plasma equipment is expressed in kW. This refers to the powder delivery rate, expressed in g / min. This refers to the carrier gas flow rate, expressed in L / min. The sheath gas H2 / Ar ratio is expressed in % (%). Powder particle size, in μm; This indicates the feed position, in mm.

[0053] 2.3 Constraints (1) Display boundary constraints (32) Based on experimental experience and equipment parameters, the constraints are as follows: 150 2 8 0 50 Implicit process constraints are as follows: To prevent powder evaporation: Ensure the raw materials are fully melted: Cooling capacity matching: ; The values ​​in the above formulas are provided by PLC data.

[0054] III. Algorithm Design: Complete Technical Solution 3.1 Stage 1 Algorithm: Quantum Dot Tagging Monitoring Quantum dot fluorescence intensity model: (1) in: To excite light intensity, For temperature-dependent quantum yield, For the emission spectral profile, Thermally activated energy, Local ambient temperature; Particle temperature inversion algorithm (based on dual-wavelength colorimetry): (2) Through experiments, obtain =600nm, =650nm, solve for the particle surface temperature. .

[0055] Melt State Index Calculation: (3) Where α=0.6, β=0, The full width at half maximum (FWHM) of the quantum dot fluorescence spectrum; The weighting coefficients for the temperature term are empirical parameters, such as: = 0.6, representing the contribution of temperature to the molten state; This refers to the solidus temperature. This is the liquidus temperature; For temperature normalization, linear interpolation represents the relative position of particle temperature within the solidus-liquidus interval; These are the weighting coefficients for the spectral terms, empirical parameters, such as: = 0, indicating that the contribution of spectral features to the molten state is 0; For reference, the measured FWHM value of the fully molten powder is usually taken.

[0056] 3.2 Stage 2 Algorithm: Spatiotemporal Prediction Using Graph Neural Networks Heterogeneous graph construction: (4) in, For process parameters, The normalized value; These are powder properties, namely, particle size, initial morphology, velocity, and position. For monitoring characteristics, including temperature field, melt index, stability index, etc.; Quality nodes include sphericity, oxygen content, particle size distribution, and flowability.

[0057] Edge type construction: setting For fully connected parametric coupling edges, The edges represent spatiotemporal evolution, with weights based on kinematic predictions. The edges representing causal relationships are determined using the Granger causality test. ST-GCN Mathematical Construction: (5) in the formula Let be the spatial adjacency matrix of the k-th order Chebyshev polynomial approximation (K=3). Features of the l-th layer nodes; The weights are learnable; the time dimension is processed by convolution with a kernel size of 3 and an expansion rate that increases with the number of layers.

[0058] To more accurately aggregate neighbor information and generate better node representations when processing graph-structured data, we constructed a graph attention mechanism (GAT): (6) (7) (8) Output decoding: (9) The four channels correspond to the probabilities of incomplete spheroidization, satellite spheres, hollow spheres, and oxidative contamination, respectively; the loss function is: (10) 3.3 Stage 3 Algorithm: Bayesian Multi-Objective Optimization The multi-output GP model is established as follows: For each target Establish independent but related GP models: (11) The Matern 5 / 2 kernel function is used to capture the nonlinear relationships between parameters: (12) in, l is the length dimension (naturally optimized by maximizing the log-likelihood). Let V be the signal variance.

[0059] Covariance structure of multi-output GP: (13) Where B is the correlation matrix between tasks.

[0060] Predicted distribution: For a given set of observation data , for new points Prediction: , (14) Desired improvement of the hypervolume acquisition function: Define the hypervolume index as: (15) in, This represents the current Pareto frontier, with r as the reference point. For Lebesgue measure.

[0061] EHI mathematical expression: (16) Monte Carlo approximation calculation: (17) Acquisition function optimization: Optimization using the L-BFGS-B algorithm. The algorithm code is as follows: def bayesian_multiobjective_optimization(): # Initialization X_init = latin_hypercube_sampling(n=10, dim=6, bounds=bounds) Y_init = evaluate_objectives(X_init) # Actual experiment or simulation # Building the initial dataset D = (X_init, Y_init) # Iterative optimization for n in range(10, max_iterations): # 1. Training a multi-output GP model gp_models = train_multioutput_GP(D, kernel='Matern52') # 2. Calculate the current Pareto front pareto_front, is_pareto = compute_pareto_front(Y_init) # 3. Define reference points (worst-case values ​​for each objective + margin) reference_point = np.max(Y_init, axis=0) + 0.1 * (np.max(Y_init, axis=0) - np.min(Y_init, axis=0)) # 4. Optimize EHI acquisition function def neg_ehi(x): return -calculate_ehi(x, gp_models, pareto_front,reference_point, n_samples=1000) x_next, _ = lbfgsb_optimization(neg_ehi, x0=random_point(),bounds=bounds) #5. Evaluate new points y_next = evaluate_objectives(x_next) # Actual experiment # 6. Update the dataset D = (np.vstack([D[0], x_next]), np.vstack([D[1], y_next])) # 7. Convergence Judgment if convergence_check(pareto_front, tolerance=0.01): break return pareto_front, D 3.4 Stage 4 Algorithm: Execution of Zoned Control The spatial division into six regions and the mapping of control parameters are as follows: Table 2. Spatial Six-Region Division and Control Parameter Mapping The mathematical mapping of regional regulation is the global parameters output by Bayesian optimization. Map to partition mapping instructions: (18) Where M is a 6*6 mapping matrix, This is the region bias vector.

[0062] 3.5 Stage 5 Algorithm: Reinforcement Learning Defect Repair The POMDP quintuple is defined as (S, A, T, R, O) for the state space. : (19) Among them, the defect type codes are: [1,0,0,0] = incomplete spheroidization, [0,1,0,0] = satellite sphere, [0,0,1,0] = hollow sphere, [0,0,0,1] = oxidation contamination; the action space A = { The physical operations of} are as follows: Table 3 Physical Operations in Motion Space State transition Approximation using a graph neural network prediction model: (20) The reward function is (twenty one) The Actor-Critic network architecture is as follows: class ActorCritic(nn.Module): def __init__(self, state_dim=16, action_dim=5): super().__init__() # Shared Feature Extraction Layer self.shared = nn.Sequential( nn.Linear(state_dim, 128), nn.ReLU(), nn.Linear(128, 128), nn.ReLU() ) # Actor: Policy Network (outputs action probabilities) self.actor = nn.Sequential( nn.Linear(128, 64), nn.ReLU(), nn.Linear(64, action_dim), nn.Softmax(dim=-1) ) # Critic: Value Network (Output State Value) self.critic = nn.Sequential( nn.Linear(128, 64), nn.ReLU(), nn.Linear(64, 1) ) def forward(self, state): features = self.shared(state) action_probs = self.actor(features) state_value = self.critic(features) return action_probs, state_value PPO algorithm training process: def ppo_train(env, actor_critic, total_timesteps=100000): # Hyperparameters clip_epsilon = 0.2 gamma = 0.95 # Discount factor lambda_gae = 0.95 # GAE parameter learning_rate = 3e-4 batch_size = 64 epochs_per_update = 10 optimizer = Adam(actor_critic.parameters(), lr=learning_rate) for iteration in range(total_timesteps / / batch_size): # 1. Collect trajectories trajectories = [] for _ in range(batch_size): state = env.reset() episode = [] done = False while not done: action_probs, value = actor_critic(state) action = sample(action_probs) next_state, reward, done, _ = env.step(action) episode.append((state, action, reward, value, action_probs[action])) state = next_state trajectories.append(episode) # 2. Compute GAE advantage estimation advantages, returns = compute_gae(trajectories, gamma,lambda_gae) # 3. PPO update (multiple epochs) for epoch in range(epochs_per_update): for batch in minibatch_generator(trajectories, batch_size): states, actions, old_probs, advs, rets = batch # Forward pass new_probs, new_values ​​= actor_critic(states) new_action_probs = new_probs.gather(1,actions.unsqueeze(1)).squeeze() # Calculate the ratio ratio = new_action_probs / old_probs # Crop Target surr1 = ratio * advs surr2 = torch.clamp(ratio, 1-clip_epsilon, 1+clip_epsilon) * advs actor_loss = -torch.min(surr1, surr2).mean() # Critic Loss critic_loss = F.mse_loss(new_values.squeeze(), rets) # Entropy Regularization (Exploration Encouraged) entropy = -(new_probs * torch.log(new_probs)).sum(dim=1).mean() # Total Loss loss = actor_loss + 0.5 * critic_loss - 0.01 * entropy # Backpropagation optimizer.zero_grad() loss.backward() optimizer.step() return actor_critic Online execution logic: def online_defect_repair(gnn_prediction, current_params, policy_network): # 1. Analyzing the defect information in GNN predictions defect_type = np.argmax(gnn_prediction.defect_class) defect_position = gnn_prediction.centroid defect_severity = gnn_prediction.probability # 2. Constructing the state vector state = np.concatenate([ one_hot(defect_type, 4), defect_position, [defect_severity], current_params, get_history_features() # History of the last 2 steps ]) # 3. Policy Network Reasoning with torch.no_grad(): action_probs, _ = policy_network(torch.FloatTensor(state)) action = torch.multinomial(action_probs, 1).item() # 4. Actions are converted into control signals control_signal = action_to_control(action, current_params) # 5. Send to the execution unit send_to_actuators(control_signal) # 6. Monitor the repair results (next cycle) return action, control_signal IV. Solution Processing: Pareto Optimal Solution Set Selection 4.1 Pareto solution set generation After Bayesian optimization iterations, an approximate Pareto optimal solution set is obtained: (22) of which A non-dominated solution.

[0063] 4.2 Final Solution Selection Strategy Strategy 1: Compromise solution selection based on preference information (weighted Chebyshev method) If the decision-maker provides the target weight (e.g., prioritizing sphericity: [0.4, 0.2, 0.2, 0.2]): (twenty three) Strategy 2: Minimize the Euclidean distance based on the ideal point The ideal point that optimizes all objective values ​​is defined as follows: : (twenty four) Strategy 3: Knee point selection based on hypervolume contribution Choose the solution that maximizes both the marginal hypervolume contribution and the curvature (i.e., the "knee point" of the Pareto front): (25) Strategy 4: Multi-attribute decision making The code is as follows: def topsis_selection(pareto_solutions, weights): # 1. Construct and standardize the decision matrix X = np.array([sol.objectives for sol in pareto_solutions]) # Np× 4 X_norm = X / np.sqrt(np.sum(X**2, axis=0)) # 2. Weighted Standardization X_weighted = X_norm * weights # 3. Determine the positive and negative ideal solutions ideal_best = np.max(X_weighted, axis=0) # Maximize the target value and minimize it. ideal_worst = np.min(X_weighted, axis=0) # 4. Calculate the Euclidean distance d_best = np.sqrt(np.sum((X_weighted - ideal_best)**2, axis=1)) d_worst = np.sqrt(np.sum((X_weighted - ideal_worst)**2, axis=1)) # 5. Calculate relative proximity closeness = d_worst / (d_best + d_worst) # 6. Select the optimal solution best_idx = np.argmax(closeness) return pareto_solutions[best_idx] 4.3 Online switching logic for strategy selection def select_final_solution(pareto_front, production_context): """ Automatically select the Pareto solution based on the production context. """ if production_context.priority == 'quality_first': # Quality Priority: Select the solution with the best overall performance among those with a spheroidization rate > 98% and an oxygen content < 500 ppm. candidates = [sol for sol in pareto_front if sol.sphericity > 0.98 and sol.oxygen_delta <500] return max(candidates, key=lambda x: x.hall_flow_rate) elif production_context.priority == 'cost_first': # Cost priority: Select a solution with power <200kW and powder feeding rate >80g / min candidates = [sol for sol in pareto_front if sol.power < 200 and sol.feed_rate > 80] return min(candidates, key=lambda x: x.oxygen_delta) elif production_context.priority == 'balanced': # Equilibrium Strategy: Using Knee Point Selection return knee_point_selection(pareto_front) else: # Default: Weighted Chebyshev, weights derived from historical preference learning weights = learn_weights_from_history() return weighted_chebyshev(pareto_front, weights) The core innovation of this invention lies in the construction of a closed-loop radio frequency plasma spheroidized powder quality control technology system that integrates sensing, prediction, optimization, regulation, and repair. Its fundamental breakthrough, which distinguishes it from existing technologies, is reflected in five mutually coupled technical levels.

[0064] At the process monitoring level, this invention creatively introduces quantum dot labeling technology to solve the problem of in-situ monitoring of powder particle state under high-temperature plasma conditions. Existing technologies rely on infrared thermometry or high-speed imaging, which can only acquire the overall temperature field of the plasma torch or the macroscopic motion trajectory of particles, but cannot track the real-time temperature history and melting state of individual powder particles. This invention selects CdSe / ZnS core-shell structured quantum dots as labeling agents, utilizing their emission peak located at 550-750nm and effectively distinguishable from the plasma background radiation spectrum. Combined with a high-speed spectral acquisition system with a frame rate of no less than 1000fps, it achieves for the first time the precise capture of the millisecond-level dynamic behavior of individual powder particles during radio frequency plasma spheroidization, increasing the fluorescence signal intensity by two to three orders of magnitude, providing a high temporal resolution and high sensitivity data foundation for subsequent intelligent prediction and control.

[0065] At the defect prediction level, this invention constructs a heterogeneous graph model based on a spatiotemporal graph convolutional network, overcoming the limitations of traditional methods in modeling the complex coupling relationships between process parameters. Existing technologies mostly employ physics-based numerical simulations or simple statistical regression models, which struggle to simultaneously characterize the multi-scale correlations and spatiotemporal evolution patterns between process parameters, powder properties, monitoring features, and product quality. This invention abstracts the radio frequency plasma spheroidization process into a heterogeneous graph structure containing process parameter nodes, powder property nodes, monitoring feature nodes, and product quality nodes. It extracts the coupling relationships and spatiotemporal evolution features between nodes through a spatiotemporal graph convolutional network architecture, and adaptively weights the critical path using a graph attention mechanism, outputting a 128×128×4 three-dimensional defect probability distribution field. This achieves early warning of the spatial distribution and evolution trends of four types of defects—incomplete spheroidization, satellite spheres, hollow spheres, and oxidation contamination—50 to 100 milliseconds in advance, with a defect prediction accuracy exceeding 96%.

[0066] At the multi-objective optimization level, this invention employs a multi-objective Bayesian optimization algorithm based on Gaussian process regression to solve the sample efficiency problem of synergistic optimization of multiple quality indicators in a high-dimensional process parameter space. Traditional empirical trial-and-error methods or orthogonal experimental design require hundreds of experiments to locate optimal parameters and cannot handle Pareto optimal trade-offs between mutually restrictive indicators such as spheroidization rate, oxygen content, particle size distribution, and powder flowability. This invention constructs a multi-output Gaussian process surrogate model based on the Matérn kernel function for six key process parameters, using expected hypervolume improvement as the acquisition function to achieve an adaptive balance between exploration and utilization. It obtains a Pareto front solution set superior to traditional methods while reducing the number of experiments by 60%, increasing the spheroidization rate from 85% to 98.5% while controlling the oxygen content increment within the range of 300 to 500 ppm.

[0067] At the process control level, this invention establishes a six-region zoned fine-grained control mechanism for the plasma torch, overcoming the shortcomings of existing global unified control that fail to consider spatial non-uniformity. Existing technologies use uniform power and gas flow settings for the entire plasma torch, failing to utilize the significant temperature gradient information along the axial direction from the central high-temperature zone to the edge cooling zone and along the radial direction from the core zone to the outer zone. This invention divides the plasma torch axially into a central high-temperature zone, a transition melting zone, and an edge cooling zone, and radially into a core zone, an intermediate zone, and an outer zone. A mapping matrix between Bayesian optimized output parameters and zoned control commands is established for the thermodynamic characteristics of different regions, enabling differentiated dynamic control of the powder feeding probe position, carrier gas injection angle, sheath gas flow ratio, and cooling chamber temperature gradient, thus optimizing the thermal history trajectory of powder particles in the plasma torch.

[0068] At the defect repair level, this invention introduces a reinforcement learning policy network based on a proximal policy optimization algorithm, filling the gap in existing technologies that can only screen residual defects offline but lack online repair capabilities. Even after optimization and control, 5% to 15% of residual defects will still be generated during the spheroidization process, which traditional methods are helpless against. This invention models defect repair as a partially observable Markov decision process, defining a state space that includes defect type, location, severity, and current process parameter configuration, and a discrete action space that includes power fine-tuning, carrier gas pulse injection, secondary powder feeding compensation, and inert gas purging. A multi-dimensional reward function is designed that comprehensively considers defect elimination rewards, operating cost penalties, and quality improvement incentives. The optimal repair action sequence is output through an Actor-Critic architecture policy network, and the execution unit performs cross-level collaborative processing on residual defects, achieving a defect repair success rate of over 85%.

[0069] The five technical aspects mentioned above are not simply superimposed, but rather form a deeply coupled closed-loop system through data flow and control flow: quantum dot monitoring data drives the update of the graph neural network model, prediction results guide Bayesian optimization decisions, optimization parameters are mapped to partitioned control instructions, the control effects are fed back to the reinforcement learning policy network, and repair experience further enriches the training dataset of the surrogate model, enabling the system to continuously learn and optimize itself. This intelligent closed-loop control architecture, which involves the collaboration of multiple technologies, elevates the preparation of radio frequency plasma spheroidized powder from traditional experience-dependent manufacturing to a data-driven next-generation intelligent manufacturing paradigm.

[0070] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0071] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.

Claims

1. A method for preparing RF plasma spheroidized powder based on multi-objective optimization, characterized in that, Includes the following steps: S1. Obtain particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and perform feature extraction and noise reduction preprocessing using signal processing algorithms; S2. Based on the particle features extracted in step S1, construct a process spatiotemporal heterogeneous graph with process parameters, powder properties, monitoring features and quality indicators as nodes, and spatiotemporal evolution relationship and parameter coupling as connection edge rules. A spatiotemporal information aggregation between nodes is realized through spatiotemporal graph convolutional network and graph attention mechanism to generate a spatial distribution field of defect probability that represents the defect evolution law. S3. Based on the spatial distribution field of the defect probability and historical data in step S2, construct a Bayesian multi-objective optimization model with process parameters as decision variables and the objective functions of maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability as the set of objective functions. Use a multi-output Gaussian process proxy model and the expected hypervolume improvement acquisition strategy to iteratively optimize the multi-objective optimization model and obtain the Pareto front solution set. S4. Based on the Pareto optimal parameter solution set output in step S3, the global variables are decoupled into differentiated control commands for six axial and radial zones through a preset spatial mapping matrix, so as to achieve fine thermal history control in accordance with the non-uniform temperature field inside the plasma torch. S5. Based on the residual defect state after adjustment in step S4 and the real-time monitoring data, a partially observable Markov decision process is modeled. An Actor-Critic network is trained using a proximal policy optimization algorithm to adaptively output power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects.

2. The method of claim 1, wherein the method is based on multi-objective optimization. Step S2, based on the particle features extracted in step S1, involves constructing a spatiotemporal heterogeneous graph of a process, with process parameters, powder properties, monitoring features, and quality indicators as nodes, and spatiotemporal evolution relationships and parameter coupling as connection rules. This graph then uses a spatiotemporal graph convolutional network and graph attention mechanism to aggregate spatiotemporal information between nodes, generating a spatial distribution field representing the defect probability and its evolutionary pattern. Specifically, this includes: S201. Heterogeneous graph construction: Construct a heterogeneous graph that includes process parameters, powder properties, monitoring features, and quality nodes; (4) in, For process parameters, The normalized value; These are powder properties, namely, particle size, initial morphology, velocity, and position. For monitoring characteristics, including temperature field, melt index, stability index, etc.; Quality indicators include sphericity, oxygen content, particle size distribution, and flowability. Edge type construction, setting For fully connected parametric coupling edges, The edges represent spatiotemporal evolution, with weights based on kinematic predictions. The edges representing causal relationships are determined using the Granger causality test. S202, Spatiotemporal Graph Convolutional Network Model Training and Feature Aggregation: Learning complex relationships between nodes using spatiotemporal graph convolutional networks; The spatiotemporal graph convolutional network model is represented as follows: (5) in, Let K be the spatial adjacency matrix of the k-th order Chebyshev polynomial approximation (K=3). The features of the l-th layer node, These are learnable weights; the time dimension is processed through convolution with a kernel size of 3 and an expansion rate that increases with the number of layers. The graph attention mechanism is represented as: (6) In the formula, For nodes With nodes Attention coefficient between them Attention weight vector transpose, The weight matrix is ​​a linear transformation matrix. For nodes The input feature vector, For nodes The input feature vector; (7) In the formula, For nodes With nodes The initial weights of the edges between them; For the first A quality objective function, Belonging to the set {1,2,3,4}, it corresponds to the sphericity, oxygen content increment, particle size distribution standard deviation, and Hall flow rate, respectively; For nodes The corresponding process parameters, For nodes Corresponding process parameters; (8) In the formula, For nodes The updated feature vector, For nodes The neighborhood set, The normalized attention coefficient. For neighboring nodes Features after linear transformation; S203, Output Decoding and Defect Prediction Output decoding: (9) In the formula, Output the tensor for the defect probability distribution field. This is the node feature matrix of the last layer of the graph neural network. Dimension labeling for real tensors, representing It is a 128*128*4 three-dimensional real tensor; It is a multilayer perceptron, which contains 2-3 layers of linear transformations and activation functions, and maps the node features output by the graph neural network to the defect probability space; The four channels correspond to the probabilities of incomplete spheroidization, satellite spheres, hollow spheres, and oxidative contamination, respectively. The loss function is: (10) In the formula, This is the total loss function value. For real labels, To predict probabilities for the model, For positive sample loss terms, is the L2 regularization coefficient, which controls the strength of the penalty for model complexity; The squared L2 norm of all learnable parameters of the model.

3. The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to claim 2, characterized in that, Step S3, based on the spatial distribution field of defect probabilities and historical data from step S2, involves constructing a Bayesian multi-objective optimization model with process parameters as decision variables and maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability as objective functions. The step of iteratively optimizing the multi-objective optimization model using a multi-output Gaussian process surrogate model and an expected hypervolume improvement acquisition strategy to obtain the Pareto front solution set specifically includes: S301, The multi-objective function is expressed as: (26) in, For sphericity, For the increase in oxygen content, The standard deviation of particle size distribution. To improve powder flowability; The process parameter vector is as follows: (31) in, The power output of the plasma equipment is expressed in kW. This refers to the powder delivery rate, expressed in g / min. This refers to the carrier gas flow rate, expressed in L / min. The sheath gas H2 / Ar ratio is expressed in % (%). Powder particle size, in μm; This indicates the feed position, in mm. S302. Constructing a multi-output Gaussian process proxy model For each target Establish independent but related Gaussian process models: (11) This is a vector of process parameters (a vector of decision variables). The Matern 5 / 2 kernel function is used to capture the nonlinear relationships between parameters: (12) in, , where l is the length dimension. The signal variance; Construct the covariance structure of the multi-output Gaussian process surrogate model using the correlation matrix B between tasks: (13) In the formula, For the extended covariance matrix of a multi-output Gaussian process, Given the input sample point set, For another set of new test input sample points, It is the spatial kernel function matrix; S303, Optimization of the Desired Supervolume Improved Acquisition Function Define the current Pareto frontier The hypervolume index relative to the reference point r is: (15) in, This represents the current Pareto frontier, with r as the reference point. Lebesgue measure; Calculate new points The expected increase in hypervolume that can be brought to the current Pareto frontier: (17) In the formula, Let S be the desired supervolume improvement value, and S be the number of Monte Carlo samplings. For the current Pareto frontier, hypervolume, This represents the extended Pareto front after adding new sampling points. Let s be the target vector sampled from the posterior of GP in the s-th iteration. Use as a reference point; Using the L-BFGS-B algorithm Perform numerical optimization to find the next candidate parameter point that maximizes EHI. ; S304, Bayesian Iterative Update and Convergence Judgment.

4. The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to claim 1, characterized in that, Step S4, based on the Pareto optimal parameter solution set output in step S3, decouples the global variables into differentiated control commands for six axial and radial zones through a preset spatial mapping matrix, in order to achieve refined thermal history control in accordance with the non-uniform temperature field inside the plasma torch. Specifically, this includes: S401. Divide the plasma torch space into six regions. The area of ​​0-300mm at the top of the torch along the axial direction is divided into a central high-temperature zone for initial rapid melting of powder, a transitional melting zone for complete melting and spheroidization of particles, and an edge cooling zone for rapid solidification of particles. The region with a radius of 0-50mm is divided radially into a core area where energy is most concentrated, a middle area, and an outer area near the cooling wall; S402, Construct the partition control parameter mapping matrix The global parameters output by Bayesian optimization Mapped to partition mapping instructions : (18) Where M is a 6*6 mapping matrix, This is the region bias vector. This refers to the operating power of the plasma equipment. To ensure powder delivery rate, Carrier gas flow rate, The ratio of sheath gas H2 / Ar is [missing information]. Powder particle size, This is the feed location.

5. The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to claim 1, characterized in that, The steps of S5, based on the residual defect state after adjustment in step S4 and real-time monitoring data, are modeled as a partially observable Markov decision process. An Actor-Critic network is trained using a proximal policy optimization algorithm to adaptively output power fine-tuning or airflow injection action commands, achieving online closed-loop defect repair. Specifically, these steps include: S501. The defect repair problem is formalized as a POMDP quintuple (S,A,T,R,O). S502. Construct an Actor-Critic network architecture, which includes: a shared feature extraction layer, a policy network, and a value network; S503. An Actor-Critic network is trained using a proximal strategy optimization algorithm to achieve online closed-loop defect repair.

6. The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to claim 5, characterized in that, The step S501, which formalizes the defect repair problem into a POMDP quintuple (S, A, T, R, O), specifically includes: (1) State space : (19) Among them, the defect type codes are: [1,0,0,0] = incomplete spheroidization, [0,1,0,0] = satellite sphere, [0,0,1,0] = hollow sphere, [0,0,0,1] = oxidation contamination; (2) Action space A A={ },in, For power fine-tuning, For carrier gas pulse, As compensation for the second batch of powder delivery, For inert purging, No operation is performed; State transition T: describes the rules that describe how the environment evolves from the current state to the next state after the agent performs an action; ; Approximation of the prediction model using graph neural networks: (20) (4) Reward function The comprehensive system includes rewards for defect elimination, incentives for quality improvement, penalties for operational costs, and penalties for deterioration. : (21) In the formula, To map to the set of real numbers, the output is a scalar reward value, with positive numbers encouraging and negative numbers penalizing. The instantaneous reward at time t; This represents the change in the severity of the defect. (5) Observation O: Based on the real-time monitoring data and defect prediction results of step S4.

7. The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to claim 5, characterized in that, S502, constructing the Actor-Critic network architecture, which includes the steps of sharing a feature extraction layer, a policy network, and a value network, specifically includes: First, a two-layer fully connected network with 128 neurons each is used as a shared feature layer to extract features from the input state, employing the ReLU activation function. Then, a policy network, based on these shared features, is constructed through a hierarchical structure from 128 neurons to 64 neurons and then to 5 neurons, using the Softmax activation function to output the probability distribution of each action, guiding action selection. The value network, also using the shared features as input, is constructed through a hierarchical structure from 128 neurons to 64 neurons and then to 1 neuron, without employing an activation function, to evaluate the value of the current state, providing a value reference for policy optimization.

8. The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to claim 1, characterized in that, The steps of S1, which involve acquiring particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and performing feature extraction and noise reduction preprocessing using signal processing algorithms, specifically include: The quantum dot fluorescence intensity model is expressed as: (1) in, To excite light intensity, For temperature-dependent quantum yield, For the emission spectral profile, Thermally activated energy, Local ambient temperature; The particle temperature inversion algorithm based on dual-wavelength colorimetry is expressed as follows: (2) Through experiments, obtain =600nm, =650nm, solve for the particle surface temperature. ; Melt State Index The calculation is as follows: (3) Among them, the weighting coefficient for the temperature term is α=0.6, and the weighting coefficient for the spectral term is β=0. The full width at half maximum (FWHM) of the quantum dot fluorescence spectrum. For reference, half-height and full width.

9. A radio frequency plasma spheroidizing powder preparation system based on multi-objective optimization, characterized in that, The method for preparing radio frequency plasma spheroidized powder based on multi-objective optimization according to any one of claims 1-8, wherein the system comprises: The sensing layer is used to acquire particle state and process parameter data through quantum dot labeling and high-speed spectral acquisition technology, and to perform feature extraction and noise reduction preprocessing using signal processing algorithms. The prediction layer is used to construct a spatiotemporal heterogeneous graph of the process based on the extracted particle features. The graph uses process parameters, powder properties, monitoring features and quality indicators as nodes and spatiotemporal evolution relationship and parameter coupling as connection edge rules. The spatiotemporal information between nodes is aggregated through spatiotemporal graph convolutional network and graph attention mechanism to generate a spatial distribution field of defect probability that represents the defect evolution law. The optimization layer is used to construct a Bayesian multi-objective optimization model based on the spatial distribution field of defect probability and historical data. The model has process parameters as decision variables and the objective functions are maximizing sphericity, minimizing oxygen content increment, minimizing particle size distribution standard deviation, and maximizing powder flowability. The multi-output Gaussian process proxy model and the expected hypervolume improvement acquisition strategy are used to iteratively optimize the multi-objective optimization model to obtain the Pareto front solution set. The control layer is used to decouple global variables into differentiated control commands for six axial and radial zones based on the Pareto optimal parameter solution set based on the output. This is achieved by using a preset spatial mapping matrix to adapt to the non-uniform temperature field inside the plasma torch and realize fine thermal history control. The repair layer is used to model the residual defect state after regulation and real-time monitoring data as a partially observable Markov decision process. The Actor-Critic network is trained using a proximal policy optimization algorithm to adaptively output power fine-tuning or airflow injection action commands to achieve online closed-loop repair of defects.

10. An application of the system according to claim 9, characterized in that, Used for quality control of radio frequency plasma spheroidized powder.