Transformer design and process adaptive optimization method and system based on deep learning

By constructing a multimodal data representation and fusion framework using deep learning methods, and combining a multiphysics proxy model and a digital twin, the problem of information silos in transformer design and manufacturing was solved. This enabled real-time optimization and adaptive control of the transformer design and manufacturing process, improving product performance consistency and manufacturing efficiency.

CN121960009APending Publication Date: 2026-05-01ZHENLAI XINYUAN COMPOSITE MATERIAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511931108.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The existing transformer design and manufacturing process suffers from information silos between the design and manufacturing stages, lacking data correlation. This results in the inability to optimize design schemes in real time, poor product performance consistency, high rework rates, and the failure of existing artificial intelligence applications to connect the entire chain.

Method used

A multimodal data representation and fusion framework based on deep learning is constructed. By combining a multiphysics proxy model and a digital twin, real-time data interaction and adaptive optimization in design and manufacturing are achieved. Decision-making is carried out through a deep reinforcement learning model to form closed-loop control.

Benefits of technology

It has enabled seamless information flow throughout the entire design and manufacturing process, improving product performance consistency and manufacturing flexibility, reducing rework rates, and enhancing design efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960009A_ABST
    Figure CN121960009A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of transformers, in particular to a transformer design and process adaptive optimization method and system based on deep learning. Through a multi-modal data fusion technology, multi-source heterogeneous data in design, material and manufacturing processes are represented in a unified manner. A physical information neural network is utilized to construct an efficient multi-physical field agent model, and an optimization design scheme is quickly generated in combination with a material knowledge graph. And the physical manufacturing process is mapped in real time and the product performance is predicted through the digital twin. And finally, on the basis of a deep reinforcement learning model, according to the state prediction and target deviation of the digital twin, adaptively adjusting subsequent process parameters or feeding back a fine tuning design scheme, and forming a'design-process-detection-optimization 'closed loop. According to the method, the problems that the design and manufacturing links of the transformer are separated, the transformer depends on experience, and online optimization is difficult are solved, data driving and self-adaptive collaboration of the whole process are achieved, and the product performance consistency, the design efficiency and the manufacturing intelligence level are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of transformer technology, and specifically to a method and system for adaptive optimization of transformer design and manufacturing process based on deep learning. Background Technology

[0002] As a core component of the power system, the performance, reliability, and manufacturing cost of transformers are of paramount importance. Traditional transformer design and manufacturing follow a sequential, experience-dependent model, resulting in numerous technical bottlenecks. During the design phase, engineers rely on empirical formulas and finite element simulations for electromagnetic and structural design. This process heavily depends on expert knowledge, and high-fidelity multiphysics simulations are costly and time-consuming, limiting the optimization space for design solutions and making it difficult to achieve a globally optimal match between performance, cost, and reliability. Material selection is also often based on fixed databases and experience, lacking dynamic and intelligent correlation with specific electromagnetic, thermal, and mechanical performance requirements.

[0003] During the manufacturing stage, process parameters (such as winding tension, lamination pressure, and curing temperature profiles) are typically set according to fixed process cards, lacking the ability to adaptively adjust to real-time disturbances such as material batch differences, equipment condition fluctuations, and environmental factors. Although modern production lines are equipped with various sensors and visual inspection equipment capable of collecting massive amounts of data, this data often exists in isolation and fails to be effectively linked to design models and performance targets, forming "data silos." Deviations generated during manufacturing (such as uneven winding gaps and insulation layer thickness fluctuations) can only be detected in the final testing stage, making real-time process diagnosis and compensation impossible, resulting in poor product performance consistency and high rework and scrap rates.

[0004] In recent years, artificial intelligence (AI) technology has provided new ideas for industrial optimization, but existing applications are mostly improvements in isolated aspects. For example, some studies use neural networks to predict transformer faults or image recognition to detect surface defects; however, these technologies do not connect the entire chain from design to manufacturing. Intelligent optimization models in the design stage (such as surrogate models) lack feedback and correction from manufacturing site data; and the inspection data in the manufacturing stage cannot drive the iteration of design solutions. How to construct a data flow and decision-making closed loop that spans the entire life cycle of a transformer, and achieve deep collaboration and adaptive optimization of design, materials, processes, and inspection, is a technical challenge that urgently needs to be solved in this field. Deep reinforcement learning shows potential in complex decision-making, but its application in industrial scenarios with high reliability and strong physical constraints, such as transformer manufacturing, still faces challenges such as poor decision interpretability, difficulty in guaranteeing safety boundaries, and difficulty in multimodal information fusion.

[0005] Therefore, the existing technology still needs further development. Summary of the Invention

[0006] The purpose of this invention is to overcome the above-mentioned technical deficiencies and provide a method and system for adaptive optimization of transformer design and process based on deep learning, so as to solve the problems existing in the prior art.

[0007] To achieve the above-mentioned technical objectives, according to a first aspect of the present invention, the present invention provides a deep learning-based adaptive optimization method for transformer design and manufacturing process, comprising: S1. Data Acquisition and Multimodal Characterization Steps: Collect multiphysics simulation data and material property data during the transformer design phase, as well as real-time process sensing data and visual inspection data during the manufacturing phase, and map the simulation data, property data, sensing data and visual data into a unified feature vector for characterization. S2. Multi-physics field collaborative optimization design steps: Based on the unified feature vector, the electromagnetic and thermal performance indicators under different combinations of design parameters are solved using a pre-trained multi-physics field proxy model. Combined with the material performance matching strategy, an electromagnetic structure design scheme and material selection scheme that meet the preset performance target are generated. S3. Process-driven and digital twin interaction steps: The electromagnetic structure design scheme and material selection scheme are converted into a sequence of key process parameters to drive the manufacturing equipment; during the manufacturing process, the state of the digital twin is updated based on the real-time process sensing data and visual inspection data, and the digital twin simulates the manufacturing process and product performance evolution. S4. Adaptive Decision-Making and Closed-Loop Optimization Steps: Based on the updated digital twin state, the deviation between the current manufacturing state and the expected target is analyzed using a preset decision model, and adjustment instructions for the subsequent process parameter sequence or fine-tuning instructions for the electromagnetic structure design scheme are generated to achieve online adaptive adjustment of the manufacturing process or re-optimization of the design scheme.

[0008] Specifically, in step S1, the real-time process sensing data includes at least one of winding tension, pressing pressure, curing temperature, and partial discharge signal; the visual inspection data includes image feature data obtained by imaging and recognizing the winding arrangement, core laminations, and insulation packaging.

[0009] Specifically, in step S2, the multiphysics proxy model is a physical information neural network, which is trained by incorporating the constraints of the electromagnetic and thermodynamic physical laws followed by the transformer.

[0010] Specifically, the material performance matching strategy is implemented through a materials knowledge graph, which links the composition, preparation process, electrical properties, thermal properties and mechanical properties of different materials.

[0011] Specifically, in step S3, the digital twin receives the sequence of key process parameters as input and simulates the resulting changes in the electromagnetic structure and material state of the transformer, thereby outputting predicted product electromagnetic performance and insulation performance indicators.

[0012] Specifically, in step S4, the decision model is a deep reinforcement learning model, which takes the state of the digital twin as the state input, the process parameter adjustment action or the design parameter fine-tuning action as the action output, and the conformity between the product performance index and the manufacturing target as the reward signal for training.

[0013] Specifically, the decision-making process of the deep reinforcement learning model integrates an attention mechanism to identify the state features that have the greatest impact on the current decision.

[0014] Specifically, the action output of the deep reinforcement learning model is subject to physical and technological constraints. The physical constraints include upper temperature limits and stress limits, while the technological constraints include the range of equipment operating parameters.

[0015] Specifically, the reward function of the deep reinforcement learning model includes: a basic reward based on product performance indicators, a penalty based on whether the physical and process constraints are violated, and an additional reward based on action smoothness and interpretability.

[0016] According to a second aspect of the present invention, a deep learning-based adaptive optimization system for transformer design and manufacturing process is provided, comprising: The data perception and fusion module is used to execute the S1 step, collect and uniformly represent multi-source heterogeneous data; The collaborative design and simulation module is used to execute the S2 step, run the multiphysics proxy model and perform material matching to generate an optimized design scheme. The process execution and digital twin module is used to execute the S3 step, convert the design scheme into process instructions, and run a digital twin that interacts synchronously with the physical manufacturing process; The intelligent decision-making and closed-loop control module is used to execute the S4 step, run the decision model, generate adaptive adjustment instructions, and send them to the manufacturing equipment or design system.

[0017] Beneficial effects: First, this invention, by constructing a unified multimodal data representation and fusion framework, maps design simulation data, material property data, real-time process sensing data, and visual inspection data to a unified feature space for the first time, completely breaking down the "information silos" between design, manufacturing, and inspection stages in the traditional model. This allows minute deviations in the manufacturing process (such as millimeter-level changes in winding gaps) to be quantified in real time as key features affecting the final electromagnetic performance and insulation reliability, and accurately compared with design targets. This provides high-quality, highly correlated input for subsequent intelligent decision-making, realizing end-to-end information connectivity and collaboration from design to manufacturing.

[0018] Secondly, this invention introduces a physical information neural network trained with constraints from physical laws as a multiphysics proxy model. While ensuring the physical plausibility of the predictions, it reduces the time required for electromagnetic-thermal coupling simulations (which previously took hours or even days) to seconds, making rapid and thorough optimization exploration possible within a vast design space. Combined with knowledge graph-based intelligent material matching, the system can efficiently generate globally optimal or Pareto optimal design and material solutions, considering multiple objectives such as performance, cost, and manufacturability, greatly improving design efficiency and innovation.

[0019] Third, this invention creatively constructs a high-fidelity digital twin that interacts synchronously with the physical manufacturing process. This twin integrates equivalent circuits, thermal networks, and aging kinetic models, and can dynamically update its state parameters based on real-time process data. This enables the system to predict the final performance indicators of the product in near real-time before the physical product is manufactured, achieving "predictive" manufacturing process. Combined with a deep reinforcement learning decision model, the system can automatically analyze the root causes of deviations based on the predicted state provided by the digital twin, and generate online adjustment instructions for subsequent process parameters or fine-tuning suggestions for the original design scheme, forming a real-time adaptive closed-loop control of "perception-prediction-decision-execution," significantly improving the flexibility of the manufacturing process and the consistency of product performance.

[0020] Fourth, the decision-making model of this invention effectively solves the interpretability and safety challenges of artificial intelligence decision-making in industrial scenarios by integrating an attention mechanism and designing a reward function with multi-objective constraints. The attention mechanism can identify the key state features on which each decision is based, enhancing trust in human-computer interaction; while the strong safety penalty term and action smoothness constraints embedded in the reward function ensure that all explorations and decisions of the agent are within the preset physical safety and process feasibility boundaries, eliminating dangerous or invalid operations, enabling this highly intelligent system to be safely and robustly applied in real industrial environments.

[0021] Fifth, the system constructed by this invention forms a self-evolving and self-learning intelligent agent. New data generated during the manufacturing process (including process parameters, test results, and final product test data) can be continuously fed back to update and optimize the physical information neural network, visual inspection model, and reinforcement learning strategy model. This enables the system to continuously accumulate experience, adapt to the introduction of new materials and processes, handle unprecedented manufacturing fluctuations, continuously improve design accuracy and process control, and ultimately achieve the comprehensive optimization of transformer product energy efficiency, reliability, and manufacturing cost, thus promoting the transformation of the transformer manufacturing industry towards a truly intelligent and adaptive production mode. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating the deep learning-based adaptive optimization method for transformer design and manufacturing process provided in a specific embodiment of the present invention. Figure 2 This is a schematic diagram of the system composition of the transformer design and process adaptive optimization system based on deep learning provided in a specific embodiment of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the embodiments in this application, other similar embodiments obtained by those skilled in the art without creative effort should all fall within the scope of protection of this application. Furthermore, directional terms mentioned in the following embodiments, such as "up," "down," "left," and "right," are only for reference to the directions in the accompanying drawings; therefore, the directional terms used are for illustrative purposes and not for limiting the invention.

[0024] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments.

[0025] Please see Figure 1 This invention provides a deep learning-based adaptive optimization method for transformer design and manufacturing process, comprising: S1. Data Acquisition and Multimodal Characterization Steps: Collect multiphysics simulation data and material property data during the transformer design phase, as well as real-time process sensing data and visual inspection data during the manufacturing phase, and map the simulation data, property data, sensing data and visual data into a unified feature vector for characterization. It should be further explained that the specific data acquisition and processing methods in step S1 are as follows: During the design phase, parametric finite element simulations are performed using commercial software such as ANSYS Maxwell and ANSYS Mechanical. The simulation data includes: for a combination of design variables (e.g., high-voltage winding turns Nhv=500, low-voltage winding turns Nlv=100, core window height Hw=400mm, window width Ww=150mm, silicon steel sheet grade 30QG120, insulation paper thickness tins=2mm), the calculated magnetic flux density B (unit: T), magnetic field strength H (unit: A / m), core loss density Pfe (unit: W / m³), winding loss density Pcu (unit: W / m³), and temperature T (unit: °C) at the nodes of the three-dimensional spatial mesh under rated load conditions. These field data are preprocessed into normalized three-dimensional tensors with dimensions, for example, 64x64x64x5 (representing a 64³ spatial mesh with 5 physical fields). Material property data is extracted from internal or public databases (such as Mat Web) and stored in structured tables, for example: {“Material ID”:“Nomex410”, “Relative Permittivity”:3.2”, “Volume Resistivity”:1E15Ω·cm”, “Thermal Conductivity”:0.15W / (m·K)”, “Breakdown Field Strength”:20kV / mm}. Real-time process sensing data is acquired through a deployed sensor network: winding tension is acquired using an S-type tension sensor (model example: ZNL-102, range 0-30N, accuracy ±0.05N) at a sampling frequency of 100Hz; core pressing pressure is acquired using a pressure transmitter (range 0-1000kN); curing temperature during vacuum casting is acquired using a Pt100 platinum resistance thermometer inserted inside the winding; partial discharge signals are acquired during winding withstand voltage tests using a Rogowski coil and a high-frequency data acquisition card (sampling rate 100MS / s). Visual inspection data is captured as JPG images by 5-megapixel industrial cameras (such as Baslerac A2440-75um) installed at key workstations under uniform backlighting. Examples include panoramic side views of completed windings and end-face views of stacked iron cores. The specific algorithm for "mapping to a unified feature vector" is as follows: For 3D simulation field data, a 3D convolutional neural network (3D-CNN) is used for feature extraction. This network contains four 3D convolutional layers with 16, 32, 64, and 128 filters, all with a kernel size of 3x3x3 and a stride of 2. Each convolutional layer is followed by a ReLU activation and a 3D batch normalization layer. Finally, a global average pooling layer converts the feature map into a 128-dimensional vector Fsim. For tabular material property data, a fully connected encoder network is used. This network has three layers with 64 and 32 neurons, respectively, and outputs a 32-dimensional vector Fmat.For one-dimensional time-series process sensing data (e.g., tension sequences over the past 10 seconds), a one-dimensional CNN (1D-CNN) is used, consisting of three one-dimensional convolutional layers (16, 32, and 64 filters, kernel size 5), followed by a Long Short-Term Memory (LSTM) layer (32 hidden units), ultimately outputting a 32-dimensional vector Fsen. For two-dimensional image data, a ResNet-18 model pre-trained on ImageNet is used, its last fully connected layer is removed, and the 512-dimensional feature vector Fvis from the penultimate layer is extracted. Finally, all feature vectors are concatenated: Funified = Concat(Fsim, Fmat, Fsen, Fvis), resulting in a unified 704-dimensional feature vector. This vector serves as the universal input interface for all subsequent models, solving the problem of fusing multi-source heterogeneous data.

[0026] S2. Multi-physics field collaborative optimization design steps: Based on the unified feature vector, the electromagnetic and thermal performance indicators under different combinations of design parameters are solved using a pre-trained multi-physics field proxy model. Combined with the material performance matching strategy, an electromagnetic structure design scheme and material selection scheme that meet the preset performance target are generated. It should be further explained that in step S2, the "multiphysics proxy model" is a physical information neural network, and its specific construction and training steps are as follows: First, define the design variable space, for example: X=[Nhv,Nlv,Hw,Ww, core material code, insulation material code, insulation thickness], etc., with a total of d variables. Generate N design samples using Latin hypercube sampling, where N is preferably 5000-10000 to ensure sufficient coverage of the design space. Perform high-fidelity finite element simulation on each sample to obtain the performance target Y, including: no-load loss P0(W), load loss Pk(W), short-circuit impedance Uk(%), top oil temperature rise Δθoil(K), and hot spot temperature rise Δθspot(K). This forms the dataset {Xi,Yi}, i=1…N. The PINN model structure is as follows: the input layer receives a subset of the unified feature vector Funified (mainly Fsim and Fmat parts) encoded in the aforementioned S1 step, passes through 5 fully connected layers (with 256, 256, 128, 128, and 64 neurons respectively), each followed by a SiLU activation function, and finally, the output layer outputs 5 performance indicators Ypred. Its loss function Ltotal is defined as: Ltotal = Ldata + λ Lphysics, where Ldata = (1 / N) Σi||Ypred_i-Ytrue_i||2Lphysics=(1 / M) Σj||R(Ypred_j,Xj)||2 Here, Ldata is the mean squared error loss of the data. Lphysics is the physical residual loss, and R is the residual operator of the physical equation. For transformers, a key physical constraint is the approximate relationship between losses and temperature rise, which can be represented by a simplified heat balance equation as the physical loss. For example, a residual can be defined as: R=Ptotal-K A ΔT, where Ptotal is the total loss (P0+Pk), K is the overall heat dissipation coefficient, A is the heat dissipation area, and ΔT is the temperature rise.

[0027] During training, M points (ideally 10% of the batch size) are randomly sampled from the computational domain, and the residual is calculated. The hyperparameter λ is used to balance the two losses, with a preferred value of 0.5. This value is chosen because: when λ=0, the model degenerates into a purely data-driven model, potentially violating physical laws in sparse regions; excessively large λ (e.g., >1) over-constrains the network, reducing its ability to fit high-precision simulation data. λ=0.5 achieves a good balance between the two. The model is trained using the Adam optimizer with an initial learning rate of 1e-3, a batch size of 32, and a total training time of 1000 epochs. After training, the surrogate model can evaluate a set of designs in approximately 0.1 seconds, a speed improvement of over four orders of magnitude compared to several hours of finite element simulation. A material performance matching strategy works in conjunction with PINN: after PINN evaluation, a performance requirement vector Rneeds is output. The material knowledge graph is matched using a graph query language (such as Cypher). For example, the query statement is: "MATCH (m:Material) WHERE m.breakdown_strength>15 AND m.thermal_conductivity>0.2 RETURN m ORDER BY (abs(m.breakdown_strength-15) + abs(m.thermal_conductivity-0.2)) LIMIT 5".

[0028] Finally, a multi-objective particle swarm optimization algorithm is used, with PINN as the fitness evaluation function, and the objectives of minimizing loss, cost, and volume are used to search for the Pareto optimal solution set in the d-dimensional design variable space, and the final design scheme is output.

[0029] S3. Process-driven and digital twin interaction steps: The electromagnetic structure design scheme and material selection scheme are converted into a sequence of key process parameters to drive the manufacturing equipment; during the manufacturing process, the state of the digital twin is updated based on the real-time process sensing data and visual inspection data, and the digital twin simulates the manufacturing process and product performance evolution. It should be further explained that in step S3, the conversion from design scheme to process parameters is achieved through a "design-process" mapping rule base, which is stored in JSON format. For example, for winding design parameters {"wire_diameter": 2.0mm, "turns_per_layer": 50, "number_of_layers": 10}, the mapping rule might be: {"tension_setpoint": 10.0N, "winding_speed": 5.0rpm, "traverse_pitch": 2.05mm}. The digital twin is a model running in a real-time environment (such as ROS2 or a dedicated simulation server), its core being a lumped-parameter equivalent circuit-thermal network model. The equivalent circuit model includes nonlinear inductance (representing the core), resistance (representing the winding), mutual inductance, etc., and its parameters (such as winding resistance Rwinding and leakage inductance Lleakage) are functions of the process state. For example, the relationship between leakage inductance Lleakage and winding gap g is represented by a pre-calibrated function: Lleakage = L0 + k g, where L0 is the design nominal leakage inductance, k is a coefficient, and g is the average gap detected visually. When real-time data (e.g., g = 0.3 mm, greater than the nominal value of 0.2 mm) is received, the digital twin immediately updates Lleakage = L0 + k. The circuit equations are resolved to obtain new predicted values ​​for short-circuit impedance and load loss. An RC model is used for the thermal network, and its thermal resistance varies with the state of the insulating material (e.g., degree of curing, calculated from the curing temperature history). The digital twin is updated and calculated synchronously with the physical manufacturing process, with a 1-second cycle.

[0030] S4. Adaptive Decision-Making and Closed-Loop Optimization Steps: Based on the updated digital twin state, the deviation between the current manufacturing state and the expected target is analyzed using a preset decision model, and adjustment instructions for the subsequent process parameter sequence or fine-tuning instructions for the electromagnetic structure design scheme are generated to achieve online adaptive adjustment of the manufacturing process or re-optimization of the design scheme. It should be further explained that in step S4, the decision model adopts a deep deterministic strategy gradient algorithm. The state space S is a multi-dimensional continuous vector, including: manufacturing stage encoding (one-hot encoding, 4-dimensional: winding, stacking, casting, curing), the deviation between the measured values ​​and target values ​​of 6 key process parameters, the severity scores of 3 types of visual defects, the deviation between the 5 key performance indicators predicted by the digital twin and the target, and a value representing the percentage of manufacturing progress. The action space A is a 6-dimensional continuous vector, each dimension corresponding to the adjustment amount of an adjustable process parameter, normalized to [-1, 1], respectively mapped to: winding tension adjustment amount (±2N), winding speed adjustment amount (±10%), oven heating rate adjustment amount (±0.5°C / min), casting pressure adjustment amount (±0.05MPa), vacuum degree adjustment amount (±10Pa), and the insulation distance fine-tuning suggestion in the design of the next product of the same type (±0.5mm). The reward function rt is designed as: rt=-[0.5 (ΔPk / Pk_target)²+0.3 (Δθspot / θspot_target)² + 0.2 (1-Rins)-1000 Iunsafe-0.01 ||at-at-1||2 Where ΔPk is the difference between the predicted load loss and the target value, Δθspot is the difference between the predicted hot spot temperature rise and the target value, and Rins is the insulation reliability index (0-1) calculated by the digital twin. Iunsafe is a safety violation indicator; it is 1 if any predicted value (such as temperature, stress) exceeds its safety threshold, and 0 otherwise. The last term is the action smoothness penalty. The coefficients are chosen for the following reasons: the performance coefficients (0.5, 0.3, 0.2) reflect the relative importance attached to loss, temperature rise, and insulation reliability; the safety penalty of -1000 is a very large negative reward, designed to absolutely prohibit unsafe operations; the smoothness penalty coefficient of 0.01 aims to avoid drastic fluctuations in control commands and ensure production stability. The decision model is trained offline in a simulation environment composed of a digital twin, using the OU process to add action noise for exploration, for a total of 2 million training steps. After training, the policy network parameters are solidified and deployed. In actual production, the system collects the state st every 10 seconds (or when a manufacturing process node is completed), inputs it into the strategy network, obtains the action at, and after inverse normalization and amplitude limiting, it is converted into actual control commands and sent to the manufacturing execution system or design software to achieve closed-loop optimization.

[0031] Specifically, in step S1, the real-time process sensing data includes at least one of winding tension, pressing pressure, curing temperature, and partial discharge signal; the visual inspection data includes image feature data obtained by imaging and recognizing the winding arrangement, core laminations, and insulation packaging.

[0032] It should be further explained that the specific monitoring point for the winding tension is located at the guide wheel before the conductor enters the winding mold, using a spoke-type tension sensor with a preferred range of 0-30N and an accuracy of ±0.1%FS. This range is chosen because the typical unwinding tension of transformer winding conductors (such as paper-insulated flat copper wire) is between 5-20N, and this range can cover this range with a margin. The sampling frequency is 100Hz, which is sufficient to distinguish rapid fluctuations in tension. The pressing pressure is monitored by four piezoelectric force sensors arranged at the four corners of the upper and lower pressure plates of the press, with a range of 0-500kN. These sensors are used to monitor the uniformity of pressure during core stacking and overall pressing, with a target of less than 10% for pressure non-uniformity (the difference between the maximum and minimum pressure divided by the average pressure). The monitoring points for curing temperature include at least: one point inside the epoxy resin mixture in the casting tank, one point inside the winding (via a pre-embedded sheathed thermocouple), and one point at the ambient temperature inside the oven. Type K thermocouples are preferred because their measurement range (-200~1260°C) fully covers the 80-180°C required for the curing process, and they offer high cost-effectiveness. Partial discharge signal detection is performed after the windings are impregnated and dried, but before assembly into the oil tank. A series pulse current method is used, connecting the test sample, the detection impedance, and the calibration pulse generator in series to the test circuit. The detection frequency band is 100kHz-1MHz, which effectively extracts the partial discharge signal and suppresses most periodic interference. The measurement threshold is set to 5pC; signals below this threshold are considered noise. This threshold is a typical value based on industry standards (such as IEC 60270) and background noise levels.

[0033] Furthermore, for visual inspection of the winding arrangement, a line scan camera is used to scan along the winding axis to acquire high-resolution (0.1mm / pixel) side images. The image processing algorithm steps are: 1) Gaussian filtering for noise reduction; 2) Canny edge detection; 3) Hough linear transform to detect straight lines on the edges of the conductors; 4) Calculate the average distance between adjacent parallel lines and compare it with the standard pitch (e.g., 2.0mm). If the deviation is greater than ±0.2mm, it is judged as uneven arrangement. ±0.2mm is chosen as the threshold because this deviation is sufficient to cause a significant change in local field strength, and industrial cameras can stably detect at this resolution. For core laminations, an area scan camera is used to photograph the end face of the laminations. Threshold segmentation is used to separate the silicon steel sheets from the background. Then, the ratio of the total length of the lamination "seam" (i.e., the gap between sheets) to the perimeter of the lamination outline is calculated and defined as the "seam ratio," with a control target of less than 1%. For surface inspection of insulating encapsulations (such as epoxy resin castings), a multi-angle LED light source is used in conjunction with a color area scan camera to acquire surface images. A U-Net semantic segmentation network was trained to identify defects such as bubbles and cracks. The training dataset contained 5000 labeled images of the surface of the casting. The network output a pixel-level mask of the defects, and an alarm was triggered when the total defect area accounted for more than 0.01% of the casting surface area. This threshold was set based on empirical data from electrical strength tests.

[0034] Specifically, in step S2, the multiphysics proxy model is a physical information neural network, which is trained by incorporating the constraints of the electromagnetic and thermodynamic physical laws followed by the transformer.

[0035] Further explanation is needed regarding the specific implementation of the Physical Information Neural Network (PINN) and the fusion method of physical constraints. Taking the prediction of the DC resistance Rdc, AC resistance Rac (considering skin effect and proximity effect), and hotspot temperature rise ΔT of a transformer winding as an example, the model input x is the encoding of design parameters (such as conductor cross-sectional area Aw, current density J, insulation thickness ti) and material properties (such as conductivity σ, thermal conductivity k). The model output y is Rdc, Rac, and ΔT. Besides the data loss Ldata, the physical loss Lphysics consists of two residual terms: Ohm's law residual and thermal equilibrium residual. Ohm's law residual: Rohm = Ploss - (I²) Rac) where Ploss is the total loss predicted by the model (which can be estimated through other subnetworks or formulas, or as an additional input), I is the operating current (known), and Rac is the AC resistance predicted by the model. This residual forces the model-predicted loss and resistance to satisfy basic electrical relationships. Thermal equilibrium residual: Rthermal = Ploss - (h A ΔT) where h is the equivalent heat dissipation coefficient (related to the cooling method and structure, and can be set as a learnable parameter or given by empirical formula), A is the effective heat dissipation area (which can be calculated from design parameters), and ΔT is the temperature rise predicted by the model. This residual forces the model to predict losses and temperature rise that satisfy the basic energy conservation law. The total physical loss Lphysics is the mean square sum of these two residuals: Lphysics = (1 / N)Σ(Rohm2 + Rthermal2). The selection of physical residual calculation points (i.e., the so-called "residual points") is crucial. In addition to calculating the residual at the input point xi of the training data, it is more important to randomly sample a large number (e.g., 1024 samples per training iteration) of unseen "residual points" xr within the domain of the input variables, input xr into the network to obtain predictions, and then calculate the physical residuals at these points. This forces the network to obey physical laws throughout the entire input space but only near the data points, greatly enhancing extrapolation ability and generalization. During network training, the total loss is Ltotal = Ldata + λphy Lphysics, with an optimal value of λphy of 0.5, is used for training with the L-BFGS optimizer, a second-order optimizer that is typically more stable for smooth optimization problems with physical constraints. Through this training, PINN can provide physically reasonable predictions even for sparsely covered design regions in the training data, which is crucial for exploring novel design schemes without historical data.

[0036] Specifically, the material performance matching strategy is implemented through a materials knowledge graph, which links the composition, preparation process, electrical properties, thermal properties, and mechanical properties of different materials.

[0037] It should be further noted that the materials knowledge graph is constructed using the Neo4j graph database. Entity types include: Material, Chemical Element, Process, and Property. Relationship types include: CONTAINS (material contains elements), PROCESSED_BY (material is prepared by a process), HAS_PROPERTY (material has properties), SIMILAR_TO (materials have similar properties), and DERIVED_FROM (material is derived from another material). For example, a specific knowledge fragment is: (:Material{name:'Epoxy Resin / Alumina Composite'})-[:CONTAINS]->(:Chemical Element{name:'C, H, O'}), (:Material)-[:CONTAINS]->(:Chemical Element{name:'Al'}), (:Material)-[:PROCESSED_BY]->(:Process{name:'Melt Blending, Hot Pressing'}), (:Material)-[:HAS_PROPERTY]->(:Property{type:'thermal_conductivity', value:0.8, unit:'W / (m·K)'}).

[0038] Furthermore, the specific algorithm steps for material matching are as follows: 1. Requirements Analysis: Receive the requirements vector from step S2, for example: Req={“breakdown_strength_min”:20,“thermal_conductivity_min”:0.5,“glass_transition_temp_min”:120}, with units of kV / mm, W / (m·K), and °C respectively.

[0039] 2. Initial screening: Execute a Cypher query to filter out material nodes that meet all minimum performance requirements.

[0040] MATCH (m:Material) WHERE (m:Material)-[:HAS_PROPERTY]->(:Property {type:'breakdown_strength'})>= 20 AND (m:Material)-[:HAS_PROPERTY]->(:Property {type:'thermal_conductivity'})>= 0.5 AND (m:Material)-[:HAS_PROPERTY]->(:Property {type:'glass_transition_temp'})>= 120 RETURN m 3. Fine Screening and Ranking: For the materials after initial screening, calculate their similarity to the ideal target (i.e., all properties are at the upper limit of the requirements). A weighted Euclidean distance is used: Distance = sqrt(Σw_i The formula is ((p_i-target_i) / range_i)²). Here, pi is the i-th property value of the material, target_i is the ideal target value for that property (which can be determined by the upper limit of demand or design goals), range_i is the range of values ​​for that property among candidate materials (used for normalization), and w_i is the weight, reflecting the importance of that property (e.g., breakdown field strength is weighted at 0.5, thermal conductivity at 0.3, and glass transition temperature at 0.2). Sort by distance from smallest to largest and return the top K materials (K=5).

[0041] 4. Related Recommendations: For materials ranked highly, further explore their related entities in the graph. For example, if material A is selected but its cost is too high, you can query: MATCH (m:Material {name:'A'})-[:SIMILAR_TO]-(similar) RETURN similar, to find a similar but lower-cost alternative material B. Alternatively, if existing materials are not ideal, you can perform knowledge reasoning: MATCH (m1:Material)-[:HAS_PROPERTY]->(p1 {type:'thermal_conductivity', value:>2.0}), (m2:Material)-[:HAS_PROPERTY]->(p2{type:'breakdown_strength', value:>30}) RETURN m1, m2, and then recommend using a high thermal conductivity material m1 (such as alumina) as a filler, combined with a high insulation strength material m2 (such as epoxy resin) to create a new material solution that meets the requirements. This reasoning process provides innovative ideas for materials engineers.

[0042] Specifically, in step S3, the digital twin receives the sequence of key process parameters as input and simulates the resulting changes in the electromagnetic structure and material state of the transformer, thereby outputting predicted product electromagnetic performance and insulation performance indicators.

[0043] It should be further explained that the specific simulation model of the digital twin is a hybrid model combining a lumped-parameter equivalent circuit, a thermal network, and an aging kinetics model. Its state updates and performance predictions are performed cyclically within a fixed time step Δt (preferably 1 second). Regarding the electromagnetic model, a T-type equivalent circuit of a two-winding transformer is used, with its parameters dynamically updated. For example, the AC resistance Rac of the winding is a function of frequency f, temperature T, and winding geometry, calculated in real-time based on the current state using a pre-trained surrogate model (such as a lightweight neural network): R... ac =f NN (f, T, wire_diameter, insulation_thickness). The magnetizing inductance Lm of the iron core is a function of the voltage U and the magnetization state of the iron core material (affected by the stacking pressure), and is realized by a lookup table method.

[0044] Furthermore, at each step, the twin receives current process parameters (such as winding tension F and pressure P), which alter the geometry and material state. For example, based on the tension F, the empirical formula ε = F / (E) is used. A cs The conductor strain ε is calculated, where E is the conductor's elastic modulus and Acs is the cross-sectional area. Strain ε causes a slight change in the conductor's cross-sectional area, which is then used to update the resistance Rac using the aforementioned surrogate model. For example, based on the visually detected core lamination joint ratio γ, the empirical formula L... m =L m0 (1-k l γ) Modify the magnetizing inductance Lm, where Lm0 is the design nominal value and kl is an empirical coefficient (e.g., 0.05). Then, solve the equivalent circuit equations to calculate the electromagnetic performance such as no-load current, load loss, and short-circuit impedance under the current state. The thermal network model adopts a 3-node (winding, core, oil) RC model. The heat capacity C of each node and the thermal resistance Rth between nodes are functions of the material and structural states. For example, the thermal conductivity kins of the insulating material varies with the degree of curing α, which is calculated from the curing temperature history using the Kamal kinetic model: dα / dt=(k1+k2) α m ) (1-α) n Where k1, k2, m, and n are temperature-dependent Arrhenius parameters. Thermal conductivity kins = kins0 + β α, where kins0 is the uncured value and β is a coefficient. The updated kins will change the thermal resistance Rth_w-o from the winding to the oil. Solving the thermal network differential equation yields the temperature rise ΔT of the winding and the core. Insulation performance is evaluated using an electric-thermal coupled aging model. First, the electric field strength E borne by the insulation layer is obtained from the voltage distribution calculated by the electromagnetic model. Then, combined with the hot spot temperature Thotspot calculated by the thermal model, the aging rate v=A is calculated using an electric-thermal joint aging rate model (such as the inverse power-law-Arrhenius model). En exp(-B / Thotspot), where A, n, and B are material constants. The cumulative aging amount D = ∫vdt, and the insulation reliability index Rins is defined as exp(-D). The digital twin outputs updated key performance indicators at each step: [P0, P... k U k ΔT hotspot R ins The deviation vector is formed by comparing it with the target value and then used for decision-making in S4.

[0045] Specifically, in step S4, the decision model is a deep reinforcement learning model, which takes the state of the digital twin as the state input, the process parameter adjustment action or the design parameter fine-tuning action as the action output, and the conformity between the product performance index and the manufacturing target as the reward signal for training.

[0046] It should be further noted that the decision-making model employs a deep deterministic policy gradient algorithm. Its specific training process is as follows: 1. Agent and Environment: The agent is the Actor network and Critic network to be trained. The environment is the digital twin simulation model described in this invention, which encapsulates the logic of the transformer manufacturing process (such as process sequence and equipment constraints).

[0047] 2. State st: At time t, the state is specifically structured as follows: Process stage: One-hot encoding is used, such as [1, 0, 0, 0] representing the winding stage; Process deviation: The normalized deviation between the current key process parameters and the target value, such as (current tension - target tension) / target tension; Visual defects: Severity rating of three types of defects (uneven distribution, large seams, air bubbles), 0-1; Performance bias: Normalized bias of five performance metrics predicted by the digital twin against the target values; Progress: Percentage of current process completed, 0-1; Historical actions: Actions taken in the past four time steps, used to provide historical information.

[0048] It should be further explained that at time t, the state st is a 35-dimensional continuous vector, and its specific structure is as follows: ① Process Stage (4-dimensional): The current manufacturing stage is represented by a unique hot code. Specifically: Phase 1: Winding, coded as [1, 0, 0, 0]; Phase 2: Core stacking, coded as [0, 1, 0, 0]; Phase 3: Insulation treatment (e.g., casting), coded as [0, 0, 1, 0]; Phase 4: Solidification, coded as [0, 0, 0, 1].

[0049] ② Key process parameter deviation (6 dimensions): The normalized relative deviation between the actual measured values ​​and the target set values ​​of the six adjustable process parameters. Calculated as (measured value - target value) / target value. Specific parameters include: 1. Winding tension deviation; 2. Deviance in winding speed; 3. Pouring pressure deviation; 4. Deviation in curing temperature rise rate; 5. Vacuum degree deviation; 6. Deviation between the measured insulation distance of the current product and the design value (obtained through visual inspection).

[0050] ③ Severity of Visual Deficits (3-dimensional): Quantitative scores for three key visual defects, ranging from [0, 1]. The scores are output by the corresponding visual detection algorithms; higher values ​​indicate more severe defects. The three types of defects are: 1. Uneven winding arrangement; 2. Core lamination joint ratio; 3. Percentage of surface defects (bubbles, cracks, etc.) on the insulating package.

[0051] ④ Key Performance Indicator Prediction Deviation (5-Dimensional): The normalized relative deviation between the five key performance indicators predicted by the digital twin based on the current state and the design target values. Calculated as (predicted value - target value) / target value. Specific performance indicators include: 1. Load loss (Pk) deviation; 2. No-load loss (P0) deviation; 3. Short-circuit impedance (Uk) deviation; 4. Hot spot temperature rise (Δθspot) deviation; 5. Insulation reliability index (Rins) deviation (calculated as Rins_pred - 1, since the target value is 1).

[0052] ⑤ Manufacturing progress (1-dimensional): The percentage of completion of the current process, ranging from [0, 1]. For example, if the winding process is 70% complete, the value is 0.7.

[0053] ⑥ Historical Action Sequence (16-dimensional): Action vectors executed at the past four decision points (t-4, t-3, t-2, t-1). Each action vector is 6-dimensional, but to control the state dimensions and focus on key adjustments, this embodiment only records the first four adjustment dimensions considered the most critical for each historical action: winding tension adjustment, winding speed adjustment percentage, pouring pressure adjustment, and curing heating rate adjustment. Therefore, 4 historical moments × 4 dimensions = 16 dimensions. These values ​​are the actual execution amounts normalized to the range [-1, 1].

[0054] 3. Action at: The output is a 6-dimensional continuous action vector, ranging from [-1, 1]. After denormalization, it is mapped to the actual adjustment amount: a real =a low +(a t +1) (a high -a low ) / 2. Where [alow, ahigh] is the actual allowable range for each adjustment amount, for example: a1: Winding tension adjustment amount, range [-1N, +1N].

[0055] a2: Winding speed adjustment percentage, range [-5%, +5%].

[0056] a3: Curing temperature rise rate adjustment, range [-0.2°C / min, +0.2°C / min].

[0057] a4: Pouring pressure adjustment range [-0.02MPa, +0.02MPa].

[0058] a5: Vacuum adjustment amount, range [-5kPa, +5kPa].

[0059] a6: Suggestions for fine-tuning the insulation distance of subsequent products of the same type, within the range of [-0.2mm, +0.2mm].

[0060] 4. Reward rt: The reward for each time step is calculated as follows: rt = -[w1(Pkerr)² + w2(Terr)² + w3 (1-Rins)2]-Cunsafe Iunsafe-λsmooth ||at-at-1||2 Where Pkerr is (predicted load loss - target loss) / target loss, Terr is (predicted hot spot temperature rise - target temperature rise) / target temperature rise, and Rins is the insulation reliability index. w1, w2, and w3 are weights, with preferred values ​​of 0.5, 0.3, and 0.2, respectively. Iunsafe is a safety indicator; if the predicted temperature exceeds the material's heat resistance rating (e.g., 155°C) or the predicted electric field strength exceeds 80% of the breakdown electric field strength, then Iunsafe = 1; otherwise, it is 0. Cunsafe is a safety penalty coefficient, set to a large number, such as -1000. λsmooth is a smoothing penalty coefficient, set to 0.01. Furthermore, a large completion bonus of +1000 is given when the entire manufacturing process is successfully completed and all final performance indicators are within allowable tolerances (e.g., ±2%).

[0061] 5. Training process: Initialization: Randomly initialize the parameters θμ and θQ of the Actor network (policy network μ) and Critic network (Q network). Initialize the target network parameters. Initialize the experience replay buffer R with a capacity of 1e6.

[0062] Interactive loop: For each training cycle (episode), reset the environment to its initial state. For each step t: a) Select the action at = μ(st|θμ) + Nt based on the current strategy and the exploration noise (using OU noise).

[0063] b) Perform action at in the digital twin environment, observe the next state st+1 and reward rt, and whether to terminate (done signal).

[0064] c) Store the empirical tuple (st, at, rt, st+1, done) into the buffer R.

[0065] d) Experience in randomly sampling a small batch (batch size=256) from R.

[0066] e) Update the Critic network: Calculate the target Q value y = r + γ Q'(st+1, μ'(st+1)), where γ=0.99 is the discount factor, and Q' and μ' are the target networks. Minimize the loss L=(Q(st, at)-y)2, and update θQ.

[0067] f) Update the Actor network: Use the policy gradient ∇θμJ≈Σ∇aQ(s, a) ∇θμμ(s), update θμ.

[0068] g) Soft update target network: θQ'<-τθQ+(1-τ)θQ', θμ'<-τθμ+(1-τ)θμ', τ=0.005.

[0069] Repeat the above steps until the total number of training steps reaches 2 million, or the policy performance converges (i.e., the average total reward per episode no longer increases significantly). After training is complete, save the final Actor network parameters for online decision-making.

[0070] Specifically, the decision-making process of the deep reinforcement learning model integrates an attention mechanism to identify the state features that have the greatest impact on the current decision.

[0071] It should be further explained that the attention mechanism is integrated into the Actor network (policy network). Specifically, the structure of the Actor network is as follows: The input state st (35 dimensions) first passes through a feature extraction layer (fully connected layer, 128 neurons, ReLU activation) to obtain an intermediate feature vector h. Then h is input into a multi-head self-attention layer. The self-attention layer operates as follows: First, h is transformed through three different linear transformations to obtain the query vector Q, the key vector K, and the value vector V. Then the attention weights are calculated: Attention(Q, K, V) = softmax(QK) T / √d k V. Here, d k This refers to the dimension of the key vector. We use four attention heads, meaning the above process is performed in parallel four times. The outputs of the four heads are then concatenated and output through a linear layer. The output of the attention weights is a matrix, where the weight value in the i-th row and j-th column represents the importance weight (normalized by softmax) assigned by the network to input feature j relative to feature i when generating the final action. During model deployment and decision-making, in addition to the output action *at*, the attention weight matrix is ​​also recorded. By visualizing this matrix, we can identify which state features (such as "large winding gap deviation" and "high predicted short-circuit force") receive the highest attention scores in specific decisions (such as "increase winding tension"). For example, in one decision, the attention weights show that the network's average attention score for the feature dimension of "winding gap deviation" is as high as 0.4, while the attention score for "ambient temperature" is only 0.01, indicating that the network's decision is mainly based on the winding gap issue. This provides an intuitive explanation for operators, enhancing human trust in AI decision-making. Furthermore, during training, a small regularization term can be added to encourage a degree of sparsity in the attention weights, for example, by adding λ to the loss function. att ||Attention_Weights||1 (L1 norm), λatt can be 0.0001, which helps the network learn a more focused and interpretable attention pattern, avoiding attention being scattered across all features.

[0072] Specifically, the action output of the deep reinforcement learning model is subject to physical and technological constraints. The physical constraints include upper temperature limits and stress limits, while the technological constraints include the range of equipment operating parameters.

[0073] It should be further explained that the specific implementation of constraint application is divided into two layers: hard constraints in the action space and soft constraints in the state / reward space, as detailed below: 1. Hard constraints on action space (process constraints): After the agent outputs an action `at`, it must be clipped before execution to ensure that it does not exceed the safe operating range of the equipment. This is specifically implemented as a clipping function `Clip(a)`. t ): at_clipped = max(amin, min(at, amax)) where amin and amax are the minimum and maximum allowed values ​​for each action dimension, determined by the device specifications. For example: a1 (tension adjustment): amin=-1.0N, amax=+1.0N (the equipment allows adjustment within ±1N near the target value). a2 (speed adjustment %): amin = -5.0, amax = +5.0; a3 (heating rate adjustment): amin = -0.2°C / min, amax = +0.2°C / min; a4 (pouring pressure adjustment): amin = -0.02MPa, amax = +0.02MPa; a5 (vacuum adjustment): amin = -5kPa, amax = +5kPa; a6 (insulation distance fine-tuning): amin=-0.2mm, amax=+0.2mm. This trimming operation is mandatory to ensure that no decision will result in instructions that prevent the equipment from executing or may cause mechanical failure.

[0074] 2. State / Reward Soft Constraints (Physical Constraints): For complex physical constraints that cannot be guaranteed by simple trimming, such as "hot spot temperature must not exceed 155°C" or "winding mechanical stress must not exceed 80% of yield strength," this is achieved by setting a large penalty term in the reward function. In the digital twin environment (during training) or the actual system (monitoring during deployment), each step checks whether the predicted or estimated physical quantity violates the constraint. Specifically, in the reward function, the safety penalty term Cunsafe... The logic for determining Iunsafe in Iunsafe is as follows: Iunsafe = 1, if (Thotspot_predicted > 155) OR (Stressestimated > 0.8) If the predicted value is (Yield_Strength) OR (any other critical physical quantity exceeding the limit), then Iunsafe = 0. Here, Thotspot_predicted comes from the digital twin model, Stressesstimated is calculated using a simplified mechanical model based on current tension, pressure, etc., and Yield_Strength is a known property of the wire material. Cunsafe is set to -1000. During training, if the agent causes Iunsafe = 1, it not only receives a large negative reward, but the training episode is usually terminated early. This strongly encourages the agent to learn to avoid any behavioral sequences that might violate physical constraints. During deployment, although the policy network has learned to avoid these violations, the system still monitors these physical quantities in real time. Once the predicted value approaches the threshold (e.g., T > 150°C), even if the policy network does not give an adjustment instruction, the safety monitoring module will trigger a level one safety alarm, and can switch to manual control mode if necessary. This multi-layered constraint mechanism of "hard pruning + soft punishment + safety monitoring" ensures that the entire system operates within absolutely safe boundaries.

[0075] Specifically, the reward function of the deep reinforcement learning model includes: a basic reward based on product performance indicators, a penalty based on whether the physical and process constraints are violated, and an additional reward based on action smoothness and interpretability.

[0076] It should be further explained that the complete and detailed formula for the reward function rt, and the rationale for the selection of coefficients for each term, are as follows: rt = Rperf + Rsafe + Rsmooth + Rexp 1. Performance bonus Rperf: Rperf = -[wpk (Pkerr)2+wtemp (Terr)2+wins (1-Rins)2+wuk [(Uk_err)2]wherein: Pkerr = (Pk_pred - Pk_target) / Pk_target, representing the relative error of load loss; Terr = (Thot_pred - Thot_target) / Thot_target, representing the relative error of hot spot temperature rise; Uk_err=(Uk_pred-Uk_target) / Uk_target, the relative error of short-circuit impedance; Rins: Insulation reliability index (0~1); wpk, wtemp, wins, and wuk are weighting coefficients. The preferred values ​​and rationale are: wpk = 0.4, wtemp = 0.3, wins = 0.2, and wuk = 0.1. The reasons are as follows: Load loss (Pk) directly relates to the transformer's operating efficiency and is a core economic indicator, thus receiving the highest weight of 0.4. Hot spot temperature rise (T) is crucial in determining insulation life and reliability, with a weight of 0.3. Insulation reliability (Rins) is the ultimate manifestation of safety, with a weight of 0.2. Short-circuit impedance (Uk) affects the system's short-circuit current, and its allowable tolerance is relatively large, hence its lowest weight of 0.1. The sum of these weights is 1, ensuring that the performance bonus is between -1 and 0 (ideally 0).

[0077] 2. Safety penalty item Rsafe: Rsafe = Cunsafe Iunsafe = 1 if and only if any of the following conditions are true: ①Thot_predicted>Tmax (Tmax=155°C, which is the heat resistance rating of the insulation material); ②Estress_predicted>0.8 Ebd (Ebd is the power frequency breakdown field strength of the insulating material, and 0.8 is the safety factor); ③Stressmech>0.9 σyield (σyield is the yield strength of the conductor, and 0.9 is the safety factor).

[0078] Any process parameter exceeding its hard constraints (this is typically prevented by action clipping, but is checked here as an additional check). Cunsafe is the safety penalty coefficient. Preferred value and rationale: Cunsafe = -1000. Choosing such a large negative value is to give the agent a clear and strong signal during training that violating safety constraints is unacceptable, and its cost far outweighs any minor performance gains. This ensures the safety of the learned policy.

[0079] 3. Motion smoothness reward term Rsmooth: Rsmooth = -λsmooth ||at-at-1||2 where ||·||2 represents the squared L2 norm of the vector, i.e., the sum of squares of the differences in adjustment amounts across dimensions. λsmooth is the smoothness coefficient. Preferred coefficient value and rationale: λsmooth = 0.01. This value needs to strike a balance between encouraging exploration and avoiding jitter. If λsmooth is too small (e.g., 0.001), the smoothness penalty is too weak, potentially leading to drastic fluctuations in the action sequence, which is unfriendly to physical devices; if it is too large (e.g., 0.1), it will overly restrict action variation, potentially causing the agent to be too conservative and unable to make necessary rapid adjustments. 0.01 has proven effective in practice.

[0080] 4. Interpretability reward item Rexp (optional but recommended): Rexp = -λexp H(Attention_Weightst) is an information entropy function used to measure the dispersion of attention weight distribution. Attention weights are obtained through the multi-head self-attention mechanism described in this invention. For each attention head, the entropy of its attention weight distribution is calculated, and then averaged over all heads. Higher entropy indicates more dispersed attention and more ambiguous decision-making criteria; lower entropy indicates more focused attention on a few features and stronger decision interpretability. λexp is the interpretability reward coefficient. The preferred coefficient value and rationale are: λexp = 0.001. This is a small coefficient, designed to subtly guide the agent to learn more focused and interpretable decision-making patterns without significantly interfering with the main rewards (performance and safety). Setting it too large may affect the final performance of the policy.

[0081] Therefore, the complete reward function is: rt = -[0.4 (Pkerr)² + 0.3 (Terr)2+0.2 (1-Rins)²+0.1 (Uk_err)2]-1000 Iunsafe-0.01 ||at-at-1||2-0.001 H(Attention_Weightst).

[0082] It is understandable that the multi-objective reward function designed in this invention systematically balances four major objectives: performance optimization, security and compliance, control stability, and decision interpretability.

[0083] Please see Figure 2 The present invention provides another embodiment, which provides a deep learning-based adaptive optimization system for transformer design and process, the deep learning-based adaptive optimization system for transformer design and process includes: The data perception and fusion module 100 is used to perform the S1 step, collect and uniformly represent multi-source heterogeneous data; The collaborative design and simulation module 200 is used to execute the S2 step, run the multiphysics proxy model and perform material matching to generate an optimized design scheme. The process execution and digital twin module 300 is used to execute the S3 step, convert the design scheme into process instructions, and run a digital twin that interacts synchronously with the physical manufacturing process; The intelligent decision-making and closed-loop control module 400 is used to execute the S4 step, run the decision model, generate adaptive adjustment instructions, and send them to the manufacturing equipment or design system.

[0084] It should be further explained that this system is a hardware and software integrated platform deployed in an industrial cloud-edge collaborative architecture, with the following specific components: 1. Data Sensing and Fusion Module 100: This module is deployed in the workshop and consists of an edge computing gateway and field devices. The hardware includes: various sensors (tension, pressure, temperature, partial discharge, current, voltage), industrial cameras, light sources, programmable logic controllers (PLCs), and industrial personal computers. The software includes a data acquisition program running on the edge IPC (based on an OPCUA client or a device-specific SDK), an image processing program (running the visual algorithms described in this invention), and a lightweight data fusion service. This service receives raw data streams and image analysis results from various sensors, and generates a unified 704-dimensional feature vector (Funified) in real time according to the encoded network structure described in step S1 (3D-CNN, 1D-CNN+LSTM, ResNet feature extractor, etc., these models have been pre-trained and deployed on the edge). This vector is then published to a specified topic in a message middleware (such as RabbitMQ) via the MQTT protocol. This module ensures the real-time nature, synchronization, and preprocessing of data, reducing the pressure on the cloud.

[0085] 2. Collaborative Design and Simulation Module 200: This module is deployed on a high-performance computing cluster in the cloud. Its core is a containerized application containing multiple microservices. It includes: Design parameter management service: Manages the design variable space and constraints.

[0086] Simulation task scheduling service: Calls commercial software such as ANSYS or open-source solvers to perform high-fidelity simulations and generate training data for proxy models.

[0087] PINN Training and Inference Service: Based on the TensorFlow or PyTorch framework, this service loads training datasets and trains physical information neural networks. After training, the service provides an API to receive design parameters and return fast performance predictions.

[0088] Materials Knowledge Graph Service: Based on the Neo4j database, it provides a material matching and recommendation interface.

[0089] Multi-objective optimization engine: Integrating algorithms such as NSGA-II, and calling PINN inference and material recommendation services, it seeks Pareto optimal design solutions under constraints. This module ultimately outputs one or more optimal electromagnetic structure design solutions (BOM and drawings) along with corresponding material lists and process parameter suggestions.

[0090] 3. Process Execution and Digital Twin Module 300: This module is deployed on a shop floor-level server. It includes two core sub-modules: Process execution engine: It receives the final design scheme from the collaborative design module, parses the process requirements, and converts them into instruction sequences (such as G-code and equipment control commands) that can be recognized by specific manufacturing equipment (such as CNC winding machines, robot stacking systems, and vacuum casting equipment) through a "process specification compiler". It then sends these instructions to the equipment controller through protocols such as OPCUA.

[0091] Digital Twin Runtime Engine: This is an independent, real-time simulation process. Internally, it encapsulates the lumped-parameter equivalent circuit-thermal network-aging dynamics coupled model described in this invention. It subscribes to the unified feature vector Funified (or relevant portions thereof) published by the data awareness and fusion module, as well as current instructions from the process execution engine. It updates its internal model state at fixed time steps (e.g., 1 second) and calculates the current predicted performance metrics. These metrics are published to the message bus in real time. This twin serves as a bridge connecting virtual design and physical manufacturing.

[0092] 4. Intelligent Decision-Making and Closed-Loop Control Module 400: This module is also deployed on the shop floor server and interacts closely with the digital twin. It includes: State awareness and feature extractor: Subscribes to the predicted performance, actual values ​​of process parameters, visual inspection results, etc. published by the digital twin from the message bus, and assembles them into a state vector st according to the state construction method described in this invention.

[0093] Reinforcement learning agent (Actor network): Loads a pre-trained policy network model (.pb or .onnx format). Upon receiving a new state st (or at a fixed time interval, such as 10 seconds), it performs a forward inference operation, outputting a 6-dimensional action vector at.

[0094] Constraint processing and command converter: Apply hard constraint trimming as described in this invention to the action at, and then denormalize it into actual process parameter adjustment amounts. Then, generate specific control commands (such as "increase the tension setting value of winding machine No. 1 by 0.5N") or design fine-tuning suggestions (such as "modify the insulation distance from 3.0mm to 3.1mm").

[0095] Command issuance interface: Sends control commands to the corresponding manufacturing equipment controller via protocols such as OPCUA to enable online adjustments. Design fine-tuning suggestions are sent back to the collaborative design and simulation module, triggering automated design revisions for the next product or the remaining parts of the current product.

[0096] All modules are connected via an enterprise service bus and an industrial IoT platform, forming a closed-loop data system. Data flows from physical devices to sensing modules, to the digital twin, to the decision-making module, and back to physical devices, achieving complete closed-loop optimization from sensing, analysis, decision-making to control. Simultaneously, new data generated during manufacturing (performance test results) is collected to retrain and update the PINN model, visual model, and reinforcement learning strategies, enabling the entire system to continuously improve itself.

[0097] In a preferred embodiment, this application also provides an electronic device, the electronic device comprising: The computer device includes a memory and a processor, wherein the memory stores computer-readable instructions that, when executed by the processor, implement the deep learning-based adaptive optimization method for transformer design and process. The computer device can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include non-volatile storage media and internal memory. The non-volatile storage media may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface and communication interface of the computer device can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the method of the present invention.

[0098] This invention can be implemented as a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the steps of the methods of embodiments of the invention to be performed. In one embodiment, the computer program is distributed across multiple network-coupled computer devices or processors, such that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be executed by one or more computer devices or processors, and one or more other method steps / operations may be executed by one or more other computer devices or processors. One or more computer devices or processors may execute a single method step / operation, or execute two or more method steps / operations.

[0099] Those skilled in the art will understand that the method steps of this invention can be performed by a computer program instructing related hardware, such as a computer device or processor, to perform the steps of this invention when executed. Depending on the context, any references herein to memory, storage, databases, or other media may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0100] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.

[0101] The specific embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention. Any other corresponding changes and modifications made in accordance with the technical concept of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A deep learning-based adaptive optimization method for transformer design and manufacturing process, characterized in that, Includes the following steps: S1. Data Acquisition and Multimodal Characterization Steps: Collect multiphysics simulation data and material property data during the transformer design phase, as well as real-time process sensing data and visual inspection data during the manufacturing phase, and map the simulation data, property data, sensing data and visual data into a unified feature vector for characterization. S2. Multi-physics field collaborative optimization design steps: Based on the unified feature vector, the electromagnetic and thermal performance indicators under different combinations of design parameters are solved using a pre-trained multi-physics field proxy model. Combined with the material performance matching strategy, an electromagnetic structure design scheme and material selection scheme that meet the preset performance target are generated. S3. Process-driven and digital twin interaction steps: Convert the electromagnetic structure design scheme and material selection scheme into a sequence of key process parameters to drive the manufacturing equipment; During the manufacturing process, the state of the digital twin is updated based on the real-time process sensing data and visual inspection data, and the digital twin simulates the manufacturing process and product performance evolution. S4. Adaptive Decision-Making and Closed-Loop Optimization Steps: Based on the updated digital twin state, the deviation between the current manufacturing state and the expected target is analyzed using a preset decision model, and adjustment instructions for the subsequent process parameter sequence or fine-tuning instructions for the electromagnetic structure design scheme are generated to achieve online adaptive adjustment of the manufacturing process or re-optimization of the design scheme.

2. The method according to claim 1, characterized in that, In step S1, the real-time process sensing data includes at least one of winding tension, pressing pressure, curing temperature, and partial discharge signal; the visual inspection data includes image feature data obtained by imaging and recognizing the winding arrangement, core laminations, and insulation encapsulation.

3. The method according to claim 1, characterized in that, In step S2, the multiphysics proxy model is a physical information neural network, which is trained by incorporating the constraints of the electromagnetic and thermodynamic physical laws followed by the transformer.

4. The method according to claim 3, characterized in that, The material performance matching strategy is implemented through a material knowledge graph, which links the composition, preparation process, electrical properties, thermal properties and mechanical properties of different materials.

5. The method according to claim 1, characterized in that, In step S3, the digital twin receives the sequence of key process parameters as input and simulates the resulting changes in the electromagnetic structure and material state of the transformer, thereby outputting the predicted electromagnetic and insulation performance indicators of the product.

6. The method according to claim 5, characterized in that, In step S4, the decision model is a deep reinforcement learning model, which takes the state of the digital twin as the state input, the process parameter adjustment action or the design parameter fine-tuning action as the action output, and the conformity between the product performance index and the manufacturing target as the reward signal for training.

7. The method according to claim 6, characterized in that, The decision-making process of the deep reinforcement learning model integrates an attention mechanism to identify the state features that have the greatest impact on the current decision.

8. The method according to claim 6, characterized in that, The action output of the deep reinforcement learning model is subject to physical and technological constraints. The physical constraints include upper temperature limits and stress limits, while the technological constraints include the range of equipment operating parameters.

9. The method according to claim 6, characterized in that, The reward function of the deep reinforcement learning model includes: a basic reward based on product performance indicators, a penalty based on whether the physical and process constraints are violated, and an additional reward based on action smoothness and interpretability.

10. A deep learning-based adaptive optimization system for transformer design and manufacturing process, characterized in that, The system for implementing the method of any one of claims 1-9 comprises: The data perception and fusion module is used to execute the S1 step, collect and uniformly represent multi-source heterogeneous data; The collaborative design and simulation module is used to execute the S2 step, run the multiphysics proxy model and perform material matching to generate an optimized design scheme. The process execution and digital twin module is used to execute the S3 step, convert the design scheme into process instructions, and run a digital twin that interacts synchronously with the physical manufacturing process; The intelligent decision-making and closed-loop control module is used to execute the S4 step, run the decision model, generate adaptive adjustment instructions, and send them to the manufacturing equipment or design system.