A method and ai-driven system for determining and monitoring parameters for automated control of additive manufacturing process

The integration of a reinforcement learning framework and a mathematical model addresses the challenges of determining optimal process parameters in additive manufacturing, achieving efficient and precise control of the AM process and enhancing product quality and adaptability.

WO2025122016A1PCT designated stage expired Publication Date: 2025-06-123D-COMPONENTS AS

Patent Information

Application Number
PCT/NO2024/050271
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The determination of optimal process parameters for additive manufacturing (AM) is challenging due to stochastic variations, requiring extensive research and development efforts, and current sensor systems face limitations in real-time data processing and decision-making.

Method used

A method utilizing a reinforcement learning (RL) framework to dynamically adjust welding parameters in real-time, combined with a mathematical model to refine process parameters, enabling efficient and precise control of the AM process.

Benefits of technology

This approach allows for quick and precise adjustments, improving product quality, reducing waste and costs, and accelerating the time to market for novel AM products by enhancing adaptability and precision in the AM process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure NO2024050271_12062025_PF_FP_ABST
    Figure NO2024050271_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A method for performing an automated additive manufacturing (AM) process comprises obtaining a current state that includes a set of first AM factors and a set of first characteristics, which are geometric or physical properties of a first AM output are generated using the first AM factors. The current state is provided to a reinforcement learning (RL) framework designed to adjust welding parameters in real-time based on the current state and reward feedback. The RL framework includes an actor network and a critic network, and safety limits are enforced through a custom loss function. The first AM factors are updated based on the critic network's evaluation, resulting in second AM factors used to perform the second AM process and generate a second AM output.
Need to check novelty before this filing date? Find Prior Art

Description

A METHOD AND AI-DRIVEN SYSTEM FOR DETERMINING AND MONITORING PARAMETERS FOR AUTOMATED CONTROL OF ADDITIVE MANUFACTURING PROCESS FIELD OF INVENTION

[0001] Aspects of the present disclosure relate to additive manufacturing processes. Specifically, but not exclusively, aspects of the present disclosure are directed to the use of artificial intelligence in combination with mathematical modelling for automation of an additive manufacturing process. Specially, but not exclusively, aspects of the present disclosure are directed to methods and systems for mathematical model-assisted selection and monitoring of additive manufacturing process parameters. BACKGROUND

[0002] Additive manufacturing (AM), such as welding, surface coating, three-dimensional (3D) printing, and directed energy deposition (DED), is a rapidly growing field that has revolutionized the manufacturing industry. More recently, AM processes use digital models to create objects by adding successive layers of material. As AM processes are highly complex, the process relies on precise control of process parameters among other factors to ensure the quality and accuracy of the final product. Even minor changes in environment of material properties can lead to substantial changes to product quality.

[0003] For example, DED involves utilizing a directed energy source, such as a laser or electric arc, to melt and fuse material onto a substrate. This process is facilitated by a delivery mechanism, such as a nozzle, which directs the flow of material, which may be in the form of powder or wire, onto the substrate surface, where it is melted and fused to create a solid object. To achieve the desired object shape and characteristics, the nozzle and energy source are precisely controlled to move in a specific three-dimensional pattern and deposit the material bead by bead. DED is a widely used technique in several industries, such as aerospace, energy, defence, and medical sectors, for creating and repairing intricate, high-performance components with complex geometries and engineered material properties.

[0004] The stability of the DED process is subject to stochastic variations, which are contingent on the tool path and type of material. Consequently, it is necessary for AM process parameters to change and be modified throughout, and it can be difficult to efficiently configure the target set of process parameters within desired time frames, especially due to presence of multiple configurable parameters.

[0005] To achieve a high level of precision in such a complex environment, sensors typically play a crucial role. Sensorial systems are required in AM industry for a multitude of tasks, including: mitigating the labour-intensive nature of finding and certifying the process parameters; and ensuring quality control. Despite significant technological advancements in sensor systems, the determination of optimal process parameters and their influence on the final properties of printed objects remains a highly challenging and resource-intensive task. To establish stable processes with given feedstock and substrate materials within a single AM technology, engineers invest hundreds of hours in research and development. Additionally, sensors in additive manufacturing help to ensure that the final product meets the required quality standards. Deviations from the desired parameters can be detected using the data from various sources including the AM machines and sensors and operators are subsequently alerted, who can then stop the process and, where possible, modify necessary parameters.

[0006] The implementation of sensorial systems is a complex task, extending beyond hardware systems and into complex integrated software systems, and necessitates a focused approach towards the generation and processing of meaningful data. The deployment of limited processing systems inhibits the ability to perform real-time data processing and presents a hindrance to the deployment of systems with high-speed decision-making capabilities. The lack of efficient and effective processing systems impedes the effectiveness of the sensorial system, leading to suboptimal outcomes. Additionally, addressing limitations of processing capacity for processing all the sensory data cannot simply be accomplished through the use of excessively large and expensive hardware, as this can have a negative impact on energy, efficiency, and use of computational resources.

[0007] Documents US 2018 / 0264553 A1 and US 2023259099 / A1 are examples of a multi- sensor-based system for monitoring AM process. The sensor data is used by machine-learning models, including classifiers such as support vector machines (SVMs) or artificial neural networks, to predict future states of AM processes or detect defects. The AM process is then monitored using the predicted future state. However, such models are difficult to reverse engineer, and the relationship between current and future AM states based on data from multiple sensors is difficult to extract, let alone understand, interpret, and correlate to a specific feature with impeccable repeatability. There is also a need for further processing to calculate optimal AM operation parameters based on deviations between a current and future state, or a future and target state, requiring additional machine learning models or human input. Additionally, the data- heavy approach necessitates large computationally complex machine learning models and hence require extensive computational resources and energy to operate, especially real-time. Therefore, such approaches are limited in both efficiency and understandability.SUMMARY OF INVENTION

[0008] According to an aspect of the present disclosure, there is provided a method for performing an automated additive manufacturing (AM) process. The method comprises obtaining a current state that includes a set of first AM factors, which consist of one or more AM process inputs associated with performing a first AM process. Additionally, the current state includes a set of first characteristics, which are geometric or physical properties of a first AM output generated by using the set of first AM factors. This current state is then provided to a reinforcement learning (RL) framework configured to learn a policy for adjusting welding parameters in real-time based on the current state and reward feedback. The RL framework includes an actor network that generates a suggested action to achieve a target AM output using a trained machine learning model, and a critic network that evaluates the suggested action by calculating an estimated reward based on its expected cumulative benefit in optimizing the first AM process. Safety limits are enforced using a loss function that penalizes exceeding predetermined limits of the AM process inputs. The set of first AM factors is updated based on the evaluation performed by the critic network, resulting in a set of second AM factors. These second AM factors are then used to perform the second AM process, thereby generating a second AM output.

[0009] According to a further aspect of the present disclosure, the method further comprises providing an updated state to the RL framework and modifying the set of second AM factors, thereby generating a set of third AM factors. The updated state includes the set of second AM factors and a set of second AM characteristics, which consist of one or more geometric or physical properties of the second AM output. A reward is calculated based on the cumulative benefit of the second AM process, which is associated with minimizing the difference between a set of target characteristics for the target AM output and the set of second characteristics. This reward is used to determine a new policy, which refines both the actor network and the critic network.

[0010] According to a further aspect of the present disclosure, the method further comprises using a mathematical model to modify the set of second AM factors if the difference between the set of target characteristics and the set of second characteristics is within a predetermined threshold. The mathematical model is associated with the relationship between the AM process inputs and the geometric or physical properties, thereby generating a set of third AM factors.

[0011] According to a further aspect of the present disclosure, there is provided a system comprising apparatus configured to perform the methods disclosed herein.

[0012] According to an additional aspect of the present disclosure, there is provided a computer- readable storage medium, which is transitory or non-transitory, comprising one or more programinstructions which, when executed by one or more processors, cause the one or more processors to perform any of the methods disclosed herein.

[0013] Beneficially, complex sensorial systems are not required, increasing both the speed and efficiency of the automated AM process. The method allows for quick and precise adjustments to be made, resulting in a higher-quality product via efficient real-time compensation of changes in AM process conditions. Furthermore, the above aspects help prevent or discover defects, resulting in reduced waste and cost, resource, and energy savings for the manufacturer. The above- described aspects also streamline the process of obtaining robust and valid process parameters, helping result in more innovative AM and welding industries and accelerating the time to market for novel AM products.

[0014] Further beneficially, the method described in the above aspects leverages a RL framework to dynamically adjust welding parameters in real-time, significantly enhancing the adaptability and precision of the automated AM process. By obtaining a current state that includes both AM process inputs and the geometric or physical properties of the AM output, the RL framework can continuously learn and optimize the process. The actor network generates suggested actions to achieve target outputs, while the critic network evaluates these actions, ensuring that the process is constantly refined for optimal performance. This iterative learning process, combined with the enforcement of safety limits through a custom loss function, ensures that the AM process remains within safe operational boundaries, thereby reducing the risk of defects and enhancing overall product quality.

[0015] In addition to the real-time adaptability and precision provided by the RL framework, the method also offers significant benefits in terms of efficiency and resource management. By continuously updating the state and modifying AM factors, the process can quickly adapt to changes, reducing downtime and increasing throughput. This dynamic adjustment capability minimizes material waste and energy consumption, leading a more sustainable manufacturing process. Furthermore, the integration of a mathematical model to fine-tune AM factors ensures that adjustments are made with high accuracy, reducing the need for extensive trial-and-error experimentation. This not only accelerates the development cycle but also enhances the consistency, explainability, and reliability of the final products. Overall, these additional advantages contribute to a more efficient, effective, and environmentally friendly manufacturing process.BRIEF DESCRIPTION OF DRAWINGS

[0016] Embodiments of the invention will now be described, by way of example only, and with reference to the accompanying drawings, in which:

[0017] Figure 1 illustrates a system architecture for performing an automated AM process;

[0018] Figure 2A and 2B illustrate a system architecture for performing an automated AM process using a predictor and a mathematical model or a RL framework;

[0019] Figure 3 shows a plurality of plots illustrating relationships between a geometric or physical property and an AM process input; and

[0020] Figures 4A and 4B show a flow-chart illustrating a method for performing an automated AM process by generating a mathematical model and / or utilising reinforcement learning;

[0021] Figure 5 shows a flow-chart illustrating a method for updating an automated AM process;

[0022] Figure 6 shows a flow-chart illustrating a method for updating a mathematical model and outputting a plot;

[0023] Figure 7 shows an example computing environment for performing any of the methods described herein. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will now be described with reference to the attached figures. It is to be noted that the following description is merely used for enabling the skilled person to understand the present disclosure, without any intention to limit the applicability of the present disclosure to other embodiments which could be readily understood and / or envisaged by the reader. In particular, whilst the present disclosure is primarily directed to the automated additive manufacturing (AM) process, the skilled person will appreciate that the methods and systems described herein are applicable to the automation and optimisation of other manufacturing processes, such as subtractive manufacturing.

[0025] In one embodiment, the present disclosure relates to a system for AM, e.g., directed energy deposition such as welding, comprising one or more sensors such as a camera, linear laser scanning profilometer, or any 3D-scanning system that can generate a representative digital reconstruction of the deposited material. situated proximate to the light radiating source to obtain geometric or physical data from the deposited material. The system further comprises a control architecture, powered by a processing unit such as a GPU-enabled computer, which processes theobtained sensor data and compares it with previously trained values using a mathematical model based on a machine learning algorithm, such as an artificial neural network. If an observed deviation is detected, the system calculates new process parameters in real-time, and applies them at the correct location, with the assistance of the robot positioning data.

[0026] Additionally, the system can be directed towards generating new process parameters, e.g., without requiring observed deviations and modification to existing parameters, such as at the start of an AM process.

[0027] The systems and methods described herein preferably focus on bead-on-plate welding rather than block-pattern strategies commonly found in prior art. A primary innovation lies in the dynamic assessment and modification of process parameters in real time during welding. Unlike prior art, which typically involves a machine learning model that simply learns relationships between welding conditions and block patterns to predict weld bead shape data, this method emphasizes real-time measurement and adjustment of process parameters. Prior art generally describes systems that creates pass data from design data, determine initial welding conditions, use the learned model to predict weld bead shapes, and adjusts welding conditions before welding if the predicted shapes do not match the desired shapes within a threshold. At best, prior art described acquiring shape data of formed weld beads for generating training data, which is not suitable for real-time control, i.e. prior art does not describe precise, efficient, and explainable real-time monitoring or adjustment. The adjustment process in prior art is part of the planning or preparation phase, where conditions are adjusted through iterations of prediction until suitable parameters are found, rather than during the actual welding process.

[0028] Additionally, while prior art may describe predictors configured to output geometric properties of a deposited bead, such prior art does not address physical or non-geometrical properties of the deposited material. In contrast, the described method and systems herein addresses the use of information from thermomechanical simulations, such as stresses, hardness, and cooling rates, to extend and enrich the design matrix used in machine learning training. This approach is a significant advancement over prior art that enables the real-time control of the process not only based on the shape and geometry, but also a cumulative set of data including the properties of the processed material.

[0029] If the machine learning capabilities are based on a reinforcement learning (RL) paradigm, the methods and systems are further improved. For example, prior art describes a machine learning device that performs supervised learning to generate a model using welding conditions and block patterns as input data and shape data of the weld bead as output data. In contrast, a RL approach optimizes welding parameters by learning a policy that selects actions based on thecurrent state to maximize cumulative rewards. Unlike with supervised learning models of the prior art, e.g. which map welding conditions and block patterns to weld bead shape data, a RL method focuses on learning optimal actions through interaction with the environment, without explicitly modelling the output shape data as a function of input conditions.

[0030] In the RL approach described herein, dimensions like height, width, or volume of the weld bead are optionally incorporated into the state representation and used to compute the reward signal. These shape data are not used as output data of a learned model derived from input welding conditions and block patterns. Instead, shape measurements guide the RL agent's learning process, avoiding conflict with prior art claims. Furthermore, RL is a distinct paradigm from supervised learning, focusing on learning optimal actions through trial and error to maximize a cumulative reward, without relying on labelled input-output pairs.

[0031] A RL-based system does not utilize a learned model that derives shape data from welding conditions and block patterns to make adjustments. Instead, it learns a policy to adjust welding parameters in real-time based on the current state and reward feedback, aiming to optimize bead characteristics directly through interaction with the environment. The computed difference between derived shape data and desired shape data is used as a reward for the neural network outputs, in contrast to prior art, where the geometries are the sole predictions of the model. Therefore, the RL method operates differently from simpler machine learning systems.

[0032] Optionally, rather than being explicitly centred on a supervised learning scope, the methods and systems herein described employ RL for machine parameter real-time control instead of only bead geometry predictions. Bead geometry information is used as a reward function for the network's exploration and exploitation of the parameter space. Additionally, the RL method includes safety limits to the neural network predictions by incorporating a custom loss function that penalizes exploration of the RL state space beyond welding and additive manufacturing machine scopes, which is not addressed in prior art.

[0033] The objective of the RL approach is to optimize welding parameters to achieve desired bead shape characteristics such as width, height, penetration, and consistency across multiple beads. Challenges include controlling a large number of parameters and potential variance in results due to environmental factors, and a solution comprises modelling the problem as a Markov Decision Process (MDP), where the state represents the current welding process parameters, bead shape and processed material properties, the action represents possible adjustments in welding parameters, and the reward represents a measure of how well the current bead shape meets the desired specifications. The RL model's learned strategy, or policy, determines which actions to take in each state to maximize the cumulative reward.

[0034] Architecturally, the RL approach optionally employs a distinct actor-critic architecture comprising two separate neural networks, where one network focuses on actions and the other on expected cumulative rewards. This dual-network setup enables the RL agent to effectively balance exploration and exploitation, providing a more robust and adaptive control mechanism compared to the single-network supervised learning model in prior art.

[0035] The steps to implement RL in welding process control include data collection and numerical simulation environment setup, defining the state and action spaces, designing the reward function, selecting an RL algorithm, training the RL agent, and conducting validation and real-world testing. The state representation includes welding parameters, bead shape metrics, and external conditions. The action representation involves adjustments in welding parameters, and the reward design measures the error between desired and actual bead dimensions. Continuous control RL methods are suitable for fine-tuned parameter adjustments.

[0036] The advantages of RL in welding include adaptability, optimization, and scalability. RL algorithms can dynamically adjust parameters in real-time based on feedback, making the system adaptive to changes in material properties or external conditions. RL can find combinations of parameters that might not be obvious through traditional experimentation or rule-based systems. Once trained, the RL model can be transferred across different welding tasks with slight modifications, improving efficiency and consistency.

[0037] Reinforcement learning offers a way to control complex welding processes by continuously adapting parameters based on feedback from bead shapes and conditions. By designing an effective state representation, action space, and reward function, the RL agent can learn to optimize welding parameters in a dynamic environment, leading to consistent and high- quality bead formation. Dynamic changes of the parameters are bound to the physically feasible and practically qualified heat input range for the given material, either based on the requirements of a job or any physical properties that can be achieved through numerical simulations. This acts as “guardrails” for the RL model.

[0038] Additional benefits include the ability to mount a sensor in front of the welding or deposition tool, with an adjustable and calibrated distance between the welding tool centre point (TCP) and the sensor TCP. This distance acts as a decision-making time for the RL model. For example, if the distance is 100 mm and the welding speed is 5 mm / sec, it will take 20 seconds for the welding tool to arrive at the location scanned by the sensor 20 seconds earlier. This time is needed to calculate the correct process parameters according to the RL algorithm, determine the full list of parameters, wrap it as a "job," and communicate it with the welding machine or any other deposition system. When the welding tool arrives at that specific location, the predefined"job" will be executed based on the determined set of parameters, making the system highly dynamic and location-specific in terms of process parameters. This approach is impractical using the supervised learning suggested by prior art.

[0039] Figure 1 illustrates a system architecture for performing an automated AM process. Specifically, Figure 1 shows a system 100 according to an example embodiment of the present disclosure, wherein system 100 is configured to perform an AM process to generate a targeted AM output based on AM process parameters identified using a mathematical model.

[0040] System 100 comprises target characteristics 102, a device 104, a processor 106, AM factors 108, AM apparatus 110, an AM process 112, an AM output 114, and AM apparatus data 116.

[0041] System 100 is configured such that target characteristics 102 of a targeted AM output are obtained from a device 104 at a processor 106. The processor 106 calculates the AM factors 108 most likely to result in the targeted AM output when used by AM apparatus 110 through use of a mathematical model. AM apparatus 110 then performs an AM process 112, resulting in the generation of an AM output 114.

[0042] By using system 100, determining the AM factors 108 required for performing an AM process 112 capable of generating a targeted AM output is optimised. System 100 therefore results in a high-quality AM output 114 based on target characteristics 102, with improved efficiency regarding time, energy, materials, and resources in comparison to alternative iterative, repetitive, and labour-intensive processes of determining the AM factors 108.

[0043] Target characteristics 102 are obtained by the processor 106 and comprise one or more geometric or physical properties of the targeted AM output, such as height, width, length, area, depth, or any combination thereof. The target characteristics 102 are associated with the shape of the targeted AM output, such as the geometric or physical properties of a 3D finished design, a single layer of a finished design, or a single weld bead of a single layer. Optionally, target characteristics 102 are obtained from a device 104, the processor 106, an external device not depicted here, or by an operator, such as a designer aided by standard design-aiding software. Device 104 is therefore not essential to system 100, and is included for illustrative purposes only. Further optionally, target characteristics 102 are associated with the quality of the targeted AM output, such as a final material quality.

[0044] The processor 106 is one of one or more processors of system 100, such as a graphics processing unit (GPU) enabled computer. The processor 106 uses a mathematical model to determine a set of AM factors 108 suitable for generating the targeted AM output using AM apparatus 110. The processor 106 is communicatively coupled to AM apparatus 110, such thatdata associated with AM apparatus 110, AM apparatus data 116, is obtained by the processor 106. The AM factors 108 are therefore specific to the AM apparatus 110, resulting in more accurate determination of AM factors 108. AM apparatus data 116 comprises data associated with the AM apparatus 110, such as operation parameters.

[0045] For example, the AM factors 108 will be within a safe operational parameter space for the AM apparatus 110. Preferably, the processor 106 is local, such as a 32 GB GPU processor communicatively coupled to the AM apparatus. Alternatively or additionally, the processor 106 is an edge device, cloud-based, or server-based processor. Optionally, device 104 comprises processor 106 and a computer-readable storage medium that can be accessed by processor 106. For example, device 104 is computing device 702 of Figure 7.

[0046] The AM factors 108 comprise one or more AM process inputs, also known as process parameters, associated with the AM apparatus 110 for performing an AM process, such as input power, speed, voltage, or current. The AM factors 108 optionally contain material specific properties, such as material composition, physical / chemical structure, or temperature. For example, where the AM apparatus 110 is configured for welding DED, the AM process inputs comprise one or more configurable process parameters, such as but not limited to any of arc current, arc voltage, travel speed, wire feeding rate, or any combination thereof. In this welding example, the AM apparatus comprises a robotic arm configured for welding, a user interface, a sensor, one or more means for communicatively connecting to processor 106 (e.g., Wi-Fi, Bluetooth, or a wired connection), a welding power source, or a laser source or a combination thereof. Optionally, the AM apparatus is further configured to generate a visualisation of the robotic arm and / or the AM process 112 to aid in monitoring and control of the AM apparatus and / or AM process 112.

[0047] By performing the AM process 112, e.g., welding a metal, based on the determined AM factors 108, AM apparatus 110 produces an AM output 114, e.g., a welded metal, based on a target AM output. As mentioned above, system 100 advantageously eliminates the need for extensive iterative AM factor determination processes, which waste time, energy, and resources. A method for determining AM factors 108 from target characteristics 102 is described in detail below in relation to method 400 of Figure 4A or method 420 of Figure 4B.

[0048] Figure 2A illustrates a system architecture for performing an automated AM process using a predictor and a mathematical model. Figure 2B illustrates a system architecture for performing an automated AM process using a RL framework. Specifically, Figures 2A and 2B are directed towards a system 200 configured to implement any of method 400, method 420, method 500, ormethod 600 described below. Additionally, system 200 comprises aspects of system 100 of Figure 1.

[0049] System 200 comprises first AM factors 202, AM apparatus 204, processor 206, AM process 208, AM output 210, sensor 212, first characteristics 214, AM data 216, machine learning model 218, predictor parameters 220, mathematical model 222, target characteristics 224, second AM factors 226, mobile unit 228, and RL framework 230. Sensor 212 is optional such that first characteristics 214 are simulated, calculated, or otherwise obtained. Additionally or alternatively to mathematical model 222, RL framework 230 is used (see method 420 below), where RL framework comprises optionally machine learning model 218. Aspects of Figures 2A and 2B can be combined, replaced, interchanged, or otherwise generalised. For example, a sensor 212 is optionally used with RL framework 230.

[0050] System 200 obtains first AM factors 202 from AM apparatus 204, which is received at a processor 206, such as using step 402 of method 400. The AM apparatus 204 performs an AM process 208, which results in the generation of an AM output 210. A sensor 212 obtains first characteristics 214 associated with geometric or physical properties of the AM output 210, such as using step 404 of method 400. AM data 216 is comprised of the first AM factors 202 and first characteristics 214, which is used to train a machine learning model 218. Optionally, the processor 206 aligns or otherwise processes the first AM factors 202 and first characteristics 214 e.g., based on one or more time points. Optionally, AM data 216 further comprises historical AM data, such as historical AM factors and corresponding historical characteristics, simulated AM data, or other AM data such as AM data associated with a different AM apparatus.

[0051] For example, the machine learning model 218 comprises a predictor, e.g. a machined learning model trained using step 406 of method 400. Predictor parameters 220 associated with internal parameters of the machine learning model 218 are extracted to generate a mathematical model 222, such as using step 408 of method 400. The mathematical model 222 uses target characteristics 224, which can be provided by the processor 206, to determine second AM factors 226, such as by using step 412 of method 400. The second AM factors 226 are then used by the AM apparatus 204 to perform further AM processes, e.g., AM process 208, such as using step 414 of method 400.

[0052] For a supervised learning process, after the machine learning model 218 is trained and the mathematical model 222 is generated, processor 206 can obtain target characteristics 224, such as using step 410 of method 400, and the mathematical model identifies second AM factors 226. The second AM factors 226 are then used by AM apparatus 204 to perform AM process 208, resulting in generation of AM output 210 associated with target characteristics 224.

[0053] Optionally, sensor 212 obtains second characteristics from AM output 210 and AM data 216 further comprises second AM factors 226 and second characteristics. The machine learning model 218 is trained using AM data 216 and the mathematical model 218 updated based on extracting updated parameters from the machine learning model 218, such as using steps 602 – 606 of method 600. Alternatively or additionally, the second characteristics are compared to the target characteristics 224 and processor 206 determines a difference. Based on the difference, mathematical model 222 is used to modify the second AM factors 226 to compensate for the difference, such as using steps 502 – 506 of method 500.

[0054] Further optionally, processor 206 is configured to output a plot of first AM factor 202 against first characteristics 214, such as using step 608 of method 600. For example, processor 206 outputs a plot to device 104 of system 100, such as plot 300-A or 300-B of Figure 3.

[0055] For example, system 200 can be used in the process of acquiring deposited bead data and using it for controlling the AM and welding process. In this example, the AM apparatus 204 is a DED or welding robot, and the AM process 208 comprises use of a deposition tool e.g., in the form of a welding torch. Sensor 212 is a linear laser scanning profilometer or seam tracker, optionally attached to the processing head of AM apparatus 204 and in close proximity to a processing location. AM output 210 comprises a substrate for deposition or weldment and AM process 208 deposits a single deposited bead under a predetermined set of process parameters. The dimensions of the deposited bead, such as its width and height, are proportional to the process parameters applied during the deposition process. Such information is utilized to determine the optimal set of process parameters required for achieving the desired dimensions of the deposited bead, e.g., second AM factors 226.

[0056] A digital representation of the deposited bead is optionally obtained through data from sensor 212, which can be processed by sensor 212 or processor 206. Such a digital representation serves as a valuable tool for analysing and evaluating the performance of the deposition process, such as by processor 206. The digital data provides a comprehensive and accurate representation of the deposited bead, facilitating the identification of any defects or irregularities that may arise during the welding process. An analysis of a digital representation obtained from the sensor 212 is performed and parameters extracted, the parameters comprising bead width, bead height and cross-section area along with any determined deviations. Preferably, the data processing is performed exclusively onboard the sensor 212, without necessitating the utilization of any external computational resources.

[0057] Optionally, no sensor is required, and data is simulated or calculated. For example, data is acquired from a trained machine learning model or mathematical model configured to generatedata associated with a current state of the system 200 and AM process 208, e.g. simulated characteristics are provided based on the AM factors. Further optionally, a sensor 212 is used initially and then simulated, calculated, or modelled data is used.

[0058] The machine learning model 218 gathers and correlates bead dimensions with the associated process parameters, which is used in generating a mathematical expression based on the compiled data. This mathematical expression forms part of mathematical model 222. Processor 206, preferably a high-performance processing unit such as a GPU-enabled computer, uses the mathematical model 222 for predictive controlling of the AM process 208. This results in a direct control channel between the material processing and sensorial readings, correcting them on demand based on the obtained geometric or physical data of the deposited material shape and use of the mathematical model 222.

[0059] As well as performing initial training, such as using method 400 of Figure 4A or method 420 of Figure 4B below, initiating the machine learning model 218 and mathematical model 222 to accurately control the AM process 208, e.g., to automate AM apparatus 204, optionally requires testing steps. For example, the first step includes using sensor 212 to obtain a set of base characteristics of a workpiece, or other pre-process material or empty workspace, before the testing AM process is performed. The second step is performing the testing AM process and obtaining a set of test AM factors associated with performing the process. The third step is obtaining a set of test characteristics of the test output generated via the testing AM process. A database of the obtained data, e.g., the set of base characteristics, the set of test AM factors, and the set of test characteristics, is formed, such that the difference between the set of base characteristics and the set of test characteristics can be determined for calibration purposes. Optionally, this process is repeated such that a plurality of sets of test AM factors is obtained, representing the safe and / or efficient operating space of the AM apparatus 204. The plurality of sets of test characteristics can be used to determine a plurality of set of predicted AM factors using the mathematical model, and the robustness of the mathematical model can be determined over the available parameter space. Additionally, or alternatively, the obtained physical and numerical data can be further used to train the predictor to increase the accuracy of the mathematical model.

[0060] Generating and testing AM factors using system 200 takes substantially less time than discovering, testing, and finalising AM factors using traditional AM systems. For example, when system 200 is applied, steps involve preparing the AM apparatus 204, generating AM factors using the mathematical model 222, performing the AM process 208, and obtaining characteristics associated with the output generated from the AM process (data collection and processing) using sensor 212 and, optionally, processor 206. When a traditional AM system is applied, steps involve preparing the AM apparatus, discovering relevant / plausible AM factors, performing an AMprocess, analysing the output of the AM process (e.g. cutting / extraction, sample preparation, grinding / polishing / isolating the sample, etching, taking microscope measurements and potentially other measurements, and other standard procedures), and reporting the analysis, which often then requires adjusting the plausible AM factors and repeating the above steps to discover increasingly optimal AM factors.

[0061] Using a sensor 212 configured to obtain geometric or physical data, such as a linear scanning sensor, is faster, more automatic, and less prone to errors than using microscopes and the like. For example, when applied to welding, geometric or physical properties of a plurality of weld beads can be obtained quickly and reliably, and such data can be used in-process for monitoring the AM process. Additionally, optimal AM factors are generated quickly and accurately via use of the mathematical model. Overall, the time and efficiency savings of using system 200 are substantial, as are potential savings to physical and computational resources. Optionally, sensor 212 is configured to determine the hardness and residual stresses (RS) of the material, and first characteristics 212 comprise hardness, RS or other material properties in addition to or as an alternative to geometric or physical properties, where hardness and RS data is associated with the quality of the AM process 208 and / or AM output 210.

[0062] System 200 optionally further comprises a database. The database comprises historical AM factors and / or simulation-based AM factors for either AM apparatus 204, AM process 208, and / or a different AM apparatus or AM processes. Alternatively, or additionally, the database comprises data associated with material properties and apparatus information. It also contains pre-trained ML models, so that models can be reused. The database is communicatively coupled to any of AM apparatus 201, sensor 212, processor 206, machine learning model 218, mathematical model 222, or a combination thereof. For example, the database is communicatively coupled via a network to processor 206, and the database is a cloud-based or otherwise accessed from an external location, such as an online repository. Alternatively, the database is local.

[0063] The database is configured to store AM data 216 or any combination of first AM factors 202, information associated with AM output 210, first characteristics 214, and information associated with AM apparatus 204 and / or AM process 208. Alternatively, the database is configured for access only. For example, the database is accessed to determine historical AM factors used to produce an AM output similar to a target AM output. E.g., when welding a first feedstock onto a first substrate, the database is accessed to determine whether historical / simulated AM factors are stored for welding the first feedstock onto the first substrate. These historical / simulated AM factors are used instead of second AM factors 226 or used by mathematical model 222 in determining second AM factors 226. In another example, information stored in the database is used to determine safe / optimal operating conditions of the AMapparatus 204, which can be used by machine learning model 218 and / or mathematical model 222 to determine boundary conditions when determining AM factors.

[0064] Optionally, system 200 comprises system 100 of Figure 1, where AM apparatus 204 is AM apparatus 110, processor 206 is processor 106, AM process 208 is AM process 112, AM output 210 is AM output 114, and the mathematical model 222 is accessibly by processor 106. For example, AM apparatus 110 comprises mobile unit 228 and / or AM apparatus 204 does not comprise mobile unit 228.

[0065] System 100 or system 200 is configured to be stationary or mobile depending on, for example, the AM apparatus, the AM process, the manoeuvrability of the AM output, the environmental conditions required for the AM process, or other consideration regarding the location. In one example, the AM apparatus 204 is in a first location while processor 206 is in a second location. In another example, the AM apparatus 204 and the processor 206 are in a first location and the system 200 is accessed or otherwise monitored from a second location. In a further example, the AM apparatus 204 is mounted on a mobile unit 228 configured to be moved and controlled such that the AM process 208 can be performed remotely. For example, mobile unit 228 comprises one or more means for mobility and / or manoeuvrability of apparatus 204, such as one or more wheels.

[0066] Figure 3 shows a plurality of plots illustrating relationships between a geometric or physical property and an AM process input. Specifically, Figure 3 shows a plot 300-A of a geometric or physical property against an AM process input and a plot 300-B of a plurality of geometric or physical properties against a plurality of AM process inputs, such as the set of first characteristics against the set of first AM factors defined in relation to method 400 of Figure 4A or method 420 of Figure 4B below.

[0067] Plot 300-A comprises an x-axis 302 representing an operable range of the AM process input A, a y-axis 304 representing a geometric or physical property B, a line 306 representing the relationship between A and B, a target geometric or physical value BTrepresented by a target geometric or physical value line 308, a target input value ATrepresented by a target input value line 310, and an intercept point 312.

[0068] Plot 300-A shows a target geometric or physical value BTcorresponding to the target geometric or physical value line 308. The target geometric or physical value BTrepresents a targeted value of geometric or physical property B, e.g., as one of the one or more geometric or physical properties that make up the set of target characteristics obtained in step 410 of method 400, or target characteristics 224 of system 200. Line 306 represents the relationship between A and B, how B changes as A changes, e.g., based on the mathematical model generated in step 408of method 400. The position where the target geometric or physical value line 308 intercepts with the line 306, the intercept point 312, represents the value for A required to generate target geometric or physical value BT, which is the target input value AT. Therefore, plot 300-A shows how the mathematical model can be used to determine an AM process input based on a geometric or physical property as in step 412 of method 400 below. Additionally, plot 300-A shows how the mathematical model can be used to determine a predicted geometric or physical property based on a known AM process input.

[0069] Optionally, geometric or physical property B is a material property B, such that plot 300- A shows a target material property BT(corresponding to the target material value line 308), such as a desired hardness or residual stress. Additionally or alternatively, BTis a minimum or maximum target value, or BTis a range of target values, where target material value line 308 would become a minimum value line, a maximum value line, or an area of target values. The target input value line 310 is then optionally a minimum input value line, a maximum input value line, an area of target input values, or, e.g., an average value line of the area of target input values.

[0070] AM processes are complex, and typically involve a plurality of targeted geometric or physical properties, e.g., a set of target characteristics, and a plurality of AM process inputs, e.g., a set of AM factors. Plot 300-B shows example relationships of each geometric or physical property (B-1, B-2, B-3) with each AM process input (A-1, A-2, A-3, A-4). For example, B-1, B-2, and B-3 represent height, width, and cross-section area, and A-1, A-2, A-3, A-4 represent speed, current, feed rate, and temperature. As in plot 300-A, a target geometric or physical value for each of B-1, B-2, and B-3 can be used to determine input values A-1, A-2, A-3, A-4 based on the relationships.

[0071] Visualising the one or more relationships offers multiple benefits, especially as one or more plots. It enhances understanding, simplifies complexity without losing essential information, supports decision-making, and facilitates effective communication of insights derived from the mathematical model. For example, plot 300-A or plot 300-B aid in identifying patterns, trends, and potential anomalies. Such outlier / anomaly detection can help maintain accuracy in the mathematical model. Moreover, the plots can be used for identifying relevant features for predictor and / or prediction function, further helping maintain accuracy in the mathematical model. Additionally, plot 300-B allows for comparing multiple variables, including analysis of trends across multiple relationships.

[0072] Figure 4A is directed towards a flow-chart illustrating a method for performing an automated AM process by generating a mathematical model. Specifically, Figure 4A illustrates a method 400 for determining a set of AM factors based on an obtained set of target characteristics using a mathematical model that estimates a relationship between AM factors and characteristics.

[0073] Any of the steps of method 400 can be performed outside of the order shown in Figure 4A, and not all steps of method 400 shown are required. Additionally, any of the steps of method 500 of Figure 5 or method 600 of Figure 6 may be performed in addition, or instead of, the steps of method 400 listed in relation to Figure 4A.

[0074] A relationship between AM factors and AM output characteristics is required to accurately determine AM factors based on target characteristics. Determining such a factors-characteristics relationship is performed in two steps: first, a predictor comprising a machine learning model is trained using known AM factors in combination with associated AM output characteristics; and second, a mathematical model is generated based on the predictor.

[0075] Using a machine learning model allows for an accurate AM apparatus-specific approximation of the relationship, without requiring extensive details about the set-up, material, or process that may not be known or fully understood in relation to base principals. A machine learning model approach therefore decreases the amount of computational resources, material resources, and time required to approximate the relationship compared to a purely simulation- based model. However, it is difficult to understand the relationship based on a machine learning model alone, due to the complexities of various algorithms and model architectures. The number of internal parameters of any one machine learning model can prohibit human understanding and can make standard quality-based assurance testing difficult or impossible. Similarly, fully reverse engineering a machine learning model can be a time and resource heavy process, and success is not guaranteed. Therefore, the second step is to generate a mathematical model based on extracting one or more predictor parameters. The predictor parameters are associated with one or more relevant internal parameters, such as weights or activation functions, from within the machine learning model. The mathematical model provides an explainable, reusable, and computationally low-resource approach to linking AM factors with geometric or physical characteristics of the output.

[0076] For example, the method 400 can be used by a combined hardware and software system for AM processes, such as system 100 of Figure 1 and system 200 of Figure 2. The system comprises a sensor, such as a linear laser scanning profilometer or an x-ray scanner. As the predictor comprises a machine learning model such as an artificial neural network, the data stream obtained from the sensor is used for training the predictor and / or used for determining AM factors. The system is designed to improve the precision and efficiency of the AM process by monitoring and analysing data from the manufacturing environment in real-time (e.g., data from the sensor is communicated to the GPU-enabled computer in under 1 second, such as at 10-100 Hz, e.g., 25 Hz), and using the mathematical model further optimizes the AM process.

[0077] The sensor, such as a profilometer, is optionally mounted in close proximity to e.g., an intensive light emitting energy source used for performing the AM process, or the sensor, such as a scanner, is manoeuvrable. The sensor further comprises a logical processing unit for compilation of the data and edge computing and communicates with a GPU-enabled server.

[0078] Converting sensor data to mathematical models can eliminate the need for extensive data storage, database maintenance, and signal jittering, resulting in the need for smaller and more efficient digital systems. Method 400 allows for the compression of large amounts of data into mathematical models, which can be easily stored and manipulated using standard digital systems. As a result, the use of mathematical models for data processing can significantly reduce the cost, computational resources, energy, and complexity of data storage and management, while also minimizing the negative impact of signal jittering on data accuracy. Therefore, method 400 provides an improved approach for data processing that leverages mathematical modelling techniques to reduce the resources and complexities associated with data storage and management.

[0079] The use of method 400 in combination with e.g., system 200 of Figure 2A above can result in improved quality control, compliance with standards, increased efficiency, enhanced safety, and reduced waste and required resources in the manufacturing process.

[0080] Method 400 comprises step 402, step 404, step 406, step 408, step 410, step 412, and step 414. Step 402 comprises obtaining a set of first AM factors; step 404 comprises obtaining a set of first characteristics; step 406 comprises training a predictor using the set of first AM factors and the set of first characteristics; step 408 comprises generating a mathematical model based on extracting one or more predictor parameters; step 410 comprises obtaining a set of target characteristics; step 412 comprises using the mathematical model to determine a set of second AM factors; and step 414 comprises using the set of second AM factors to perform a second AM process.

[0081] Step 402 comprises obtaining a set of first AM factors, wherein the set of first AM factors comprises one or more AM process inputs associated with performing a first AM process, such as from the AM apparatus 204 or processor 206. The one or more AM process inputs comprise process inputs / process parameters necessary for performing the first AM process, such as the heat input.

[0082] For example, the first AM process is performed by an AM apparatus, such as AM apparatus 110 of Figure 1, and the first AM process is a welding process. Therefore, the set of first AM factors comprises weld speed, weld current, and wire feed rate, or any other configurable factor and combination thereof. In this example, the AM apparatus is suitable for welding, e.g., the AMapparatus comprises a welding torch, and the AM process comprises welding a metal. Alternatively, the AM process comprises welding a plastic, ceramic, glass, composite, or biological material. Optionally, the AM process inputs are associated with limits of the AM apparatus performing the first AM process, such as the minimum and maximum heat input.

[0083] Step 404 comprises obtaining a set of first characteristics from a sensor, the first set of characteristics comprising one or more geometric or physical properties of a first AM output. Preferably, the sensor is configured to process sensor data into one or more geometric or physical properties. Alternatively, the sensor data is processed by a processor, such as processor 106 of Figure 1. For example, the sensor is an optical scanning sensor, e.g., a linear laser profilometer, or a high-speed x-ray scanner, e.g., configured for diffraction-based hardness and residual stress measurements.

[0084] Returning to the welding process example, the one or more geometric or physical properties comprise weld bead characteristics, such as any of the height, length, depth, width, and / or area of a plurality of weld beads that make up at least part of a first AM output generated from the first AM process, e.g., a welding workpiece generated from a welding process performed using the set of first AM factors. Preferably, the sensor is configured to operate even in bright arc condition, such that the set of first characteristics can be obtained while the AM process is ongoing. Other example geometric or physical properties for one or more weld beads include bead width, heat affected zone (HAZ) width, bead height, HAZ depth, layer thickness, hatch spacing, roundness, hardness, residual stress, or other relevant property.

[0085] In addition or as an alternative to one or more sensors configured to obtain topographical data, either directly or indirectly e.g., from processing point-cloud data, one or more sensors are optionally configured to obtain data associated with monitoring the AM apparatus, the AM process, the surrounding conditions, or the output from the AM process, such as temperature, acoustic wavelengths, motion data, one or more electromagnetic wavelengths, deviations, or a combination thereof. For example, the one or more sensors are configured to obtain hardness and / or residual stress measurements based on x-ray scanning. Optionally, the one or more sensors are configured to obtain geometric properties, physical properties, or a combination thereof, such that the first set of characteristics comprise one or more geometric properties, one or more physical properties, or one or more geometric and physical properties.

[0086] Optionally, the sensor is configured to obtain and process data in real-time. For example, a processor, such as processor 106 of system 100 or processor 206 of system 200, receives data from the set of first characteristics within 1 second. Preferably, the processor comprises a GPU-enabled computer configured to obtain geometric or physical properties at a rate of 1-100 Hz, e.g., 25 Hz.

[0087] Step 406 comprises training a predictor using the set of first AM factors and the set of first characteristics. The predictor comprises a machine learning model configured to generate the set of first characteristics based on the set of first AM factors.

[0088] The machine learning model is configured to output one or more first characteristics (of the set of first characteristics) from an input comprising one or more first AM factors (of the set of first AM factors). The set of first AM factors and the set of first characteristics from step 402 and step 404 respectively form training data, and the machine learning model is training using standard supervised, unsupervised, or semi-supervised learning methods depending on machine learning model. For example, the machine learning model is a supervised artificial neural network (ANN) trained using backpropagation to minimise differences between outputs and the set of first characteristics. Alternatively, the machine learning model is an unsupervised algorithm, such as an autoencoder. Alternatively or additionally, the machine learning method is a reinforcement learning algorithm that controls the first and second characteristics of the process in real-time.

[0089] Returning to the welding process example, a predictor comprising an ANN configured to output the average height, the average width, and the average depth of a plurality of weld beads could be beneficial, based on an input comprising weld speed, weld current, and wire feed rate. Optionally, the ANN is configured to output the height, width, and depth of each weld bead of the plurality of weld beads, such that the AM factors resulting in the lowest standard deviation in height, width, or cross-section area can be determined for a consistent and high-quality weld. Alternatively, the ANN is configured to output the total height, total width, and total area e.g., for a given stage of the AM process, representing the part-completed or finished 3-dimensional shape of the weld.

[0090] Optionally, the predictor comprises a plurality of machine learning models. For example, the predictor comprises a first machine learning model configured to generate the set of first characteristics based on the set of first AM factors, as discussed above, and a second machine learning model configured to identify deviations in the set of first characteristics, such as an autoencoder. Deviations can then be compensated for, such as using step 412 or method 500 below, and / or prevented from being used in training the first machine learning model. Alternatively, deviations can be detected by the sensor of step 404 above, or by the mathematical model of step 408 and step 412 below.

[0091] Further optionally, the predictor comprises a reinforcement learning (RL) framework. For example, the predictor is a model / agent of the RL framework, and training the predictor using theset of first AM factors and the set of first characteristics comprises using the set of first characteristics as at least part of a reward function configured to guide an ANN or alternative suitable machine learning model to adjust process parameters, e.g. AM factors, to achieve a desired bead geometry, e.g. characteristics.

[0092] For example, step 406 or method 400 comprises any of the following processes: model section; data collection and / or preprocessing; initial model training; evaluation and validation; pilot testing and calibration; deployment and real-time or near real-time operation; monitoring and continuous improvement; or expansion to full scale and automation.

[0093] For model selection, one or more RL algorithms suitable for continuous action spaces are chosen, such as Soft Actor-Critic (SAC) or Twin Delayed Deep Deterministic Policy Gradient (DDPG) (TD3). Preliminary evaluations are preferably conducted to determine the best- performing model architecture for welding control. For data collection and preprocessing, structured data is collected from welding and profilometer sensors through design of experiments (DoE). Preprocessing is performed, comprising outlier removal, normalization, and feature engineering. The data collection is also optionally performed through a set of numerical analyses of the AM processes, e.g. the weld, such as pre-calibrated and / or validated finite element simulations, or similar. For initial model training, offline training of the actor-critic RL model is conducted using historical and simulated data. The model is trained to adjust AM factors, e.g., welding parameters, with simulated rewards, optimizing for bead geometry, process efficiency, and safety, e.g. characteristics. For evaluation and validation, the model is validated in a simulated environment using evaluation metrics, e.g. cumulative rewards and Q-value stability. The adaptability across various processing scenarios and materials is preferably also assessed.

[0094] For pilot testing and calibration, the model is deployed in a controlled pilot environment. Calibrated with real-world process settings, adjustments are performed for any variations not captured in the training data. For deployment and (near) real-time operation, the model is implemented on the edge hardware to control welding parameters, e.g. AM factors, in real time. Integrated sensors help initiate a feedback loop for real-time or near real-time adjustments based on continuous welding conditions. For monitoring and continuous improvement, performance is monitored during deployment, and optionally feedback gathered, to improve the model e.g. with additional data. Periodic or on-demand updates to the reward function and / or retraining the model results in better precision. For expansion to full scale operation and automation, the model’s deployment is expanded to additional machines and materials, possible due to its generalization capabilities, e.g. once proven in pilot testing. Integration with robotics results in autonomous material processing applications, allowing for wide-scale industrial use.

[0095] In a more detailed example, the main type of data used in the RL process is structured data from the welding robot and a non-contact profilometer sensor. The metadata about welding and AM processes is also preferable, comprising information about the consumable materials (e.g. wire or powder) and the substrate material, temperatures distribution during the process, and physical and mechanical properties of the materials used, such as tensile and yield strengths, hardness, thermal conductivity, or similar suitable parameters. Incorporating material properties helps the model generalize across different materials and conditions, enhancing its predictive capabilities over known methods.

[0096] Examples of data used in the process include process parameters such as, but not limited to, current, voltage, wire feed rate, and / or welding speed. Sensor data examples comprise temperature, deposited material profile and geometry and / or other non-contact sensor readings such as residual stresses and hardness, which are collected by the implemented sensors. As mentioned above, while deposition profiles are not essential as inputs for prediction, they preferably play a role in the reinforcement learning (RL) framework by serving as part of the reward function. This guides the machine learning model, e.g. NN, to adjust process parameters to achieve the desired bead geometry.

[0097] Several preprocessing stages are optionally performed, depending on the problem(s) at hand. General techniques for data cleaning include outlier removal using methods such as Inter- Quartile Range (IQR) filtering and noise reduction via low-pass filters or moving averages. For machine learning models, data centring and normalization are beneficial, e.g. involving normalizing both the data and its standard deviations to specific ranges or values. Common practices include scaling data between -1 and 1 and standardizing it to have a mean of 0 and a standard deviation of 1. Additional optional preprocessing steps involve Principal Component Analysis (PCA) for dimensionality reduction and feature engineering to aid the models in learning more effectively. Feature engineering optionally comprises creating new features from existing data, such as the rate of change of certain parameters, to capture more complex patterns.

[0098] The amount of data needed depends on whether the system is in the development phase or in the deployment phase. In the deployment stage, the goal is to have a set of pre-trained models, which can be enriched and tuned for specific purposes on a case-by-case basis, e.g. depending on the AM process, available equipment, or factors such as location. Data collection preferably occurs for both tuning and more generalised pre-training. Preferably, all data is structured in terms of time-series data, where a temporal flag is attributed to every sensor measurement or control parameter. This temporal alignment ensures synchronization between the states, actions, and rewards in the RL framework, which is beneficial for effective learning. Inthe deployment phase, the sensor is optional and the system can adapt the machine learning model and / or reinforcement learning framework to determine the process parameters.

[0099] Data collection is conducted using a DoE, following designs such as full factorial, Latin Hypercube Sampling (LHS) or Taguchi methods. Such a collection approach is beneficial to create a representative design space that limits the large number of involved parameters and the extensive experimental ranges yet retains statistically significant. For context, an experiment with two process parameters and 10 levels each would require 100 experiments in a full factorial design but can be reduced to around 30 significant ones using LHS. Optionally, new DoEs are required for different welding and AM hardware setup and materials used to result in more precise prediction and operation. In practice, a well-designed experiment for a combination of deposition and substrate materials optionally consists of a table with around 3,000 rows and 10 columns before preprocessing.

[0100] The system where method 400 or 420 is deployed, e.g. system 100, comprises an environment with continuous action space. Therefore, algorithms like Deep Deterministic Policy Gradient (DDPG), Twin Delayed DDPG (TD3), and Soft Actor-Critic (SAC) are among the suitable algorithms available to the skilled person. DDPG is an off-policy, model-free algorithm that combines actor-critic methods with deterministic policy gradients, allowing precise control actions. TD3 improves upon DDPG by addressing function approximation errors through twin critics and delayed policy updates, reducing overestimation bias and enhancing learning stability. SAC, on the other hand, is an off-policy, entropy-regularized actor-critic algorithm that encourages exploration by maximizing a trade-off between expected return and policy entropy, providing robust performance in continuous control tasks. Given the complexities of AM and welding processes, SAC or TD3 are preferable candidates due to their stability and performance.

[0101] Actor-critic RL methods mentioned above, having two main components, an actor and a critic network, are preferable. For example, the actor network's architecture begins with an input layer processing multiple data streams: sensor readings (e.g., temperature, hardness, residual stresses, bead geometry etc.), instantaneous process parameters, and material properties such as, but not limited to, tensile and yield strengths, hardness, and thermal conductivity. Machine or and material properties that are categorical variables are handled through embedding layers that convert categorical data into numerical representations. The hidden layers consist of fully connected layers with an arbitrary set of activations functions such as ReLU or Leaky ReLU. For handling temporal dependencies, the architecture supports optional LSTM or GRU layers. The output layer produces continuous values representing process parameter adjustments, employing tanh or sigmoid activation functions to ensure outputs remain within desired ranges. The critic network complements the actor by evaluating action quality. It processes both the stateinformation (matching the actor's input) and the proposed action, outputting a scalar value (e.g. Q-value) that estimates the expected cumulative reward. This network provides feedback for policy improvement, helping the actor refine its decision-making process over time.

[0102] For example, a welding environment is a physical welding system comprises the welding and AM machinery, sensors, and models. This environment generates real-time feedback through various sensors measuring bead geometry, possibly temperature distributions, and other process parameters. After each control action, the environment provides both the new state of the system, and a reward signal based on weld quality metrics. The actor network functions as the primary controller, receiving instantaneous welding state information (including welding robot position, process parameters, temperature readings, and material properties) and determining appropriate adjustments to welding parameters. For example, when the state indicates excessive heat input, the actor might adjust travel speed or reduce current. This network essentially encodes the welding parameter control policy, mapping sensor readings to parameter adjustments. The critic network evaluates the quality of the actor's parameter adjustments. It processes both the current welding state and the proposed parameter adjustments, estimating how these changes will affect future weld quality. For instance, if the actor suggests increasing travel speed, the critic estimates how this will impact bead geometry and overall weld quality over the next several timesteps. The advantage (Adv) calculation quantifies how much better (or worse) a specific parameter adjustment is compared to the average adjustment for that welding state. This is calculated as: Advantage = Q(state_t, adjustment_t) - V(state_t), where Q(state, adjustment) represents the expected quality of making a specific parameter adjustment in the current welding state, at specified time “t”, and V(state) represents the expected quality of average parameter adjustments in that state.

[0103] For training, the model uses data from either real or simulated welding sessions. This data forms what is called an “experience replay buffer” in RL, allowing the agent to sample past experiences to break training correlations when needed, and stabilize learning. The model alternates between exploration, where it tries new parameter adjustments, and exploitation, where it applies what it has learned to optimize welding quality. The choice of reward function here plays an important role in guiding the RL agent towards desired welding outcomes. Preferably, a custom reward function is designed to prioritize bead geometry accuracy, process efficiency, and safety.

[0104] Geometry Deviation Penalty encourages the agent to minimize bead geometry error by penalizing deviations from the target geometry. For example, the reward at time t can be reduced by^௧ = ^௧ − ^ × (Geometry Error)௧where α is a scaling factor that adjusts the weight of geometry error in the overall reward. An additional or alternative is a safety penalty making the reward function include penalties for violating safety constraints or operating beyond equipment limits: ^௧ = ^௧ − ^ × (Safety Error)௧where β is a scaling factor that controls the impact of safety violations on the reward. The functions relating to what constitutes the safety errors or geometry errors are here unspecified. For the geometry error, several options such as root mean squared error could follow.

[0105] Rewards can also be given for efficient welding practices, such as minimizing energy consumption, reducing material waste, environmental footprint and stabilizing process parameters over time. Balancing the reward components involves fine-tuning these scaling factors (e.g. α or β) and finding optional metrics to measure the deviations. During training, the agent's performance is periodically evaluated based on cumulative rewards from test episodes. These episodes reflect typical or challenging welding conditions.

[0106] Additional metrics include the Q-value accuracy and stability of critic predictions, which are preferable for meaningful feedback to the actor. This evaluation helps confirm the agent’s ability to generalize across different welding scenarios, providing insights into its adaptability to varied material properties and process parameters.

[0107] Returning to the above example, the actor and critic networks are updated through gradient steps based on specific loss functions. In each training step, the critic’s parameters are adjusted to reduce the error between the predicted Q-value and a target Q-value. By performing a gradient descent step on this loss, the critic becomes better at estimating the impact of adjustments, like modifying travel speed or current, providing more accurate feedback to the actor. Similarly, the actor is trained to choose actions that maximize the critic’s Q-value, meaning it learns to make parameter adjustments that are likely to improve the stability and quality of process and properties. Through gradient ascent, the actor’s parameters adapt to produce actions with higher expected rewards. Both networks are updated with gradients from a mini-batch of previous experiences, where the critic uses gradient descent to minimize Q-value prediction errors, and the actor uses gradient ascent to refine its choices. In TD3, the actor updates are delayed avoiding overestimations, while SAC introduces an entropy term to foster exploration. These gradient steps gradually optimize both action selection and evaluation, enabling precise and adaptive control over the process parameters.

[0108] The hardware architecture of the system, e.g. system 100, is preferably centred on edge computing utilizing a GPU-enabled computational resource, with optional integration of cloudinfrastructure. The GPU-enabled computer serves as the primary edge computing platform, interfacing directly with welding equipment and robots to manage real-time inference, online training, and process control decisions through its hybrid CPU / GPU-accelerated computing capabilities. This setup allows for efficient handling of computationally intensive tasks at the edge, reducing latency and improving responsiveness.

[0109] An example data acquisition system comprises high-precision weld seam track sensors, including but not limited to a laser profilometer, which feed data directly into the GPU-enabled resource. This system achieves low latency for control decisions through optimized local processing loops. Communication within the local environment relies on high-speed industrial ethernet, while encrypted channels are used for cloud connectivity when implemented. This ensures secure and efficient data transfer between local and cloud resources.

[0110] Cloud integration through established providers can offer supplementary storage and computational resources for intensive tasks such as large-scale model training, data analytics, and system monitoring. The cloud can also store pretrained models specific to different materials or machines, which can serve as initial steps for user-specific models. In this hybrid workflow, the GPU-enabled resource manages continuous online learning and model updates based on specific materials, use cases, and machines, while the cloud handles more extensive training using aggregated data from multiple sources when authorized. Importantly, users have the flexibility to maintain and interact with all data, including trained models and operational parameters, entirely locally on the edge device if preferred, ensuring complete data sovereignty. The modular architecture supports the integration of additional sensors and processing capabilities, with flexible resource allocation for expanding data and model requirements, regardless of whether a local or hybrid storage approach is chosen.

[0111] Any machine learning models of the RL framework can be improved in various orthogonal directions. One possible route is through refinements of reward components such as geometry deviation, efficiency, and safety penalties, ensuring that the RL agent adapts to different scenarios and making the reward function more representative of desired process outcomes. Additional machine learning techniques can also be experimented with, such as data augmentation. Through a process called domain randomization, synthetic data variations are generated by altering conditions in simulation (e.g., temperature fluctuations or material variability). This extends the model's exposure to diverse processing conditions, supporting generalization across new materials and conditions. Another goal is to develop a model that can extrapolate to different materials based on material properties and a set of experiments conducted for some materials, similar to "Model-Agnostic Meta-Learning" (MAML). This could be achieved by incorporating material properties into the neural network's input layer, allowing the model to generalize itslearning across various materials and conditions. Lastly, the technology can be optionally improved by incorporating real-time user feedback on weld quality and system performance into the training pipeline, allowing the system to learn and improve from practical applications and adjust based on feedback in deployment.

[0112] To summarise, a trained RL-based control method is preferably applied within an industrial welding and / or AM environment, where it operates in real-time or near real-time to control welding parameters. The application process includes a setup phase where the RL agent undergoes a calibration run. During this phase, the system aligns with the specific welding machine settings, material types, and environmental conditions. Once calibrated, the RL agent monitors the welding process, adjusting parameters such as current, travel speed, and wire feed rate in real time. It continuously adapts to variations in material properties and environmental changes, ensuring optimal welding conditions, and can respond to desired weld geometries such as specific widths, heights, and penetration depths. Although the process parameters are controlled and adjusted adaptively in real-time, the initial objective defined by the end-user, which is normally in compliance with industry standards, will be set as barriers to limit the parameter variations within a given range. This ensures a straightforward translation to the welding and AM procedure specifications. Accordingly, steps 408 to 414 are considered optional or performed within the RL-based control method.

[0113] Preferably, in deployment, the RL agent receives continuous feedback on welding quality and geometry, enabling it to make immediate adjustments for welding precision. The system also logs additional metrics and potential anomalies, which can be used in post-process analysis to identify areas for improvement. Beyond parameter control, the system’s insights are used in quality control, where it identifies optimal settings for different welding conditions. Additionally, the model can support automation processes by enabling robotics to execute complex welding tasks or design of experiments autonomously, with reduced human intervention and, beneficially, with much faster process parameter experimentation and welding specification times.

[0114] Step 408, which is optional, comprises generating a mathematical model based on extracting one or more predictor parameters. The mathematical model optionally comprises a prediction function, which mathematically represents an estimated / approximate relationship between one or more AM process inputs and one or more geometric or physical properties. The one or more predictor parameters are associated with internal parameters of the predictor of Step 406, such as weights, biases, or activation function. Internal parameters of the predictor are the values used to transform input data (e.g., AM factors) into output data (e.g., geometric or physical properties of the set of characteristics). These values change during training to transform the input data more accurately into the output data, such as by using backpropagation. Therefore, theinternal parameters are associated with the relationship between the AM factors and geometric or physical properties of the set of characteristics.

[0115] If there is a predictor of first and second characteristics, it is an ANN algorithm where the prediction function mathematically represents an approximated relationship between i) example AM factors comprising the weld speed, weld current, and wire feed rate, and ii) example characteristics comprising the height, width, and cross-section area of weld beads. For example, the prediction function is a nonlinear mathematical expression comprising the one or more AM process inputs as variables, with coefficients based on the one or more predictor parameters, such that the one or more geometric or physical properties can be calculated. The coefficients depend on the specific AM apparatus, resulting in a mathematical model tailored to a specific AM process or AM apparatus

[0116] Optionally, unlike with complex “black-box” machine learning models, the mathematical model can be scrutinised and the approximate relationship can be analysed. An explainable mathematical model therefore results in a clear link between AM process inputs and geometric or physical properties of AM outputs, which can be modified, and reverse engineered to generate greater understanding of a highly complex relationship. Beneficially, the mathematical model is less resource heavy than a standard machine learning model.

[0117] Alternatively, the present invention comprises a hierarchical control system wherein a mathematical model provides a foundational framework for determining and monitoring automated additive manufacturing process parameters, and a Reinforcement Learning (RL) framework enables dynamic parameter adjustments. The mathematical model maintains scrutiny and analysis capabilities of the approximate relationships between additive manufacturing process inputs and geometric or physical properties of additive manufacturing outputs, while the RL framework provides real-time adaptive control. The combined system ensures explainability through the mathematical model while enabling dynamic optimization through the RL framework, thereby maintaining efficiency in computational resources and energy consumption while providing enhanced control capabilities. The mathematical model can therefore be updated and used with less energy consumption, computational resources, and time compared to many machine learning models. For example, the mathematical model, once generated, can be updated and operated using a smaller / more efficient GPU or a CPU-based processor, instead of needing the GPU requirements necessary for training and running a machine learning model.

[0118] In one example, the predictor comprises a single ANN with an input layer, and output layer, and a hidden layer where the number of nodes is determined using step 604 of method 600 below. The hidden layer is a dense layer, and the ANN is trained until a calculated R-Squared valueis above 0.999. All nodes use a TanH activation function, and a boosting algorithm is also applied, such as AdaBoost or gradient boost. In this example, the AM process is welding, and the AM apparatus comprises a welding power source. The first AM factors comprise travel speed, power source voltage, power source current, and wire speed, and the first characteristics comprise average cross-section area, average width, and average height.

[0119] Internal parameters of the trained ANN, e.g., weights and activation functions, are extracted to form a prediction formula using standard mathematical methods. The mathematical model comprises one or more prediction formula, and each prediction formula has either a predefined structure or is tailored to each predictor. An example one-dimensional prediction formula to determine height, ℎ, is provided below:

[0120] ℎ = TanH(0.2 + 0.1 ∗ log((ିଽଽ.^ ା ^௨^^^^௧)(ଶଶଶ.^ ି^௨^^^^௧)) − 0.4 ∗ loglog((ିଷ,^ ା ^^^^ ௌ^^^ௗ)(ଽ,ସ ି^^^^ ௌ^^^ௗ)

[0121] As is typical, the activation function TanH is defined as:

[0123] In the example of a node within the hidden layer of the above ANN, ^ would be the sum of weighted inputs to the node.

[0124] The prediction formula can also be multi-dimensional, utilising vectors, tensors, and / linear algebra to mathematically express the factors-characteristics relationship.

[0125] Optionally, alternative algorithms, structures, thresholds, statistical measures, activation functions, training methods, or combinations thereof are used. For example, alternative activation functions include sigmoid, linear, Gaussian, rectified linear unit (ReLU), or other activation functions not listed here. Alternatively or additionally, multiple machine learning models are used, e.g., depending on the obtained sensor data. For example, multiple machine learning models are trained in step 406, but only internal parameters from the most accurate machine learning model(s) are extracted in step 408 to generate the mathematical model. In another example, internal parameters from multiple machine learning models are extracted in step 408 to generate the mathematical model, which comprises one or more prediction formulae. Additionally, the mathematical model is optionally generated using standard reverse engineering approaches to extracting a relationship between inputs and outputs from machine learning models, e.g., recovering the ratio between weights and biases. For example, the predictor comprises a RL framework as described above in relation to step 406.

[0126] Further optionally, the mathematical model comprises multiple prediction formulae. For example, a first prediction formula is generated based on extracting internal parameters from a first machine learning model of the predictor, and the second prediction formula is generated based on extracting internal parameters from a second machine learning model of the predictor. Either the best prediction formula, e.g., the most relevant for the specific AM process or the most accurate, is selected or one or more prediction formulae are combined within the mathematical model.

[0127] By generating the mathematical model from a machine learning algorithm, less time, energy, equipment, material, and computational resources are required compared to generating a mathematical model from base principals. For example, generating a mathematical model from base principals (e.g., known physics-, chemistry-, electronics-, and mechanics-based equations and theories) would require immense quantities of data and a complex multiple-sensor system. Substantial data processing would also be required, as the particle-level data necessary for an accurate mathematical model cannot be directly obtained from readily available sensors. Additionally, all the interactions between base principals are still not fully understood due to the complexity of AM processing. Therefore, even if it were possible, such a model would be large and prohibitively complex, requiring immense processing capabilities to operate accurately. In contrast, the computational requirements of extracting predictor parameters and formulating a mathematical model are lower without compromising on accuracy.

[0128] Step 410 comprises obtaining a set of target characteristics, wherein the set of target characteristics comprise one or more geometric or physical properties of a target AM output. For example, the set of target characteristics are target characteristics 102 of system 100 above.

[0129] Returning to the above welding process example, the target AM output is a high-quality weld with a desired average height, width, depth and cross-section area for each weld bead. The set of target characteristics therefore comprise the instantaneous height, width, depth, and cross- section area of the target AM output.

[0130] The set of target characteristics are obtained by a processor (e.g., processor 106 of system 100), such as from storage or from an input device, or are obtained by determining the heigh quality output that can be generated from the AM apparatus and AM process, e.g., based on the available materials and parameter space. Optionally, the target characteristics are predefined. Alternatively, the target characteristics are based on training requirements for the predictor, such as if more training data is required due to, e.g., the accuracy being below a predetermined threshold.

[0131] Step 412 comprises using the mathematical model to determine a set of second AM factors based on the set of target characteristics. The set of second AM factors comprise one or more AM process inputs associated with performing a second AM process. For example, the set of second AM factors are AM factors 108 of system 100 above.

[0132] Returning to the above welding process example, the mathematical model comprises an approximate relationship between AM process inputs and geometric or physical properties. To obtain the average height, width, and cross-section area of the target AM output as defined by the set of target characteristics, a specific set of inputs must be used by the AM apparatus. These specific inputs can be determined based on the relationship between AM process inputs and geometric or physical properties defined by the mathematical model, allowing the required process parameters, e.g., travel speed, current, and wire feeding rate, to be identified.

[0133] Preferably, the operational limits and boundary conditions defined within the system apply concurrently to both the mathematical model and the RL framework. The RL system operates within these predefined constraints while continuously monitoring real-time process data. When approaching operational boundaries, the RL framework implements adaptive responses while maintaining compliance with the defined operational limits. This dual-constraint system ensures process stability while enabling dynamic parameter optimization, thereby preventing process disruptions while maximizing manufacturing efficiency.

[0134] Optionally, the mathematical model is tailored to each AM apparatus and / or AM process in steps 406 and 408, such that limits to available parameters are enforced by boundary conditions. For example, the maximum current, voltage, wire speed, etc. are incorporated into the mathematical model, and the set of second AM factors identified by the mathematical model are within safe operational limits of the AM apparatus and / or AM process. Optionally, the limits are obtained from the AM apparatus and, further optionally, these limits are adjustable, such as by an operator of the AM apparatus or other safety system to increase safety and efficiency of the AM process.

[0135] Optionally, determining a set of second AM factors comprises calculating one or more process inputs of the set of second AM factors. Further optionally, determining a set of second AM factors comprises identifying one or more relevant AM process inputs to be included in the set of second AM factors.

[0136] Step 414 comprises using the set of second AM factors to perform the second AM process, thereby generating a second AM output. For example, referencing system 100 above, the AM apparatus 110 uses the AM factors 108 identified using the mathematical model to perform an AM process 112 to generate an AM output 114. Ideally, the second AM output comprises the targetcharacteristics of the target AM output, such that the second AM output is a generated version of the target AM output. Any differences between the second AM output and the target AM output can be used to update and, preferably, improve the mathematical model, as described in more detail in reference to method 500 below.

[0137] As referenced above, any of the steps of method 400 can be performed alone or repeated. For example, once a mathematical model has been generated, step 410 and step 412 can be repeated to determine AM factors required to generate specific target AM outputs (as in system 100 above) or to continually update the AM process to maintain quality and efficiency. In another example, step 402 and step 404 can be repeated and, optionally, the obtained data stored in order to generate a large training data set for training the predictor. In another example, step 414 does not need to be performed, as completing the second AM process may no longer be necessary, or it may be performed by an external AM apparatus.

[0138] Beneficially, method 400 provides an efficient and reliable means for controlling the deposition and manufacturing processes, thereby improving the quality and accuracy of the final product. The disclosed steps of method 400 can be performed for an unknown set of materials or be used for a material that is previously trained and found in a system database, which can be accessed by a processor and / or the predictor and mathematical model. The apparatus for performing method 400 comprises a processor and, optionally, a memory that stores a set of instructions for executing the automatic selection of AM process parameters and controlling the deposition in real-time. The processor is configured to execute the set of instructions stored in the memory to automatically select AM process parameters for the target materials and control the deposition in real-time based on the selected parameters.

[0139] Figure 4B is directed towards a flow-chart illustrating a method for controlling an automated AM process via an RL system. Specifically, Figure 4B illustrates a method 420 for controlling a set of AM factors based on set of obtained target characteristics. Any of the steps of method 500 of Figure 5 or method 600 of Figure 6 may be performed in addition, or instead of, the steps of method 420 listed in relation to Figure 4B. Similarly, steps of method 420 may be incorporated or replaced with steps of method 400 where beneficial.

[0140] In combination with RL, a relationship between AM factors and AM output characteristics is required to accurately determine AM factors based on target characteristics. Determining such a factors-characteristics relationship is performed in two steps: first, a predictor comprising a machine learning model is trained using known AM factors in combination with associated AM output characteristics; and second, a mathematical model is generated based on the predictor.

[0141] Method 420 comprises step 422, step 424, step 426, step 428, step 430, step 432, and step 434. Step 422 comprises collecting a state of the system; step 424 comprises feeding the state to actor and critic networks; step 426 comprises the actor network providing a suggested action to the critic network; step 428 comprises updating process parameters; step 430 comprises generating a new state and calculating a reward; step 432 comprises creating a new policy using the reward and feeding the new state into the critic network; and step 434 comprises calculating a new reward and updating using the new reward. The method 420 is designed to loop and repeat, allowing for continuous optimization and real-time adjustments based on the evolving state of the system.

[0142] Step 422 comprises collecting the current state of the system, which serves as the foundational data for the subsequent steps in method 420. One or more sensors or data acquisition devices are strategically placed throughout the system to capture a comprehensive set of parameters and conditions. These sensors may include, but are not limited to, temperature sensors, pressure sensors, motion detectors, and other relevant measurement tools capable of capturing real-time data about the system's performance and environmental conditions. The collected data comprises relevant physical measurements such as temperature, pressure, humidity, and / or other environmental factors, and optionally operational parameters such as machine settings, process variables, or system outputs. This data collection ensures that relevant aspects of the system's state are accurately captured, providing a detailed snapshot of the system's current operational status.

[0143] Preferably, the data collection process is designed to be thorough and precise, involving continuous monitoring and recording of data over a specified period to account for any fluctuations or variations in the system's performance. This may include high-frequency data sampling to capture transient events and ensure that the collected data reflects the true state of the system. The collected data is then compiled and formatted in a manner suitable for further analysis and processing. Preprocessing steps such as filtering, normalization, and outlier removal are often employed to enhance the quality and reliability of the data. Filtering helps in eliminating noise, normalization ensures that the data is on a consistent scale, and outlier removal addresses any anomalies that could skew the analysis. This preprocessing is beneficial for ensuring that the data fed into the subsequent steps is clean, accurate, and representative of the system's true state.

[0144] By ensuring that relevant data is accurately captured and preferably pre-processed, step 422 enables the effective functioning of the actor and critic networks in step 424, the generation of suggested actions in step 426, and the subsequent optimization and policy creation steps. This foundational step ensures that the system operates based on the most current and accurate information, leading to improved performance and efficiency.

[0145] Step 424 comprises feeding the collected state of the system into the actor and critic networks, which are integral components of the reinforcement learning framework employed in method 420. The actor and critic networks are configured to process the state information and generate actions and evaluations that guide the optimization of the system's performance. The state data, collected in step 422, is preferably pre-processed and formatted to be compatible with the input requirements of these one or more neural networks.

[0146] The actor network receives the state data and processes it through multiple layers of neurons, each layer applying a series of transformations to extract relevant features and patterns. The actor network's primary function is to propose a set of actions that are expected to optimize the system's performance based on the current state. These actions are generated by the network's output layer, which produces continuous values representing adjustments to process parameters. The architecture of the actor network may include fully connected layers, activation functions such as ReLU or tanh, and optional recurrent layers like LSTM or GRU to handle temporal dependencies in the data.

[0147] Simultaneously or in parallel, the critic network also receives the same state data, along with the actions proposed by the actor network (see step 426 below). The critic network evaluates the quality of these actions by estimating the expected cumulative reward, which reflects the long- term benefit of the actions in optimizing the system's performance. The critic network processes the combined state and action data through its own set of layers, producing a scalar value known as the Q-value. This value provides feedback to the actor network, helping it refine its action proposals in future iterations. The critic network's architecture typically mirrors that of the actor network, with layers designed to capture the complex relationships between state, action, and reward.

[0148] By feeding the state data into the actor and critic networks, step 424 enables the reinforcement learning framework to generate informed actions and evaluations that drive the continuous optimization of the system. This step ensures that the system operates based on the most current and accurate information, allowing for real-time adjustments and improvements in performance.

[0149] Step 426 comprises the actor network generating a suggested action based on the current state of the system and providing this action to the critic network for evaluation, e.g. as described in step 424. The actor network, having processed the state data received in step 424, utilizes its trained neural architecture to propose an optimal action aimed at improving the system's performance. This action is represented as a set of continuous values corresponding toadjustments in process parameters, such as temperature, pressure, speed, or other relevant operational variables.

[0150] The suggested action is derived from the output layer of the actor network, which applies activation functions to ensure the action values are within acceptable ranges. The actor network's design, which may include fully connected layers, activation functions like ReLU or tanh, and optional recurrent layers such as LSTM or GRU, allows it to capture complex patterns and dependencies in the state data. By leveraging these patterns, the actor network can propose actions that are expected to yield the highest rewards based on its training and the current state of the system.

[0151] Once the actor network generates the suggested action, it is fed into the critic network along with the current state data. The critic network's role is to evaluate the quality of the proposed action by estimating its expected cumulative reward. This evaluation helps determine how beneficial the action will be in achieving long-term optimization of the system's performance. The critic network processes the combined state and action data through its layers, producing a Q-value that reflects the expected reward.

[0152] By providing the suggested action to the critic network, step 426 facilitates a beneficial feedback loop within the reinforcement learning framework. This loop allows the system to continuously refine its actions based on the critic's evaluations, leading to iterative improvements and real-time optimization of the system's performance.

[0153] Step 428 comprises updating the process parameters, e.g. AM factors, based on the evaluations provided by the critic network and the mathematical model. When deviations are within predetermined limits, e.g. one or more predefined thresholds, the system comprises the mathematical model of method 400 for systematic, long-term improvements based on historical data analysis. However, if deviations exceed these limits, the RL-based control system is triggered to make immediate, real-time adjustments. This dual mechanism ensures that the system can respond promptly to process variations while maintaining the accuracy and stability provided by the mathematical model.

[0154] When using RL-based control, step 428 comprises updating the process parameters based on the evaluations provided by the critic network. After the critic network has assessed the suggested action from the actor network in step 426, it generates a Q-value that reflects the expected cumulative reward of the proposed action. This Q-value serves as a critical piece of feedback, guiding the adjustment of the process parameters to optimize the system's performance.

[0155] The process parameters are updated by incorporating the feedback from the critic network into the actor network's decision-making process. This involves adjusting the weights and biases within the actor network through a process known as backpropagation, where the network learns to improve its action proposals based on the critic's evaluations. The goal is to minimize the difference between the predicted Q-value and the actual reward observed from the system's response to the action. By iteratively refining the actor network's parameters, the system becomes better at proposing actions that maximize the expected reward.

[0156] In practical terms, the updated process parameters are applied to the system, altering its operational settings to reflect the optimized values suggested by the actor network. These parameters may include adjustments to variables such as temperature, pressure, speed, or other relevant factors that influence the system's performance. The continuous updating of process parameters ensures that the system adapts in real-time to changing conditions and maintains optimal performance.

[0157] By updating the process parameters based on the critic network's evaluations, step 428 enables the system to dynamically optimize its operations. This iterative process of action proposal, evaluation, and parameter adjustment forms the core of the reinforcement learning framework, driving continuous improvement and real-time adaptation in the system's performance, while optionally maintaining alignment with the long-term goals defined by the mathematical model.

[0158] Step 430 comprises generating a new state of the system and calculating the reward based on the updated process parameters applied in step 428. This step is beneficial for assessing the impact of the adjustments made to the system and for providing feedback that will guide further optimization. In step 430, the system operates under the updated parameters, leading to a new state. This new state is captured using sensors, and the reward is calculated based on the system's performance. The reward function considers both the immediate performance metrics and, preferably, the long-term trends identified by the mathematical model. This dual consideration ensures that the system's performance is optimized in real-time while

[0159] Once the updated process parameters are applied, the system operates under these new conditions, leading to a new state. This new state is captured using the same sensors and data acquisition devices employed in step 422. The one or more sensors collect real-time data reflecting the system's current operational status and environmental conditions, including any changes resulting from the updated parameters. This data is then compiled and formatted to represent the new state of the system.

[0160] Additionally, the reward is calculated based on the new state and the system's performance under the updated parameters. The reward function is designed to quantify the effectiveness of the applied changes in achieving the desired outcomes. It takes into account various performance metrics, such as efficiency, quality, and stability, and assigns a numerical value representing the overall benefit of the adjustments. The reward function may include penalties for deviations from target values or for violating operational constraints, ensuring that the system remains within safe and optimal operating conditions.

[0161] By generating a new state and calculating the reward, step 430 provides beneficial feedback for the reinforcement learning framework. This feedback is used to evaluate the success of the applied changes and to guide further adjustments in subsequent iterations. The continuous loop of state generation, reward calculation, and parameter updating enables the system to adapt in real-time, optimizing its performance through iterative improvements.

[0162] Step 432 comprises creating a new policy based on the reward calculated in step 430 and feeding the new state into the critic network for further evaluation. The actor network's parameters are updated using the reward, which reflects both real-time performance and long- term trends. The new state is preferably also evaluated by the critic network to ensure that the immediate adjustments align with the systematic improvements suggested by the mathematical model. This step ensures that the system continuously refines its policies based on both real-time feedback and historical data analysis. Additionally, this step is beneficial for refining the decision- making process of the actor network and ensuring continuous optimization of the system's performance.

[0163] The new policy is created by updating the actor network's parameters using the reward feedback. The reward provides a measure of how effective the previous action was in achieving the desired outcomes. By incorporating this feedback, the actor network adjusts its internal parameters, such as weights and biases, to improve its future action proposals. This process, known as policy optimization, involves using algorithms like gradient ascent to maximize the expected reward. The goal is to enhance the actor network's ability to propose actions that lead to higher rewards, thereby improving the overall performance of the system.

[0164] Simultaneously, the new state generated in step 430 is fed into the critic network. The critic network evaluates this new state along with the updated action proposed by the actor network, producing a new Q-value that reflects the expected cumulative reward. This evaluation helps in assessing the quality of the new policy and provides further feedback for refining the actor network. The critic network's continuous evaluation ensures that the actor network's policy remains aligned with the system's performance goals and adapts to changing conditions.

[0165] By creating a new policy using the reward and feeding the new state into the critic network, step 432 facilitates an iterative learning process. This process enables the system to continuously refine its actions and policies based on real-time feedback, leading to ongoing improvements in performance and efficiency.

[0166] Step 434 comprises calculating a new reward based on the updated state and action, and using this reward to further refine the actor and critic networks. This step is beneficial for closing the feedback loop in the reinforcement learning framework, ensuring that the system continuously learns and adapts to optimize its performance.

[0167] The new reward is calculated by evaluating the system's performance under the updated state and action generated in step 432. The reward function assesses various performance metrics, such as efficiency, quality, and stability, and assigns a numerical value that quantifies the overall benefit of the applied changes. This reward reflects how well the updated action has achieved the desired outcomes and provides a measure of the system's current performance.

[0168] Once the new reward is calculated, it is used to update both the actor and critic networks. The actor network uses the reward to adjust its internal parameters, such as weights and biases, through a process known as policy gradient optimization. This involves using algorithms like gradient ascent to maximize the expected reward, thereby improving the network's ability to propose actions that lead to higher rewards in future iterations. The critic network, on the other hand, uses the reward to refine its evaluation of the Q-value, ensuring that its assessments of the expected cumulative reward remain accurate and reliable.

[0169] By calculating the new reward and updating the actor and critic networks, step 434 ensures that the system continuously learns from its performance and adapts its actions to achieve optimal results. This iterative process of reward calculation and network updating drives the continuous improvement of the system, enabling real-time optimization and enhanced performance over time.

[0170] Preferably, after completing step 434, the method 420 loops back to step 422 to ensure continuous optimization and real-time adaptation of the system's performance. This looping process allows the system to iteratively refine its actions and policies based on the most current state and feedback, leading to ongoing improvements. Upon calculating the new reward and updating the actor and critic networks in step 434, the system is now equipped with refined policies and evaluations. These updates are based on the latest performance metrics and feedback, ensuring that the system's decision-making process is optimized for the current conditions. The method then transitions back to step 422 to begin a new iteration of the loop. In this new iteration, step 422 involves collecting the current state of the system once again. The sensors and dataacquisition devices capture real-time data reflecting the system's operational status and environmental conditions under the newly updated parameters. This data collection is crucial for providing an accurate and up-to-date snapshot of the system's state, which serves as the foundation for the subsequent steps.

[0171] The collected state data is then fed into the actor and critic networks in step 424. The actor network processes this data to propose a new set of actions, while the critic network evaluates these actions based on the updated state. This step ensures that the system's decision-making process is informed by the most current state information. In step 426, the actor network generates a suggested action based on the current state and provides it to the critic network for evaluation. The critic network assesses the quality of this action by estimating its expected cumulative reward, providing critical feedback for refining the actor network's proposals. Based on the critic network's evaluation, the process parameters are updated in step 428. These updated parameters are applied to the system, altering its operational settings to reflect the optimized values suggested by the actor network. This step ensures that the system adapts in real-time to the proposed actions.

[0172] In step 430, the system operates under the updated parameters, leading to a new state. This new state is captured using the one or more sensors, and the reward is calculated based on the system's performance under these conditions. The reward provides a measure of the effectiveness of the applied changes. The reward calculated in step 430 is used to create a new policy in step 432. The actor network's parameters are updated based on this reward, and the new state is fed into the critic network for further evaluation. This step ensures that the system continuously refines its policies based on real-time feedback. Finally, in step 434, a new reward is calculated based on the updated state and action. This reward is used to further refine the actor and critic networks, closing the feedback loop and preparing the system for the next iteration. By looping back to step 422 after completing step 434, method 420 ensures continuous optimization and real-time adaptation of the system's performance. This iterative process allows the system to learn from its performance, refine its actions, and achieve ongoing improvements in efficiency and effectiveness.

[0173] Method 420 represents an enhancement of method 400 by incorporating a RL framework to enable continuous optimization and real-time adaptation of the system's performance. While method 400 focuses on determining a set of AM factors based on target characteristics using a mathematical model, method 420 builds upon this foundation by integrating actor and critic networks to dynamically adjust process parameters. This integration allows method 420 to not only determine the optimal AM factors but also to iteratively refine these factors based on real- time feedback, leading to improved efficiency and precision in the manufacturing process.

[0174] In method 400, the process involves obtaining a set of first AM factors and characteristics, training a predictor using these data sets, generating a mathematical model, and using this model to determine a set of second AM factors for performing a second AM process. Method 420 adapts this approach by introducing steps that leverage reinforcement learning to continuously update and optimize the process parameters. Specifically, method 420 comprises steps such as feeding the state of the system to actor and critic networks, generating suggested actions, updating process parameters, and calculating rewards based on the system's performance. These additional steps enable the system to learn from its actions and make real-time adjustments, thereby enhancing the overall performance and adaptability of the manufacturing process.

[0175] A key benefit of method 420 lies in its ability to create a feedback loop that continuously refines the decision-making process. By incorporating the actor and critic networks, method 420 can evaluate the effectiveness of the applied changes and adjust the process parameters accordingly. This iterative learning process ensures that the system remains responsive to changing conditions and can optimize its performance in real-time. The use of reinforcement learning in method 420 allows for a more dynamic and adaptive approach compared to the static optimization provided by method 400. As a result, method 420 offers significant advantages in terms of efficiency, precision, and adaptability, making it a more robust solution for complex manufacturing environments.

[0176] Method 400 and method 420 can be seamlessly integrated to create a comprehensive approach for optimizing and adapting the manufacturing process in real-time. Initially, method 400 is employed to obtain a set of first AM factors (step 402) and first characteristics (step 404). These data sets are used to train a predictor (step 406), which is integrated into the actor network of the RL framework in method 420. The predictor, now part of the actor network, processes the state of the system, which comprises both the AM factors and characteristics, to generate suggested actions.

[0177] In the integrated method, the state of the system is collected in step 422, capturing real- time data on the current AM factors and characteristics using various sensors and data acquisition devices. This state data is then fed into the actor and critic networks in step 424. The actor network, which includes the predictor from method 400, processes the state data to propose a new set of actions aimed at optimizing the system's performance. The critic network evaluates these actions by estimating their expected cumulative reward, providing feedback to the actor network.

[0178] The suggested action generated by the actor network in step 426 is then evaluated by the critic network, which produces a Q-value reflecting the expected reward. Based on this evaluation,the process parameters are updated in step 428. This step incorporates a dual mechanism: if deviations are within a predetermined threshold, the system relies on the mathematical model for systematic, long-term improvements based on historical data analysis. If deviations exceed the threshold, the RL-based control system is triggered to make immediate, real-time adjustments. The system then generates a new state and calculates a reward in step 430, considering both immediate performance metrics and long-term trends identified by the mathematical model.

[0179] The reward calculated in step 430 is used to create a new policy in step 432. The actor network's parameters are updated based on this reward, and the new state is fed into the critic network for further evaluation. This step ensures that the system continuously refines its policies based on both real-time feedback and historical data analysis. Finally, in step 434, a new reward is calculated based on the updated state and action, and both the actor and critic networks are refined using this reward. This closes the feedback loop and prepares the system for the next iteration. By looping back to step 422, the integrated method ensures continuous optimization and real-time adaptation of the system's performance. The optional generation of a mathematical model in step 408 of method 400 provides an additional layer of explainability to the RL framework, allowing for a clear understanding of the relationship between AM factors and characteristics. This combined approach leverages the initial parameter determination and training from method 400, while the RL framework in method 420 ensures ongoing real-time improvements, resulting in a highly efficient and adaptive manufacturing process.

[0180] Method 420 offers several significant advantages over traditional supervised machine learning models, such as neural networks, for predicting future states and correcting the system accordingly. Traditional supervised machine learning models are typically trained on static datasets and do not adapt in real-time. Once trained, these models can predict future states based on historical data but may not effectively respond to new, unforeseen changes in the system or environment. Method 420 continuously learns and adapts in real-time by interacting with the environment. The actor and critic networks iteratively refine their actions based on real-time feedback, allowing the system to dynamically adjust to changing conditions and optimize performance continuously. Supervised models can predict future states and suggest corrections, but they often require retraining with new data to adapt to significant changes. This retraining process can be time-consuming and may not be feasible for real-time applications. Method 420 enables real-time optimization by continuously updating process parameters based on the latest state and reward feedback. This ensures that the system remains responsive and can make immediate adjustments to maintain optimal performance. Supervised models rely on historical data and may not explore new actions beyond the training dataset. This limitation can result in suboptimal performance if the model encounters scenarios not represented in the training data.Method 420 balances exploration (trying new actions) and exploitation (using known actions that yield high rewards) to discover the most effective strategies for optimizing the system. This balance allows the system to explore a wider range of actions and find better solutions over time. Supervised models may struggle with complex and dynamic environments, especially if the training data does not capture all possible variations. These models may require extensive feature engineering and large datasets to achieve comparable performance. Method 420 is well-suited for complex and dynamic environments where the relationships between inputs and outputs are not fully understood or are constantly changing. The RL framework can adapt to these complexities by continuously learning from the environment. Supervised models do not inherently include a feedback loop for policy improvement. While they can be updated with new data, this process is not as seamless or continuous as the RL framework's iterative learning. Method 420 incorporates a feedback loop where the reward signal guides the improvement of the actor and critic networks. This iterative process ensures that the system's policies are continuously refined based on real- time performance metrics.

[0181] In summary, method 420 offers significant benefits over traditional supervised machine learning models by enabling continuous learning and adaptation, real-time optimization, effective exploration and exploitation, handling complex and dynamic environments, and incorporating a feedback loop for ongoing policy improvement. These advantages make method 420 a more robust and adaptive solution for optimizing manufacturing processes and other applications requiring real-time decision-making and performance optimization.

[0182] Additionally, the dual mechanism of combining a mathematical model with a RL framework offers substantial efficiency savings and enhanced system performance by leveraging the strengths of both approaches. The mathematical model provides a systematic, long-term perspective based on historical data analysis, ensuring that the system maintains accuracy and stability over time. This model is particularly beneficial for making informed decisions about process parameters when deviations are within predetermined limits, as it can predict outcomes based on well-understood relationships between AM factors and characteristics. By periodically updating the mathematical model with accumulated process data, the system can continuously improve its baseline predictions, leading to more efficient and reliable operations. This approach reduces the need for constant real-time adjustments, thereby saving computational resources and minimizing the risk of overfitting to transient variations.

[0183] On the other hand, the RL-based control system excels in handling immediate, real-time adjustments when deviations exceed the predefined thresholds. This capability is crucial for responding to unexpected process variations and maintaining optimal performance under dynamic conditions. The RL framework's ability to explore and exploit different actions allows thesystem to adapt quickly to new scenarios, ensuring that it can optimize performance even in the face of unforeseen changes. By continuously monitoring key performance indicators and switching to RL control when necessary, the system can achieve a balance between long-term stability and real-time adaptability. This dual mechanism not only enhances the overall efficiency and effectiveness of the manufacturing process but also provides a layer of explainability through the mathematical model. Operators can understand the underlying relationships and trust the system's decisions, while the RL framework ensures that the system remains responsive and agile. This combination results in a robust, adaptive, and efficient solution that maximizes performance while maintaining system stability and explainability.

[0184] Example thresholds optionally include a temperature deviation threshold of ±2°C from the target value, a pressure deviation threshold of ±5% from the optimal operating range, or a material deposition rate deviation threshold of ±3% from the desired rate. If the system detects that the temperature remains within ±2°C of the target, the mathematical model can be used to make systematic, long-term adjustments based on historical data. However, if the temperature deviation exceeds ±2°C, the RL-based control system is triggered to make immediate, real-time adjustments to bring the temperature back within the acceptable range. Similarly, if the pressure or material deposition rate deviations exceed their respective thresholds of ±5% and ±3%, the RL framework takes over to ensure rapid correction. These example thresholds ensure that the system can maintain optimal performance by leveraging the strengths of both the mathematical model for stability and the RL framework for adaptability, providing a balanced and efficient approach to process control.

[0185] Figure 5 is directed to a flow-chart illustrating a method for updating an automated AM process. Specifically, Figure 5 is directed towards using the mathematical model of method 400 to generate new / modified AM factors, such that an AM process can be updated, resulting in an improved AM output. Preferably, method 500 is used to update AM factors during an AM process, such that the AM process in monitored and updated to increase the quality of the AM output.

[0186] Any of the steps of method 500 can be performed outside of the order shown in Figure 5, and not all steps of method 500 shown are required in order to perform method 500. Additionally, any of the steps of method 400 of Figure 4A, method 420 of Figure 4B, or method 600 of Figure 6 may be performed in addition or instead of the steps of method 500 listed here.

[0187] As discussed above, differences between the second AM output and the target AM output can be detected. AM processes are highly complex, and minute changes in material properties, environmental factors, and other conditions can result in alternations to generated AM outputs. Therefore, it is important to use the mathematical model to compensate for such changes bydetermining new AM factors. By using method 500 throughout the AM process, the end-product (the final AM output) will be high quality without necessarily requiring manual intervention.

[0188] Method 500 comprises step 502, step 504, step 506, and step 508. Step 502 comprises obtaining a set of second characteristics; step 504 comprises determining a difference between a set of target characteristics and the set of second characteristics; step 506 comprises using a mathematical model to generate a set of third AM factors; and step 508 comprises using the third set of AM factors to perform a third AM process.

[0189] Step 502 comprises obtaining, such as from the sensor, a set of second characteristics. The set of second characteristics comprises one or more geometric or physical properties of a second AM output, such as the second AM output generated in step 414 of method 400 above.

[0190] Returning to the aforementioned welding process example, the sensor determines the average height, width, and cross-section area of weld beads of the second AM output.

[0191] Step 504 comprises determining a difference between a set of target characteristics, such as the set of target characteristics obtained in step 410 of method 400 above, and the set of second characteristics. For example, the difference is a collated difference of each corresponding geometric or physical property of the set of target characteristics and the set of second characteristics.

[0192] Returning to the above welding process example, the difference comprises: a difference in the average height of the second AM output weld beads and the average height of the target AM output weld beads; a difference in the average width of the second AM output weld beads and the average width of the target AM output weld beads; and a difference in the average cross-section area of the second AM output weld beads and the average cross-section area of the target AM output weld beads.

[0193] Alternatively, step 504 comprises determining if deviations are present in the set of second characteristics, such as by use of the mathematical model, the predictor, a processor, or the sensor. Deviation detection, also known as anomaly detection, can be achieved without comparison to the set of target characteristics. For example, the sensor can be configured to determine whether a geometric or physical property exceeds a predefined range. In another example, the predictor comprises a deviation detection algorithm, such as an autoencoder.

[0194] Step 506 comprises using a mathematical model, such as the mathematical model generated in step 408 of method 400, to modify one or more AM process inputs of a set of second AM factors, such as the set of second AM factors identified in step 412 of method 400. Modificationis based on the difference determined in step 504 above, and results in the generation of a set of third AM factors. Therefore, the set of third AM factors is a modified set of second AM factors.

[0195] Returning to the above welding process example, the set of second AM factors are modified using the mathematical model to, e.g., reduce the welding speed, and hence the third AM factors comprise a reduced welding speed in respect to the second AM factors and a maintained welding current and wire feed rate. The welding speed was reduced as the mathematical model showed the determined difference could be compensated by (in this example) adjusting the welding speed by a set quantity equal to the reduction.

[0196] Step 508 comprises using the third AM factors to perform a third AM process. For example, the AM apparatus performs an AM process using the AM process inputs of the third AM factors. The third AM process results in the generation of a third AM output. Due to the modifications made using the mathematical model in step 506, any difference between the target AM output and the third AM output are decreased in comparison to the difference determined in step 504, such that the third AM output is closer to the target AM output in relation to associated geometric or physical properties.

[0197] Returning to the above welding speed example, the reduction in welding speed resulted an average height, width, and cross-section area of welding beads for the third AM output within a quantity threshold associated with the set of target characteristics, and hence the third AM output is considered high quality.

[0198] Method 500 can be repeated such that AM process inputs are continually modified and improved to result in high quality AM outputs. For example, step 502 is repeated such that a set of third characteristics are obtained comprising one or more geometric or physical properties of the third AM output (generated as a result of performing the third AM process in step 508). Step 504 is repeated such that a difference between the set of third characteristics and the set of target characteristics is determined. Step 506 is optionally repeated as, for example, step 506 may not be required if the difference is within a quality threshold. If step 506 is repeated, the mathematical model is used to generate a set of fourth AM factors, which can then be used in step 508 to perform a fourth AM process.

[0199] Preferably, the first, second, third, or fourth AM processes are one or more continuous AM processes. For example, the first, second, third and fourth AM processes are a single AM process at multiple subsequent time points, where: the first AM process is the start of the single AM process used to train the predictor and mathematical model such that it is tailored for the specific single AM process; the second AM process is a continuation of the first AM process and represents a correction and / modification of the first AM process based improved AM factors obtained usingthe mathematical model; the third AM process is a correction and / or modification of the second AM process based on improved AM factors obtained using the mathematical model (e.g. using method 500); and the fourth AM process is a correction and / or modification of the third AM process to result in even higher quality outputs. Therefore, the mathematical model can be used to continually improve and update the AM process.

[0200] Figure 6 shows a flow-chart illustrating a method for updating a mathematical model and outputting a plot. Specifically, Figure 6 is directed towards multiple interchangeable methods for training the predictor, updating the mathematical model, and outputting a plot based on the relationship(s) defined by the mathematical model.

[0201] Any of the steps of method 600 can be performed outside of the order shown in Figure 6, and not all steps of method 600 shown are required in order to perform method 600. Each step of method 600 may be performed individually. Additionally, any of the steps of method 400 of Figure 4A, method 420 of Figure 4B, or method 500 of Figure 5 may be performed in addition or instead of the steps of method 600 listed here.

[0202] Training the predictor using additional data can help improve the accuracy of characteristics generated by the machine learning model. Data associated with a particular AM apparatus and / or AM process tailors the predictor to more accurately generate characteristics of AM outputs generated by the particular AM apparatus / process. Data associated with a different AM apparatus and / or AM process generalises the predictor, so generated characteristics from additional AM apparatus / processes are more accurate.

[0203] The mathematical model of the present invention maintains accuracy through periodic updates based on accumulated process data. When process parameters require adjustment, the system employs two distinct but complementary mechanisms: (1) direct mathematical model updates for systematic, long-term improvements based on historical data analysis, and (2) RL- based control for immediate, real-time adjustments responding to process variations. The transition between these mechanisms is governed by one or more defined performance thresholds, wherein deviations beyond predetermined limits trigger the RL control system while maintaining the mathematical model's baseline predictions. The system continuously monitors key performance indicators to determine the appropriate control mechanism, ensuring optimal process outcomes while maintaining system stability.

[0204] However, if differences between generated characteristics and target characteristics are substantial, e.g., a difference determined in step 504 of method 500 is above a predefined threshold, updating the predictor may be required. For example, there may be a change in AM process and / or AM apparatus that cannot be accurately compensated for by updating themathematical model directly. After the predictor is trained / re-trained / updated, the one or more predictor parameters extracted in generating the mathematical change. Therefore, the mathematical model can be updated based on extracting one or more updated parameters from the retrained predictor.

[0205] Method 600 comprises step 602, step 604, step 606, and step 608, where any of the steps can be performed individually or in combination. Step 602 comprises training a predictor using a second set of AM factors and a set of second AM characteristics; step 604 comprises changing the hyperparameters within the predictor until residuals fall below a predefined threshold; step 606 comprises updating a mathematical model based on extracting one or more updated parameters from the predictor; and step 608 comprises outputting, based on a prediction function of the mathematical model, a plot of a geometric or physical property against an AM process input.

[0206] Step 602 comprises training a predictor using a set of second AM factors and a set of second characteristics, such as the predictor and sets of method 600. For example, the predictor is initially trained on the set of first AM factors and the set of first characteristics. Training the predictor further on the set of second AM factors and the set of second characteristics then improves the predictor, as additional training data results in more accurately generated characteristics.

[0207] Returning to the aforementioned welding process example, the predictor comprises an ANN trained on first AM factors and corresponding first characteristics associated with the first AM output generated using the first AM factors. The second characteristics are associated with the second AM output generated using the second AM factors, where the same AM apparatus is used. Therefore, the re-trained ANN is tailored to provide more accurate characteristics for the specific AM apparatus.

[0208] Step 604 comprises changing hyperparameters within the predictor until residuals fall below a predefined threshold. This step may be performed in addition to step 602, in addition to alternative steps of method 400.

[0209] Returning to the above welding process example, the ANN is trained for 10 epochs using the set of second characteristics and the set of second AM factors. The difference between characteristics generated from the predictor and the set of second characteristics is above the predefined threshold, so the network architecture and / or other parameters will change. The altered ANN comprising a different architecture and set of hyperparameters is then further trained for additional epochs using the same data. The difference between characteristics generated from the predictor and the set of second characteristics is then below the predefined threshold, and the ANN is accurate enough to not require additional modifications. Beneficially,this method of changing the network architecture based on residuals allows for preservation of computational resources, as the ANN is as large as is necessary to generate outputs with the required accuracy.

[0210] If increasing the number of nodes is not relevant, such as in a scenario where the predictor does not comprise a relevant machine learning model relative such as a neural network, alternative hyperparameters can be increased in the place of nodes. Additionally, residuals (the difference between generated and target outputs from the machine learning model) may not be readily available, such as when using unsupervised algorithms, and alternative threshold-based accuracy determination can be used in place of residuals falling below a predefined threshold.

[0211] Step 606 comprises updating a mathematical model based on extracting one or more updated parameters from the predictor. The one or more updated parameters are associated with internal parameters of the predictor after training / re-training / updating. For example, step 606 can follow step 602, 604, in addition or in combination of steps of method 400.

[0212] Returning to the above welding process example, the ANN is retrained and modified with an additional node, and hence the internal parameters have changed. The extracted one or more updated parameters are then used to update e.g., the coefficients of the prediction function, and hence the mathematical model is updated. Beneficially, updating the mathematical model based on an updated predictor allows for increased accuracy in determining AM factors based on targeted characteristics without impacting the increased explainability and efficiency from using a mathematical model.

[0213] Step 608 comprises outputting, based on the prediction function, a plot of a geometric or physical property of the set of first characteristics against an AM process input of the set of first AM factors. For example, step 608 comprises outputting plot 300-A of Figure 3, described below.

[0214] Returning to the above welding process example, step 608 comprises outputting a plot of e.g., height against e.g., welding speed based on the relationship between heigh, and welding speed represented by the prediction function.

[0215] Alternatively, a plurality of geometric or physical properties is plotted against a plurality of AM process inputs, whether individually or in combination, such as shown by plot 300-B of Figure 3. If step 608 is performed after step 606, step 608 may instead comprise outputting, based on the updated prediction function, a plot of a geometric or physical property of either the first characteristics or second characteristics against an AM process input of the first AM factors or second AM factors.

[0216] Figure 7 shows an example computing system. Specifically, Figure 7 shows a block diagram of an embodiment of a computing system 700 according to example embodiments of the present disclosure.

[0217] Computing system 700 can be configured to perform any of the operations disclosed herein such as, for example, any of the operations discussed with reference to method 400 of Figure 4. Computing system 700 includes one or more computing devices 702. Computing device 702 of computing system 700 comprises one or more processors 704 and memory 706. device(s) 702 of computing system 700 comprise one or more processors 704 and memory 706. One or more processors 704 can be any general-purpose processor(s) configured to execute a set of instructions. For example, one or more processors 704 can be one or more general-purpose processors, one or more field programmable gate array (FPGA), and / or one or more application specific integrated circuits (ASIC). In one embodiment, one or more processors 704 include one processor. Alternatively, one or more processors 704 include a plurality of processors that are operatively connected. One or more processors 704 are communicatively coupled to memory 706 via address bus 708, control bus 710, and data bus 712. Memory 706 can be a random-access memory (RAM), a read-only memory (ROM), a persistent storage device such as a hard drive, an erasable programmable read-only memory (EPROM), and / or the like. Computing device(s) 702 further comprise input / output (I / O) interface 714 communicatively coupled to address bus 708, control bus 710, and data bus 712.

[0218] Memory 706 can store information that can be accessed by one or more processors 704. For instance, memory 706 (e g., one or more non-transitory computer-readable storage mediums, memory devices) can include computer-readable instructions (not shown) that can be executed by one or more processors 704. The computer-readable instructions can be software written in any suitable programming language or can be implemented in hardware. Additionally, or alternatively, the computer-readable instructions can be executed in logically and / or virtually separate threads on one or more processors 704. For example, memory 706 can store instructions (not shown) that when executed by one or more processors 704 cause one or more processors 704 to perform operations such as any of the operations and functions for which computing system 700 is configured, as described herein. In addition, or alternatively, memory 706 can store data (not shown) that can be obtained, received, accessed, written, manipulated, created, and / or stored. The data can include, for instance, the data and / or information described herein in relation to Figures 1 to 6. In some implementations, computing device(s) 702 can obtain from and / or store data in one or more memory device(s) that are remote from the computing system 700.

[0219] Computing system 700 further comprises storage unit 716, network interface 718, input controller 720, and output controller 722. Storage unit 716, network interface 718, inputcontroller 720, and output controller 722 are communicatively coupled to central control unit or computing devices 702 via I / O interface 714.

[0220] Storage unit 716 is a computer readable medium, preferably a non-transitory computer readable medium, comprising one or more programs, the one or more programs comprising instructions which when executed by one or more processors 704 cause computing system 700 to perform the method steps of the present disclosure. Alternatively, storage unit 716 is a transitory computer readable medium. Storage unit 716 can be a persistent storage device such as a hard drive, a cloud storage device, or any other appropriate storage device.

[0221] Network interface 718 can be a Wi-Fi module, a network interface card, a Bluetooth module, and / or any other suitable wired or wireless communication device, e.g. a module configured for 5G, 6G or future communication and data transfer technology. In an embodiment, network interface 718 is configured to connect to a network such as a local area network (LAN), or a wide area network (WAN), the Internet, or an intranet.

[0222] Figure 7 illustrates one example computer system 700 that can be used to implement the present disclosure. Other computing systems can be used as well. Computing tasks discussed herein as being performed at and / or by one or more functional unit(s) (e.g., as described in relation to Figure 2) can instead be performed remote from the respective system, or vice versa. Such configurations can be implemented without deviating from the scope of the present disclosure. The use of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. Computer-implemented operations can be performed on a single component or across multiple components. Computer-implemented tasks and / or operations can be performed sequentially or in parallel. Data and instructions can be stored in a single memory device or across multiple memory devices.

[0223] Regarding the above disclosure, references to items in the singular should be understood to include items in the plural, and vice versa, unless explicitly stated otherwise or clear from the context. Additionally, grammatical conjunctions are intended to express any and all disjunctive and conjunctive combinations of conjoined clauses, sentences, words, and the like, unless otherwise stated or clear from the context. Thus, the term “or” should generally be understood to mean “and / or” and so forth. The use of any and all examples, or exemplary language (“e.g.,” “such as,” “including,” or the like) provided herein, is intended merely to better illuminate the embodiments and does not pose a limitation on the scope of the embodiments or the claims.

[0224] Methods described herein may relate to a computer storage product with a non-transitory computer-readable medium (also can be referred to as a non-transitory processor-readablemedium) having instructions or computer code thereon for performing various computer- implemented operations, such that the methods are performed. The computer-readable medium (or processor-readable medium) is non-transitory in the sense that it does not include transitory propagating signals per se (e.g., a propagating electromagnetic wave carrying information on a transmission medium such as space or a cable). The media and computer code (also can be referred to as code) may be those designed and constructed for the specific purpose or purposes. Examples of non-transitory computer-readable media include, but are not limited to, magnetic storage media such as hard disks, floppy disks, and magnetic tape, optical storage media such as Compact Disc / Digital Video Discs (CD / DVDs), Compact Disc-Read Only Memories (CD-ROMs), and holographic devices; magneto-optical storage media such as optical disks; carrier wave signal processing modules; and hardware devices that are specially configured to store and execute program code, such as Application-Specific Integrated Circuits (ASICs), Programmable Logic Devices (PLDs), Read-Only Memory (ROM) and Random-Access Memory (RAM) devices. Other embodiments described herein relate to a transitory computer program product, which can include, for example, the instructions and / or computer code discussed herein.

[0225] Some embodiments and / or methods described herein can be performed by software (executed on hardware), hardware, or a combination thereof. Hardware modules include, for example, a general-purpose processor, a field programmable gate array (FPGA), and / or an application specific integrated circuit (ASIC). Software modules (executed on hardware) can be expressed in a variety of software languages (e.g., computer code), including C, C++, Java, Ruby, Visual Basic, Python, and / or other object-oriented, procedural, or other programming language and development tools. Examples of computer code include, but are not limited to, micro-code or micro-instructions, machine instructions, such as produced by a compiler, code used to produce a web service, and files containing higher-level instructions that are executed by a computer using an interpreter. For example, embodiments can be implemented using imperative programming languages (e.g., C, Fortran, etc.), functional programming languages (Haskell, Erlang, etc.), logical programming languages (e.g., Prolog), object-oriented programming languages (e.g., Java, C++, etc.) or other suitable programming languages and / or development tools. Additional examples of computer code include, but are not limited to, control signals, encrypted code, and compressed code. Optionally, the embodiments and / or methods described herein are implemented using an operating system such as Robot Operating System (ROS).

Claims

CLAIMS 1. A method for performing an automated additive manufacturing (AM) process, the method comprising: obtaining a current state comprising: a set of first AM factors, wherein the set of first AM factors comprises one or more AM process inputs associated with performing a first AM process; a set of first characteristics comprising one or more geometric or physical properties of a first AM output, wherein the first AM output is generated by performing the first AM process using the set of first AM factors; providing the current state to a reinforcement learning (RL) framework configured to learn a policy to adjust welding parameters in real-time based on the current state and reward feedback, wherein the RL framework comprises: an actor network for generating, based on the current state, a suggested action associated with obtaining a target AM output, wherein the actor network comprises a trained machine learning model; a critic network for performing an evaluation of the suggested action by calculating an estimated reward based on an expected cumulative benefit of the suggested action in optimizing the first AM process; wherein safety limits are enforced using a loss function on that penalizes exceeding predetermined limits of the one or more AM process inputs; updating the set of first AM factors, thereby generating a set of second AM factors, based on the evaluation performed by the critic network; and using the set of second AM factors to perform the second AM process, thereby generating a second AM output.

2. The method of claim 1, further comprising providing an updated state to the RL framework and modifying the set of second AM factors, thereby generating a set of third AM factors, wherein: the updated state comprises the set of second AM factors and a set of second AM characteristics comprising one or more geometric or physical properties of the second AM output; a reward is calculated based on the cumulative benefit of the second AM process, wherein the cumulative benefit is associated with minimising a difference between a set of target characteristics associated with the target AM output and the set of second characteristics; a new policy is determined using the reward, wherein the new policy the reward and are used to refine the actor network and the critic network.

3. The method of claim 1, further comprising, if the difference between the set of target characteristics and the set of second characteristics is within a predetermined threshold, using a mathematical model to modify the set of second AM factors, wherein the mathematical model is associated with a relationship between the one or more AM process inputs and the one or more geometric or physical properties, thereby generating a set of third AM factors.

4. The method of claim 2 or claim 3, further comprising using the set of third AM factors to perform a third AM process.

5. The method of any preceding claim, wherein the RL framework comprises a plurality of artificial neural networks.

6. The method of any preceding claim, further comprising: selecting a plurality of RL algorithms suitable for continuous action spaces, and conducting preliminary evaluations to determine the best-performing model architecture for AM process control.

7. The method of any preceding claim, wherein obtaining a current state comprises collecting data from a sensor, a machine or system for performing the first AM process, or other data source.

8. The method of any preceding claim, wherein obtaining the current state comprises preprocessing data through design of experiments (DoE), comprising any of outlier removal, normalization, and feature engineering.

9. The method of any preceding claim, wherein the AM process comprises depositing a plurality of weld beads, and wherein one or more geometric or physical properties comprises any of height, width, cross-section area, hardness or residual stresses for at least one of the plurality of weld beads.

10. The method of any preceding claim, further comprising: conducting offline training of the predictor using historical or simulated data, and training the predictor to adjust AM factors with simulated rewards, optimizing for bead geometry, process efficiency, or safety.

11. The method of any preceding claim, wherein the target AM output has a predefined geometry and is associated with an optimised second AM output.

12. The method of any preceding claim, further comprising validating the RL framework in a simulated environment using evaluation metrics, such as cumulative rewards and Q-value stability, and assessing adaptability across various processing scenarios and materials.

13. The method of any preceding claim, further comprising implementing the RL framework on edge hardware to control welding parameters in at least near real-time, and using one or more integrated sensors comprising the sensor to initiate a feedback loop for real-time or near real- time adjustments based on continuous AM process conditions.

14. A system comprising: apparatus configured to perform one or more additive manufacturing (AM) processes; a sensor configured to determine one or more geometric or physical properties of an output of the one or more AM processes; a GPU-enabled computer configured to obtain data from the apparatus and the sensor; and a memory storing instructions that, when executed, cause the system to perform the method of any of claims 1-13.

15. A computer-readable medium storing instructions for performing the method of any of claims 1-13.

Citation Information

Patent Citations

  • Method and system for monitoring additive manufacturing processes

    US20180264553A1

  • Machine learning device, additive manufacturing system, machine learning method for welding condition, method for determining welding condition, and a non-transitory computer readable medium

    US20230259099A1

Cited By

  • Processing parameter intelligent adaptation method and system for multi-layer false tooth

    CN120686726A

  • Ship steel plate digital cutting and typesetting optimization method and system

    CN120909215A

  • Friction plate formula design method and system based on deep reinforcement learning

    CN121093761A

  • Task processing method and device, electronic equipment and computer storage medium

    CN121234198A

  • Spring steel wire drawing control method based on reinforcement learning

    CN121277098A