Machine learning framework incorporating thermodynamic models for chemical process simulation

WO2026177756A2PCT designated stage Publication Date: 2026-08-27SIEMENS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/036646
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-14
Filing Date
2025-07-07
Publication Date
2026-08-27

Smart Images

  • Figure US2025036646_27082026_PF_FP_ABST
    Figure US2025036646_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for chemical process simulation integrates thermodynamic models into machine learning frameworks to enhance predictive accuracy and optimize chemical separation processes. The system employs a variety of methods, such as using a physics- informed neural network (PINN)-like approach and methods that embed thermodynamic equations directly into the architecture, ensuring adherence to physical laws such as conservation of energy and mass. Additionally, experimental data, virtual sampling points derived from thermodynamic models, and auxiliary predictions capturing interdependencies between physical and chemical variables are combined into an augmented dataset for training. A hybrid loss function balances contributions from experimental data and thermodynamic insights, dynamically adjusted via a scaling hyperparameter. The system predicts target properties, such as adsorption efficiency, under diverse operating conditions while maintaining compliance with thermodynamic constraints.
Need to check novelty before this filing date? Find Prior Art

Description

202418740MACHINE LEARNING FRAMEWORK INCORPORATING THERMODYNAMIC MODELS FOR CHEMICAL PROCESS SIMULATIONSTATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENTThis invention was made with Government support under U.S. Department of Energy (DOE) under Award No.: DE-EE0009768.PRIORITY CLAIMThis application claims priority to U.S. Provisional Patent application Serial No. 63 / 706,822, filed October 14, 2024, the disclosure of which is incorporated by reference herein in its entirety.BACKGROUND

[0001] The present disclosure pertains to the field of machine learning and chemical engineering, specifically addressing the integration of thermodynamic models within machine learning frameworks to enhance the optimization of chemical processes.

[0002] Chemical process modeling and simulation play a significant role in the development, design, and optimization of a wide range of industrial operations. Over the years, computational techniques have enhanced the ability to predict and control complex separation processes that are essential to industries involved in chemical, pharmaceutical, and energy production. Historically, simulation efforts have combined experimental measurements with theoretical formulations based on well-established physical principles. However, accurately representing intricate chemical interactions while managing the limitations imposed by data availability continues to challenge conventional approaches. The evolving landscape of industrial processing highlights the importance of establishing robust simulation frameworks that integrate diverse sources of information, thereby providing deeper insights into process behavior and performance.

[0003] In pursuit of improved operational efficiency and process reliability, there is a strong impetus toward achieving more precise and streamlined simulation tools. Enhanced prediction capabilities contribute to better resource management, energy conservation, and optimization of process configurations. Achieving these objectives ensures that process202418740designs are both economically viable and aligned with safety and regulatory standards. The drive will further aid in the effective design, scale-up, and control of equipment, reducing production costs and minimizing environmental impacts. As industrial systems become more sophisticated, the integration of multiple analytical perspectives into simulation methodologies becomes increasingly important.

[0004] Many established simulation methods rely heavily on either extensive collections of experimental data or isolated theoretical constructs. When dealing with systems characterized by limited or noisy datasets, these techniques can struggle to capture the dynamic interactions fundamental to chemical processes. Moreover, conventional modeling approaches may fall short when attempting to consistently reflect the balance of significant physical parameters under varying operational conditions. Such shortcomings can lead to discrepancies between modeled outcomes and actual process performance, posing challenges to achieving reliable forecasts and efficient process management. This general limitation has prompted ongoing efforts to explore more comprehensive frameworks that can better negotiate the trade-offs between data scarcity and complex process behavior.

[0005] Particularly in highly specialized separation processes, accurately modeling system dynamics becomes a significant challenge. Specific issues arise when attempting to reconcile foundational principles with variable operating conditions and limited quantitative information about the substrates in a mixture to separate. This is especially relevant in scenarios where maintaining adherence to physical constraints, such as energy and mass balances, is necessary throughout the simulation of chemical interactions. The challenges associated with representing these factors highlight the need for more adaptable simulation methodologies. Addressing these issues is necessary for the effective design and operation of advanced processing systems, ensuring that emerging computational tools can meet the stringent demands of modern industrial applications while reducing resource waste and improving overall efficiency.SUMMARY

[0006] In one embodiment, the disclosure provides a method for predicting a target property associated with a chemical separation process. The method comprises obtaining experimental data corresponding to a plurality of operating conditions for the chemical separation process; generating a plurality of virtual sampling points by applying a202418740thermodynamic model to the experimental data, the thermodynamic model being configured to derive the virtual sampling points in accordance with thermodynamic principles; incorporating auxiliary predictions related to the target property into a training dataset, the auxiliary predictions capturing interdependencies between physical and chemical variables; training a machine learning model on an augmented dataset comprising the experimental data, the virtual sampling points, and the auxiliary predictions, the training including utilization of a hybrid loss function that combines a loss based on the experimental data, a loss based on the virtual sampling points, and a scaling hyperparameter to balance the losses; and predicting the target property using the trained machine learning model, the prediction adhering to predetermined thermodynamic constraints.

[0007] In addition to one or more of the features described herein, generating the plurality of virtual sampling points further comprises applying thermodynamic equations to the experimental data and deriving material-specific parameters.

[0008] In addition to one or more of the features described herein, incorporating auxiliary predictions comprises generating predictions of one or more additional chemical process variables and correlating the additional chemical process variables with the target property.

[0009] In addition to one or more of the features described herein, training the machine learning model further comprises employing a neural network architecture and configuring the neural network architecture to integrate physical laws and thermodynamic models.

[0010] In addition to one or more of the features described herein, the hybrid loss function comprises a weighted sum of a loss based on the experimental data, a loss based on the virtual sampling points, and a weighting determined by the scaling hyperparameter.

[0011] In addition to one or more of the features described herein, predicting the target property further comprises enforcing predetermined thermodynamic constraints, including conservation of energy and conservation of mass.

[0012] In addition to one or more of the features described herein, the augmented dataset further comprises a discrepancy model and capturing differences between the experimental data and the virtual sampling points.202418740

[0013] In addition to one or more of the features described herein, the scaling hyperparameter is dynamically adjusted during training based on a monitored performance metric of the machine learning model.

[0014] In addition to one or more of the features described herein, the virtual sampling points are generated to represent operating conditions extending beyond those included in the experimental data.

[0015] In addition to one or more of the features described herein, the machine learning model is implemented as a physics-informed neural network, with thermodynamic equations integrated directly into an architecture of the neural network.

[0016] In another embodiment, the disclosure provides a method for predicting a target property associated with a chemical separation process. The method comprises obtaining experimental data corresponding to a plurality of operating conditions for the chemical separation process; constructing a physics-informed neural network that integrates thermodynamic principles by embedding thermodynamic equations directly into an architecture of the neural network; training the physics-informed neural network on the experimental data, the training including utilization of a hybrid loss function that combines a loss based on the experimental data, a loss derived from thermodynamic principles, and a scaling hyperparameter to balance the losses; and predicting the target property using the trained physics-informed neural network, the prediction enforcing predetermined thermodynamic constraints including conservation of mass and energy.

[0017] In addition to one or more of the features described herein, constructing the physics-informed neural network further comprises integrating thermodynamic equations into the network and replacing at least one activation function with a thermodynamically based function.

[0018] In addition to one or more of the features described herein, training the physics-informed neural network further comprises dynamically adjusting the scaling hyperparameter and basing the adjustment on a monitored performance metric of the physics-informed neural network.202418740

[0019] In addition to one or more of the features described herein, the experimental data corresponds to operating conditions that include variations in temperature, pressure, and pH for the chemical separation process.

[0020] In addition to one or more of the features described herein, the loss derived from thermodynamic principles comprises a loss calculated based on deviations from conservation of energy and conservation of mass.

[0021] In addition to one or more of the features described herein, the physics-informed neural network is further configured to incorporate auxiliary predictions related to one or more additional chemical process variables.

[0022] In addition to one or more of the features described herein, training the physics-informed neural network further comprises utilizing a discrepancy model and capturing differences between the experimental data and predictions provided by the embedded thermodynamic equations.

[0023] In addition to one or more of the features described herein, constructing the physics-informed neural network further comprises calibrating material-specific parameters using a subset of the experimental data and employing the calibrated parameters within the embedded thermodynamic equations.

[0024] In addition to one or more of the features described herein, the hybrid loss function further comprises a weighted sum of the loss based on the experimental data, the loss derived from thermodynamic principles, and the weighting determined by the scaling hyperparameter.

[0025] In addition to one or more of the features described herein, predicting the target property further comprises filtering the predicted values to reduce noise originating from variations in the experimental data and enhancing prediction accuracy.

[0026] Additional technical features and benefits are realized through the techniques of the present disclosure. Embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed subject matter. For a better understanding, refer to the detailed description and to the drawings.BRIEF DESCRIPTION OF THE DRAWINGS202418740

[0027] FIG. 1 illustrates a block diagram of a processing system in accordance with an embodiment.

[0028] FIG. 2 illustrates a schematic block diagram of a chemical process simulation system integrating thermodynamic models and machine learning frameworks in accordance with an embodiment.

[0029] FIG. 3 illustrates a block diagram of a physics-informed neural network model integrating thermodynamics-informed loss into a neural network framework in accordance with an embodiment.

[0030] FIG. 4 illustrates a schematic block diagram of an integrated model combining thermodynamic equations and neural network architecture for chemical separation predictions in accordance with an embodiment.

[0031] FIG. 5 illustrates a schematic block diagram of a neural network integrating thermodynamics-based descriptors and relationships for predicting target and related thermodynamical properties in accordance with an embodiment.

[0032] FIG. 6 illustrates a schematic block diagram of a discrepancy modeling system integrating thermodynamic models and neural networks for chemical process simulation in accordance with an embodiment.

[0033] FIG. 7 illustrates a schematic flow chart diagram of a method for predicting a target property in a chemical separation process using thermodynamic-informed machine learning in accordance with an embodiment.

[0034] FIG. 8 illustrates a schematic flow chart diagram of a method for predicting a target property in a chemical separation process using a physics-informed neural network in accordance with an embodiment.

[0035] In the accompanying figures and following detailed description of the disclosed embodiments, the various elements illustrated in the figures are provided with three-digit reference numbers. In some instances, the leftmost digits of each reference number corresponds to the figure in which its element is first illustrated.DETAILED DESCRIPTION202418740

[0036] The present detailed description provides illustrative embodiments of the disclosed subject matter, which pertains to the integration of thermodynamic models into machine learning frameworks for chemical process simulation and optimization. The disclosed subject matter is broadly applicable to the field of chemical engineering, particularly in the modeling and prediction of separation processes, such as organic acid separations, where thermodynamic principles significantly influence outcomes. By embedding domain-specific scientific knowledge into machine learning methodologies, the disclosed subject matter addresses challenges associated with limited experimental datasets and improves predictive accuracy across a range of industrial applications.

[0037] The examples and embodiments described herein are provided for illustrative purposes only and are not intended to limit the scope of the described subject matter. Certain conventional techniques, components, and processes that are well-known to those skilled in the art may not be described in detail to avoid obscuring the described subject matter.Furthermore, various modifications, rearrangements, and substitutions of the described elements and methodologies may be made without departing from the spirit and scope of the subject matter, as defined by the claims.

[0038] The field of chemical process simulation and optimization has long relied on conventional methods that integrate experimental data with theoretical models based on established physical principles. While these approaches have proven effective in many scenarios, they face significant limitations when applied to complex separation processes, particularly in cases where experimental datasets are sparse or noisy. For example, in organic acid separation processes, accurately modeling system dynamics is challenging due to the intricate interplay of physical and chemical properties, variable operating conditions, and the need to maintain adherence to thermodynamic constraints such as energy and mass balances. Conventional simulation techniques often struggle to capture these dynamic interactions, leading to discrepancies between predicted outcomes and actual process performance. This limitation hinders the ability to design efficient, scalable, and cost-effective processes, especially for novel chemicals or separation technologies.

[0039] The present system addresses these challenges by integrating thermodynamic models and domain-specific scientific knowledge directly into machine learning frameworks, creating a hybrid approach that enhances predictive accuracy and reduces reliance on extensive experimental datasets. Unlike traditional methods that either depend heavily on202418740experimental data or rely solely on theoretical constructs, the disclosed system employs multiple strategies to improve machine learning performance. First, auxiliary predictions related to the target property are incorporated into the training process, enabling the model to uncover interdependencies between physical and chemical variables. Second, virtual sampling points generated from thermodynamic models are used to supplement the experimental data, enriching the training dataset and expanding the range of operating conditions represented. A hybrid loss function balances the contributions of experimental data and thermodynamic insights, ensuring the model prioritizes real-world data while leveraging the predictive power of thermodynamic principles.

[0040] By embedding thermodynamic equations into machine learning models, the described approach ensures that predictions comply with established physical laws, such as conservation of energy and mass, even in scenarios where experimental data is limited or unavailable. This methodology not only enhances the precision of predictions but also expedites the simulation process, allowing for rapid testing of various process configurations to determine favorable solutions. The approach has been validated through improved R2values in adsorption and membrane separation models, demonstrating its capability to achieve greater predictive accuracy compared to traditional machine learning techniques. Additionally, this methodology supports cost-efficient process design and optimization by minimizing the reliance on extensive experimental datasets, making it particularly beneficial for industries such as pharmaceuticals, food production, and biofuel manufacturing.

[0041] FIG. 1 illustrates an example of a processing system 100 that can be used to implement the computer-based components described herein. The processing system 100 includes an exemplary computing device (“computer”) 102 configured for performing various aspects of the operations described herein in accordance with aspects of the invention. In addition to computer 102, exemplary processing system 100 includes network 114, which connects computer 102 to additional systems (not depicted) and can include one or more wide area networks (WANs) and / or local area networks (LANs) such as the Internet, intranet(s), and / or wireless communication network(s). Computer 102 and the additional system are in communication via network 114, e.g., to communicate data between them. In exemplary embodiments, the system 100 may be embodied in a processing system 100.

[0042] Exemplary computer 102 includes processor cores 104, main memory (“memory”) 110, and input / output component(s) 112, which are in communication via bus202418740103. Processor cores 104 include cache memory (“cache”) 106 and controls 108, which include branch prediction structures and associated search, hit, detect, and update logic, which will be described in more detail below. Cache 106 can include multiple cache levels (not depicted) that are on or off-chip from processor 104. Memory 110 can include various data stored therein, e.g., instructions, software, routines, etc., which, e.g., can be transferred to / from cache 106 by controls 108 for execution by processor 104. Input / output component(s) 112 can include one or more components that facilitate local and / or remote input / output operations to / from computer 102, such as a display, keyboard, modem, network adapter, etc. (not depicted).

[0043] A cloud computing system 120 is in wired or wireless electronic communication with the processing system 100. The cloud computing system 120 can supplement, support, or replace some or all of the functionality (in any combination) of the processing system 100. Additionally, some or all of the functionality of the processing system 100 can be implemented as a node of the cloud computing system 120.

[0044] Referring now to FIG. 2, a block diagram of a chemical process simulation system 200 in accordance with an embodiment is shown. The chemical process simulation system 200 is designed to integrate thermodynamic principles into machine learning methodologies for enhanced prediction and optimization of chemical separation processes. The chemical process simulation system 200 includes multiple interconnected components, each contributing to the overall functionality and performance of the simulation framework.

[0045] In exemplary embodiments, the data acquisition module 202 serves as the entry point for collecting experimental data and other relevant information required for the simulation process. The data acquisition module 202 interfaces with external data sources, such as laboratory equipment, sensors, and databases, to gather data corresponding to various operating conditions, including temperature, pressure, pH, and material-specific parameters. In exemplary embodiments, the data acquisition module 202 is configured to preprocess the acquired data, ensuring compatibility with downstream components. Preprocessing may include normalization, outlier detection, and noise reduction. The data acquisition module 202 interacts directly with the thermodynamic model(s) 204 to provide foundational data necessary for generating virtual sampling points and calibrating material-specific parameters.202418740

[0046] In exemplary embodiments, the thermodynamic model(s) 204 embed domainspecific scientific knowledge into the simulation framework. The thermodynamic model(s) 204 utilize thermodynamic equations, such as conservation of energy and mass, equilibrium equations, and adsorption isotherms, to derive virtual sampling points and enrich the dataset. The thermodynamic model(s) 204 are configured to simulate scenarios beyond the scope of experimental data, enabling the chemical process simulation system 200 to generalize better across a wider range of operating conditions. In one embodiment, the thermodynamic model(s) 204 interact with the discrepancy analysis module 206 to identify deviations between thermodynamic predictions and experimental data, and with the augmented dataset generator 208 to supplement the training dataset with virtual sampling points.

[0047] In exemplary embodiments, the discrepancy analysis module 206 evaluates the differences between experimental data and predictions generated by the thermodynamic model(s) 204. In one embodiment, the discrepancy analysis module 206 employs statistical techniques and machine learning algorithms to quantify discrepancies and identify patterns that may indicate limitations in the thermodynamic models or experimental inaccuracies. The discrepancy analysis module 206 outputs discrepancy metrics that can be used by the augmented dataset generator 208 to refine the dataset and by the hybrid loss function optimizer 212 to adjust the loss function parameters during model training.

[0048] In exemplary embodiments, the augmented dataset generator 208 creates enriched datasets by combining experimental data, virtual sampling points generated by the thermodynamic model(s) 204, and auxiliary predictions related to the target property. The augmented dataset generator 208 ensures that the training dataset represents a diverse range of operating conditions, improving the robustness and accuracy of the machine learning model(s) 210. In one embodiment, the augmented dataset generator 208 employs algorithms to balance the contributions of experimental and thermodynamic data, ensuring that the dataset aligns with the requirements of the hybrid loss function optimizer 212.

[0049] In exemplary embodiments, the machine learning model(s) 210 are the computational engines of the chemical process simulation system 200, and the machine learning model(s) 210 are trained to predict target properties associated with chemical separation processes. In one embodiment, the machine learning model(s) 210 are implemented using advanced neural network architectures, such as physics-informed neural networks (PINNs), which integrate thermodynamic principles directly into their structure. In202418740addition to providing virtual sampling points, these equations can also be directly integrated into the machine learning architecture. These additions can provide a soft influence on the model, can be used to strictly enforce relationships between variables, or can be used as combination of both. This differs from the PINNS approach, where the loss function is used to guide the model towards physical behavior but not necessarily strictly enforce the physical constraints. The machine learning model(s) 210 leverage the augmented dataset provided by the augmented dataset generator 208 and are optimized using the hybrid loss function optimizer 212. The machine learning model(s) 210 are configured to enforce physical constraints, such as conservation laws, during the prediction process, ensuring adherence to established scientific principles.

[0050] In exemplary embodiments, the hybrid loss function optimizer 212 governs the training process of the machine learning model(s) 210. The hybrid loss function optimizer 212 employs a hybrid loss function that combines losses derived from experimental data and thermodynamic model-generated virtual sampling points. A scaling hyperparameter is used to balance the contributions of these losses, allowing the hybrid loss function optimizer 212 to prioritize real-world data while leveraging thermodynamic insights. The hybrid loss function optimizer 212 interacts with the scaling hyperparameter adjuster 214 to dynamically adjust the weighting of the loss components based on monitored performance metrics.

[0051] In exemplary embodiments, the scaling hyperparameter adjuster 214 dynamically modifies the scaling hyperparameter used in the hybrid loss function during the training process. The scaling hyperparameter adjuster 214 monitors performance metrics, such as predictive accuracy and adherence to physical constraints, to determine the optimal balance between experimental and thermodynamic losses. The scaling hyperparameter adjuster 214 ensures that the machine learning model(s) 210 achieve high accuracy while maintaining compliance with thermodynamic principles.

[0052] In exemplary embodiments, the prediction engine 216 utilizes the trained machine learning model(s) 210 to predict target properties associated with chemical separation processes. The prediction engine 216 is configured to enforce predetermined thermodynamic constraints, such as conservation of energy and mass, during the prediction process. The prediction engine 216 interacts with the physical constraints enforcer 218 to validate the predictions and ensure their alignment with established scientific laws. The202418740output of the prediction engine 216 is made accessible to users via the user interface module 220.

[0053] In exemplary embodiments, the physical constraints enforcer 218 validates the predictions generated by the prediction engine 216, ensuring compliance with thermodynamic principles and physical laws. The physical constraints enforcer 218 employs algorithms to enforce constraints such as conservation of energy and mass, filtering out predictions that violate these principles. The physical constraints enforcer 218 interacts with the prediction engine 216 to refine the output and with the user interface module 220 to provide users with confidence metrics for the predictions.

[0054] In exemplary embodiments, the user interface module 220 provides a user-friendly interface for interacting with the chemical process simulation system 200. The user interface module 220 allows users to input experimental data, configure simulation parameters, and visualize prediction results. The user interface module 220 integrates with the data storage unit 222 to retrieve historical data and with the prediction engine 216 to display real-time predictions. Advanced visualization tools are included to help users interpret the results and make informed decisions regarding process optimization.

[0055] In exemplary embodiments, the data storage unit 222 serves as the repository for all data associated with the chemical process simulation system 200. The data storage unit 222 includes experimental data, virtual sampling points, auxiliary predictions, model parameters, and historical simulation results. The data storage unit 222 is designed to support efficient data retrieval and storage, ensuring seamless interaction with other components such as the data acquisition module 202, augmented dataset generator 208, and user interface module 220.

[0056] FIG. 3 illustrates one embodiment of a physics-informed neural network model 300 designed to integrate thermodynamic principles into machine learning frameworks for chemical separation process predictions. The physics-informed neural network model 300 includes a neural network 310, thermodynamics-based descriptors 312, intermediate output 314, and a thermodynamic informed loss model 320, such as gBET. As used herein gBET refers to a thermodynamic model used to generalize adsorption isotherms for mixed-gas adsorption equilibria. It is derived from the BET (Brunauer-Emmett-Teller) theory,202418740which is a foundational model for describing the physical adsorption of gas molecules on solid surfaces.

[0057] The neural network 310 functions as the computational centerpiece of the physics-informed neural network model 300. The neural network 310 processes input data, including thermodynamics-based descriptors 312, to generate predictions related to chemical separation processes. The neural network 310 is designed to integrate domain-specific scientific knowledge, ensuring that predictions align with established physical laws. The thermodynamics-based descriptors 312 include parameters such as apparent concentration (CA ,t), initial pH, acid identity, and column identity, which can used for effectively modeling the chemical separation process. These descriptors supply the neural network 310 with foundational information required for producing reliable predictions.

[0058] The intermediate output 314 represents the results generated by the neural network 310 after processing the thermodynamics-based descriptors 312. This output includes variables such as apparent concentration (CA,t) and amount of adsorbed species (nA), which are indicative of the chemical separation process's behavior under specific operating conditions. The intermediate output 314 serves as a bridge between the neural network 310 and the thermodynamics informed loss model 320, ensuring that the predictions align with thermodynamic principles. The thermodynamics-informed loss model 320 assesses the predictions produced by the neural network 310 in relation to thermodynamic constraints, including conservation of energy and mass. The thermodynamics-informed loss model 320 ensures that the loss calculations are based on thermodynamic foundations, allowing the physics-informed neural network model 300 to enhance its predictions and achieve greater precision.

[0059] By integrating these components, the physics-informed neural network model 300 provides a robust framework for predicting target properties associated with chemical separation processes. This integration ensures that the predictions comply with thermodynamic constraints, enhancing the reliability and applicability of the model in industrial scenarios.

[0060] FIG. 4 illustrates an integrated model 400 designed to predict target properties associated with chemical separation processes by combining thermodynamics-based descriptors 412, thermodynamic models 414, and outputs 416. The integrated model 400202418740leverages a neural network architecture to embed thermodynamic principles directly into the prediction framework.

[0061] The thermodynamics-based descriptors 412 serve as input parameters to the integrated model 400. These descriptors include apparent concentration, initial pH, acid identity, and column identity (type of adsorbent), which are important for characterizing the chemical separation process. These inputs provide fundamental information that allows the neural network to model the complex interactions present in the process.

[0062] The thermodynamic models 414 are embedded within the neural network architecture of the integrated model 400. These models include gBET, Langmuir Isotherm, and Henderson-Hasselbalch equations, which represent foundational thermodynamic principles governing adsorption equilibria, chemical interactions, and pH-dependent behavior. The thermodynamic models 414 interact with the neural network to ensure that predictions align with established physical laws, such as conservation of energy and mass, partially established physics (such as empirical and semi-empirical models), and other derived thermodynamic and physics-based models.

[0063] The outputs 416 generated by the integrated model 400 include variables such as apparent concentration ( ,t) and the amount of adsorbed species (nA). These outputs represent the predicted properties of the chemical separation process under specific operating conditions. The outputs 416 are derived from the neural network's processing of the thermodynamics-based descriptors 412 and the embedded thermodynamic models 414, ensuring that the predictions are both accurate and physically consistent.

[0064] FIG. 5 illustrates a thermodynamics knowledge model 500 designed to predict related thermodynamical properties by leveragingthermodynamics-based descriptors 512 and a neural network 514. The thermodynamics knowledge model 500 integrates domainspecific scientific knowledge into the prediction framework, ensuring that outputs 516 adhere to established thermodynamic principles. The thermodynamics-based descriptors 512 serve as input parameters to the thermodynamics knowledge model 500. These descriptors include apparent concentration, initial pH, acid identity, and column identity, which are necessary for characterizing the chemical separation process. These inputs provide foundational information that enables the neural network 514 to model the complex interactions present in the process.202418740

[0065] The neural network 514 processes the thermodynamics-based descriptors 512 to generate predictions related to the chemical separation process. The neural network 514 is configured to incorporate thermodynamic relationships and constraints, ensuring that the predictions align with physical laws such as conservation of energy and mass. The neural network 514 interacts with the thermodynamics knowledge model 500 to enhance the reliability and accuracy of the predictions. The outputs 516 generated by the thermodynamics knowledge model 500 include related thermodynamical properties such as true acid concentration and loading. These outputs represent the predicted properties of the chemical separation process under specific operating conditions. The outputs 516 are derived from the neural network's processing of the thermodynamics-based descriptors 512, ensuring that the predictions are both accurate and physically consistent.

[0066] FIG. 6 illustrates a discrepancy modeling system 600 designed to integrate thermodynamic models 610 and neural networks 620 for chemical process simulation. The system 600 includes a thermodynamics model 610, a neural network 620, thermodynamicsbased descriptors 612, simulated data 614, and experimental data 622.

[0067] The thermodynamics model 610 utilizes thermodynamic principles to process the thermodynamics-based descriptors 612, which include parameters such as apparent concentration, initial pH, acid identity, and column identity. These descriptors provide foundational information for the thermodynamics model 610 (one example is using the gBET model) to generate simulated data 614. The simulated data 614 includes predicted values such as apparent concentration ( ,t) and the amount of adsorbed species (nA), which are derived using thermodynamic equations and models.

[0068] The neural network 620 processes experimental data 622, which represents real-world measurements of chemical separation processes under various operating conditions. The experimental data 622 includes values for apparent concentration (C4,t) and the amount of adsorbed species (nA), serving as a benchmark for comparison against the simulated data 614. The discrepancy model system 600 evaluates differences between the simulated data 614 generated by the thermodynamics model 610 and the experimental data 622 processed by the neural network 620. This evaluation identifies deviations and quantifies discrepancies, enabling the refinement of the thermodynamics model 610 and the neural network 620 to improve predictive accuracy. The discrepancy model system 600 ensures that202418740the integrated system adheres to thermodynamic principles while accounting for variations in experimental observations.

[0069] Referring now to FIG. 7, a flowchart of a computer-implemented method 700 for predicting a target property associated with a chemical separation process according to one or more embodiments is shown. In exemplary embodiments, the method 700 is performed by the chemical process simulation system 200 depicted in FIG. 2.

[0070] The method 700 begins at block 702 with obtaining experimental data corresponding to a plurality of operating conditions for the chemical separation process. This step involves collecting data from laboratory experiments, sensors, or databases, including measurements such as temperature, pressure, pH, and material-specific parameters. For example, experimental data might include the concentration of organic acids in a mixture, adsorption rates, or membrane separation efficiencies under varying conditions. For example, the experimental data may include measurements of organic acid concentration at temperatures ranging from 25°C to 55°C, pressures between 3 and 28 bar, and pH values from 2 to 12.

[0071] Next, as shown in block 704, the method 700 involves generating a plurality of virtual sampling points by applying a thermodynamic model to the experimental data. The thermodynamic model is configured to derive virtual sampling points in accordance with thermodynamic principles, such as conservation of energy and mass, equilibrium equations, and adsorption isotherms. These virtual sampling points represent operating conditions beyond those directly measured in the experimental dataset, enriching the training data and enabling the system to generalize across a wider range of scenarios.

[0072] In one embodiment, generating a plurality of virtual sampling points by applying a thermodynamic model to the experimental data can be performed applying a thermodynamic model to the experimental data to derive virtual sampling points that represent operating conditions beyond those directly measured. This can be achieved by using thermodynamic equilibrium equations, such as adsorption isotherms or conservation laws, to extrapolate material-specific parameters. For instance, the thermodynamic model may calculate equilibrium concentrations of adsorbed species under hypothetical conditions, such as a temperature of 30°C and a pressure of 20 bar, even if these conditions were not part of the original experimental dataset.202418740

[0073] For example, the thermodynamic model may utilize the Langmuir adsorption isotherm to predict the amount of adsorbed species (nA) based on the apparent concentration (C4,t) and material-specific adsorption parameters derived from the experimental data. By fitting the experimental data to the Langmuir equation, the model can generate virtual sampling points that reflect adsorption behavior under a broader range of operating conditions. These virtual sampling points are then added to the training dataset, enriching it with additional data that captures the underlying thermodynamic relationships. This expanded dataset enables the machine learning model to generalize better across scenarios, improving its predictive accuracy for chemical separation processes. For example, the enriched dataset may allow the model to predict adsorption efficiency at a pH of 6.5 and a temperature of 40°C, even if these specific conditions were not experimentally measured.

[0074] At block 706, the method 700 proceeds with incorporating auxiliary predictions related to the target property into a training dataset. Auxiliary predictions capture interdependencies between physical and chemical variables, such as acid concentration, adsorption rates, or pH-dependent behavior captured by the dissociation constant. For example, the system may predict additional chemical process variables, such as the amount of adsorbed species or the true acid concentration, which are correlated with the target property. These auxiliary predictions enhance the training dataset by providing additional insights into the relationships between variables.

[0075] In one embodiment, incorporating auxiliary predictions related to the target property into the training dataset involves generating predictions of additional chemical process variables that are interdependent with the target property. For example, in the context of organic acid separation processes, the target property may be the adsorption efficiency of an acid, while auxiliary predictions may include the true acid concentration, the amount of adsorbed species (nA), or pH-dependent behavior. These auxiliary predictions are derived using domain-specific scientific knowledge and thermodynamic relationships, such as equilibrium equations or adsorption isotherms. Even if the auxiliary predictions are not highly precise, they provide valuable insights into the underlying relationships between physical and chemical variables, enabling the machine learning model to uncover patterns and dependencies that enhance the accuracy of the target property predictions. By incorporating these auxiliary predictions into the training dataset, the system leverages202418740interdependencies between variables to improve model performance, particularly in scenarios where experimental data is sparse or noisy.

[0076] Following this, at block 708, the method 700 involves training a machine learning model on an augmented dataset comprising the experimental data, the virtual sampling points, and the auxiliary predictions. The training process utilizes a hybrid loss function that combines losses derived from experimental data and thermodynamic insights. A scaling hyperparameter is employed to balance the contributions of these losses, ensuring that the model prioritizes real-world data while leveraging the predictive power of thermodynamic principles. In exemplary embodiments, the machine learning model is implemented as a physics-informed neural network, which integrates thermodynamic equations directly into its architecture to enforce physical constraints during training.

[0077] In one embodiment, training the machine learning model involves employing a neural network architecture specifically configured to integrate physical laws of conservation of mass and energy. For example, the neural network may be implemented as a physics-informed neural network, which embeds thermodynamic equations directly into its structure. This integration ensures that the model adheres to established scientific principles during both training and prediction phases. The neural network architecture may include specialized layers or activation functions that enforce thermodynamic constraints, such as conservation of energy and mass, throughout the learning process. By embedding these physical laws, the neural network is able to produce predictions that are not only accurate but also physically consistent, even when experimental data is limited or noisy. This approach enhances the reliability of the model and ensures that its outputs align with the fundamental principles governing chemical separation processes.

[0078] In one embodiment, the hybrid loss function employed during the training of the machine learning model includes a weighted sum of multiple loss components, including a loss based on experimental data, a loss derived from virtual sampling points generated by thermodynamic models, and a weighting determined by a scaling hyperparameter. The loss based on experimental data ensures that the model prioritizes real-world observations, while the loss derived from virtual sampling points allows the model to leverage thermodynamic insights to generalize across a broader range of operating conditions. The scaling hyperparameter dynamically adjusts the relative importance of these loss components, enabling fine-tuning of the training process to achieve optimal performance. For example,202418740during early stages of training, the scaling hyperparameter may prioritize experimental data to establish a strong foundation, while later stages may increase the contribution of thermodynamic losses to refine predictions and enforce physical consistency. This hybrid loss function ensures that the machine learning model achieves high predictive accuracy while maintaining adherence to established thermodynamic principles.

[0079] Finally, as shown in block 710, the method 700 involves predicting the target property using the trained machine learning model. The prediction adheres to predetermined thermodynamic constraints, such as conservation of energy and mass, ensuring that the outputs are both accurate and physically consistent. For example, the system may predict the adsorption efficiency of an organic acid under specific operating conditions, providing actionable insights for optimizing chemical separation processes.

[0080] In one embodiment, predicting the target property further includes enforcing predetermined thermodynamic constraints, including conservation of energy and conservation of mass. For example, during the prediction phase, the machine learning model is configured to validate its outputs against physical laws to ensure compliance with established scientific principles. This enforcement may involve applying thermodynamic equations, such as energy balance equations or mass conservation laws, to filter out predictions that violate these constraints. By incorporating these checks, the system ensures that the predicted target property, such as adsorption efficiency or separation yield, is both accurate and physically consistent. This approach enhances the reliability of the predictions, particularly in scenarios where experimental data is sparse or noisy, and provides actionable insights for optimizing chemical separation processes while maintaining adherence to fundamental thermodynamic principles.

[0081] In one embodiment, the augmented dataset further includes a discrepancy model that captures differences between the experimental data and the virtual sampling points generated by the thermodynamic model. The discrepancy model is configured to evaluate deviations between the predictions produced by the thermodynamic model and the observed experimental data, identifying patterns or inconsistencies that may arise due to limitations in the thermodynamic model or inaccuracies in the experimental measurements. By incorporating these discrepancies into the augmented dataset, the machine learning model is able to learn from both the strengths and weaknesses of the thermodynamic model, improving its ability to generalize across diverse operating conditions. This approach ensures that the202418740machine learning model accounts for real-world variations while maintaining adherence to thermodynamic principles, ultimately enhancing the accuracy and reliability of the predicted target property.

[0082] In one embodiment, the scaling hyperparameter is dynamically adjusted during the training process based on monitored performance metrics of the machine learning model. For example, the system may evaluate metrics such as predictive accuracy, adherence to thermodynamic constraints, and generalization performance across diverse operating conditions. The scaling hyperparameter adjuster continuously monitors these metrics and modifies the relative weighting of the loss components in the hybrid loss function to optimize the training process. During early stages of training, the scaling hyperparameter may prioritize losses derived from experimental data to establish a strong foundation, while later stages may increase the contribution of thermodynamic losses to refine predictions and enforce physical consistency. This dynamic adjustment ensures that the machine learning model achieves high accuracy while maintaining compliance with thermodynamic principles, resulting in reliable and physically consistent predictions.

[0083] In one embodiment, the virtual sampling points are generated to represent operating conditions extending beyond those included in the experimental data. For example, the thermodynamic model may extrapolate data to simulate scenarios that were not directly measured during experimentation, such as higher or lower temperature ranges, varying pressures, or altered pH levels. This extrapolation is achieved by applying thermodynamic principles, such as equilibrium equations, adsorption isotherms, or conservation laws, to derive material-specific parameters under hypothetical conditions. For instance, the thermodynamic model may predict the adsorption behavior of an organic acid at a temperature of 60°C and a pressure of 25 bar, even if these conditions were not part of the original experimental dataset. By generating virtual sampling points that expand the scope of operating conditions, the system enriches the training dataset, enabling the machine learning model to generalize better and improve predictive accuracy across a wider range of scenarios. This approach is particularly beneficial for novel processes or chemicals where experimental data is limited.

[0084] In one embodiment, the machine learning model is implemented as a physics-informed neural network that integrates thermodynamic equations directly into the architecture of the neural network. This implementation ensures that the model adheres to202418740established physical laws, such as conservation of mass and energy, during both training and prediction phases. For example, specific layers within the neural network may be designed to enforce thermodynamic constraints by embedding equations governing equilibrium, adsorption, or energy balances. Additionally, conventional activation functions may be replaced with thermodynamically derived functions to ensure that intermediate outputs align with physical principles. By incorporating these thermodynamic relationships into the neural network architecture, the model achieves higher predictive accuracy and physical consistency, even in scenarios where experimental data is sparse or noisy. This approach enhances the reliability of predictions and provides actionable insights for optimizing chemical separation processes.

[0085] The method 700 enables the chemical process simulation system 200 to achieve high predictive accuracy while reducing reliance on extensive experimental datasets, making it particularly beneficial for industries such as pharmaceuticals, food production, and biofuel manufacturing.

[0086] Referring now to FIG. 8, a flowchart of a computer-implemented method 800 for predicting a target property associated with a chemical separation process using a physics-informed neural network according to one or more embodiments is shown. In exemplary embodiments, the method 800 is performed by the chemical process simulation system 200 depicted in FIG. 2.

[0087] The method 800 begins at block 802 with obtaining experimental data corresponding to a plurality of operating conditions for the chemical separation process. This step involves collecting data from laboratory experiments, sensors, or databases, including measurements such as temperature, pressure, pH, and material-specific parameters. For example, experimental data may include the concentration of organic acids in a mixture, adsorption rates, or membrane separation efficiencies under varying conditions.Representative operating conditions may include temperatures ranging from 25°C to 55°C, pressures between 3 and 28 bar, and pH values from 2 to 12.

[0088] Next, as shown at block 804, the method 800 involves constructing a physics-informed neural network that integrates thermodynamic principles by embedding thermodynamic equations directly into the architecture of the neural network. This step ensures that the neural network adheres to established physical laws, such as conservation of202418740mass and energy, during both training and prediction phases. For example, specific layers within the neural network may be designed to enforce thermodynamic constraints by embedding equations governing equilibrium, adsorption, or energy balances. Additionally, conventional activation functions may be replaced with thermodynamically derived functions to ensure that intermediate outputs align with physical principles.

[0089] At block 806, the method 800 proceeds with training the physics-informed neural network on an augmented dataset comprising the experimental data, auxiliary predictions, and any additional thermodynamic insights. The training process utilizes a hybrid loss function that combines losses derived from experimental data and thermodynamic principles. A scaling hyperparameter is employed to balance the contributions of these losses, ensuring that the model prioritizes real-world data while leveraging the predictive power of thermodynamic relationships. This approach enhances the reliability of the model and ensures that its outputs align with the fundamental principles governing chemical separation processes.

[0090] Finally, as shown at block 808, the method 800 involves predicting the target property using the trained physics-informed neural network. The prediction adheres to predetermined thermodynamic constraints, such as conservation of energy and mass, ensuring that the outputs are both accurate and physically consistent. For example, the system may predict the adsorption efficiency of an organic acid under specific operating conditions, providing actionable insights for optimizing chemical separation processes. This step enables the chemical process simulation system 200 to deliver reliable predictions that are aligned with scientific principles, making it particularly beneficial for industries such as pharmaceuticals, food production, and biofuel manufacturing.

[0091] In one embodiment, constructing the physics-informed neural network includes integrating thermodynamic equations directly into the network architecture and replacing at least one conventional activation function with a thermodynamically based function. For example, the neural network may include specialized layers that enforce thermodynamic constraints, such as conservation of mass and energy, by embedding equations governing equilibrium, adsorption, or energy balances. Additionally, activation functions traditionally used in neural networks, such as ReLU or sigmoid functions, may be replaced with thermodynamically derived functions that reflect the physical relationships inherent in the chemical separation process. This integration ensures that intermediate202418740outputs generated by the neural network align with established physical principles, enhancing the model's ability to produce predictions that are both accurate and physically consistent. By embedding thermodynamic equations and replacing activation functions, the physics-informed neural network achieves higher predictive reliability, particularly in scenarios where experimental data is sparse or noisy.

[0092] In one embodiment, training the physics-informed neural network further includes dynamically adjusting the scaling hyperparameter based on monitored performance metrics of the neural network. For example, during the training process, the system evaluates metrics such as predictive accuracy, adherence to thermodynamic constraints, and generalization performance across diverse operating conditions. The scaling hyperparameter adjuster continuously monitors these metrics and modifies the relative weighting of the loss components in the hybrid loss function to optimize the training process. During early stages of training, the scaling hyperparameter may prioritize losses derived from experimental data to establish a strong foundation, while later stages may increase the contribution of thermodynamic losses to refine predictions and enforce physical consistency. This dynamic adjustment ensures that the physics-informed neural network achieves high accuracy while maintaining compliance with thermodynamic principles, resulting in reliable and physically consistent predictions.

[0093] In one embodiment, the experimental data corresponds to operating conditions that include variations in temperature, pressure, and pH for the chemical separation process. For example, the experimental data may include measurements of organic acid concentration at temperatures ranging from 25°C to 55°C, pressures between 3 and 28 bar, and pH values from 2 to 12. These operating conditions are critical for characterizing the behavior of the chemical separation process under diverse scenarios. By collecting data across a wide range of conditions, the system ensures that the physics-informed neural network is trained on a comprehensive dataset that captures the complex interactions between physical and chemical variables. This diversity in operating conditions allows the model to generalize effectively, improving its predictive accuracy for scenarios that may not have been directly measured during experimentation.

[0094] In one embodiment, the loss derived from thermodynamic principles includes a loss calculated based on deviations from conservation of energy and conservation of mass. During the training process, the physics-informed neural network evaluates its predictions202418740against thermodynamic constraints to ensure compliance with established physical laws. For example, the loss function may penalize predictions that violate energy balance equations or mass conservation laws, thereby guiding the model to produce outputs that adhere to these principles. This thermodynamic loss component is integrated into the hybrid loss function alongside the loss derived from experimental data, ensuring that the model prioritizes real-world observations while leveraging thermodynamic insights. By incorporating this thermodynamic loss, the system enhances the physical consistency of the predictions, particularly in scenarios where experimental data is sparse or noisy, and ensures that the outputs align with the fundamental principles governing chemical separation processes.

[0095] In one embodiment, the physics-informed neural network is further configured to incorporate auxiliary predictions related to one or more additional chemical process variables. For example, in the context of organic acid separation processes, the target property may be the adsorption efficiency of an acid, while auxiliary predictions may include the true acid concentration, the amount of other adsorbed species, or pH-dependent behavior. These auxiliary predictions are derived using domain-specific scientific knowledge and thermodynamic relationships, such as equilibrium equations or adsorption isotherms. By incorporating these auxiliary predictions into the training process, the neural network is able to uncover interdependencies between physical and chemical variables, enhancing its ability to identify patterns and dependencies that improve the accuracy of the target property predictions. This approach is beneficial in scenarios where experimental data is sparse or noisy, as the auxiliary predictions provide additional insights that strengthen the model's performance and reliability.

[0096] In one embodiment, training the physics-informed neural network further includes utilizing a discrepancy model to capture differences between the experimental data and predictions provided by the embedded thermodynamic equations. The discrepancy model evaluates deviations between the outputs of the thermodynamic model and the observed experimental data, identifying patterns or inconsistencies that may arise due to limitations in the thermodynamic model or inaccuracies in the experimental measurements. By incorporating these discrepancies into the training process, the physics-informed neural network is able to learn from both the strengths and weaknesses of the thermodynamic model, improving its ability to generalize across diverse operating conditions. This approach ensures that the neural network accounts for real-world variations while maintaining202418740adherence to thermodynamic principles, ultimately enhancing the accuracy and reliability of the predicted target property.

[0097] In one embodiment, constructing the physics-informed neural network further includes calibrating material-specific parameters using a subset of the experimental data and employing the calibrated parameters within the embedded thermodynamic equations. For example, the experimental data may be used to determine adsorption coefficients, equilibrium constants, or other material-specific properties that are critical for accurately modeling the chemical separation process. These calibrated parameters are then integrated into the thermodynamic equations embedded within the neural network architecture, ensuring that the model reflects the unique characteristics of the materials and operating conditions involved. By incorporating calibrated parameters, the physics-informed neural network achieves higher predictive accuracy and physical consistency, enabling it to produce reliable predictions even in scenarios where experimental data is limited or noisy. This approach enhances the model's ability to generalize across diverse chemical systems and operating conditions, making it particularly valuable for optimizing separation processes in industrial applications.

[0098] In one embodiment, the hybrid loss function further includes a weighted sum of the loss based on the experimental data, the loss derived from thermodynamic principles, and the weighting determined by a scaling hyperparameter. The loss based on experimental data ensures that the model prioritizes real-world observations, while the loss derived from thermodynamic principles enforces adherence to established physical laws, such as conservation of energy and mass. The scaling hyperparameter dynamically adjusts the relative importance of these loss components during the training process, allowing the model to balance accuracy with physical consistency. For example, during the initial stages of training, the scaling hyperparameter may emphasize experimental data to establish a strong foundation, while later stages may increase the contribution of thermodynamic losses to refine predictions and enforce compliance with physical constraints. This hybrid loss function enables the physics-informed neural network to achieve high predictive accuracy while maintaining alignment with thermodynamic principles, ensuring reliable and physically consistent outputs.

[0099] In one embodiment, predicting the target property further includes filtering the predicted values to reduce noise originating from variations in the experimental data and enhancing prediction accuracy. For example, the system may employ post-processing202418740techniques to validate the predictions against thermodynamic constraints, such as conservation of energy and mass, ensuring that the outputs are both physically consistent and reliable. This filtering process may involve applying statistical methods or thermodynamic equations to identify and remove outliers or inconsistencies in the predictions. By refining the predicted values, the system ensures that the outputs align with established scientific principles and provide actionable insights for optimizing chemical separation processes. This approach is particularly beneficial in scenarios where experimental data is sparse or noisy, as it enhances the robustness and reliability of the predictions while maintaining adherence to thermodynamic laws.

[0100] As used herein the virtual sampling points refer to data points generated by applying thermodynamic models to experimental data to simulate operating conditions beyond those directly measured. These points are derived using thermodynamic principles, such as equilibrium equations, adsorption isotherms, and conservation laws, to extrapolate material-specific parameters under hypothetical conditions. For example, a thermodynamic model may predict the equilibrium concentration of adsorbed species at a temperature and pressure not included in the experimental dataset. Virtual sampling points enrich the training dataset by expanding the range of operating conditions represented, enabling the machine learning model to generalize better across diverse scenarios.

[0101] Auxiliary predictions are additional outputs generated by the machine learning model during training that are related to, but distinct from, the target property being predicted. These predictions capture interdependencies between physical and chemical variables, such as acid concentration, adsorption rates, or pH-dependent behavior. For instance, in an organic acid separation process, auxiliary predictions may include the true acid concentration or the amount of adsorbed species (nA). Even if these predictions are not highly precise, they provide valuable insights into the relationships between variables, enhancing the accuracy of the target property predictions. Auxiliary predictions are derived using domain-specific scientific knowledge and thermodynamic relationships embedded in the model.

[0102] In one embodiment, a discrepancy model is a computational framework designed to evaluate differences between predictions generated by thermodynamic models and observed experimental data. This model quantifies deviations and identifies patterns or inconsistencies that may arise due to limitations in the thermodynamic model or inaccuracies202418740in the experimental measurements. The discrepancy model outputs a discrepancy signal, which is incorporated into the augmented training dataset to refine the machine learning model's predictions. By learning from these discrepancies, the machine learning model improves its ability to generalize across diverse operating conditions while maintaining adherence to thermodynamic principles.

[0103] In one embodiment, a physics-informed neural network (PINN) is a specialized neural network architecture that integrates thermodynamic principles directly into its structure. This integration is achieved by embedding thermodynamic equations, such as conservation of mass and energy, equilibrium equations, and adsorption isotherms, into the network layers or activation functions. For example, conventional activation functions, such as ReLU or sigmoid, may be replaced with thermodynamically derived functions to ensure intermediate outputs align with physical laws. PINNs enforce physical constraints during both training and prediction phases, ensuring that the model produces outputs that are both accurate and physically consistent.

[0104] In one embodiment, the training process for the physics-informed neural network involves using an augmented dataset comprising experimental data, virtual sampling points, auxiliary predictions, and discrepancy signals. The neural network is optimized using gradient-based techniques based on the hybrid loss function. In one embodiment, the hybrid loss function is a weighted combination of multiple loss components, including losses derived from experimental data, virtual sampling points, and thermodynamic principles. Each loss component is calculated independently and combined using a scaling hyperparameter to balance their contributions. For example, the loss based on experimental data ensures the model prioritizes real-world observations, while the loss derived from virtual sampling points leverages thermodynamic insights to improve generalization. The scaling hyperparameter is dynamically adjusted during training based on monitored performance metrics, such as predictive accuracy and adherence to physical constraints.

[0105] In one embodiment, the scaling hyperparameter is a dynamic parameter used to balance the relative importance of different loss components in the hybrid loss function. During training, the scaling hyperparameter is adjusted based on performance metrics, such as the model's ability to adhere to thermodynamic constraints and its predictive accuracy. For example, in the early stages of training, the scaling hyperparameter may prioritize losses derived from experimental data to establish a strong foundation. As training progresses, the202418740hyperparameter may increase the contribution of thermodynamic losses to refine predictions and enforce physical consistency.

[0106] For the sake of brevity, conventional techniques related to making and using the disclosed embodiments may or may not be described in detail herein. In particular, various aspects of computing systems and specific computer programs to implement the various technical features described herein are well known. Accordingly, in the interest of brevity, many conventional implementation details are only mentioned briefly or are omitted entirely without providing the well-known system and / or process details.

[0107] The various components / modules / models of the systems illustrated herein are depicted separately for ease of illustration and explanation. In embodiments of the invention, the functions performed by the various components / modules / models can be distributed differently than shown without departing from the scope of the various embodiments of the invention described herein unless it is specifically stated otherwise.

[0108] Aspects of the invention can be embodied as a system, a method, and / or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.

[0109] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element components, and / or groups thereof.

[0110] While the present invention has been described with reference to an exemplary embodiment or embodiments, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted for elements thereof without departing from the scope of the present invention. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present invention202418740without departing from the essential scope thereof. Therefore, it is intended that the present invention not be limited to the particular embodiment disclosed as the best mode contemplated for carrying out this present invention, but that the present invention will include all embodiments falling within the scope of the claims.

Claims

202418740CLAIMSWhat is claimed is:

1. A method for predicting a target property associated with a chemical separation process, comprising:obtaining experimental data corresponding to a plurality of operating conditions for the chemical separation process;generating a plurality of virtual sampling points by applying a thermodynamic model to the experimental data, the thermodynamic model being configured to derive the virtual sampling points in accordance with thermodynamic principles;incorporating auxiliary predictions related to the target property into a training dataset, the auxiliary predictions capturing interdependencies between physical and chemical variables;training a machine learning model on an augmented dataset comprising the experimental data, the virtual sampling points, and the auxiliary predictions, the training including a utilization of a hybrid loss function that combines:a loss based on the experimental data;a loss based on the virtual sampling points; anda scaling hyperparameter to balance the loss based on the experimental data and the loss based on the virtual sampling points; andpredicting the target property using the trained machine learning model, the prediction adhering to predetermined thermodynamic constraints.

2. The method according to claim 1, wherein generating the plurality of virtual sampling points further comprises:applying thermodynamic equations to the experimental data; andderiving material-specific parameters.

3. The method according to claim 1, wherein incorporating auxiliary predictions comprises:generating predictions of one or more additional chemical process variables; and correlating the additional chemical process variables with the target property.

4. The method according to claim 1, wherein training the machine learning model further comprises:202418740employing a neural network architecture; andconfiguring the neural network architecture to integrate physical laws and thermodynamic models.

5. The method according to claim 1, wherein the hybrid loss function comprises: a weighted sum of a loss based on the experimental data;a loss based on the virtual sampling points; anda weighting determined by the scaling hyperparameter.

6. The method according to claim 1, wherein predicting the target property further comprises:enforcing predetermined thermodynamic constraints;including a conservation of energy; andincluding a conservation of mass.

7. The method according to claim 1, wherein the augmented dataset further comprises a discrepancy model and capturing differences between the experimental data and the virtual sampling points.

8. The method according to claim 1, wherein the scaling hyperparameter is dynamically adjusted during training based on a monitored performance metric of the machine learning model.

9. The method according to claim 1, wherein the virtual sampling points are generated to represent operating conditions extending beyond those included in the experimental data.

10. The method according to claim 1, wherein the machine learning model is implemented as:a physics-informed neural network; andthermodynamic equations integrated directly into an architecture of the neural network.

11. A method for predicting a target property associated with a chemical separation process, comprising:obtaining experimental data corresponding to a plurality of operating conditions for202418740the chemical separation process;constructing a physics-informed neural network that integrates thermodynamic principles by embedding thermodynamic equations directly into an architecture of the neural network;training the physics-informed neural network on the experimental data, the training including a utilization of a hybrid loss function that combines:a loss based on the experimental data; anda loss derived from thermodynamic principles,a scaling hyperparameter to balance the loss based on the experimental data and the loss derived from thermodynamic principles; andpredicting the target property using the trained physics-informed neural network, the prediction enforcing predetermined thermodynamic constraints including conservation of mass and energy.

12. The method according to claim 11, wherein constructing the physics-informed neural network further comprises:integrating thermodynamic equations into the network; andreplacing at least one activation function with a thermodynamically based function.

13. The method according to claim 11, wherein training the physics-informed neural network further comprises:dynamically adjusting the scaling hyperparameter; andbasing the adjustment on a monitored performance metric of the physics-informed neural network.

14. The method according to claim 11, wherein the experimental data corresponds to operating conditions that include variations in temperature, pressure, and pH for the chemical separation process.

15. The method according to claim 11, wherein the loss derived from thermodynamic principles comprises a loss calculated based on deviations from conservation of energy and conservation of mass.

16. The method according to claim 11, wherein the physics-informed neural network is further configured to incorporate auxiliary predictions related to one or more additional chemical process variables.20241874017. The method according to claim 11, wherein training the physics-informed neural network further comprises:utilizing a discrepancy model; andcapturing differences between the experimental data and predictions provided by the embedded thermodynamic equations.

18. The method according to claim 11, wherein constructing the physics-informed neural network further comprises:calibrating material-specific parameters using a subset of the experimental data; and employing the calibrated parameters within the embedded thermodynamic equations.

19. The method according to claim 11, wherein the hybrid loss function further comprises:a weighted sum of the loss based on the experimental data;the loss derived from thermodynamic principles; andthe weighting determined by the scaling hyperparameter.

20. The method according to claim 11, wherein predicting the target property further comprises:filtering the predicted values to reduce noise originating from variations in the experimental data; andenhancing prediction accuracy.