Physical information intelligent calculation model based on sensor data

By simulating and modifying sensor response data using a process-based model, a training dataset is generated, which solves the problem of insufficient training data for machine learning models and improves the accuracy and robustness of the model in predicting complex physical systems.

CN122070552APending Publication Date: 2026-05-19TRANSCEND ENG & TECH LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480052634.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-20
Filing Date
2024-06-20
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, machine learning models struggle to accurately predict the behavior of complex physical systems, especially in natural environments, when there is a lack of sufficient quantity and quality of training data. Furthermore, synthetic data often fails to accurately reflect unobservable or difficult-to-observe variables in physical systems, resulting in insufficient model generalization ability.

Method used

By simulating physical systems using process-based models, sensor response data is generated and modified to produce simulated data that more closely approximates the actual sensor output. This simulated data serves as the training dataset for training machine learning models. The training dataset is generated by simulating the imperfections of real physical sensors by utilizing sensor performance characteristics and noise bias.

Benefits of technology

It improves the accuracy and generalization ability of machine learning models in predicting complex physical systems, ensures that training data reflects the complexity and imperfections of sensors, and enhances the robustness and prediction accuracy of the models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122070552A_ABST
    Figure CN122070552A_ABST
Patent Text Reader

Abstract

Computational modeling of physics-based, intelligent machine learning-based complex natural phenomena using sensor data as input is disclosed. Computational modeling includes calculating sensor performance characteristics of a physical sensor used in measuring attributes of a physical system. The modeling further includes simulating the physical system according to the process-based model to produce simulated data corresponding to one or more physical state variables of the natural system, and applying the calculated sensor performance characteristics of the physical sensor to the simulated data to destroy the simulated data, thereby generating one or more simulated sensor responses closer to the actual output of the physical sensor. A training dataset is generated from the simulation data, the training dataset reflecting the simulation sensor response and input parameters of the process-based model to train a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 521,974, filed June 20, 2023, entitled “Intelligent Computational Model Based on Sensor Data,” the disclosure of which is incorporated herein by reference in its entirety.

[0003] Statement regarding federally sponsored research or development

[0004] All aspects of this disclosure were carried out with the support of the U.S. government under a contract awarded to the U.S. Army Corps of Engineers (Contract #W913E519C0003). The government retains certain rights to this disclosure. Technical Field

[0005] This disclosure generally relates to intelligent computing models, and more specifically, to developing machine learning-based computing models for complex physical phenomena that use sensor data as input. In some implementations, this disclosure relates to developing machine learning-based computing models for the properties of fluid flow in unsaturated porous media that use sensor data as input. Background Technology

[0006] In machine learning, training data is used to train machine learning models. Obtaining training data for machine learning may require some human input. The quantity and quality of training data can vary depending on the machine learning technique and the type of model being trained.

[0007] A machine learning model is trained by feeding it training data and modifying its parameters to reduce the error between the expected and actual outputs. This process can be repeated multiple times until the model reaches an acceptable level of accuracy, driven by optimization algorithms such as gradient descent using backpropagation or stochastic gradient descent. Summary of the Invention

[0008] According to one aspect of this disclosure, a method for computational modeling based on intelligent machine learning to describe complex physical phenomena is disclosed. The method uses sensor data as input. The method includes calculating sensor performance characteristics of physical sensors used in measuring properties of a physical system. The method also includes simulating the physical system according to a process-based model to generate simulated data corresponding to one or more physical state variables of a natural system, and applying the calculated sensor performance characteristics of the physical sensors to the simulated data to corrupt the simulated data, thereby generating one or more simulated sensor responses that more closely approximate the actual output of the physical sensors. In some cases, the physical phenomenon may be a natural phenomenon. A training dataset is generated to train the machine learning model, using the simulated sensor responses, the simulated data, and / or the process-based model as input. In some embodiments, the simulated sensor responses may replace the simulated data from which the simulated sensor responses are generated in the training dataset. In other words, the simulated sensor responses are generated by applying at least one sensor performance characteristic to a state variable represented in the simulated data; this state variable may be replaced by the simulated sensor responses in the training data, or it may be replaced in the simulated data before the simulated data is stored in the training dataset. The training dataset may include multiple training examples. Therefore, as used herein, the simulated data stored in the training dataset reflects any simulated sensor responses generated from the simulated data. Each training example can identify a property of the physical system as a training objective. The training objective can be a property of the physical system. The training objective can be a state variable. In some implementations, simulated sensor responses can be identified as training objectives. A training example can identify multiple state variables as target training objectives. Each training example can include a time series of simulated data for the training objectives(s). Each training example can include an instance of a time series of simulated data. A process-based model can be a virtual copy of the physical system that mimics the behavior of the physical system using real-world data and / or other input data. A physical sensor can be one or more sensors of the same type used to measure a given state variable of the physical system, or it can be one of multiple sensors of different types used to measure one or more state variables of the physical system.

[0009] In one implementation, the method includes performing simulation steps and application steps multiple times, each of which corresponds to a scenario. Each scenario can be defined by multiple input parameters representing properties of the physical system being simulated. Properties of the physical system can be invariant, meaning that the properties do not change during the simulation. Such invariant properties are called physical system properties or simply properties. An example of a physical system property is a domain definition. A domain definition specifies the arrangement of physical objects or materials within a physical system, including intrinsic and extrinsic physical properties. Examples of properties include the properties of materials present in a physical system. For example, the spatial variability of the intrinsic properties of a fixed material within a simulation domain is one aspect of a domain definition. The order and thickness of each layer in a multi-layered profile of soil, along with the porosity, permeability, and other physical properties of each layer, are examples of domain definitions.

[0010] Properties of a physical system can also be time-varying (e.g., time-dependent). Such time-varying properties are also called physical state variables or simply state variables. When physical state variables are used as inputs to a model, simulation, or another process, they can also be called physical state parameters or state parameters. Examples of state variables include temperature, pressure, flux, chemical concentration, or any other physical property that can be measured by physical sensors. However, state variables are not limited to properties that can be measured by sensors. For example, a state variable could include the in-situ permeability of a soil region. Another example of a state variable is the flux rate of groundwater in a particular direction. Yet another example of a state variable is the presence or absence of underground tunnels within the simulation domain. Some state variables can be initial conditions for simulating a physical system. Initial conditions are state variables that are used as input parameters to the physical system at the start of the simulation. In other words, input conditions specify the conditions (state variable values) of the time-varying properties of the physical system at the start of the simulation. For example, initial conditions could include initial temperature, pressure, or chemical concentration. State variables can also include variable boundary conditions. Boundary conditions represent how the system behaves at the boundaries of the domain explicitly represented in the simulation. Variable boundary conditions can also be considered state variables of the physical system. Boundary conditions can also be static. Static boundary conditions can be considered as properties of the physical system. Multiple scenarios can be generated, each corresponding to a simulation. This simulation produces an unmodified set of simulated data and a set of simulated sensor responses obtained through disruption, based on sensor response characteristics of at least some unmodified simulated data. A machine learning model is trained based at least on the simulated sensor responses used as training input.

[0011] According to embodiments of this disclosure, a computer program product is disclosed. The computer program product includes one or more computer-readable storage devices and program instructions stored on at least one of the one or more tangible storage devices. The program instructions are executable by a processor and include program instructions for intelligent modeling of complex physical phenomena. The computer program product includes program instructions for calculating sensor performance characteristics of physical sensors used in measuring properties of a physical system. The computer program product also includes program instructions for simulating the physical system according to a process-based model to generate simulated data corresponding to one or more physical state variables of the natural system. The computer program product includes program instructions for applying the calculated sensor performance characteristics of the physical sensors to the simulated data to corrupt the simulated data, thereby generating one or more simulated sensor responses that more closely approximate the actual output of the physical sensors. The computer program product includes program instructions for generating a training dataset to train a machine learning model, the training dataset being generated using the simulated data. The simulated data stored in the training dataset can be modified such that the simulated data reflects simulated sensor responses rather than the state variables used to generate the simulated sensor responses. In other words, a simulated sensor response is generated by applying at least one sensor performance characteristic to a state variable represented in simulated data; this state variable may be replaced by the simulated sensor response in the training data, or it may be replaced in the simulated data before the simulated data is stored in the training dataset. In either case, the simulated data stored in the training dataset reflects the generated simulated sensor response (if any). The training dataset may identify properties (or more properties) of the physical system as training targets (or more targets).

[0012] According to embodiments of this disclosure, a non-transitory computer-readable storage medium tangibly embodying computer-readable program code is disclosed. The computer-readable program code includes computer-readable instructions that, when executed, cause a processor to perform a method comprising calculating sensor performance characteristics of physical sensors used in measuring state variables (variable attributes) of a physical system, and simulating the physical system according to a process-based model to generate simulated data corresponding to one or more physical state variables of the natural system. The processor applies the calculated sensor performance characteristics of the physical sensors to the simulated data to modify the simulated data, thereby generating one or more simulated sensor responses that more closely approximate the actual output that a physical sensor in an actual physical system corresponding to the simulated physical system would produce. A training dataset is generated by the processor to train a machine learning model. The training dataset includes the simulated data and at least some input parameters of the process-based model. At least one attribute of the physical system (e.g., at least one physical system property or at least one state variable) is identified as a training objective in the training dataset. The training dataset may include several training examples, each training example associated with a corresponding training objective and simulated data of the training objective.

[0013] groundwater

[0014] According to one aspect of this disclosure, a method is disclosed that describes computational modeling of the vertical flux of fluids in unsaturated porous media, as well as other movement and storage properties (such as groundwater flux through unsaturated soil, also referred to herein as complex natural phenomena) based on intelligent machine learning. The method uses sensor data as input. The method includes using physical sensors to measure state variables (i.e., variable properties) of a physical system, such as the water content of the physical system, which is an unsaturated porous medium. The method also includes simulating the physical system according to a process-based unsaturated groundwater flow model to generate simulated data corresponding to one or more physical state variables of the natural system, to generate one or more simulated sensor responses that approximate the actual outputs of the physical sensors. A training dataset is generated to train the machine learning model. The training dataset is generated using the simulated data and identifies at least one property of the physical system as a training objective. The training objective may correspond to boundary conditions. The training objective may correspond to a domain definition. The training objective may correspond to a state variable that cannot be measured by physical sensors. The training objective may be one of multiple training objectives. Each of the multiple training objectives may represent a different property of the physical system. Process-based unsaturated groundwater flow models can be virtual replicas of physical systems, using real-world data and / or other input data to mimic the behavior of the physical system.

[0015] In one implementation, the method includes performing simulation and application steps multiple times, each of which corresponds to a scenario. Each scenario can be defined by multiple input parameters representing a domain definition, initial conditions, and / or boundary conditions of a physical system. For example, the domain definition may include a spatial distribution of numerical values ​​of physical properties characterizing soil hydraulic behavior, the initial conditions may include an initial spatial distribution of soil water content, and the boundary conditions may include time-varying water pressure or “head” specified at the upper surface of the soil. Another boundary condition may include time-varying water flux from a vertical region of the soil, representing plant root uptake from the rhizosphere. Additional initial and boundary conditions may be specified. Multiple scenarios can be generated, each corresponding to a simulation that produces a set of simulated data and a set of simulated sensor responses obtained by computationally placing one or more virtual sensors in the simulation domain. Each scenario may correspond to training examples stored in a training dataset. Each scenario may correspond to multiple training examples stored in the training dataset, where each training example represents a different time period during the scenario. A machine learning model is trained based at least on the simulated sensor responses used as training inputs.

[0016] According to embodiments of this disclosure, a computer program product is disclosed. The computer program product includes one or more computer-readable storage devices and program instructions stored on at least one of the one or more tangible storage devices. The program instructions are executable by a processor and include program instructions for intelligently modeling the time-varying flow of fluid through a porous medium. The computer program product includes program instructions for calculating the response of a sensor performance characteristic of a physical sensor used to measure water-based variant properties (state variables) of the porous medium. The computer program product also includes program instructions for simulating a physical system based on a process-based unsaturated groundwater flow model to generate simulated data corresponding to one or more physical state variables of the natural system. The computer program product includes program instructions for generating one or more simulated sensor responses that approximate the actual output of the physical sensor. The computer program product includes program instructions for generating a training dataset to train a machine learning model, as described herein. Attached Figure Description

[0017] The accompanying drawings are illustrative embodiments. They do not show all embodiments. Other embodiments may be used additionally or alternatively. Details that may be obvious or unnecessary may be omitted to save space or for more efficient description. Some embodiments may be practiced using additional components or steps and / or without all components or steps shown. When the same number appears in different drawings, it refers to the same or similar components or steps.

[0018] Figure 1A block diagram depicts a data processing environment, illustrating a network in which an illustrative implementation of a data processing system can be carried out;

[0019] Figure 2 A block diagram of a data processing system in which illustrative embodiments can be implemented is depicted;

[0020] Figure 3 A data synthesis configuration in which illustrative implementation methods can be achieved is described;

[0021] Figure 4 A machine learning engine in which illustrative implementations can be achieved is described;

[0022] Figure 5 A block diagram depicts an example training architecture for a model of a complex natural phenomenon in which illustrative implementations can be achieved.

[0023] Figure 6 A block diagram depicting the configuration of an intelligent computing model in which illustrative implementation methods can be realized;

[0024] Figure 7 A block diagram depicts a tunnel detection configuration that can implement the illustrative implementation.

[0025] Figure 8 Examples of illustrative implementation methods are shown.

[0026] Figure 9 A data synthesis configuration in which illustrative implementation methods can be achieved is described;

[0027] Figure 10 A graph depicting simulated data of volumetric water content in which illustrative embodiments can be implemented is presented.

[0028] Figure 11 A block diagram depicts a vertical soil water flux estimation configuration in which an illustrative implementation can be achieved;

[0029] Figure 12 The routines in which illustrative implementations can be carried out are described;

[0030] Figure 13A A first graph depicting the predicted flux and the actual flux according to an illustrative embodiment is shown.

[0031] Figure 13B A second graph depicting the predicted flux and the actual flux according to an illustrative embodiment is shown.

[0032] Figure 13C A third graph depicting the predicted flux and the actual flux according to an illustrative embodiment is shown.

[0033] Figure 13DA fourth graph depicting the predicted flux and the actual flux according to the illustrative implementation is presented. Detailed Implementation

[0034] In the following detailed description, many specific details are illustrated by way of examples to provide a thorough understanding of the relevant teachings. However, it should be apparent, however, that these teachings can be practiced without these details. In other cases, well-known methods, processes, components, and / or circuits have been described at a relatively high level without detail in order to avoid unnecessarily obscuring aspects of these teachings.

[0035] In machine learning, obtaining a sufficient amount of labeled training data through regular physical observations is often impractical. Insufficient labeled training data is a technical problem because a lack of data negatively impacts the quality of model predictions; that is, it causes the model to fail to generalize accurately beyond the specific examples or combinations of input variables provided to it during training. A model that fails to generalize accurately (a poorly generalizing model) will be limited in its applicability (by an insufficiently wide range of input variability) or overfit to the limited training data, preventing it from generalizing well even within the range of input variability presented during training. This limits the ability to use data-scarce supervised machine learning techniques to predict the behavior of complex physical systems based on sensor data, especially in natural environments.

[0036] While synthetic data can be frequently used to train machine learning models, a technical challenge in using synthetic data for training is that data generated by physical systems rarely follows well-ordered parametric probability distributions. This makes the synthesis of such data a highly complex task, where it is crucial that the synthetic data accurately reflects the unobservable or difficult-to-observe variables of the physically determined system. For example, because data generated by physical systems rarely follows ordered parametric probability distributions, basic sampling and stratification methods for obtaining training data often produce questionable results, negatively impacting the prediction quality of models trained using such data.

[0037] For example, if a model is trained on samples from a well-ordered parameter probability distribution, it will perform poorly when using physical sensor data as input (i.e., in inference mode) because the physical sensor data reflects the imperfections and variable quality of the physical sensors involved. In other words, physical sensors may distort the physical reality of the physically determined system in which the sensors are placed in one or more ways. This potential distortion is referred to in this paper as the sensor's performance characteristics or limiting characteristics. A model trained on synthetic sensor data that does not reflect these limiting characteristics will not take these characteristics into account, resulting in poor model output. Therefore, the technical problem lies not only in obtaining sufficient training data, but also in ensuring that the training data reflects the complex, imperfect, and erratic quality of the sensors involved, and / or approximates measurements of the sensors involved, especially when collecting sufficient real-world measurement data to induce robust and accurate performance from data-scarce machine learning algorithms may simply be impractical or impossible.

[0038] This illustrative implementation provides a technical solution to the insufficient quantity and quality of labeled training data for machine learning models supporting physical processes. More specifically, the implementation relates to the generation and preparation of training data for applications involving classification and regression in physical systems, such as natural systems, where the input to the supervised machine learning model is sensor data. This illustrative implementation can synthesize a large amount of representative training data using a combination of one or more process-based models and inputs from measurements and synthesis of the process-based models. This illustrative implementation modifies the theoretically perfect output of the process-based model to simulate the imperfections of real-world physical sensors by applying the modification. The modification may include transfer functions and probabilistic representations of noise and bias obtained by characterizing the performance of actual sensors in response to known or controlled experimental conditions. This illustrative implementation can modify the theoretically perfect output of the process-based model to approximate the destruction of physical information inherent in the output from real-world physical sensors. The modified output of the process-based model is then used as input and / or a training objective to train the machine learning model. The training objective is the desired predictive output of the model being trained. In the disclosed implementation, the training objective may represent a property of the physical system, such as a domain definition or variable condition (state variable) used as input to a process-based model.

[0039] As discussed in this paper, a process-based model can be a simulation or mathematical description depicting how a system or process behaves over time. This model can represent the processes that occur within the system and how these processes interact with each other. The model typically involves a set of equations or algorithms describing the relationships between different model properties, such as inputs, outputs, and internal states. This model enables the understanding and prediction of the behavior of complex systems, such as weather patterns, soil systems, and biological systems. Process-based models can be a combination of physics-based models and / or empirical models.

[0040] Physics-based models can use physical laws to represent physical processes occurring within a physical system. These models can be based on mathematical descriptions representing physical processes, such as the conservation of mass and energy, Newton's laws of motion, the laws of thermodynamics, and the laws of electromagnetism. However, empirical models can be based on observations and measurements of the system, rather than on first principles or underlying physical laws. These models can employ statistical techniques to identify patterns and relationships within the data and can be used to predict the future behavior of the system or to understand its past behavior. Process-based models, which can include process-based models of unsaturated groundwater flow, can also be hybrid models, which can be a combination of at least two of physics-based models, empirical models, and any other models. In a non-limiting example, an empirical model could be a modified St. Venant model that correlates flow velocity, water depth, and velocity in a river or open channel with channel geometry, riverbed slope, friction, and other factors affecting flow dynamics. In another non-limiting example, an empirical model could be the Penman-Monteith model, which estimates (predicts) evapotranspiration based on energy balance and aerodynamic concepts, taking into account factors such as net radiation, air temperature, humidity, wind speed, and vegetation characteristics. Another non-limiting example of an empirical model could be the Bishop, Sandberg, and Tong (BST) correlation used in nuclear power applications to predict the critical heat flux in nuclear fuel rods, which is the maximum heat flux that can be removed by boiling before a vapor film forms on the rod surface, leading to a rapid decrease in heat transfer efficiency. An example of an empirical model in the context of soil water flux modeling (which can be a mixture of physics-based and empirical modeling techniques) could be the Van Genuchten-Mualem equations, which can be used to describe soil water properties. In this context, a physics-based model could be the Richardson-Richards equations, which can represent the movement of water in unsaturated soils.

[0041] In a non-restrictive example, the physics-based model could be a numerical implementation of point dynamics equations, which are a set of first-order differential equations used in nuclear engineering to predict the time-dependent behavior of neutron swarms in nuclear reactors.

[0042] In a non-limiting example, a mixing process-based model could be one that combines a physics-based numerical implementation of point dynamics equations with physics-based flow equations (such as the Navier-Stokes equations) and empirical heat transfer relationships (such as the Bishop, Sandberg, and Tong (BST) correlations), as well as other models and thermodynamic relationships, to form a comprehensive model of a nuclear reactor and its power generation and cooling systems. Another non-limiting example of a mixing process-based model could be the Variable Saturated Flow (VS2D / VS2DT) model developed by the U.S. Geological Survey (USGS), which uses the empirical Van Genuchten model to predict (estimate) the hydraulic properties of variable saturated soils and uses a physics-based numerical implementation of the Richardson-Richards equations to solve for unsaturated flow.

[0043] In one aspect, a method for generating training data for supervised learning about a physical system using a machine learning model can be disclosed. This method may include simulating the physical system using at least one set of input parameters and a process-based model that generates outputs corresponding to state variables. The physical system may be a natural system, and the process-based model may be a virtual copy of the physical system that mimics its behavior using real-world data (such as sensor data) and / or other input data (such as the domain definition of the physical system, initial conditions of the physical system, boundary conditions of the physical system, etc.). The physical state variables may or may not be measurable by physical sensors. In the example method, multiple measurable outputs of the process-based model may be corrupted by added uncertainty, which represents an inherent defect in the measurement of the corresponding real state variables by real physical sensors. Uncertainty reflects an overall lack of precision and accuracy in the measurement. Uncertainty can be represented by noise. Noise refers to random variability in data that cannot be attributed to any particular cause. In other words, noise represents unpredictable fluctuations. Uncertainty can be represented by bias. Bias is a systematic error that skews the results in a particular direction. Bias may occur due to assumptions or methods that consistently distort the measurement. Uncertainty can reflect both noise and bias. In one respect, manipulation can capture the chaos, limitations, or characteristics that sensors may exhibit during use, rather than the chaos about the real world that the model fails to capture. Therefore, manipulation can replace one or more pure values ​​of the process-based model output with one or more corresponding simulated sensor responses, as measured by a virtual sensor with the same characteristics as a real physical sensor. Thus, the simulated data stored in the training dataset is understood to reflect the corresponding simulated sensor responses. Therefore, the manipulated values ​​are more realistic versions of the pure values / model outputs. In another example approach, multiple measurable outputs of a process-based unsaturated groundwater flow model can be modified or selected from the physical domain to generate or represent simulated sensor responses that approximate sensor data.

[0044] In the example method, simulated sensor responses, whether modified or unmodified outputs of a process-based model, can be used as inputs to train a machine learning model, as described below. Inputs and / or outputs of the process-based model, such as time-varying state properties (state variables representing initial and boundary conditions), time-invariant properties (domain definitions, static boundary conditions), synthetic sensor readings (e.g., state variables corrupted using sensor performance characteristics), etc., can be adopted as training targets depending on the purpose of the model to be trained. In implementations using unsaturated groundwater flow models, different types of sensors can be used, including, for example, water content sensors, temperature sensors, and pressure sensors (e.g., tensiometers).

[0045] In one training method, the machine learning model can be provided with simulated outputs from sensors that measure properties or state variables of a physical system. Optionally, the simulated outputs can be combined with the actual outputs from the sensors for training. In another training method, the machine learning model can be provided with simulated outputs from sensors that measure properties or state variables of a physical system. Optionally, the simulated outputs can be combined with the actual outputs from the sensors for training. Typically, state variables can represent time-series (time-varying) data in a dynamic system. In other words, state variables have values ​​that fluctuate over time as the simulation (process-based model) proceeds. Properties can be time-invariant attributes of a natural physical domain or environment. In other words, attributes represent data that remains unchanged over time. To use the trained machine learning model during inference, input data including actual sensor data (such as previously unseen sensor data) can be used to predict the state of one or more variables or one or more properties of a physical system.

[0046] On the other hand, synthetic data approaches can be applied to a range of architectures, including Convolutional Neural Networks (CNNs), Transformer Neural Networks (TNNs), Visual Transformer (ViT) Neural Networks, Autoencoders (AEs, a form of CNN), Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs, a form of RNN), and non-ANN (non-artificial neural network) machine learning models and architectures such as Random Forests (RFs), other Classification and Regression Trees (CART) methods, Partial Least Squares Regression (PLSR) Gradient Boosting Regression, Support Vector Regression (SVR), and so on. While the description presented in this paper may be beneficial in all supervised machine learning applications of natural physical systems whose input data includes sensor data, this technique may be particularly helpful for ANNs or RFs in solving multi-objective regression problems (predicting multiple objects simultaneously) because the process of synthesizing input data often involves process-based models where multiple state variables are computed as part of a simulation process and are therefore available as outputs of the simulation. These outputs can be converted into simulated sensor responses that provide physically consistent inputs and corresponding training objectives, thus providing appropriate inference bias for multi-objective prediction.

[0047] Some operations are described as occurring at a specific component or location in the implementation. This locality of operation is not intended to limit the illustrative implementation. Any operation described herein as occurring at or performed by a specific component can be implemented in such a way that a component-specific function causes the operation to occur or be performed at another component (e.g., at a local or remote machine learning (ML) engine).

[0048] The illustrative embodiments are described by way of example only, relating to certain types of data, functions, algorithms, equations, model configurations, implementation locations, additional data, devices, data processing systems, environments, components, and applications. Any particular manifestation of these and other similar products is not intended to limit this disclosure. Any suitable manifestation of these and other similar products may be chosen within the scope of the illustrative embodiments.

[0049] Furthermore, illustrative embodiments can be implemented with respect to any type of data, data source, or access to a data source via a data network. Within the scope of this disclosure, any type of data storage device can provide data to embodiments of this disclosure either locally within a data processing system or via a data network. Within the scope of the illustrative embodiments, when embodiments are described using mobile devices, any type of data storage device suitable for use with mobile devices can provide data to such embodiments either locally within the mobile device or via a data network.

[0050] The illustrative implementations are described using specific code, designs, architectures, protocols, layouts, diagrams, and tools as examples only, and are not limited to. Furthermore, for clarity, specific software, tools, and data processing environments are used in some instances as examples only. The illustrative implementations can be used in conjunction with other equivalent or similar structures, systems, applications, or architectures. For example, within the scope of this disclosure, other equivalent mobile devices, structures, systems, applications, or architectures can be used in conjunction with this implementation of the disclosure. The illustrative implementations can be implemented in hardware, software, or a combination thereof.

[0051] The examples in this disclosure are for clarity of description only and are not limited to illustrative embodiments. Additional data, operations, actions, tasks, activities, and manipulations are conceived from this disclosure and are contemplated within the scope of the illustrative embodiments.

[0052] Any advantages listed herein are merely illustrative and are not intended to limit the illustrative implementations. Additional or different advantages may be achieved through specific illustrative implementations. Furthermore, specific illustrative implementations may have some, all, or none of the advantages listed above.

[0053] Example Architecture

[0054] Refer to the accompanying drawings, and specifically refer to... Figure 1 and Figure 2 These figures are example diagrams of data processing environments in which illustrative implementation methods can be carried out. Figure 1 and Figure 2This is merely an example and is not intended to assert or imply any limitation regarding the environment in which different implementations may be carried out. Specific implementations may be based on the following description and many modifications may be made to the depicted environment.

[0055] Figure 1 A block diagram depicts a network in which an illustrative implementation of a data processing system can be implemented. Data processing environment 100 is a computer network in which the illustrative implementation can be implemented. Data processing environment 100 includes network / communication infrastructure 104. Network / communication infrastructure 104 is a medium for providing communication links between various devices, databases, and computers connected together within data processing environment 100. Network / communication infrastructure 104 may include connections such as wired, wireless communication links, or fiber optic cables.

[0056] The client or server are merely example roles of certain data processing systems connected to network / communication infrastructure 104 and are not intended to exclude other configurations or roles of these data processing systems. Servers 106 and 108 are coupled to network / communication infrastructure 104 along with storage unit 110. Software applications can execute on any computer in data processing environment 100. Clients 112 and 114 are also coupled to network / communication infrastructure 104. Client 112 can be a remote computer with a display. Client 114 can be a mobile device configured with an application to send or receive information, such as receiving information from server 106. Data processing systems such as server 106 or server 108, clients (client 112, client 114), data synthesis engine 102, and sensing system 124 can contain data and may have software applications or software tools executing thereon.

[0057] This is merely an example and does not imply any limitations on this architecture. Figure 1 Some components of an example implementation that can be used in the implementation are depicted. For example, the server and client are merely examples and do not imply a limitation on the client-server architecture. As another example, as shown, the implementation can be distributed across several data processing systems and data networks, while another implementation can be implemented on a single data processing system within the scope of the illustrative implementation. The data processing systems (server 106, server 108, client 112, client 114, data synthesis engine 102, sensing system 124) also represent example nodes, partitions, and other configurations suitable for implementing the implementation in a cluster.

[0058] The data synthesis engine 102 may include configuration and code for simulating a physical system based on a process-based model and generating simulated data corresponding to one or more physical state variables of the physical system. In one example, the process-based model represents a saturated or unsaturated (variably saturated) groundwater flow model. In another example, the process-based model represents a vehicle suspension model. In yet another example, the process-based model represents a soil respiration model. These examples are non-limiting, and the disclosed methods can be applied to other process-based models. In some implementations, the engine may manipulate the simulated data by applying calculated sensor performance characteristics of physical sensors to the simulated data to generate one or more simulated sensor responses that more closely approximate the actual outputs of the physical sensors. In some implementations, the engine may generate simulated sensor responses from the simulated data. In some implementations, the engine may generate simulated sensor responses without manipulating the simulated data. The engine may also use the simulated data (including those incorporating simulated sensor responses) and the inputs and / or outputs (simulated data) of the process-based model as training targets to generate a training dataset to train a machine learning model. The engine may further use the trained machine learning model to predict unknown properties of the physical system.

[0059] Sensing system 124 may include one or more physical sensors 122 and a configuration for experimentally determining sensor performance characteristics, which may include one or more combinations of experimentally determined sensor transfer function, impulse response, sensitivity, selectivity, repeatability, uncertainty (noise and / or bias), and spatial weighting. Sensitivity is the minimum value or change in value of a physical quantity that a sensor can detect or resolve. For example, the minimum concentration of nitrate that an ion-selective electrode can record represents the sensitivity of that sensor. Selectivity is the ability of a sensor to distinguish between two physical effects to which it may be sensitive. For example, if the aforementioned ion-selective electrode also has some sensitivity to sulfate, this would represent insufficient selectivity for nitrate.

[0060] Sensor performance characteristics can be any quantitative representation of the accuracy with which a sensor represents or fails to represent physical reality. In other words, sensor performance characteristics describe the potential distortion of the sensor to the physical reality of the system in which it is placed. Sensor performance characteristics can be calculated by characterizing the physical sensor performance in response to one or more controlled experimental conditions to generate one or more transfer functions and probabilistic representations of the physical sensor response. One or more of the simulated data can then be corrupted by modifying the simulated data using the parameter descriptions of the transfer functions and probabilistic representations.

[0061] Client application 120 or any other application such as server application 116 implements the implementation described herein. Any application can synthesize training data or use data from data synthesis engine 102 and predict one or more physical state variables and / or physical system properties of the physical system. The application can also obtain data for predictive analysis from storage unit 110. In some implementations, the data can be stored in an indexable manner, such as in database 118. The application can also execute in any data processing system, such as server 106 or server 108, client 112, client 114, data synthesis engine 102, sensing system 124.

[0062] Server 106, server 108, storage unit 110, client 112, client 114, data synthesis engine 102, and sensing system 124 can be coupled to network / communication infrastructure 104 using wired connections, wireless communication protocols, or other suitable data connections. Clients 112 and 114 can be, for example, mobile phones, personal computers, or network computers.

[0063] In the depicted example, server 106 can provide data, such as boot files, operating system images, and applications, to other data processing systems. Clients 112 and 114 can include their own data, boot files, operating system images, and applications. Data processing environment 100 can include additional servers, clients, and other devices not shown.

[0064] In the depicted example, data processing environment 100 can be the Internet. Network / communication infrastructure 104 can represent a collection of networks and gateways that communicate with each other using Transmission Control Protocol / Internet Protocol (TCP / IP) and other protocols. The core of the Internet is the backbone of data communication links between master nodes or master computers, including thousands of commercial, government, educational, and other computer systems that route data and messages. Of course, data processing environment 100 can also be implemented as multiple different types of networks, such as, for example, intranets, local area networks (LANs), or wide area networks (WANs). Figure 1 This is intended as an example, not as an architectural limitation on different illustrative implementations.

[0065] Among other uses, the data processing environment 100 can be used to implement a client-server environment in which illustrative implementations can be carried out. The client-server environment enables software applications and data to be distributed across a network, allowing applications to function through the interoperability between client data processing systems and server data processing systems. The data processing environment 100 can also adopt a service-oriented architecture, where interoperable software components distributed across a network can be packaged together as a consistent business application. The data processing environment 100 can also take the form of a cloud and employ a service-delivered cloud computing model to enable convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with service providers.

[0066] refer to Figure 2 The figure depicts a block diagram of a data processing system in which illustrative embodiments can be implemented. Data processing system 200 is an example of a computer, such as... Figure 1 Server 106 or server 108, client 112, client 114, data synthesis engine 102, sensing system 124, or computer-usable program code or instructions for implementing the process may be located therein for use in another type of device for illustrative implementation.

[0067] The data processing system 200 is described as a computer by way of example only and is not limited thereto. Without departing from the general description of the operation and function of the data processing system 200 described herein, Figure 1 Other implementations of the device form may modify the data processing system 200, such as by adding a touch interface, and even by removing certain depicted components from the data processing system 200.

[0068] In the depicted example, the data processing system 200 employs a hub architecture including a Northbridge and Memory Controller Hub (NB / MCH) 202 and a Southbridge and Input / Output (I / O) Controller Hub (SB / ICH) 204. A processing unit 206, main memory 208, and a graphics processor 210 are coupled to the Northbridge and Memory Controller Hub (NB / MCH) 202. The processing unit 206 may contain one or more processors and may be implemented using one or more heterogeneous processor systems. The processing unit 206 may be a multi-core processor. In some embodiments, the graphics processor 210 may be coupled to the Northbridge and Memory Controller Hub (NB / MCH) 202 via an Accelerated Graphics Port (AGP).

[0069] In the depicted example, a local area network (LAN) adapter 212 is coupled to the Southbridge and input / output (I / O) controller hub (SB / ICH) 204. An audio adapter 216, a keyboard and mouse adapter 220, a modem 222, a read-only memory (ROM) 224, a universal serial bus (USB) and other ports 232, and a PCI / PCIe device 234 are coupled to the Southbridge and input / output (I / O) controller hub (SB / ICH) 204 via bus 218. A hard disk drive (HDD) or solid-state drive (SSD) 226a and a CD-ROM 230 are coupled to the Southbridge and input / output (I / O) controller hub (SB / ICH) 204 via bus 228. The PCI / PCIe device 234 may include, for example, an Ethernet adapter, an add-in card, and a PC card for a notebook computer. PCI uses a card bus controller, while PCIe does not. The read-only memory (ROM) 224 may be, for example, a flash binary input / output system (BIOS). The hard disk drive (HDD) or solid-state drive (SSD) 226a and CD-ROM 230 can use, for example, an integrated drive electronics (IDE), a serial advanced technology accessory (SATA) interface, or variants such as external SATA (eSATA) and micro SATA (mSATA). The super I / O (SIO) device 236 can be coupled to the southbridge and input / output (I / O) controller hub (SB / ICH) 204 via bus 218.

[0070] Memory such as main memory 208, read-only memory (ROM) 224, or flash memory (not shown) are some examples of computer-usable storage devices. Hard disk drive (HDD) or solid-state drive (SSD) 226a, CD-ROM 230, and other similar available devices are some examples of computer-usable storage devices that include computer-usable storage media.

[0071] The operating system runs on the processing unit 206. The operating system coordinates and provides services to... Figure 2 The data processing system 200 controls various components within it. The operating system can be a commercially available operating system for any type of computing platform, including but not limited to server systems, personal computers, and mobile devices. Object-oriented or other types of programming systems can work in conjunction with the operating system and make calls to the operating system from programs or applications executing on the data processing system 200.

[0072] For use in operating systems, object-oriented programming systems, and applications or programs (such as...) Figure 1The instructions for the server application 116 and client application 120 reside on a storage device, such as a hard disk drive (HDD) or solid-state drive (SSD) 226a in the form of data synthesis code 126, and may be loaded into at least one of one or more memories (such as main memory 208) for execution by the processing unit 206. The processes of the illustrative implementation can be executed by the processing unit 206 using computer-implemented instructions, which may reside in memory, such as, for example, main memory 208, read-only memory (ROM) 224, or one or more peripheral devices.

[0073] Furthermore, in one scenario, the data-synthesized code 126 can be downloaded from the remote system 214c via network 214a, where code 214e is stored on storage device 214g. In another scenario, the data-synthesized code 126 can be pushed to the remote system 214c via network 214a, where code 214e is stored on storage device 214g.

[0074] Figure 1 and Figure 2 The hardware can vary depending on the implementation method. Besides... Figure 1 and Figure 2 The hardware described in the text, or its replacement Figure 1 and Figure 2 The hardware described herein can be replaced with other internal hardware or peripheral devices, such as flash memory, equivalent non-volatile memory, or optical disc drives. Furthermore, the processes described in the illustrative embodiments can be applied to multiprocessor data processing systems.

[0075] In some illustrative examples, the data processing system 200 may be a personal digital assistant (PDA), which is typically configured with flash memory to provide non-volatile memory for storing operating system files and / or user-generated data. The bus system may include one or more buses, such as a system bus, I / O bus, and PCI bus. Of course, the bus system can be implemented using any type of communication structure or architecture that provides data transfer between different components or devices attached to that structure or architecture.

[0076] The communication unit may include one or more devices for sending and receiving data, such as a modem or network adapter. Memory may be, for example, main memory 208 or a cache, such as the cache found in the Northbridge and Memory Controller Hub (NB / MCH) 202. The processing unit may include one or more processors or CPUs.

[0077] Figure 1 and Figure 2The examples depicted and those described above are not intended to imply architectural limitations. For example, in addition to taking the form of a mobile or wearable device, the data processing system 200 could also be a tablet computer, a laptop computer, or a telephone device.

[0078] When a computer or data processing system is described as a virtual machine, virtual device, or virtual component, the virtual machine, virtual device, or virtual component operates in a manner similar to that of the data processing system 200, using virtualized representations of some or all of the components depicted in the data processing system 200. For example, in a virtual machine, virtual device, or virtual component, processing unit 206 is represented as a virtualized instance of all or some of the hardware processing units 206 available in the host data processing system, main memory 208 is represented as a virtualized instance of all or some portions of the main memory 208 available in the host data processing system, and hard disk drive (HDD) or solid-state drive (SSD) 226a is represented as a virtualized instance of all or some portions of the hard disk drive (HDD) or solid-state drive (SSD) 226a available in the host data processing system. In this case, the host data processing system is represented by data processing system 200.

[0079] Figure 3 The data synthesis configuration 300 has been disclosed, which can form Figure 1 Part of the data compositing engine 102 or Figure 1 The data synthesis engine 102. The data synthesis configuration 300 may include a sensor performance module 306, a sensor response simulator 308, a process-based simulator 318, a machine learning engine 312, and a data storage 322. The data synthesis configuration 300 can be used to synthesize training and validation data, and to perform training and testing of machine learning models. In one aspect, the sensor performance module 306 can calculate sensor performance characteristics 304 of the physical sensor 122 used in measuring the state variables of the physical system. The process-based simulator 318 can be used to simulate the physical system according to the process-based model 302 to generate simulated data 320 corresponding to one or more physical state variables of the physical system. In one aspect, the physical system can be a natural system. A natural system is a system that can be studied using natural sciences and appears in nature. These systems can vary in size and can include systems involving the motion of physical bodies, heat transfer, material behavior, etc., where various physical, chemical, materials science, mathematical models, experiments, and observations can be used to understand how these systems operate and how their properties change over time and physical space. Examples may include the solar system, atmosphere, ocean, weather, etc. On the other hand, a physical system is a natural system that has been altered by human intervention, an artificial system, or a combination of natural systems, natural systems that have been altered by human intervention, and artificial systems.

[0080] The computational sensor performance characteristics 304 of the physical sensor can be applied to one or more simulated data 320 to disrupt the simulated data, thereby generating one or more simulated sensor responses 310 that can more closely reflect the actual output of the physical sensor 122. In response to generating sufficient simulated sensor responses 310 (which would otherwise be impractical in the real world due to, for example, physical limitations or a lack of a sufficient number of physical sensors), a machine learning engine 312 can participate in training a machine learning model (not shown) based on the simulated sensor responses 310. Training inputs may include at least one or more of the simulated sensor responses 310, and training objectives may include one or more of the simulated data 320 and / or the inputs to the process-based model 302. This can be helpful because simulated sensor responses are often easier to obtain in the real world for a specific location, as they more readily represent the cumulative effects of real-world sensor limitations and surrounding space on the measured property, despite their limited breadth and applicability. Inputs to simulated data 320 and / or process-based models 302 may be more useful because they are functions of a specific time and / or space, free from unwanted noise / bias and other limitations from sensors, and therefore more difficult to manually measure at multiple locations and / or time periods. This is also because deploying one or more sensors to measure training targets may alter the physical system, thus changing its behavior, or because the training targets may be unmeasurable, for example, due to the absence of sensors capable of measuring them. Therefore, physical and natural systems can be studied accurately and / or precisely, and their properties measured via the machine learning techniques described herein, without being limited by the availability and unavailability of sensors on the measured properties.

[0081] In one aspect, the process-based simulator 318 includes one or more process-based models 302. The process-based model 302 can be a physics-based model, an empirical model, or a combination of both. For example, in an example application concerning soil respiration, the process-based model could be the porous media flow module of the COMSOL multiphysics software model, which includes functionality for modeling single-phase flow in porous media based on Darcy's law. Typically, models such as these and other models with extensive validation, trust, and support can be utilized.

[0082] The process-based model 302 can receive data representing independent variables used to predict the response of the physical system. The data can be the actual input parameters 316 of the physical system, determined based on the problem definition (problem definition 314). Such data may include the domain definition of the physical system, the initial conditions of the physical system, and the static or time-varying boundary conditions of the physical system. In soil respiration phenomena (tunnel detection, see...) Figure 7In the example, the process-based model input data (real input parameter 316) can include spatially variable pressure and flow conductivity properties of subsurface soil (although spatially variable, this property does not change in a single simulation and can therefore be considered an invariant property of the physical system for a specific simulation), depth to impermeable boundaries (e.g., groundwater level), and surface atmospheric pressure variations driving the subsurface response. These input data can be derived from time series measurements by atmospheric pressure sensors, field surveys of subsurface material distribution, or synthetic implementations based on features of observed natural variations. To generate a realistic domain definition, for example, several soil layers can be represented, each with a separate mean air conductivity. Further realism can be represented by adding spatial variability to flow-related soil properties in the model domain using correlated random fields, as is often observed in the field. When specifying correlated random fields to represent the variability of soil properties, different correlation lengths in the vertical and horizontal directions can be used to generate more autocorrelation in the horizontal direction than in the vertical direction, which is typical in real loose geological sediments. In another example, the response of a vehicle suspension to disturbances crossing speed bumps can be modeled, and the vehicle's overload state can be predicted based on the detection of vehicle motion characteristics. In this example, the actual input parameters 316 of the physical system can be determined based on common design goals of passenger vehicle suspensions, such as the amplitude, period, and decay rate of pitch oscillations constrained by road variability and driver input.

[0083] The output data of the process-based model 302 can be simulation data 320 representing the physical state variables of the physical system. These state variables may or may not be measurable by physical sensors 122 within the physical system. Unmeasurable state variables may be unmeasurable due to, for example, a lack of existing sensors capable of measuring properties. Measurable state variables may be measurable by physical sensors, but generating measurable state variables from the physical system may be impractical due to, for example, physical limitations or a lack of a sufficient number of physical sensors to generate enough measurements to train the machine learning model, or because obtaining measurements by installing sensors or acquiring material samples is impractical, or would disrupt the physical system to the point of altering its behavior.

[0084] For example, Figure 3The simulated data X or Y can be unmeasurable and can represent hidden or imperceptible variables of the physical system that a machine learning model might ultimately want to predict or infer. Such data can be referred to as unmeasurable simulated data or unmeasurable state variables. For example, simulated data X could be the water temperature in a river downstream of a nuclear power plant's heat emissions into the river without the plant exerting any influence, allowing for an assessment of compliance with regulatory limits on temperature increases caused by the river. In this example, the temperature increase caused by the power plant might not be obtainable by subtracting the river temperature when the plant is not operating from the river temperature when the plant is operating, since two temperatures cannot be obtained simultaneously at the compliance point in the river, regardless of whether the plant is operating or not. Therefore, simulated data X can represent an additional training objective that can be used during training to provide inference constraints. On the other hand, simulated data Z can be measurable and represent, for example, the temperature of water leaving a cooling tower, which existing sensors are capable of measuring. Simulated data Z can also be referred to as measurable simulated data or measurable state variables. However, in the case of simulated data Z, the measurement may be impractical, or there may be physical limitations to measuring it. For example, in the nuclear power plant example, measuring the amount of water lost through evaporation in the cooling tower or the temperature of the main reactor cooling water entering the main heat exchanger is impractical because a temperature sensor cannot be practically maintained in the radioactive environment of the reactor. Instead, the temperature can be modeled using physical principles applied to reactor operation knowledge. Furthermore, until and unless subsequent limitations of the physical sensor are applied to the simulated data Z, the data may be too pure and does not represent the actual measurements reported by the physical sensor 122 due to its limitations or characteristics. Therefore, the sensor response simulator 308 can use the simulated data Z to generate a corresponding simulated sensor response 310 based on the sensor performance characteristics 304, which is more representative of the measurements that can be generated by the physical sensor 122.

[0085] Back Figure 3 The sensor performance module 306 can be used to characterize the physical sensor 122 and establish its sensor performance characteristics 304. The sensor performance module 306 can experimentally or computationally determine the sensor response characteristics by exposing one or more physical sensors 122 to known controlled conditions and measuring their responses. From this set of measurements, statistical descriptions of the sensor response characteristics, such as sensitivity, selectivity, repeatability, uncertainty (noise and / or bias), and spatial weighting, can be developed. These parametric descriptions of the sensor transfer function can be used to modify the output of a theoretically perfect process model to simulate the quantitative state reported by real sensors as system variables.

[0086] More generally, a theoretically perfect output can be obtained for the process-based model 302, and when physical sensors 122 are used in the real world to measure input or output state variables as input to the ML model, the corresponding theoretical output of the process-based model may be corrupted to simulate the defects that actual physical sensors would introduce in the reported measurements, whether due to transfer characteristics such as the sensor's averaging time or effective probe volume, or due to uncertainties (noise and / or biases) inherent in the conversion of physical properties to electrical signals. This can provide the desired "message" of the input data, which is essential for making the machine learning solution robust to the degree of uncertainty, inaccuracy, and non-reproducibility inherent in real-world inputs.

[0087] The values ​​of physical state variables reported by sensors may deviate from the actual values, or true values, of the corresponding state variables in the physical system in several ways. Deviations between sensor readings and true values ​​can include noise, drift, and limitations on the sensor's ability to accurately resolve the measured property in either space or time.

[0088] Noise is a relatively high frequency variation in the long-term response of a sensor to physical conditions that are frequency-stable relative to noise. The sensor adds unpredictable (i.e., random) residuals to the representation of the true value of the property it measures, and the noise may or may not be autocorrelated over time.

[0089] Bias is the relatively constant long-term residual between the noise that represents the central tendency of sensor values ​​(e.g., mean, median, or mode) and the actual value of the sensed physical property. Bias can also vary gradually over time. The amount by which bias changes over time is also called drift.

[0090] Limitations on resolution can be temporal and / or spatial. Temporal resolution limitations typically manifest as a time lag in the sensor's response to changes in the state of the variable it is measuring. For example, the thermal mass of a temperature sensor may limit its ability to respond to rapid changes in the ambient temperature it is measuring. Similarly, scanning sensor techniques can have a finite scan time associated with each reported value, making the measurement of the instantaneous state of variable properties such as spectral reflectance not instantaneously available.

[0091] Spatially resolution-limited sensors can respond to properties within a finite detection volume, on which the sensor synthesizes, integrates, or averages the values ​​of its reported state variables. The synthesis, integration, or averaging of properties within the detection volume can occur if the spatial distribution of the properties within the detection volume is equal or spatially variable-weighted. For example, the volume of material whose moisture content affects the readings acquired by a dielectric-based moisture sensor will be defined by the geometry of the electric field induced in the material by the sensor during measurement. Because the field strength decreases further away from the sensor, the contribution of material directly in contact with the sensor will have a greater impact on the obtained sensor value than the contribution of material further away from the sensor but still within the electric field.

[0092] The way a sensor averages, integrates, synthesizes, or otherwise combines stimulus variations over a finite time period into a single reading is called the sensor's temporal transfer characteristic. The way a sensor averages, integrates, synthesizes, or otherwise combines spatial variations of properties within its detection volume into a single measurement is called the sensor's spatial transfer characteristic. An interaction may occur between the sensor's temporal and spatial transfer characteristics. The sensor's transfer function is a mathematical model of its temporal and / or spatial transfer characteristics. A sensor may have both a temporal and spatial transfer function, or it may have a spatiotemporal transfer function.

[0093] The time transfer function of a sensor can be established by exposing the sensor to a step change in the state of the physical system property (state variable) being measured by the sensor over time. An example is to suddenly move the sensor from one environment to another, such as from air to water, and observe the rate at which the sensor reading changes in response to the step change. A mathematical model of the sensor's time transfer function is defined by the time constant of the sensor's response. A common definition of the sensor's time constant is the time taken for the sensor output to reach the proportion (1-1 / e) of the instantaneous change (i.e., the step change) in the state variable measured by the sensor.

[0094] The spatial transfer function of a sensor can be established by exposing the sensor to step changes in material properties occurring at a series of discrete distances from the sensing element and recording the changes in the sensor output as a function of the distance to the material change. For example, the spatial transfer characteristics of a dielectric-based humidity sensor can be characterized by progressively repositioning the sensor from a position completely surrounded by the first fluid to a position completely surrounded by a second fluid with a different dielectric constant, by passing the sensing element through an interface formed by two fluids with contrasting dielectric constants (e.g., methanol and vegetable oil). Since the sensor senses the dielectric constant of the material within a finite probe volume defined by the geometry of the electric field emitted by the sensor element, the influence of the second fluid on the dielectric constant measured by the sensor increases as the distance between the sensing element and the fluid interface decreases, until the sensor is completely immersed in the second fluid, at which point the dielectric constant of the second fluid dominates. Furthermore, the influence of the dielectric constant of the first fluid continues to decrease as the sensor is repositioned further away from the fluid interface and into the second fluid. The spatial transfer function of the sensor can be defined as any mathematical function that appropriately transposes the mapping of the fluid dielectric constant to the sensor position into the sensor response of the fluid dielectric constant to the sensor position.

[0095] One approach to establishing (i.e., empirically characterizing and modeling) sensor noise is to expose the sensor to controlled, constant physical conditions while acquiring repeated measurements from it. The variability of the measurements with respect to their average value forms a probability distribution that can be randomly sampled to simulate noise added to a theoretically perfect analog variable. For example, a dielectric sensor can be immersed in a liquid bath with a known dielectric constant, and the sensor output can be repeatedly sampled. The average value of the samples and the residuals from the average value are calculated. The deviation of the time average of the samples from the controlled value of the variable measured by each sensor is the bias of that sensor. The probability distribution of potential sensor biases in a single group of sensors, each an instance of a given sensor model, can be defined by exposing multiple identically manufactured sensors to the same controlled conditions or the same environment and examining the variation in the average values ​​produced between the sensors. For example, one way to determine the probability density function of the bias of an atmospheric pressure sensor is to place multiple identical pressure sensors in a pressure-stable room or chamber, average multiple readings for each sensor, compare the average of each sensor output to obtain the probability distribution of the bias, and examine the residuals of each sensor measurement to its respective average value to obtain the probability distribution of noise. The probability distribution of noise can be obtained by summing the deviation of the average value of each sensor across all sensors, or by summing the deviation of a subset of sensors evaluated in the same or different environments.

[0096] To simulate the effects of real-world sensor bias, a sensor-specific value, randomly drawn once from the bias distribution for each sensor, is added to the process-based model output, which will be measured by the sensor in the physical world. To simulate the effects of real-world sensor noise, for each simulated sensor measurement, a value randomly drawn from the sensor noise distribution is added to each output value generated by the process-based model, which will be measured by the sensor in the physical world. Noise and bias are not mutually exclusive and can be simulated simultaneously by adding both to the process-based model output. The probability distribution of bias can be a Gaussian distribution or any other parametric or non-parametric distribution sufficient to describe the distribution of bias for empirical measurements. The probability distribution of noise can be a Gaussian distribution or any other parametric or non-parametric distribution sufficient to describe the distribution of noise for empirical measurements, and / or noise can be generated in a not-completely random but time-correlated manner that matches or closely mimics the time correlation observed in empirically characterized noise.

[0097] A set of realistic input parameters 316 can correspond to a simulation run, also known as a scenario, and can be used to generate simulation data 320 and simulated sensor responses 310. For example, if simulating soil respiration (physical system) to obtain an output including soil gas pressure (physical state variable) at multiple depths in a porous medium, the realistic input to the driving boundary conditions at the medium surface could be a time series of actual atmospheric pressure measurements obtained from a weather monitoring station. In such a process-based model scenario, the realistic input to the transmissivity of the medium and its spatial distribution within the model domain could be estimates (predictions) of soil porosity and soil permeability obtained by applying a geotechnical transfer function (PTF) to a material classification curve obtained from a well log generated from a geotechnical survey. In a non-limiting example where the physical system is a power plant coupled with a cooling tower and a river as a radiator for cooling the power plant, the realistic input to the process-based model could include the geometry of the watershed (domain definition), historical data of river flow velocity based on upstream measurements (physical state variable), historical data of thermal power generation at the power plant (physical state variable), and historical data of observed weather conditions affecting the performance of the cooling tower (physical state variable). These inputs can be combined in different ways from different time periods to simulate a wide range of operational scenarios that are impractical to accumulate solely from direct experience. Simulated data 320, simulated sensor responses 310, and real input parameters 316 for each scenario can form training examples in a synthetic dataset that can be stored in data storage 322. Each training example has at least one physical system property identified as a training objective. In some implementations, one or more boundary conditions can be identified as training objectives. In some implementations, a domain definition can be identified as a training objective. In some implementations, one or more state variables can be identified as one or more training objectives. Each training example can represent a time series of state variables associated with a training objective. In some implementations, a training example can represent an instance of a time series (e.g., a snapshot of the time series). In some implementations, a scenario can result in multiple training examples, each representing a period of time within the simulated scenario. In some implementations, each training example can represent a separate scenario (e.g., with different initial conditions / boundary conditions / sensor placements, etc.). Data storage 322 can be sampled and used by machine learning engine 312. When a sufficient set of simulations is obtained, machine learning engine 312 can be used to train a machine learning model based on the synthetic dataset, as described below.

[0098] Figure 4 A block diagram of a machine learning engine 400 is shown. The machine learning engine 400 can be... Figure 3 Machine learning engine 312 or Figure 9An example of machine learning engine 904. Machine learning engine 400 can extract data from a synthetic dataset in data storage 422 (e.g., data storage 322 or data storage 916) via data extraction module 402 for training ML model 410. The synthetic dataset can be a combination of inputs to a process-based model, the pure output of a process-based model, and simulated sensor output responses from a sensor response simulator, forming a training, testing, and / or validation dataset for training one or more ML models 410. ML model 410 can be Figure 9 Example of machine learning engine 904. ML model 410 can be... Figure 3 Example of a machine learning engine 312. A trained ML model can predict or estimate properties (domain definitions) or physical state variables (e.g., initial conditions, boundary conditions) that are not measurable (or unmeasurable) in real-world / physical systems.

[0099] Therefore, the data extraction module 402 can, for example, extract data from the data storage 422 and divide the data into training data 406 and validation data 408. The training data 406 and validation data 408 can be stored in the data partition 404. In some embodiments, at least some of the measurable simulated data may have been replaced by the corresponding simulated sensor responses in the data storage 422. In some embodiments, the data extraction module 402 can replace at least some of the measurable simulated data with the corresponding simulated sensor responses before dividing the data. In either case, the training dataset represented by the data partition 404 reflects one or more of the simulated sensor responses 411 (e.g., simulated sensor response 310 and / or simulated sensor response 906). Therefore, the simulated sensor responses can be used as at least a portion of the input 412 (training or validation input). A training objective 414 can be identified in the training dataset represented by the data partition 404. The training objective 414 can be any property of the physical system represented during simulation. The training objective 414 can represent physical system properties, such as domain definitions or static boundary conditions. The training objective 414 can represent physical state variables. Training objective 414 can represent a combination of physical state variables, i.e., multi-objective prediction. Training objective 414 can also represent unmeasurable simulated data. More generally, training data can be used to train the model, while validation data can be used to tune the hyperparameters of the machine learning model and make decisions about the model architecture, such as choosing between different architectures. Furthermore, test data can be used to evaluate the final performance of the machine learning model and estimate its ability to generalize to new, unseen data. Thus, supervised machine learning models of any architecture can be developed and well maintained at an exponentially cheaper and more time-efficient rate, requiring a large enough amount of data not only for training but also for a much more diverse range of sensor training data than is affordable in the real world and in physical environments.

[0100] Now go to Figure 5 An example training architecture 502 is disclosed. The ML model can be a neural network ML model and can be trained using various types of training datasets. According to the illustrative implementation, the training architecture 502 can be configured for machine learning-based recommendation generation. Program code (such as data synthesis code 126) can extract various features 506 from the training data 504. The training data 504 is included in... Figure 4Examples of data in data partition 404. Components of training data 504 have labels L. These features are used to develop a predictor function H(x) or hypothesis, which the program code uses as the ML model 410. In identifying various features in training data 504, the program code may utilize various techniques, including but not limited to mutual information, which is an example of a method that can be used to identify features in the implementation. Other implementations may utilize different techniques to select features, including but not limited to principal component analysis, diffusion mapping, random forests, and / or recursive feature elimination (a powerful method for feature selection). “P” is an available output (e.g., flux, the presence of tunnels in the ground, etc.), which, upon receipt, may further trigger the data processing environment 100 to perform other steps, such as storing instructions. The program code may utilize machine learning ML algorithm 510 to train the ML model 410, including providing weights to the output, so that the program code can prioritize various changes based on the predictor function including the ML model 410. The output can be evaluated using quality metric 508.

[0101] By selecting a different set of training data 504, the program code trains the ML model 410 to identify and weight various features of the physical system. To utilize the ML model 410, the program code obtains (or derives) input data or features to generate an array of at least one or more simulated sensor responses, which are then fed into the input neurons of the neural network. In response to these inputs, the output neurons of the neural network produce an array that includes, for example, temporal and / or spatially relevant properties of the physical system (physical state variables and / or physical system properties) to be simultaneously presented or used. In particular, hidden or imperceptible properties of the physical system may be most useful.

[0102] refer to Figure 6 This figure depicts an example configuration 600 for predicting the properties of physical or natural phenomena using sensor data as input, based on intelligent machine learning. In other words, Figure 6 An example configuration 600 for using an ML model trained using the disclosed techniques in inference mode is described. It can be used... Figure 6 Application 604 is used to make predictions. Application 604 can be any application related to the attributes (physical state variables, properties) of the ML model 410 for which it was trained. Non-limiting and non-exhaustive examples of application 604 include applications for analyzing saturated or unsaturated groundwater flow, applications for analyzing vehicle suspension, applications for analyzing soil respiration (e.g., tunnel detection), etc. Application 604 is Figure 1Examples of server application 116 or client application 120 in the example. Application 604, for example, receives or monitors a set of input data 602 in real time. Input data 602 is relevant to the purpose of application 604. Input data 602 may include actual sensor responses 618, such as actual pressure, actual flow rate through cooling system pipes, actual upstream river temperature, actual volumetric water content of soil columns, etc. Due to the use of actual physical sensors, sensor performance characteristics have already been considered in the actual sensor response 618. In other words, it is not necessary to disrupt the actual sensor response 618.

[0103] In one or more non-limiting embodiments, configuration 600 includes a feature selection component 612. Feature selection component 612 can be configured to drive feature selection of the actual sensor response 618. In a particular embodiment, feature selection component 612 can select, for example, all or part of the actual sensor response 618. In another embodiment, the system (e.g., feature selection component 612) can prioritize certain features over others. In yet another embodiment, feature selection component 612 is configured to generate relevant features. Relevant features are those desired by the ML model 410. In some embodiments, relevant features can be based on content requested from application 604. For example, feature selection component 612 can isolate features from one or more of the actual sensor responses 618. This process of isolating features can be referred to as engineered feature extraction. For example, color saturation or hue can be extracted from an image because those features of the image may be more important to the problem the model is solving. As another example, diurnal variation can be extracted from a time series of subsurface pressure. In other words, feature selection component 612 can use feature extraction to reduce the raw sensor output to features or a set of features used by the model. Therefore, feature selection component 612 enables prediction module 614 to use extracted features from sensor data (e.g., from actual sensor responses 618) instead of raw sensor data. Although illustrated as part of application 604, in some embodiments, feature selection component 612 may operate on the actual sensor responses 618 before they are stored or provided as input data 602.

[0104] Using the extracted features and a trained ML model 410 that has been trained on a large number of different datasets (e.g., datasets in data stores 322 or 916), the prediction module 614 determines the model output 610, such as physical state variables and domain definitions. Output 610 represents any data item output by the prediction module 614. Therefore, the prediction module 614 can also predict one or more of the physical state variables (representing initial or boundary conditions) or domain definitions of the physical system. Boundary conditions can be time-varying. Of course, these examples are not intended to be limiting, and any combination of these and other examples is possible given this description. As an example, the prediction module can be configured to predict at least one of the following: (a) unsaturated groundwater flux based on time-series sensor measurements of soil water content at one or more depths; (b) unsaturated groundwater content based on time-series sensor measurements of soil temperature at one or more depths; (c) unsaturated groundwater flux based on time-series sensor measurements of soil temperature at one or more depths; (d) unsaturated groundwater pressure based on sensor measurements of soil water content at one or more depths; or (e) unsaturated groundwater pressure based on time-series sensor measurements of soil temperature at one or more depths.

[0105] The prediction module 614 may be based on a neural network such as a CNN or RNN, although this is not intended to be limiting. In an illustrative embodiment, the model output 610 may be presented or used by the presentation component or the usage component 606 of the application 604. For example, the presentation component or the usage component 606 may use the model output 610 to inform actions performed on a physical system. Non-limiting examples of actions performed on a physical system include controlling an irrigation system, providing an alert for the suspected presence of a tunnel, logging information related to environmental regulations, etc. In some embodiments, the adaptation component 608 may optionally be configured to receive input from a user to adapt the model output 610 (e.g., t-state variables, domain definitions, conditions), if desired. For example, changing the domain definition proposed by the prediction module 614 causes a recalculation of the output to take the new domain definition into account.

[0106] Feedback component 616 optionally collects user feedback relative to the prediction model output 610. In one embodiment, application 604 is configured not only to compute model output 610 (e.g., state variables, physical system properties) but also to provide input feedback to the user, where the feedback indicates the accuracy of the prediction. Feedback component 616 applies feedback from machine learning techniques to, for example, ML model 410, to modify ML model 410 for better predictions. In an illustrative embodiment, application 604 analyzes the feedback input and strengthens the ML model 410 of prediction module 614. If the feedback is satisfactory or unsatisfactory for the accuracy of the prediction, the parameters of ML model 410 are strengthened or weakened, respectively. In other words, the feedback can be used to further train or refine ML model 410 to improve the quality of model outputs in future predictions.

[0107] refer to Figure 7 The figure depicts a block diagram of an example configuration 708 for tunnel detection based on soil respiration analysis, according to an illustrative embodiment. Configuration 708 is... Figure 3 This is a non-limiting example of configuration 300. In this implementation, the trained ML model is configured to predict tunneling information 710. Of course, this is an example implementation and is not intended to be limiting, as other examples are possible given the description herein. For example, process-based models for natural systems (such as rivers, oceans, space / sky / atmosphere, or other natural systems) can be obtained to compute their temporal and / or spatially relevant properties for training corresponding machine learning models. Process-based models for physical systems can also be obtained accordingly.

[0108] Cross-border tunnels pose a significant threat to national security and are proliferating as physical barriers and surveillance along borders improve. Most cross-border tunnels are discovered by artificial intelligence because reliably detecting tunnels by technical means has previously proven elusive and prone to false positives. The machine learning method disclosed in this paper utilizes a trained machine learning model to identify the presence of nearby tunnels based on changes in subsurface pressure, overcoming the shortcomings of other available techniques, such as active geophysical surveying and imaging techniques, and passive acoustic monitoring. Specifically, the disclosed technique generates training data that can be used to train the model to identify the presence or absence of tunnels based on changes in subsurface pressure. This signal is unaffected by non-tunnel activities (e.g., cultural noise). The effectiveness of using subsurface pressure changes increases with tunnel depth and requires very low communication bandwidth and computational power for collection and processing. Furthermore, this model leverages infrastructure that is inexpensive to build and maintain. The disclosed training model can also adapt to the geological environment and does not require extensive pre-characterization of the local subsurface environment.

[0109] The tunnel detection (soil respiration) configuration 708 employs a model that directly correlates pressure patterns in the ground with tunnel information 710 using machine learning-based smart sensor technology. The process-based model can receive model input data, including, for example, spatially variable pressure and flow conduction properties of the subsurface soil 704 (parameter state values), voids (or not) representing the subsurface space of the tunnel (or not) 712 (domain definition), and atmospheric boundary conditions, such as variations in atmospheric pressure 702 above the ground that drive the subsurface response. This input data can be derived from time series measurements by atmospheric pressure sensors, field surveys of subsurface material distribution, or synthetic implementations based on observed natural variations. This implementation recognizes that atmospheric pressure in the ground can vary from one location to another, and the pressure at a first location M closer to the atmosphere may differ from the pressure at a second location N below the first location. Pressure variations may decrease with increasing depth until a relaxation depth, where virtually no detectable changes occur. If the tunnel is near the second location N, the pressure in the tunnel may be closer to atmospheric pressure, which could affect the pressure at location N due to the movement of air from the tunnel through the ground to location N. By feeding the pressure into a machine learning algorithm (tunnel detection module 706), tunnel information 710 can be predicted, such as whether one or more tunnels exist nearby, or the distance to a tunnel can be predicted as a regression, and / or the probability that a tunnel exists nearby.

[0110] In one implementation, a virtual observation training / validation pair can be generated from data storage 322 using simulated sensor responses 310 and inputs and / or outputs of a process-based simulator 318 via a tunnel detection (soil breathing) configuration 708. Inputs may include atmospheric pressure 702 and spatially variable pressure and flow conduction properties 704 of the subsurface soil. Because spatially variable pressure and flow conduction properties do not change during a particular simulation, they can be considered as properties used for simulation (e.g., domain definitions). Other simulations can be run with different spatial distributions (different domain definitions). The pure, undestructed output of the process-based simulator 318 (e.g., physical state variables in simulation data 320) may include, for example, pressures in the ground at various depths. The simulation outputs (e.g., pressures in the ground at various depths) can be used to generate the simulated sensor response 310. The simulated sensor response 310 can better match values ​​that will be measured by actual physical sensors placed at various depths in the ground. Based on the virtual observation training / validation pair, a tunnel detection module 706 can be trained to predict tunnel information 710.

[0111] Now go to Figure 8A routine 800 for modeling complex natural phenomena is disclosed. In routine 800, a data synthesis engine 102 experimentally determines or computes sensor performance characteristics of a physical sensor 122 used in measuring state variables (variable properties) of a physical system in block 802. In block 804, the data synthesis engine 102 simulates the physical system according to a process-based model (such as model 302) to generate simulated data (e.g., simulated data 320) corresponding to one or more physical state variables of the physical system. In block 806, the data synthesis engine 102 applies the experimentally determined or computed sensor performance characteristics of the physical sensor (e.g., sensor performance characteristics 304) to the simulated data to corrupt the simulated data, thereby generating one or more corresponding simulated sensor responses that more closely approximate the actual output that the physical sensor will provide. Simulated sensor response 310 is an example of such a simulated sensor response. In block 808, the data synthesis engine 102 generates a training dataset to train a machine learning model, such as... Figure 4 The ML model 410. The training dataset can be generated from simulated data, which includes (reflects) simulated sensor responses. In some implementations, one or more state variables in the simulated data can be identified (used) as training objectives. In some implementations, one or more properties of the physical system, such as measurements obtained from actual physical sensors and / or human-generated or human-obtained inputs to a process-based model, can be identified (used) as training objectives. In some implementations, the training objectives can represent time series of the properties. In some implementations, each training objective can be considered a training example.

[0112] In one aspect of routine 800, simulations of simulated data are performed for multiple scenarios, and sensor response characteristics are applied to the simulated data, where each scenario is defined by multiple input parameters representing the domain definition, initial conditions, and / or boundary conditions of the physical system. This creates a collection of simulated data and simulated sensor responses along with their corresponding input parameters, which together form a synthetic dataset that can be used to train a machine learning model.

[0113] In one aspect of routine 800, the calculated sensor performance characteristics include properties selected from a list consisting of experimentally determined sensor transfer functions, sensitivity, selectivity, repeatability, uncertainties (noise and / or bias), and spatial weighting. The calculated sensor performance characteristics can be calculated by characterizing the physical sensor performance in response to one or more controlled experimental conditions to generate one or more transfer functions and probabilistic representations of the physical sensor's response. The calculated sensor performance characteristics can be used to manipulate simulated data. For example, simulated data can be corrupted by modifying the parameters of one or more transfer functions and probabilistic representations. More specifically, the simulation data can be corrupted as follows: First, a transfer function is applied to the simulation data. This transfer function, in mathematical terms, describes the sensor's sensitivity to its surrounding medium, which may be homogeneous or non-homogeneous relative to the property the sensor is measuring. This results in the sensor value not being exactly the value of the sensed property of the medium at the sensor's exact location, but rather the spatial average of the property values ​​in the medium near the sensor's location. Then, an realization of the sensor bias, derived from a probability distribution of biases associated with multiple sensors and invariant for the specific sensor being simulated, is added. Furthermore, an realization of the noise, derived from a probability distribution of noise exhibited by a particular model or type of sensor, is added to each simulation data point. This process derives a comprehensive realization of the sensor uncertainty.

[0114] In another aspect of routine 800, the analog sensor response includes physical properties of state or mass transfer, which may include one or more of temperature, flow rate, pressure, concentration, and measurements of wave phenomena (such as amplitude, audio frequency, radio waves, light waves, etc.).

[0115] In one aspect of routine 800, a machine learning model is trained for multi-objective prediction based on multiple simulated properties representing physical state variables. A subset of the multiple simulated physical state variables is used as a training objective to train the machine learning model for inference, and a training loss function is formulated to penalize estimation errors across multiple training objectives. In other words, the machine learning model...

[0116] In another aspect of routine 800, a trained machine learning model is used to infer patterns to predict one or more real-world variable properties (physical state variables) of a physical system based on inputs from responses to one or more real sensors. Furthermore, the machine learning model and training method can be designed to predict one or more properties of the domain and initial and / or boundary conditions.

[0117] In another aspect of routine 800, the output of the machine learning model can be implemented or further used in practical applications to measure the temporal and / or spatially relevant properties (physical state variables) of physical or natural phenomena (e.g., tunnels, power plants, rivers, and vehicle systems) that would otherwise be extremely difficult or impossible to measure in the real world with physical sensors. This is due to (i) the lack of any commercially available or even feasible sensors configured to accurately measure said properties, (ii) the impracticality of using available sensors at multiple locations (e.g., thousands of locations) or at locations that are difficult or impossible to reach (e.g., at certain depths underground), and (iii) the time and computational costs that would otherwise require a manual solution. For example, as a result of the prediction (an estimate of the properties of the physical system), routine 800 may include notifications (e.g., initiating, providing, etc.) of actions performed on the physical system. In some implementations, the action may relate to providing a user interface related to the prediction. For example, the action may include providing information related to environmental regulations in the user interface, or providing an alert for a suspicious tunnel, etc. Administrators can use the user interface to take further action. In some implementations, the action may include the execution of one or more additional aspects of routine 800, such as initiating a remedial action as a result of a prediction. Therefore, the physical system can be studied accurately and / or precisely, and its properties can be measured without being limited by the availability or unavailability of the measured properties by sensors.

[0118] groundwater

[0119] The disclosed implementations can be used to quantify soil water fluxes, which is important for understanding hydrobiogeochemical processes, irrigation planning and sustainable water use, climate monitoring and forecasting, and regulatory applications. Past methods have not provided practical solutions that can be used at scale. For example, weighing lysimeters are invasive, expensive, and maintenance-intensive. As another example, thermal pulse technology is power-intensive and presents calibration challenges. As yet another example, capillary cores require soil water properties and are maintenance-intensive. Current time-series techniques for volumetric water content cannot quantify fluxes or subsidence because water can flow without changing volumetric water content, subsidence terms (root uptake and deep drainage) are not directly observable (i.e., cannot be measured by physical sensors), and soil hydraulic behavior, a key physical system property, is lacking. The implementations can be used to generate training data that allows machine learning models to learn soil hydraulic behavior by observing stored curves over time.

[0120] Figure 9 The data synthesis configuration 900 has been disclosed, which can form Figure 1 Part of the data compositing engine 102 or Figure 1 The data synthesis engine 102. The data synthesis configuration 900 may include a process-based unsaturated groundwater flow simulator 912, a machine learning engine 904, and a data storage 916. The data synthesis configuration 900 can be used to synthesize training and validation data, and to perform training and testing of machine learning models. The process-based unsaturated groundwater flow simulator 912 can be used to simulate a physical system according to a process-based unsaturated groundwater flow model 902 to generate simulation data 914 corresponding to one or more physical state variables of the physical system (unsaturated porous medium). In one aspect, the physical system can be a natural system. A natural system is a system that can be studied using natural sciences and that appears in nature. These systems can vary in size and can include systems involving the motion of physical bodies, heat transfer, fluid flow, material behavior, etc., where various physical, chemical, materials science, mathematical models, experiments, and observations can be used to understand how these systems operate and how their properties change over time and physical space. Examples may include the solar system, soil, atmosphere, ocean, weather, etc. In another aspect, a physical system is a man-made natural system, an artificial system, or a combination of natural systems, man-made natural systems, and artificial systems.

[0121] In order to generate sufficient simulated sensor responses 906, which would otherwise be impractical in the real world due to, for example, physical limitations or a lack of a sufficient number of physical sensors, the machine learning engine 904 may participate in training a machine learning model based on the simulated sensor responses 906. The training input may include at least one or more of the simulated sensor responses 906, and the training objective may include at least the input of simulated data 914 and / or a process-based unsaturated groundwater flow model 902.

[0122] In one aspect, the process-based unsaturated groundwater flow simulator 912 includes one or more process-based unsaturated groundwater flow models 902. The process-based unsaturated groundwater flow model 902 can be a physics-based model, an empirical model, or a combination of both. For example, in an example application of soil respiration, the process-based model could be a version of the MODFLOW (Modular Three-Dimensional Finite Difference Groundwater Flow Model) groundwater flow model that solves the Darcy equations, modified to model airflow dynamics rather than saturated groundwater flow. In another example of unsaturated groundwater flux, it could be the generally accepted HYDRUS (i.e., a hydrological model analyzing water flow and solute transport in variable saturated media, e.g., “HYDRUS-1D” simulating the movement of water, heat, and solutes in a one-dimensional variable saturated medium), which solves the Richardson-Richards equations. Yet another example of a groundwater flow model is the COMSOL Multiphysics Groundwater Flow Module. Typically, models with extensive validation, trust, and support can be utilized.

[0123] The process-based unsaturated groundwater flow model 902 can receive data representing independent variables used to predict the response of the physical system. The data can be the actual input parameters 910 of the physical system, determined based on the problem definition (problem definition 908). Such data may include the domain definition of the physical system, the initial conditions of the physical system, and the static or time-varying boundary conditions of the physical system. For example, in the soil groundwater flux estimation problem, the input data for the process-based model may include, for example, the soil type, and a general record modified to realistically represent the irrigation schedule of the simulated natural physical system (see [link to relevant documentation]). Figure 11 More specifically, input data can include various static vertical distributions of soil hydraulic properties obtained from the Natural Resources Conservation Service Soil Survey Geography (NRCSSSURGO) database, laboratory analysis from soil cores, and measurements from soil pits or in-situ profiling sensors; combined with various measured, interpolated, or synthetic time-variable irrigation or precipitation patterns; and further combined with crop root uptake data or crop definitions as supplementary and integrated inputs to various user-defined root depth intervals and crop root uptake models. To generate realistic domain definitions, for example, several soil layers, each with its own independent soil hydraulic properties, can be represented. Further realism can be achieved by adding spatial variability to the flow-related soil properties in the model domain using correlated random fields, as is often observed in the field. When specifying correlated random fields to represent the variability of soil properties, different correlation lengths in the vertical and horizontal directions can be used to generate more autocorrelation in the horizontal direction than in the vertical direction, which is typical in real-world stratified soil environments.

[0124] The output data of the process-based unsaturated groundwater flow model 902 can be simulation data 914, representing the temporal and / or spatially relevant properties (physical state variables) of the physical system. These properties may or may not be measurable by physical sensors 122 within the physical system. Unmeasurable properties may be unmeasurable due to, for example, a lack of existing sensors capable of measuring the property. Measurable properties may be measurable by physical sensors, but generating measurable properties from the physical system may be impractical due to, for example, physical limitations or a lack of a sufficient number of physical sensors to generate enough measurements to train a machine learning model, or because obtaining measurements by installing sensors or acquiring material samples would be so disruptive to the physical system that it alters the research behavior of the physical system.

[0125] For example, Figure 9 Simulated data X or Y can be unmeasurable and can represent hidden or imperceptible properties of a physical system that a machine learning model may ultimately want to predict or infer. Such data can be referred to as unmeasurable simulated data or unmeasurable state variables. For example, simulated data X could be the volumetric flux of water through a unit volume of soil at a given time and a given location in the soil, which no existing sensor can measure. Thus, simulated data X can represent additional training objectives, such as pore water matrix potential, i.e., pressure, that can be used to provide inference constraints during training. On the other hand, simulated data Z can be measurable and represents, for example, the volumetric water content per unit volume of soil in a physical system, which can be measured by existing sensors. Simulated data Z can also be referred to as measurable simulated data or measurable state variables. However, in the case of simulated data Z, measurement may be impractical, or there may be physical limitations to measuring it, such as the impracticality of measuring the volumetric water content of a given volume of soil at several depths and locations (e.g., hundreds or thousands of these).

[0126] exist Figure 10 The image shows a graph illustrating the exemplary volumetric water content attribute 1002 as purely simulated data 914 using a HYDRUS process model. This can be converted into one or more other simulated sensor responses that approximate actual sensor responses for training machine learning models. This is an example output of a process-based model (e.g., Figure 12 (Output of 1202). In Figure 10 In the example, each trace on the graph represents a time series of the volumetric water content 1002 (VWC) (state variable) at a given depth in the soil during a portion of the entire simulation (e.g., a 45-day simulation). 1004a corresponds to a depth of 10 cm, 1006a to a depth of 20 cm, 1008a to a depth of 40 cm, and so on, with 1022a corresponding to a depth of 90 cm. Figure 10The simulation data has not been corrupted, therefore it represents a theoretically perfect output, for example, corresponding to Figure 9 The output of process-based unsaturated groundwater flow model 902 and / or Figure 11 The state variables of the process-based unsaturated groundwater flow simulator 1120. Figure 10 The example represents a time series of state variables (volume water content at different depths at different instances of the time series). At the beginning of the time series, Figure 10 It is shown that for all sensors, the volumetric water content is low, and when water is applied (in a later instance of the time series), the water first reaches a shallow depth (1004a) and finally travels to the deepest depth (1022a). Implementations can generate training examples using the entire time series, a portion of the time series, or an instance of the time series. As described herein, the simulated data can be corrupted when used in training examples.

[0127] A set of real-world input parameters 910 may correspond to a simulation or scenario and may be used to generate simulation data 914 and simulation sensor responses 906. Examples of real-world inputs for a process-based model of unsaturated groundwater infiltration and drainage may include irrigation scheduling records (e.g., as part of a domain definition). Another real-world input for a process-based model of unsaturated groundwater infiltration and drainage may be soil hydraulic properties obtained from SSURGO (Soil Survey Geographic Database) by cross-referencing USDA Natural Resources Defense Council (NRCS) soil maps (e.g., as part of a domain definition). The simulation data 914, simulation sensor responses 906, and real-world input parameters 910 for each scenario may be training examples from a synthetic dataset stored in a data store 916. Each training example has at least one physical system property identified as a training objective. In some embodiments, one or more boundary conditions may be identified as training objectives. In some embodiments, a domain definition may be identified as a training objective. In some embodiments, one or more state variables may be identified as training objectives. Each training example may represent a time series of state variables associated with a training objective. In some embodiments, a scenario may result in multiple training examples, each representing a period of time within the simulated scenario. In some implementations, each training example can represent a separate scene (e.g., with different initial conditions / boundary conditions / sensor placements, etc.). Data storage 916 can be sampled and used by machine learning engine 904. When a sufficient set of simulations is obtained, machine learning engine 904 can be used to train a machine learning model based on a synthetic dataset, as described below.

[0128] refer to Figure 11The figure depicts a block diagram of an example configuration 1112 for estimating fluid properties in unsaturated porous media (e.g., vertical soil moisture estimation in the ground), according to an illustrative embodiment. Configuration 1112 is... Figure 3 A non-restricted example of configuration 300. In Figure 11 In the example, a trained ML model (e.g., fluid property estimation module 1108) is configured to predict the movement or rate of movement of fluids (such as water through soil) as fluid property 1116. Fluid property 1116 can be one or more extrinsic properties of the fluid in a porous medium, as well as a set of properties describing its movement and storage. The illustrative implementation recognizes that the movement of water through soil provides more information about the intrinsic properties of the soil than the water saturation level alone, which is an extrinsic property. The illustrative implementation recognizes that, conventionally, there is no means to directly measure fluid properties (such as water flux) under real-world conditions. Groundwater flux is not a measurement for which there is any operationally practical equipment. Instead, groundwater flux is typically studied indirectly through complex mathematical models or by utilizing the relationship between alternative soil characteristics and moisture. It has been recognized that neither of these methods is a practical solution for applications requiring flux estimation at high temporal resolution. Because soil is able to conduct water flow without proportionally changing its stored water content, directly equating flux to changes in water content is unreliable. In addition to estimating flux states, other states of fluids (and similar substances) passing through unsaturated porous media, such as storage states, pressure states, temperature states, and / or nutrient states, can be predicted, as discussed herein. Therefore, examples of significant outputs of the fluid property estimation module 1108, which can be trained individually or via multi-objective prediction for prediction and can be either the output or input of the process-based unsaturated groundwater flow model 902, may include one or more combinations of: (i) groundwater flux, (ii) storage, (iii) pressure, or (iv) nutrients, which will be discussed in further detail below.

[0129] Groundwater flux can be represented as the volume of water passing through or across a vertical reference surface per unit time. Important examples of fluxes used to inform management decisions and / or assess environmental impacts include: deep drainage flux (water flowing downwards through the bottom of the root zone), infiltration (water entering the soil at the surface), and root uptake (the flux of water flowing out of the soil through plants at depth intervals corresponding to the root zone, and which can alternatively be calculated as the difference between infiltration, deep drainage, and storage changes).

[0130] Water storage can be expressed in several ways, including: volumetric water content (VWC) at one or more points in an unsaturated porous medium; a storage curve of VWC as a function of depth; and storage length, which is the VWC integrated over a depth interval, typically corresponding to the root region. Storage length can be expressed in feet or meters (i.e., volume per acre per unit surface area - feet, hence "feet").

[0131] Pressure (matrix potential of water) can also be estimated by the fluid property estimation module 1108. Additional properties (if modeled in the process-based modeling step) can also be predicted, such as nutrient-related properties, including one or more nutrient fluxes and one or more nutrient concentrations. These can together form at least a portion of the possible fluid properties 1116, such as... Figure 11 As shown.

[0132] The inputs to the fluid property estimation module 1108, which can be measured by sensors, include VWC, pressure, temperature, and any combination of other inputs as described herein. VWC at one or more discrete depths and at one or more time points can be measured using, for example, a time domain reflectometer (TDR), a capacitance-based sensor, or other dielectric-based sensors. Pressure (also known as matrix potential) at one or more depths and at one or more time points, measured by a tensiometer, can be used as an input to the fluid property estimation module 1108. Furthermore, temperature at one or more depths and at one or more time points, as measured by any kind of temperature sensor (including, but not limited to, thermistors, thermocouples, and solid-state temperature sensors), can serve as an input to the fluid property estimation module 1108, and the fluid property estimation module 1108 is trained accordingly to predict fluid property 1116.

[0133] Additional inputs to the ML model may optionally include rainfall as a function of time, irrigation as a function of time, air temperature as a function of time, irrigation or precipitation temperature as a function of time, and solar radiation as a function of time (air temperature and solar radiation may affect plant root uptake, which is a component of the mass balance of deep drainage and therefore may provide inference bias when used as training targets in multi-objective prediction).

[0134] Any one or more of the inputs listed above can be used alternatively as outputs in multi-objective training of a machine learning model to provide inference bias to the training. Even if such variables will be measured in the real system of interest and therefore do not need to be predicted, and / or if the estimation or prediction of such variables will not be directly used for any subsequent actions or decisions based on the output of the machine learning model, their presence as outputs during model training can help improve the accuracy and performance of the model in estimating or predicting other output properties on which actions or decisions may be based.

[0135] Examples of process-based models that can be used to simulate the flow (e.g., flux) and storage of groundwater in unsaturated (i.e., seepage) zones of soil, and also to simulate the water-coupled transport of heat and / or chemical components (e.g., nutrients) depending on the installed and invoked modules, include HYDRUS-1D, HYDRUS-3D, MODFLOW Unsaturated Zone Flow Package (UZF1), and COMSOL Multiphysics Porous Media Flow Module. The terms “HYDRUS,” “HYDRUS-1D,” “HYDRUS-3D,” “MODFLOW,” and “COMSOL Multiphysics” may be subject to trademark rights in various jurisdictions worldwide and are used herein only to refer to products or services correctly named as trademarks, provided such trademark rights are possible.

[0136] like Figure 11 As shown, the assumption of configuration 1112 for estimating the vertical fluid properties of unsaturated porous media is that soil hydraulic properties can be implicitly inferred using machine learning based on the observed distribution of vertical fluid over time in response to surface wetting events of varying intensities and durations. Another assumption is that infiltration and depth drainage can be accurately inferred across a range of soils without providing soil-specific hydraulic properties, and this forms the basis for estimating root uptake. It should be recognized that this presents a complex problem with high degrees of freedom, and the method described in this paper can be used to provide a generalized solution.

[0137] Virtual observation training / validation pairs can be generated from data storage 1122 using a vertical fluid property estimation configuration 1112 for unsaturated porous media, simulating sensor responses 1110, and the inputs and / or outputs 1114 of a process-based unsaturated groundwater flow simulator 1120. Inputs may include, for example, records 1102 of irrigation schedules, soil type 1104, crop type 1106, and root absorption pattern 1118. Records 1102 of irrigation schedules, soil type 1104, crop type 1106, and root absorption pattern 1118 are examples of real input parameters 910. Based on the virtual observation training / validation pairs, a fluid property estimation module 1108 can be trained to predict fluid properties 1116.

[0138] In specific experimental tests, the industry-standard, physics-based unsaturated soil flow code (HYDRUS-1D) was used as the process-based model to generate virtual observation training / validation pairs for approximately 1,350 different soil types that fully represent the USDA soil texture triangle. Data was generated under various surface boundary conditions. The dataset introduced as input to the ML algorithm consisted of volumetric water content and its temporal and spatial gradients at multiple depth increments up to 1 meter. It was found that pre-computed gradients added important information to the inference flux, significantly improving performance. Water content measurements were assumed to be performed at 1 cm depth and 1 minute time intervals. Five neural network architectures were tested, with all neural networks optimized using Bayesian hyperparameter optimization via Gaussian processes. Hyperparameter configuration and fine-tuning were performed via Bayesian optimization. Results from a single-objective, rootless surface flux (5 cm depth) estimation neural network demonstrate that the model is adept at capturing flux profiles of various soil types at a 1-minute time resolution. The model achieved a coefficient of determination of 0.82. Using a multi-task model, the coefficient of determination was increased to an impressive 0.88. These results support the project hypothesis that analysis of time-varying water content profiles based on ML can support flux estimation without requiring soil-specific information or knowledge of surface boundary conditions.

[0139] In another specific experimental test, virtual observation training / validation pairs were generated for approximately 45,000 different physical scenarios involving all soils in Fresno, King, and Tuller counties, California, listed in the USDA SSURGO database and associated with California almond orchards. HYDRUS-1D modeled diurnal and seasonal variations in root uptake using a modeled reference evapotranspiration demand and published Feddes parameters to account for the effects of soil water potential, and employed the University of California Cooperative Extension's almond irrigation recommendations for Kern County (with variations imposed around the recommendations across different scenarios) to specify time-varying surface boundary conditions. Based on previous laboratory characterizations of the noise, bias, and spatiotemporal transfer characteristics of each sensor type in response to controlled laboratory conditions, each model in the ensemble simulated 56 days of unsaturated groundwater dynamics, with the final 28 days used to simulate the sensor responses of several commercially available multi-stage VWC sensors. Commercial sensors included the SoilVUE10 (Campbell Scientific, Logan, Utah), Drill and Drop (Sentek, Australia), and EnviroScan (Sentek, Australia). The terms SoilVUE10, Drill and Drop, and EnviroScan may be subject to trademark rights in various jurisdictions around the world and are used herein only to refer to the products or services correctly named by the trademarks, for as long as such trademark rights may exist. Training / validation datasets generated from simulations of each model from the commercial probe were used to optimize hyperparameters and train individual LSTM models to predict, as output, the instantaneous vertical groundwater flux at a 120 cm deep vertical datum (also known as deep drainage), from inputs including previous 24-hour measurements of VWC at each depth supported by the corresponding sensor model. The LSTM models achieved impressive performance, for example, exhibiting a coefficient of determination of 0.91 for a model trained using EnviroScan data as input. See also Figures 13A-13D The diagram shows four graphs comparing the predicted flux 1302 (i.e., depth discharge—flux at a depth of 120 cm) of a machine learning model using a simulated sensor as input to form an LSTM, with the corresponding true flux 1304. These graphs demonstrate the accuracy of the model trained using the disclosed technique. The true flux 1304 was determined in a very limited, controlled environment. There is no way to actually measure the flux in a real-world environment.

[0140] Therefore, through Figure 11The configuration disclosed presents a method / system for estimating fluid properties 1116 based on time series measurements of soil water or fluid content in a vertical profile. Furthermore, the machine learning model / fluid property estimation module 1108 is not necessarily used to predict soil hydraulic properties as an endpoint, merely replacing traditional iterative forward model inversion, but rather infers soil hydraulic characteristics internally, unconstrained by the approximations and assumptions inherent in the derivation of the "governing equations." This more comprehensively overcomes some of the most documented and elusive limitations of conventional methods for simulating groundwater phenomena.

[0141] Now go to Figure 12 The document discloses routines 1200 for modeling complex natural phenomena. In block 1202, a data synthesis engine 102 simulates a physical system as an unsaturated porous medium based on a process-based unsaturated groundwater flow model 902 to generate simulated data corresponding to one or more physical state variables of the physical system. In block 1204, the data synthesis engine 102 generates one or more corresponding simulated sensor responses that approximate the actual outputs that physical sensors would provide. In block 1206, the data synthesis engine 102 generates a training dataset to train a machine learning model, which uses simulated sensor responses as training inputs and simulated data, measurements obtained from actual physical sensors, and / or artificially generated or artificially obtained process-based model inputs as training targets.

[0142] In one aspect of routine 1200, simulations of simulated data and simulated sensor responses are performed for multiple scenarios, each scenario defined by multiple input parameters representing the domain definition, initial conditions, and / or boundary conditions of the physical system. This creates a collection of simulated data and simulated sensor responses, along with the corresponding input parameters, which together form a synthetic dataset that can be used to train a machine learning model.

[0143] In one aspect of routine 1200, a machine learning model is trained for multi-objective prediction based on multiple simulated time and / or space-related attributes (physical state variables), wherein a subset of the multiple simulated time and / or space-related physical state variables is used as training objectives to train the machine learning model for inference, and a training loss function is formulated to penalize estimation errors across multiple training objectives.

[0144] In another aspect of routine 1200, a trained machine learning model is used in inference mode to predict one or more real-time and / or spatially relevant (variable) properties (i.e., physical state variables) of a physical system based on inputs from one or more real-world sensor responses. Furthermore, the machine learning model and training method can also be designed to predict (estimate) one or more properties of a domain (e.g., properties of subsurface soil) as well as initial and / or boundary conditions. For example, in flux estimation, irrigation schedules can be predicted based on the same inputs used to predict deep drainage (i.e., properties of the physical system can be estimated). As another example, the machine learning model can predict the spatial distribution of soil permeability. As another example, the machine learning model can predict (estimate) patterns of water applied in an irrigation system. As another example, the machine learning model can predict patterns of surface infiltration.

[0145] In another aspect of routine 1200, the output of the machine learning model can be implemented or further used in practical applications to measure the time- and / or spatially relevant properties (physical state variables) of unsaturated porous media that would otherwise be extremely difficult or impossible to measure in the real world with physical sensors, due to, for example (i) the lack of any commercially available or even feasible sensors configured to accurately measure said properties, (ii) the impracticality of measuring physical state variables at multiple locations (e.g., thousands of locations) or at locations that are difficult or impossible to reach with available sensors (e.g., at certain depths underground), and (iii) the time and computational costs that would otherwise require a manual solution. For example, as a result of a prediction (an estimate of the properties of the physical system), routine 1200 may include notifications (e.g., initiation, provision, etc.) of actions performed on the physical system. In some embodiments, the action may relate to providing a user interface related to the prediction. For example, the action may include providing an irrigation schedule in the user interface or providing information related to environmental regulations in the user interface. In some embodiments, the action may include performing one or more additional aspects of routine 1200, such as controlling the irrigation system. Therefore, unsaturated porous media such as soil can be studied accurately and / or precisely, and their properties can be measured in agricultural applications without being limited by the availability and unavailability of sensors for the properties being measured.

[0146] Various embodiments of this teaching have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

[0147] While the content and / or other examples considered to be in their best form have been described above, it should be understood that various modifications may be made therein, and the subject matter disclosed herein can be implemented in various forms and examples, and the teachings can be applied to many applications, only a few of which have been described herein. The appended claims are intended to claim protection for any and all applications, modifications, and variations that fall within the true scope of this teaching.

[0148] The components, steps, features, purposes, benefits, and advantages discussed herein are merely illustrative. None of them, and the discussions relating to them, are intended to limit the scope of protection. While various advantages have been discussed herein, it should be understood that not all embodiments are necessarily required to include all advantages. Unless otherwise stated, all measurements, values, grades, positions, amplitudes, dimensions, and other specifications set forth in this specification (including in the appended claims) are approximate, not precise. They are intended to have a reasonable range consistent with the functions they pertain to and the conventions of the art to which they belong.

[0149] Many other implementations are also envisioned. These include implementations with fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. They also include implementations in which components and / or steps are arranged and / or ordered differently.

[0150] This document describes aspects of the disclosure with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0151] These computer-readable program instructions may be provided to a processor of a computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, the instructions create components for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to function in a certain way, such that the computer-readable medium storing the instructions includes an article of manufacture comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0152] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module, segment, or instruction containing one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions mentioned in the blocks may not occur in the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified functions or actions, or using a combination of dedicated hardware and computer instructions.

[0154] While the foregoing has been described in conjunction with exemplary embodiments, it should be understood that the term "exemplary" means only as an example, and not the best or optimal. Nothing stated or described above is intended or should be construed as causing any component, step, feature, object, benefit, advantage, or equivalent to be offered to the public, whether or not it is recited in the claims.

[0155] It should be understood that the terms and expressions used herein have the general meaning consistent with those in the respective fields of investigation and research, unless otherwise specified herein. Relational terms such as "first" and "second" may be used merely to distinguish one entity or action from another, without necessarily requiring or implying any actual such relationship or order between these entities or actions. The terms "comprise," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but may also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further constraints, an element beginning with "a" or "an" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes that element.

[0156] An abstract of this disclosure is provided to allow the reader to quickly determine the nature of this technical disclosure. It should be understood at the time of submission that it is not intended to interpret or limit the scope or meaning of the claims. Furthermore, as can be seen in the foregoing detailed description, various features are combined in various embodiments for the purpose of simplifying this disclosure. The method of this disclosure should not be construed as reflecting an intention to have more features than expressly recited in each claim. Rather, as reflected in the following claims, the subject matter lies in fewer than all features of a single disclosed embodiment. Therefore, the following claims are incorporated herein by reference, wherein each claim is independently claimed as a separate subject matter.

[0157] Clause 1. A method comprising: establishing sensor performance characteristics of a physical sensor used in measuring properties of a physical system; simulating the physical system by providing input parameters to a process-based model to generate simulated data corresponding to state variables of the physical system; replacing at least one state variable with a simulated sensor response in the simulated data by applying the sensor performance characteristics of the physical sensor to at least one state variable corresponding to the property of the physical system to disrupt the at least one state variable; and generating a training dataset comprising the simulated data and input parameters of the process-based model, wherein at least one physical system property or at least one state variable of the input parameters is identified as a training objective.

[0158] Clause 2. The method according to Clause 1 further includes using the training dataset to train a machine learning model to predict the value of the training objective as output.

[0159] Clause 3. The method described in Clause 2, wherein the machine learning model is configured to use sensor measurements of underground pressure as input to predict tunnel information.

[0160] Clause 4. The method according to Clause 2 further includes: using the machine learning model in inference mode after the training; and using the output to notify the action to be performed on the physical system.

[0161] Clause 5. The method described in Clause 4, wherein the actions performed include controlling the irrigation system.

[0162] Clause 6. The method described in Clause 4, wherein the actions performed include warning of the suspected presence of a tunnel.

[0163] Clause 7. The methods described in Clause 4, wherein the actions performed include recording information relating to environmental regulations.

[0164] Clause 8. The method according to any one of Clauses 1 to 7, wherein the physical system property is a static boundary condition.

[0165] Clause 9. The method according to any one of Clauses 1 to 7, wherein the physical system property is a domain definition.

[0166] Clause 10. The method according to any one of Clauses 1 to 7, wherein the at least one state variable identified as the training objective is a variable boundary condition.

[0167] Clause 11. The method according to any one of Clauses 1 to 7, wherein the at least one state variable identified as the training target is a property of the physical system that cannot be measured by physical sensors.

[0168] Clause 12. The method according to any one of Clauses 1 to 11, wherein the process-based model is a physical model or an empirical model.

[0169] Clause 13. The method according to any one of Clauses 1 to 12, wherein the training dataset comprises a plurality of training examples, each training example comprising a time series of the simulated data and a corresponding training objective.

[0170] Clause 14. The method according to Clause 13, wherein the corresponding training target is an instance in the time series.

[0171] Clause 15. The method according to any one of Clauses 1 to 14 further includes performing simulations and replacements multiple times, each of the multiple times corresponding to a scenario in a plurality of scenarios, wherein each scenario is defined by corresponding input parameters representing the domain definition, initial conditions and / or boundary conditions of the physical system, wherein each scenario generates at least one corresponding training example in the training dataset.

[0172] Clause 16. The method according to any one of Clauses 1 to 15, wherein the training dataset further includes measurements from actual physical sensors as training targets.

[0173] Clause 17. The method according to Clause 1, wherein the physical system property of the input parameter is artificially generated or artificially obtained.

[0174] Clause 18. The method according to any one of Clauses 1 to 17, wherein the sensor performance characteristics include experimentally determined sensor transfer function, repeatability, uncertainty, or spatial weighting.

[0175] Clause 19. The method according to any one of Clauses 1 to 17, wherein the sensor performance characteristics include one or more transfer functions calculated by characterizing the physical sensor performance in response to one or more controlled experimental conditions, and the simulation data is corrupted by applying the one or more transfer functions to the simulation data.

[0176] Clause 20. The method of any one of Clauses 1 to 17, wherein the sensor performance characteristics include one or more probabilistic representations of the response of a physical sensor calculated by characterizing the physical sensor performance in response to one or more controlled experimental conditions, and the simulation data is corrupted by applying a synthetic implementation of sensor uncertainties derived from the parameter descriptions of the one or more probabilistic representations to the simulation data.

[0177] Clause 21. The method according to any one of Clauses 1 to 20, wherein the simulated sensor response includes temperature, flow rate, pressure, concentration, flux, flux rate, wave amplitude, or wave frequency.

[0178] Clause 22. The method according to any one of Clauses 1 to 21, the method further comprising: using the training dataset to train a machine learning model to simultaneously estimate multiple training objectives.

[0179] Clause 23. A method comprising: simulating the interaction between a fluid and an unsaturated porous medium by providing input parameters to a process-based model, the simulation generating simulated data corresponding to physical state variables of the unsaturated porous medium, the physical state variables representing properties of the fluid in the unsaturated porous medium; generating simulated sensor responses from the simulated data, the simulated sensor responses approximating actual outputs of physical sensors; generating a training dataset comprising the simulated data reflecting the simulated sensor responses and physical system properties, wherein at least one physical system property or at least one state variable represented in the simulated data is identified as a training objective; and providing a training dataset for training a machine learning model.

[0180] Clause 24. The method according to Clause 23, wherein the physical state variable includes one or more of flux state, storage state, pressure state, or nutrient state.

[0181] Clause 25. The method of Clause 23 or Clause 24, wherein the simulated sensor response comprises: a) volumetric water content, b) fluid pressure measured at one or more depths in an unsaturated porous medium and / or at one or more time points, c) fluid temperature measured at one or more depths in an unsaturated porous medium and / or at one or more time points, d) rainfall as a function of time, e) irrigation as a function of time, f) air temperature as a function of time, or g) solar radiation as a function of time.

[0182] Clause 26. The method according to any one of Clauses 23 to 25, wherein the simulation data includes data corresponding to physical state variables that cannot be measured by physical sensors.

[0183] Clause 27. The method according to any one of Clauses 23 to 26, wherein the simulated data and simulated sensor responses are used for scenarios in a plurality of scenarios, and the method further comprises: performing simulations and generating simulated sensor responses multiple times, wherein each scenario is defined by corresponding input parameters representing a domain definition, initial conditions and / or boundary conditions of a physical system, wherein each scenario generates at least one corresponding training example in a training dataset.

[0184] Clause 28. The method according to any one of Clauses 23 to 27 further includes: generating the training dataset using measurements obtained from actual physical sensors as an additional training objective.

[0185] Clause 29. The method according to any one of Clauses 23 to 28, wherein the machine learning model is configured to predict at least one of: (a) unsaturated groundwater flux measured by time-series sensors based on soil water content at one or more depths; (b) unsaturated groundwater content measured by time-series sensors based on soil temperature at one or more depths; (c) unsaturated groundwater flux measured by time-series sensors based on soil temperature at one or more depths; (d) unsaturated groundwater pressure measured by sensors based on soil water content at one or more depths; or (e) unsaturated groundwater pressure measured by time-series sensors based on soil temperature at one or more depths.

[0186] Clause 30. A method comprising: receiving state variables of a physical system, the state variables corresponding to responses from physical sensors placed in the physical system; receiving physical system properties of the physical system; and obtaining predictions of attributes of the physical system by providing the state variables and physical system properties to a machine learning model, the machine learning model providing predictions of attributes, wherein the predictions are used to inform actions performed on the physical system, and wherein the machine learning model has been trained with a training dataset to provide estimates given the state variables and physical system properties, the training dataset comprising simulated data including simulated sensor responses, input parameters of a process-based model, and attributes identified as training targets.

[0187] Clause 31. The method described in Clause 30, wherein the prediction involves soil properties or soil hydraulic properties.

[0188] Clause 32. The method according to Clause 30, wherein the prediction involves the spatial distribution of soil permeability, the pattern of water applied in the irrigation system, or the pattern of surface permeability.

[0189] Clause 33. The method according to any one of Clauses 30, wherein the actions performed include controlling the irrigation system.

[0190] Clause 34. The method described in Clause 30, wherein the actions performed include warning of the suspected presence of a tunnel.

[0191] Clause 35. The methods described in Clause 30, wherein the actions performed include recording information relating to environmental regulations.

[0192] Clause 36. A computer system comprising: at least one processor; and a memory storing instructions that, when executed by said at least one processor, perform the methods described in Clauses 1 to 35.

Claims

1. A method comprising: The sensor performance characteristics of the physical sensors used when measuring the properties of a physical system; The physical system is simulated by providing input parameters to a process-based model to generate simulation data corresponding to the state variables of the physical system. In the simulated data, the at least one state variable is replaced with a simulated sensor response by applying the sensor performance characteristics of the physical sensor to at least one state variable corresponding to the property of the physical system to disrupt the at least one state variable. and A training dataset is generated, the training dataset including the simulation data and the input parameters of the process-based model, wherein at least one physical system property or at least one state variable of the input parameters is identified as a training target.

2. The method of claim 1, further comprising using the training dataset to train a machine learning model to predict the value of the training target as output.

3. The method according to claim 2, wherein, The machine learning model is configured to use sensor measurements of underground pressure as input to predict tunnel information.

4. The method according to claim 2, further comprising: After the training, the machine learning model is used in inference mode; and The output is used to notify the actions to be performed on the physical system.

5. The method according to claim 4, wherein, The actions performed include controlling the irrigation system.

6. The method according to claim 4, wherein, The actions taken included warning of the suspected existence of a tunnel.

7. The method according to claim 4, wherein, The actions performed include recording information related to environmental regulations.

8. The method according to claim 1, wherein, The physical system properties are static boundary conditions.

9. The method according to claim 1, wherein, The properties of the physical system are defined as domains.

10. The method according to claim 1, wherein, The at least one state variable identified as the training target is a variable boundary condition.

11. The method according to claim 1, wherein, The at least one state variable identified as the training target is a property of the physical system that cannot be measured by physical sensors.

12. The method according to claim 1, wherein, The training dataset includes multiple training examples, each of which includes a time series of the simulated data and a corresponding training objective.

13. The method according to claim 12, wherein, The corresponding training target is an instance in the time series.

14. The method according to claim 1, further comprising: The simulation and the replacement are performed multiple times, each of which corresponds to a scenario in a plurality of scenarios, wherein each scenario is defined by corresponding input parameters representing the domain definition, initial conditions, and / or boundary conditions of the physical system. In this context, at least one corresponding training example is generated for each scenario in the training dataset.

15. The method according to claim 1, wherein, The training dataset also includes measurements from actual physical sensors as training targets.

16. The method according to claim 1, wherein, The sensor performance characteristics include experimentally determined sensor transfer function, repeatability, uncertainty, or spatial weighting.

17. The method according to claim 1, wherein, The sensor performance characteristics include one or more transfer functions calculated by characterizing the physical sensor performance in response to one or more controlled experimental conditions, and the simulation data is corrupted by applying the one or more transfer functions to the simulation data.

18. The method according to claim 1, wherein, The sensor performance characteristics include one or more probabilistic representations of the response of the physical sensor, calculated by characterizing the physical sensor performance in response to one or more controlled experimental conditions, and the simulation data is corrupted by applying a synthetic implementation of sensor uncertainties derived from the parameter descriptions of the one or more probabilistic representations to the simulation data.

19. The method according to claim 1, wherein, The simulated sensor response includes temperature, flow rate, pressure, concentration, flux, flux rate, wave amplitude, or wave frequency.

20. The method according to claim 1, further comprising: The training dataset is used to train a machine learning model to estimate multiple training objectives simultaneously.

21. A method comprising: The interaction between a fluid and an unsaturated porous medium is simulated by providing input parameters to a process-based model. The simulation produces simulation data corresponding to physical state variables of the unsaturated porous medium, which represent the properties of the fluid in the unsaturated porous medium. A simulated sensor response is generated from the simulated data, the simulated sensor response being approximately equal to the actual output of a physical sensor; Generate a training dataset comprising simulated data reflecting the simulated sensor responses and physical system properties, wherein at least one physical system property or at least one state variable represented in the simulated data is identified as a training objective; and The training dataset is provided for training the machine learning model.

22. The method according to claim 21, wherein, The physical state variables include one or more of flux state, storage state, pressure state, or nutrient state.

23. The method according to claim 21, wherein, The simulated sensor response includes: a) Volumetric water content, b) The pressure of the fluid measured at one or more depths in the unsaturated porous medium and / or at one or more time points. c) The temperature of the fluid measured at one or more depths in the unsaturated porous medium and / or at one or more time points. d) Rainfall as a function of time, e) Irrigation as a function of time f) Air temperature as a function of time, or g) Solar radiation as a function of time.

24. The method according to claim 21, wherein, The simulation data includes data corresponding to physical state variables that cannot be measured by physical sensors.

25. The method according to claim 21, wherein, The simulated data and the simulated sensor response are used for scenarios in multiple scenarios, and the method further includes: The simulation is performed multiple times to generate simulated sensor responses, where each scenario is defined by corresponding input parameters representing the domain definition, initial conditions, and / or boundary conditions of the physical system. In this context, at least one corresponding training example is generated for each scenario in the training dataset.

26. The method of claim 21, further comprising: The training dataset is generated using measurements obtained from actual physical sensors as additional training objectives.

27. The method according to claim 21, wherein, The machine learning model is configured to predict at least one: (a) Unsaturated groundwater flux measured by time-series sensors based on soil water content at one or more depths; (b) Unsaturated groundwater content measured by time-series sensors based on soil temperature at one or more depths; (c) Unsaturated groundwater flux measured by time-series sensors based on soil temperature at one or more depths; (d) Unsaturated groundwater pressure measured by sensors based on soil water content at one or more depths; or (e) Unsaturated groundwater pressure measured by time-series sensors based on soil temperature at one or more depths.

28. A method comprising: Receive state variables of a physical system, the state variables corresponding to responses from physical sensors placed in the physical system; Receive the physical system properties of the physical system; and Predictions of the properties of the physical system are obtained by providing the state variables and the properties of the physical system to a machine learning model, wherein the machine learning model provides the predictions of the properties. The prediction is used to notify the actions to be performed on the physical system, and The machine learning model has been trained using a training dataset to provide estimates given state variables and the properties of the physical system. The training dataset includes simulated data, which includes simulated sensor responses, input parameters of the process-based model, and the attributes identified as training targets.

29. The method according to claim 28, wherein, The predictions involve soil properties or soil hydraulic properties.

30. The method according to claim 28, wherein, The predictions involve the spatial distribution of soil permeability, the pattern of water applied in the irrigation system, or the pattern of surface infiltration.

31. The method according to claim 28, wherein, The actions performed include controlling the irrigation system.

32. The method according to claim 28, wherein, The actions taken included warning of the suspected existence of a tunnel.

33. The method according to claim 28, wherein, The actions performed include recording information related to environmental regulations.

34. A computer system, comprising: At least one processor; and A memory for storing instructions that, when executed by the at least one processor, perform the method according to claims 1 to 33.