Method for determining one or more biological and / or chemical properties
By generating synthetic sample data and adjusting the model using quality metrics, the problem of unreliable data-driven model outputs is solved, enabling more reliable monitoring and control of chemical and biological products.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BASF SE
- Filing Date
- 2024-09-25
- Publication Date
- 2026-04-24
AI Technical Summary
Existing data-driven models have unreliable outputs when monitoring and controlling chemical and biological products, and users find it difficult to interpret their predictive driving mechanisms, leading to barriers to use.
By providing sample data to the data-driven classification model, synthetic sample data is generated, and the model is adjusted using classification quality metrics and synthetic quality metrics to ensure the reliability of the output.
It improves the interpretability and output reliability of data-driven models, and reduces the risk of erroneous predictions, especially in monitoring and control in critical areas.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of reliable artificial intelligence for controlling and / or monitoring chemical and / or biological products. It also relates to methods and apparatus for determining biological and / or chemical properties associated with sample data related to the biological and / or chemical product. Background Technology
[0002] Chemical and / or biological products can be monitored in various ways related to one or more chemical and / or biological properties. These (multiple) chemical and / or biological properties can include quality, application characteristics, technical characteristics, physicochemical properties, tissue characteristics, etc. To monitor sample data measured from chemical and / or biological products, data-driven models trained on the sample data can be used.
[0003] However, data-driven models may not be interpretable. The outputs of these models may be unreliable, and users of the system may still not understand what drives a particular output prediction. This can pose a barrier to the use of data-driven models, especially in critical areas such as monitoring chemical and / or biological products. Summary of the Invention
[0004] On the one hand, a method, particularly a computer-implemented method, is disclosed for determining one or more biological and / or chemical properties associated with sample data related to biological and / or chemical products, the method comprising:
[0005] - Provide sample data associated with the biological and / or chemical product.
[0006] - The sample data is provided to a data-driven classification model, which determines one or more biological and / or chemical attributes. The model is then parameterized based on historical sample datasets and the corresponding biological and / or chemical attributes.
[0007] - Synthetic sample data is generated by providing sample data to a data-driven synthesis data generator, and the synthetic sample data is determined based on classified chemical and / or biological properties, wherein the data-driven synthesis data generator is configured to transform the sample data and generate synthetic sample data with respect to one or more biological and / or chemical properties.
[0008] - Provide at least one classification quality metric related to the classification determination and / or at least one synthetic quality metric related to the generation of the synthetic data.
[0009] - Based on the provided classification quality metric and / or synthetic quality metric, provide trigger signals for adjusting the classification model.
[0010] - Optionally, an adapted classification model is provided to classify sample data into one or more biological and / or chemical attributes.
[0011] On another front, the present invention relates to a method, particularly a computer-implemented method, for obtaining and / or adjusting a data-driven classification model for determining one or more biological and / or chemical attributes associated with a biological product and / or a chemical product, wherein the (multiple) biological and / or chemical attributes can be obtained from sample data associated with the biological and / or chemical product, the method comprising:
[0012] - Provide sample data associated with the biological and / or chemical product.
[0013] - The sample data is provided to a data-driven classification model, which determines one or more biological and / or chemical properties. This classification model is parameterized based on historical sample datasets and the corresponding biological and / or chemical properties.
[0014] - Synthetic sample data is generated by providing sample data to a data-driven synthesis data generator, and the synthetic sample data is determined based on classified chemical and / or biological properties, wherein the data-driven synthesis data generator is configured to transform the sample data and generate synthetic sample data with respect to one or more biological and / or chemical properties.
[0015] - Provide at least one synthetic quality metric related to the generation of the synthetic data and, optionally, at least one classification quality metric related to the classification determination, and / or
[0016] - Based on the provided classification quality metrics and / or synthesis quality metrics, trigger signals are provided for adjusting the classification model. The data-driven classification model can be configured to determine one or more biological and / or chemical properties based on sample data.
[0017] On the other hand, this disclosure relates to an apparatus for obtaining and / or adjusting a data-driven classification model for determining one or more biological and / or chemical properties associated with a biological product and / or a chemical product, wherein the biological and / or chemical properties can be obtained from sample data associated with the biological and / or chemical product, the apparatus comprising: a processor configured to perform any of the methods described herein. The data-driven classification model can be configured to determine one or more biological and / or chemical properties based on sample data.
[0018] In another aspect, an apparatus for determining one or more biological and / or chemical properties associated with sample data related to biological and / or chemical products is disclosed, the apparatus comprising:
[0019] - A sample data provision interface configured to provide sample data associated with the biological and / or chemical product.
[0020] - A classification generator configured to feed the sample data to a data-driven classification model, which then determines one or more biological and / or chemical attributes, wherein the classification model is parameterized based on a historical sample dataset and the corresponding one or more biological and / or chemical attributes.
[0021] - A synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator, and to determine the synthetic sample data based on classified chemical and / or biological properties, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to one or more biological and / or chemical properties to generate synthetic sample data.
[0022] - A quality metric provider configured to provide at least one classification quality metric related to the classification determination and / or at least one synthetic quality metric related to the synthetic data generation.
[0023] - A trigger signal provider configured to provide trigger signals for adjusting the classification model based on the provided classification quality metric and / or the synthetic quality metric.
[0024] - Optionally, a model provider is configured to provide the classification generator with an adjusted classification model for classifying sample data into one or more biological and / or chemical attributes.
[0025] In another aspect, an apparatus for determining one or more biological and / or chemical properties associated with sample data related to biological and / or chemical products is disclosed, the apparatus comprising: a processor configured to perform any of the methods described herein.
[0026] In another aspect, a method, particularly a computer-implemented method, for monitoring and / or controlling biological and / or chemical products based on one or more biological and / or chemical properties is disclosed, comprising:
[0027] - Provide sample data associated with the biological and / or chemical product.
[0028] - The sample data is provided to a data-driven classification model, which determines one or more biological and / or chemical properties. This classification model is parameterized based on historical sample datasets and the corresponding biological and / or chemical properties.
[0029] - Synthetic sample data is generated by providing sample data to a data-driven synthesis data generator, and the synthetic sample data is determined based on classified chemical and / or biological properties, wherein the data-driven synthesis data generator is configured to transform the sample data and generate synthetic sample data with respect to one or more biological and / or chemical properties.
[0030] - Provide at least one classification quality metric related to the classification determination and / or at least one synthetic quality metric related to the generation of the synthetic data.
[0031] - Based on the provided classification quality metric and / or the synthesis quality metric, the generated synthetic data is provided as associated with one or more biological and / or chemical properties, wherein the classification is adjusted by training the classification model with the synthetic data and the associated biological and / or chemical properties.
[0032] -In this context, a modified data-driven classification model is provided to classify sample data into one or more biological and / or chemical attributes.
[0033] In another aspect, an apparatus for monitoring and / or controlling biological and / or chemical products based on one or more biological and / or chemical properties is disclosed, the apparatus comprising:
[0034] - A sample data provision interface configured to provide sample data associated with the biological and / or chemical product.
[0035] - A classification generator configured to feed the sample data to a data-driven classification model, which then determines one or more biological and / or chemical attributes, wherein the classification model is parameterized based on a historical sample dataset and the corresponding one or more biological and / or chemical attributes.
[0036] - A synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator, and to determine the synthetic sample data based on classified chemical and / or biological properties, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to one or more biological and / or chemical properties to generate synthetic sample data.
[0037] - A training module configured to provide generated synthetic data as associated with one or more biological and / or chemical properties based on a provided classification quality metric and / or the synthetic quality metric, wherein the classification is adjusted by training the classification model with the synthetic data and the associated biological and / or chemical properties.
[0038] - A model provider configured to supply the classification generator with a tuned data-driven classification model for classifying sample data into one or more biological and / or chemical attributes.
[0039] In another aspect, the use of one or more biological and / or chemical properties identified according to the methods disclosed herein for monitoring and / or controlling biological and / or chemical products is disclosed.
[0040] In another aspect, the use of one or more biological and / or chemical properties identified according to the methods disclosed herein for monitoring and / or controlling the production and / or treatment of biological and / or chemical products is disclosed.
[0041] In another aspect, the use of synthetic data generated according to the methods disclosed herein for training classification models is disclosed. In yet another aspect, a method for generating synthetic sample data according to the methods disclosed herein for controlling and / or monitoring chemical and / or biological products is disclosed.
[0042] In another aspect, the use of synthetic data generated according to the methods disclosed herein for training classification models is disclosed. In yet another aspect, a method for generating synthetic sample data according to the methods disclosed herein for controlling and / or monitoring the production and / or treatment of chemical and / or biological products is disclosed.
[0043] In another aspect, the use of trigger signals generated according to the method disclosed herein for initiating the training of a classification model is disclosed. In yet another aspect, a method for training a classification model according to the method disclosed herein is disclosed, wherein the classification model is configured to classify sample data associated with biological and / or chemical products based on biological and / or chemical properties.
[0044] In another aspect, an adapted classification model generated according to the methods disclosed herein is disclosed for the purpose of determining one or more biological and / or chemical properties and / or for monitoring and / or controlling biological and / or chemical products. In yet another aspect, a method is disclosed for classifying sample data associated with biological and / or chemical products based on biological and / or chemical properties using an adapted classification model.
[0045] In another aspect, a computer element having instructions, such as a computer program product or a machine-readable medium, is disclosed that, when executed on one or more computing nodes or processors, are configured to perform steps of the methods disclosed herein or are configured to be executed by the means disclosed herein.
[0046] On another front, this disclosure relates to a method, particularly a computer-implemented method, for determining biological and / or chemical properties associated with sample data related to biological and / or chemical properties. The method, particularly the computer-implemented method, includes: providing sample data, particularly via a user interface; providing the sample data to a classification model to determine whether the sample data is associated with a biological and / or chemical property, wherein the classification model is trained to provide the sample data and the associated biological and / or chemical property, wherein, in response to providing the classification model with additional sample data associated with a first biological and / or chemical property and receiving a second biological and / or chemical property, the classification model is further trained based on the additional sample data; and in response to determining that the sample data is associated with the biological and / or chemical property, the biological and / or chemical property is provided.
[0047] On the other hand, this disclosure relates to the use of a classification model for determining biological and / or chemical properties, the classification model being trained to be provided with sample data and to provide biological and / or chemical properties associated with the sample data, and to be further trained based on the additional sample data in response to providing the classification model with additional sample data associated with a first biological and / or chemical property and receiving a second biological and / or chemical property, to determine biological and / or chemical properties associated with sample data related to the biological and / or chemical property.
[0048] On the other hand, this disclosure relates to the use of the classification models described herein for determining biological and / or chemical properties.
[0049] On the other hand, this disclosure relates to the use of additional sample data associated with the first biological and / or chemical attribute for further training of the classification model as described herein for determining the biological and / or chemical attribute.
[0050] On the other hand, this disclosure relates to a method, particularly a computer-implemented method, for classifying sample data associated with biological and / or chemical properties based on biological and / or chemical attributes. The method includes: providing sample data, particularly via a user interface; providing the sample data to a classification model to determine whether the sample data is associated with a biological and / or chemical attribute, wherein the classification model is trained to provide the sample data and the biological and / or chemical attribute associated with the sample data; and further training the classification model based on the additional sample data in response to providing the classification model with further sample data associated with a first biological and / or chemical attribute and receiving a second biological and / or chemical attribute; and providing the biological and / or chemical attribute in response to determining that the sample data is associated with the biological and / or chemical attribute.
[0051] On the other hand, this disclosure relates to a method, particularly a computer-implemented method, for generating additional sample data associated with a first biological and / or chemical attribute for a classification model, the method comprising: providing historical sample data; providing the historical sample data to a data generation model to generate the additional sample data, the data generation model being trained to provide sample data in response to receiving the historical sample data; and providing the additional sample data to further train the classification model based on the additional sample data in response to providing the additional sample data to the classification model and receiving a second biological and / or chemical attribute, wherein the classification model is trained to provide the sample data and provide the biological and / or chemical attribute associated with the sample data.
[0052] On the other hand, this disclosure relates to a method, particularly a computer-implemented method, for training a classification model to determine biological and / or chemical attributes associated with sample data related to the same biological and / or chemical attributes. The method includes: providing additional sample data associated with a first biological and / or chemical attribute to the classification model, wherein the additional sample data is generated by: providing historical sample data and generating the additional sample data based on the historical sample data; in response to providing the additional sample data to the classification model and receiving a second biological and / or chemical attribute, training the classification model based on the additional sample data associated with the first biological and / or chemical attribute; optionally, providing the classification model.
[0053] On the other hand, this disclosure relates to a method, particularly a computer-implemented method, for adjusting a data-driven classification model for determining one or more biological and / or chemical attributes associated with a biological and / or chemical product, wherein the (multiple) biological and / or chemical attributes can be obtained from sample data associated with the biological and / or chemical product, and wherein the classification model is parameterized based on a historical sample dataset and corresponding one or more biological and / or chemical attributes, the method comprising:
[0054] - Provide sample data associated with the biological and / or chemical product.
[0055] - Provide the sample data to a data-driven classification model, which will then determine one or more biological and / or chemical properties.
[0056] - Synthetic sample data is generated by providing sample data to a data-driven synthesis data generator, and the synthetic sample data is determined based on classified chemical and / or biological properties, wherein the data-driven synthesis data generator is configured to transform the sample data and generate synthetic sample data with respect to one or more biological and / or chemical properties.
[0057] - Provide at least one synthetic quality metric related to the generation of the synthetic data and, optionally, at least one classification quality metric related to the classification determination, and / or
[0058] - Provide trigger signals for adjusting the classification model based on the provided classification quality metric and / or synthetic quality metric.
[0059] Any disclosures, embodiments, and examples described herein relate to the methods, systems, apparatuses, uses, and computer elements listed above and below. Advantageously, the benefits provided by any embodiments and examples also apply to all other embodiments and examples. Example
[0060] Embodiments of this disclosure will be outlined below through examples and / or embodiments. It should be understood that this disclosure is not limited to the embodiments and / or examples described.
[0061] This disclosure relates to the field of trustworthy artificial intelligence in the chemical and pharmaceutical industries. In this environment, the safe use of AI is particularly relevant, as false positives or unexplained black-box functionality can lead to adverse effects. Specifically, transforming sample data with regard to one or more biological and / or chemical properties to generate synthetic sample data produces a second set of data, which may be counterfactual or adversarial relative to the original sample data, and thereby generates a second instance for examining the results. Further, by providing at least one classification quality metric related to classification determination and / or at least one synthesis quality metric related to the generation of synthetic data, these quality metrics can be used to evaluate the classification results and ensure the reliability of the results. This is particularly advantageous for classification models that classify the chemical and / or biological properties of chemical and / or biological products, as these operations are particularly sensitive, and erroneous results can lead to adverse effects when monitoring and / or controlling chemical and / or biological products.
[0062] One or more biological and / or chemical properties associated with sample data may be related to the characteristics or properties of a chemical and / or biological product. One or more biological and / or chemical properties may be related to the characteristics or properties of a production process and / or treatment of a chemical and / or biological product. One or more biological and / or chemical properties may be related to the characteristics or properties of treatment with or on a chemical and / or biological product. One or more biological and / or chemical properties may be related to characteristics or properties directly or indirectly associated with a chemical and / or biological product. The characteristics of a chemical and / or biological product may be chemical, physical, biological, and / or environmental characteristics. A chemical characteristic may be a characteristic that can only be established by changing one or more chemical structures associated with at least one chemical and / or biological product. Examples of chemical characteristics may be acidity, oxidation state, or reactivity. A physical characteristic may be one of the following: mechanical, electrical, optical, thermal, etc. For example, physical characteristics may include one or more of the following: density, scratch resistance, electrical conductivity, color, absorbency, heat capacity, etc. Biological and / or chemical properties can be obtained from the sample data. Biological and / or chemical properties can be measured from sample data.
[0063] In embodiments, the environmental characteristics of chemical and / or biological products may include at least one of the following: emission data of the chemical and / or biological product, the recoverable content of the chemical and / or biological product, the bio-based content of the chemical and / or biological product, the renewable content of the chemical and / or biological product, chemical and / or biological product declaration data, chemical and / or biological product safety data, the share of chemical and / or biological product to be recycled, the biodegradable share of the chemical and / or biological product, the non-degradable share of the chemical and / or biological product, or a combination thereof. Emission data may include any data related to the environmental footprint. An environmental footprint may refer to an entity and its associated environmental footprint. An environmental footprint may be entity-specific. For example, an environmental footprint may be associated with: chemical and / or biological products, companies, processes such as manufacturing processes, raw materials or base substances, chemical and / or biological products or materials, components, component assemblies, final products, combinations thereof, or additional entity-specific relationships. Emission data may include data related to the carbon footprint of the chemical and / or biological product. Emissions data may include data related to greenhouse gas emissions, such as those released during the production of chemical and / or biological products. Greenhouse gas emissions may include, for example, emissions of carbon dioxide (CO2), methane (CH4), nitrous oxide (N2O), hydrofluorocarbons (HFCs), perfluorocarbons (PFCs), sulfur hexafluoride (SF6), nitrogen trifluoride (NF3), and combinations thereof, as well as other emissions. Emissions data may include data related to greenhouse gas emissions generated by the entity or company's own operations (production, power plants, and waste incineration). Scope 2 includes emissions generated from the production of energy supplied externally. Scope 3 includes all other emissions generated along the value chain. Specifically, this includes greenhouse gas emissions from raw materials obtained from suppliers. Product carbon footprint (PCF) is the sum of greenhouse gas emissions and removals generated by consecutive and interrelated process steps associated with a specific product. Cradle-to-gate PCF aggregates greenhouse gas emissions based on selected process steps: from resource extraction to the product leaving the company's plant gates. This type of PCF is called a partial PCF. In order to achieve such aggregation, each company that supplies any product must be able to provide the contribution of each of its products to the scope 1 and scope 2 of the PCF as accurately as possible, and obtain reliable and consistent PCF data for the energy purchased (scope 2) and its raw materials (scope 3).
[0064] Sample data associated with biological and / or chemical products can be correlated with measurement data representing the characteristics or properties of the production process of the chemical and / or biological products. Sample data can be correlated with measurement data representing the characteristics or properties of the production process of the chemical and / or biological products. Sample data can be correlated with measurement data representing the characteristics or properties of the treatment with or involving chemical and / or biological products. Sample data can be correlated with measurement data representing the characteristics or properties directly or indirectly associated with the chemical and / or biological products.
[0065] A data-driven classification model is used to determine one or more biological and / or chemical attributes. One or more biological and / or chemical attributes can be used as classifiers in the data-driven classification model. The data-driven classification model can be configured to ingest sample data at an input layer and determine one or more classifiers associated with one or more biological and / or chemical attributes. A biological and / or chemical attribute can be used as a binary classifier, i.e., a true classifier or a false classifier representing the presence or absence of the biological and / or chemical attribute. At least two or more biological and / or chemical attributes can be used as classifiers. The classification model can be parameterized based on a historical sample dataset and corresponding one or more biological and / or chemical attributes. The classification model can be trained using a historical sample dataset and corresponding one or more biological and / or chemical attributes. This historical dataset can include multiple sets of input-output pairs, which include sample data and corresponding one or more biological and / or chemical attributes. The classification model can be trained based on input-output pairs including sample data and corresponding one or more biological and / or chemical attributes. The classification model can include a neural network or a deep neural network architecture, which includes multiple layers, such as an input layer, an output layer, and hidden layers. Classification models can include specialized models that can be trained to monitor and / or control specific types of chemical and / or biological products based on one or more biological and / or chemical properties. Classification models can be based on neural network architectures suitable for a specific purpose. Since the essence of this disclosure applies to any classifier model suitable for classifying one or more biological and / or chemical properties based on sample data associated with chemical and / or biological products, any suitable neural network architecture (such as CNN, RNN, LSTM, etc.) can be applied. It should be understood that those skilled in the art can select an appropriate architecture based on the task to be performed by the classification model.
[0066] Data-driven synthetic data generators can be associated with counterfactual or adversarial synthetic data generators. A synthetic data generator can determine synthetic sample data by transforming sample data with respect to one or more biological and / or chemical properties based on classified chemical and / or biological attributes, and by generating synthetic sample data. A data-driven synthetic data generator can include a generative model adapted to transform sample data with respect to one or more biological and / or chemical attributes and generate synthetic sample data. Data-driven synthetic data generators can include flow-based models, autoencoders, or generative adversarial networks (GANs), particularly generative models associated with GANs. Furthermore, data-driven synthetic data generators can include diffusion models. Diffusion models can be based on adding noise to training data. Diffusion models can include latent variable models that use fixed Markov chains to map to a latent space. Diffusion models can map inputs to a latent space and then to an output space via diffusion. GANs, particularly generative models associated with GANs, can be trained to generate synthetic sample data by mapping numerical representations associated with synthetic sample data, particularly noise associated with synthetic sample data, to synthetic sample data. Synthetic sample data can be fed to a detector associated with a GAN. The detector can be trained, particularly in conjunction with a generative model, to determine whether sample data is likely real or synthetic. GANs, and especially generative models associated with GANs, can be trained to generate synthetic sample data that the detector identifies as real. Stream-based models can include normalized streams.
[0067] Classification quality metrics related to classification determination can be statistical measures of classifier performance or prediction accuracy. Classification quality metrics can represent confidence intervals, classification error, or classification accuracy. Classification quality metrics can be correlated with classification accuracy or classification error as a proportion or ratio. Classification quality metrics can represent the proportion of correct or incorrect predictions made by a classification model.
[0068] Synthetic quality metrics associated with synthetic data generation can be related to synthetic sample data, particularly counterfactual sample data. Adjusting the classification model can include training the model using synthetic sample data and the corresponding classifier.
[0069] In one embodiment, one or more biological and / or chemical attributes are associated with characteristic features embedded in sample data. In this context, characteristic features can refer to features that a classification model can generate or select as a classifier based on one or more biological and / or chemical attributes. In the same context, characteristic features can refer to features that a human expert can select as a classifier based on one or more biological and / or chemical attributes. In the context of interpretable artificial intelligence, the characteristic features that a classification model can use to select a classifier may be related to or overlap with the characteristic features that a human expert can use to select a classifier.
[0070] In another embodiment, sample data is provided via a measurement interface configured to acquire sample data as measured by a sensor. The sensor measuring the sample data can include any sensor suitable for measuring signals that include characteristics or properties directly or indirectly related to the chemical and / or biological product. Chemical and / or biological products can be monitored and / or controlled based on classified chemical and / or biological properties. Such monitoring and / or control can be directly or indirectly related to the chemical and / or biological product. For example, the production process of the chemical and / or biological product, the properties of the chemical and / or biological product, and / or the treatment of the chemical and / or biological product or the handling of the chemical and / or biological product can be controlled and / or monitored.
[0071] In another embodiment, generating synthetic sample data includes providing the sample data to a data-driven counterfactual generator. The counterfactual generator can be configured to transform the sample data into a representation (e.g., a latent space representation) encoding characteristic features embedded within the sample data. The counterfactual generator can also be configured to transform the sample data into a numerical representation associated with the sample data. The numerical representation and / or the representation encoding characteristic features can be associated with a space different from that of the sample data. Therefore, the format of the numerical representation associated with the sample data and / or the format of the representation encoding characteristic features can differ from the format of the sample data. In particular, the numerical representation and / or the representation encoding characteristic features can be associated with a dimension and / or a smaller value than the dimension and / or the value associated with the sample data.
[0072] A counterfact generator can be configured to determine counterfactual sample data based on changes in classified biological and / or chemical properties, preferably associated with synthetic sample data generated from sample data. Changes in classified biological and / or chemical properties can refer to variations in the classified biological and / or chemical properties. Specifically, changes in classified biological and / or chemical properties can refer to chemical and / or biological properties associated with synthetic sample data that differ from those associated with the same sample data. Therefore, changes in classified biological and / or chemical properties can refer to alterations in biological and / or chemical properties based on sample data, particularly by generating synthetic sample data from sample data.
[0073] Counterfactual generators can include pre-trained generative models. For transformations, the counterfactual generator can be configured to map sample data leading to a specific classification of biological and / or chemical properties to a latent representation, such as a lower-dimensional representation. The counterfactual generator can be configured to determine, in the latent space, a counterfactual latent representation corresponding to the minimum distance that causes a change in the classifier's prediction. In this context, the change in the classifier includes a change from the original biological and / or chemical property to at least one different biological and / or chemical property. In other words, in the latent space, a counterfactual latent representation corresponding to at least one different biological and / or chemical property and a minimum distance to the decision boundary between the original biological and / or chemical property and at least one different biological and / or chemical property can be determined. The counterfactual generator can be configured to map the counterfactual latent representation to counterfactual sample data. This allows the generation of counterfactual sample data that resembles the characteristic features embedded in the sample data rather than noise, compared to adversarial sample data. Thus, the counterfactual sample data can be used to model the decision strategy of the classification model.
[0074] Counterfactual sample data can be generated by transforming the sample data into a representation encoding characteristic features using a counterfactual generator trained on sample data and corresponding biological and / or chemical properties, and an augmented data-driven classification model initialized with random weights. The augmented data-driven classification model initialized with random weights can include a data-driven classification model trained on sample data with different weights. The augmented data-driven classification model can be trained and / or configured based on the sample data and the corresponding biological and / or chemical properties determined by the data-driven classification model, particularly after initializing the augmented data-driven classification model with random weights. Therefore, the architecture of the augmented data-driven classification model can correspond to and / or be the architecture of the data-driven classification model. The augmented data-driven classification model can be configured to determine changes in the classified biological and / or chemical properties based on variations in the sample data, particularly stepwise variations. In particular, determining changes can preferably include determining the degree of change associated with the determination of the biological and / or chemical properties associated with the sample data, preferably by the augmented data-driven classification model. In this context, the sample data can include provided sample data and / or synthetic sample data based on and / or obtained from the provided sample data. Enhanced data-driven classification models can be configured to receive sample data and / or synthetic sample data, particularly progressively, and to determine changes in classified biological and / or chemical attributes associated with the sample data and / or synthetic sample data. Based on these changes in classified biological and / or chemical attributes, particularly progressive changes, the data-driven classification model can be configured to determine the confidence level or class probability of the determined biological and / or chemical attribute. Once the determined change in biological and / or chemical attribute corresponds to a predefined confidence level and / or target confidence level, the sample data associated with the determined change in biological and / or chemical attribute or class change can be counterfactual sample data and / or can be provided. Once the predefined confidence level and / or target confidence level for the change can be reached, the sample data associated with the determined and / or predicted change in biological and / or chemical attribute or class change can be counterfactual sample data and / or can be provided. Counterfactual determination can be achieved with greater reliability by using a counterfactual generator trained on sample data for confidence or class probability determination and an enhanced data-driven classification model initialized with random weights for gradient determination. This achieves a quality metric that allows positive feature evaluation by assessing the correctness (positive) of features, rather than negative feature evaluation by assessing the incorrectness (non-negative) of features. For the interpretability of AI, the enhanced reliability of counterfactual determination makes positive evaluations superior to non-negative evaluations, which are easier and faster for human operators.
[0075] In another embodiment, generating synthetic sample data includes generating and / or storing a set of synthetic sample data from a set of sample data, and providing each synthetic sample data in the set with one or more chemical and / or biological properties, at least one synthetic quality metric, and / or at least one classification quality metric. The generated and / or stored set of sample data may be selected based on the synthetic quality metric and / or classification quality metric. The generated and / or stored set of sample data may include sample data that results in inadequacy of at least one synthetic quality metric and / or classification quality metric during classification and / or synthetic data generation. In other words, the generated and / or stored set of sample data may include sample data that is identified during classification and / or synthetic data generation to trigger adjustments to the classifier model. If the synthetic quality metric or classification quality metric represents a confounding factor or outlier, the synthetic sample data may be stored during classification. Such synthetic data may be provided to the tuning process to focus the tuning of the classification model and related training on the data subspace where model performance is insufficient. In other words, such synthetic data may be provided to the tuning process for more efficient and targeted training of the classification model.
[0076] In another embodiment, the at least one synthesis quality metric includes the generated synthetic data and / or at least one counterfactual interpretation, which is related to characteristic features embedded in the sample data that enable a classification model to determine one or more chemical and / or biological properties. Generating synthetic sample data may further include determining biological and / or chemical properties associated with the synthetic sample data. The at least one synthesis quality metric may include an indication of whether the biological and / or chemical properties associated with the synthetic sample data are target biological and / or chemical properties. The synthetic sample data may be associated with target biological and / or chemical properties. The target biological and / or chemical properties may be biological and / or chemical properties that are correctly associated with the synthetic sample data. For example, biological and / or chemical properties determined by a data-driven synthesis data generator to be associated with the synthetic sample data may be correctly or incorrectly associated with the synthetic sample data. Human experts may be qualified to determine whether biological and / or chemical properties can be correctly or incorrectly associated with the synthetic sample data. The at least one synthesis quality metric may include an indication of whether the synthetic sample data can be associated with biological and / or chemical properties different from those of the sample data.
[0077] Counterfactual interpretations can be correlated with the differences between the original sample data and the counterfactual sample data. They can also be correlated with human expert input, which represents the consistency or difference between the characteristics on which the classification model classifies the sample data and the characteristics on which a human expert would classify the sample data. Through counterfactual interpretations, specific characteristics of the decision-making strategy of the classification model can be identified and analyzed. This provides human users or systems used for monitoring and / or control with a reliable and robust option to check the performance of the classification model and ensure the reliability of monitoring and / or control specifically based on (multiple) chemical and / or biological properties.
[0078] In another embodiment, an adjustment to the classification model is triggered if the classification quality metric indicates a correlation unrelated to the mapping of sample data to chemical and / or biological properties expected by human experts. This correlation may include a non-causal correlation with the mapping of sample data to chemical and / or biological properties expected by human experts. The chemical and / or biological properties expected by human experts may be referred to as expected chemical and / or biological properties. Expected chemical and / or biological properties may be provided, for example, via a data provision interface (e.g., a user interface). An adjustment to the classification model may be triggered if the classification quality metric indicates confounding factors or outliers. The classification quality metric may be outside a predefined range or exceed a predefined threshold, indicating poor performance of the classification model. In other words, the classification quality metric may indicate an incorrect classification made by the classification model based on a statistical probability metric. A trigger signal for adjusting the classification model may be provided in response to the classification model determining that a biological and / or chemical property not equal to the target biological and / or chemical property can be associated with the synthetic data. The target biological and / or chemical property may be a ground truth associated with the synthetic sample data. Benchmark facts associated with synthetic sample data and / or sample data can be determined by human experts and / or provided via a user interface.
[0079] In embodiments, any of these methods may further include adjusting the classification model in response to providing a trigger signal for adjusting the classification model. Providing a trigger signal for adjustment may include adjusting the classification model. Further, a trigger signal for determining one or more biological and / or chemical properties by the adjusted classification model may be provided specifically in response to providing a trigger for adjusting the classification model. In embodiments, a trigger signal for determining one or more biological and / or chemical properties by the adjusted classification model may be provided while adjusting the classification model. Providing a trigger signal for determining one or more biological and / or chemical properties by the adjusted classification model may include and / or may cause the adjusted classification model to determine one or more biological and / or chemical properties. Adjustment of the classification model may be necessary when a misclassification is detected. A misclassification may refer to a deviation between the classification determined by the classification model and the classification of a human expert. The classification of the human expert may be referred to as the expected classification. The expected classification may be provided, for example, via a data provision interface (such as a user interface). Therefore, this feature allows for improvement of the decision boundary of the classification model through specific and interpretable misclassifications.
[0080] In embodiments, any of these methods may further include displaying an indication that the classification model has been adjusted, particularly in response to providing a trigger signal for adjustment and a trigger signal for determining one or more biological and / or chemical properties by the adjusted classification model. Preferably, providing the trigger signal for adjustment and / or adjusting the classification model may further provide sample data and / or synthetic sample data, particularly along with the indication that the classification model has been adjusted.
[0081] In another embodiment, an adjustment to the classification model is triggered if a counterfactual interpretation relating to characteristic features embedded in the sample data (which enable the classification model to determine one or more chemical and / or biological properties) indicates a correlation unrelated to the mapping of the sample data to the chemical and / or biological properties expected by human experts. This correlation may include a non-causal correlation with the mapping of the sample data to the chemical and / or biological properties expected by human experts. An adjustment to the classification model may also be triggered if a counterfactual interpretation relating to characteristic features embedded in the sample data (which enable the classification model to determine one or more chemical and / or biological properties) indicates a confounding factor or outlier. In this case, the results of the classification model may not be interpretable with reference to human expert evaluation.
[0082] In another embodiment, based on the trigger signal, a sample dataset and / or a corresponding synthetic sample dataset associated with the classification quality metric and / or synthetic quality metric that triggers the adjustment are provided. If the synthetic quality metric or classification quality metric represents a confounding factor or outlier, the sample dataset and / or the corresponding synthetic sample dataset can be stored during classification. This sample dataset and / or the corresponding synthetic sample dataset can be provided to the adjustment process to focus the adjustment of the classification model and related training on the data subspace where model performance is insufficient. In other words, this sample dataset and / or the corresponding synthetic sample dataset can be provided to the adjustment process for more efficient and targeted training of the classification model.
[0083] In another embodiment, synthetic data is generated for one or more sample datasets by providing them to a data-driven synthetic data generator based on a trigger signal. Based on the trigger signal, synthetic data for one or more sample datasets can be provided for tuning a classification model. This generation can be based on sample datasets and / or corresponding synthetic sample datasets provided in relation to a classification quality metric and / or synthetic quality metric that triggers the tuning. The sample datasets and / or corresponding synthetic sample datasets can be used as anchors to generate additional synthetic sample data in the vicinity of the provided synthetic sample data. This allows mapping of a sub-data space related to confounding factors, and enables targeted training of the classification model.
[0084] In another embodiment, the generated synthetic data is provided as associated with one or more biological and / or chemical attributes, for example, by providing synthetic sample data via a user interface and / or receiving one or more biological and / or chemical attributes via a user interface. Further, the synthetic sample data can be provided for adjusting a classification model. In embodiments, the synthetic sample data can be provided along with one or more corresponding biological and / or chemical attributes, particularly one or more corresponding target biological and / or chemical attributes. Since human experts are the benchmark for trusted artificial intelligence, one or more biological and / or chemical attributes can be provided by human experts. This process can be referred to as human expert labeling. One or more biological and / or chemical attributes can be provided via a user interface to label the synthetic data or associate one or more biological and / or chemical attributes with the generated synthetic data.
[0085] In this embodiment, the synthetic data may be synthetic sample data.
[0086] In another embodiment, synthetic data and associated biological and / or chemical properties are provided to the classification model, and the classification is adjusted by training the classification model with the synthetic data and associated biological and / or chemical properties. Adjustment may include changing the weights of the classification model through the training process.
[0087] In another embodiment, a modified data-driven classification model is provided for classifying sample data into one or more biological and / or chemical properties. This allows for a classification model that enables more reliable and robust monitoring and / or control of chemical and / or biological products.
[0088] Chemical and biological processes require a high degree of determinism and involve significant resource demands. Typically, these processes involve multiple interdependent steps, meaning errors in chemical and biological processes can propagate along the value chain. Since these products form the basis of final consumer goods, ensuring high product quality is crucial.
[0089] In response to providing the classification model with additional sample data associated with a first biological and / or chemical attribute, and upon receiving a second biological and / or chemical attribute, further training the classification model based on that additional sample data, robust and reliable identification and / or classification of sample data is permitted. This ensures efficient resource utilization and a safe process at large scale.
[0090] The classification model was customized according to the use case specification by further training it in response to additional sample data and receiving a second biological and / or chemical attribute. Thus, the performance of the classification model can be tuned to the needs of biological and / or chemical processes. Ultimately, this leads to an improved product-to-waste ratio, which helps reduce the environmental impact of chemical and biological processes.
[0091] In embodiments, biological and / or chemical properties can be first biological and / or chemical properties, and / or second biological and / or chemical properties, and / or third biological and / or chemical properties, and / or fourth biological and / or chemical properties. Biological and / or chemical properties can be related to biological characteristics and / or chemical characteristics and / or physical properties, can indicate biological characteristics and / or chemical characteristics and / or physical properties, and can be part of or a combination of biological and / or chemical objects. Biological objects can include living organisms, such as cells, animals, bacteria, viruses, etc. Chemical objects can include one or more compounds. Chemical characteristics can be properties defined by the structure of at least one chemical substance. Chemical characteristics can be properties that can be established by changing the structure of at least one chemical substance. Examples of chemical characteristics can be acidity, oxidation state, or reactivity. Physical characteristics can indicate physical parameters associated with a sample. Physical characteristics can be properties that can be established by changing the physical state of a sample. Physical characteristics can be one of the following: mechanical properties, electrical properties, optical properties, thermal properties, etc. For example, physical characteristics can include one or more of the following: density, scratch resistance, electrical conductivity, color, absorption, heat capacity, etc. Specifically, end-of-use characteristics can refer to the characteristics of a product at the end of its use phase. Biological characteristics can indicate biological parameters associated with a sample. Biological characteristics can be those that can be established through biological processes. Biological characteristics can refer to toxicity, bioactivity, biodegradability, growth rate, bioaccumulation, etc.
[0092] In embodiments, the classification model may be adapted to classify sample data, additional sample data, historical sample data, and / or additional historical sample data based on biological and / or chemical attributes, particularly first and / or second biological and / or chemical attributes associated with sample data, additional sample data, historical sample data, and / or additional historical sample data. The classification model may determine confidence scores indicating biological and / or chemical attributes, particularly first and / or second biological and / or chemical attributes. Confidence scores may specify biological and / or chemical attributes, particularly first and / or second biological and / or chemical attributes associated with sample data, additional sample data, historical sample data, and / or additional historical sample data.
[0093] In an embodiment, the data generation model may be trained and / or parameterized to provide additional sample data in response to receiving an instruction for additional sample data. The instruction for additional sample data may include at least one of the following: at least a portion of the additional sample data, a representation of the additional sample data, a random tensor, and / or a combination thereof. The representation of the additional sample data may include a textual representation of the additional sample data to be generated and / or a visual representation of the additional sample data. The textual representation of the additional sample data may include a textual description of the additional sample data. The visual representation of the additional sample data may include a sketch of the additional sample data.
[0094] In embodiments, additional sample data may be additional sample image data, additional sample numerical data, particularly additional sample tabular data, additional sample text data, additional sample audio data, etc. The additional sample data may be associated with a first biological and / or chemical attribute. Specifically, the additional sample data may represent the first biological and / or chemical attribute. Additionally or alternatively, the first biological and / or chemical attribute may be derived from, preferably based on, the additional sample data. Specifically, the first biological and / or chemical attribute may be derived from the additional sample data based on features of the additional sample data. Where the additional sample data is additional sample image data, the feature may be one or more portions of the image. Where the additional sample data is additional sample numerical data, the feature may be a numerical value and / or a combination of two or more numerical values. Where the additional sample data is additional sample tabular data, the feature may be at least a portion of the table and / or a combination of two or more portions of the table. Where the additional sample data is additional sample text data, the feature may be a portion of a word and / or a combination of two or more portions of a word. In cases where the additional sample data can be additional sample audio data, the features can be a sound and / or a combination of two or more sounds.
[0095] Additional sample data may include synthetic data. Additional sample data may indicate biological and / or chemical properties associated with the sample (e.g., a synthetic sample). Additional sample data may be generated based on historical sample data. At least a portion of the historical sample data may be associated with a first biological and / or chemical property. Providing at least a portion of the historical sample data to a classification model allows a second biological and / or chemical property to be received from the classification model. Thus, the classification model can provide biological and / or chemical properties other than those associated with the historical sample data. At least a portion of the historical sample data associated with the biological and / or chemical property, preferably the first biological and / or chemical property, may be part of a validation training dataset and / or a test training dataset. Additional sample data may be generated from at least a portion of the historical sample data associated with the biological and / or chemical property, preferably the first biological and / or chemical property, by manipulating the representation of the historical sample data in the model space to generate a representation of additional sample data in the model space associated with biological and / or chemical properties other than those associated with the representation of the historical sample data, particularly the second biological and / or chemical property. Additional sample data can be generated from at least a portion of historical sample data associated with biological and / or chemical properties, preferably a first biological and / or chemical property, by selecting a representation of additional sample data associated with biological and / or chemical properties other than those associated with the historical sample data. The representation of the historical sample data can be a dimensionality-reduced representation of the historical sample data. This representation can be obtained by providing the historical sample data to the encoder. The representation of the historical sample data can be a tensor. Similarly, the representation of additional sample data can be a dimensionality-reduced representation of the additional sample data. This additional representation can be obtained by providing the additional sample data to the encoder. The additional sample data can be a tensor.
[0096] In embodiments, historical sample data can be historical sample image data, historical sample numerical data, particularly historical sample tabular data, historical sample text data, historical sample audio data, or a combination thereof. Historical sample data can be associated with historical biological and / or chemical properties. Specifically, historical sample data can represent historical biological and / or chemical properties. Additionally or alternatively, historical biological and / or chemical properties can be derived from, and preferably from, historical sample data. Classification models can be trained based on historical sample data indicating historical biological and / or chemical properties. Specifically, historical biological and / or chemical properties can be features derived from, and / or obtained from, the historical sample data. Where historical sample data is historical sample image data, features can be one or more portions of an image. Where historical sample data is historical sample numerical data, features can be numerical values and / or combinations of two or more numerical values. Where historical sample data is historical sample tabular data, features can be at least a portion of a table and / or combinations of two or more portions of a table. Where historical sample data is historical sample text data, features can be a portion of a word and / or combinations of two or more portions of a word. In cases where historical sample data can be historical sample audio data, the features can be sounds and / or combinations of two or more sounds.
[0097] The biological and / or chemical properties of historical samples associated with historical sample data may be available, and / or can be provided, and / or can be received.
[0098] In this embodiment, historical sample data can be obtained using sensors. Therefore, historical sample data can be sensor data. Historical sample data can be associated with historical samples. Historical samples can be samples for which biological and / or chemical properties are available and / or known and / or provided.
[0099] In an embodiment, the predefined range may be a first predefined range and / or a second predefined range. The second predefined range may indicate a numerical range associated with a second biological and / or chemical attribute. Additionally or alternatively, the second predefined range may indicate a numerical range associated with a first biological and / or chemical attribute, and / or a third biological and / or chemical attribute, and / or a fourth biological and / or chemical attribute. The first predefined range may indicate a numerical range associated with a similarity score.
[0100] In embodiments, sample data may be sample image data, sample numerical data, particularly sample tabular data, sample text data, sample audio data, or a combination thereof. Sample data may be associated with and / or indicate biological and / or chemical properties. Specifically, sample data may represent biological and / or chemical properties. Additionally or alternatively, biological and / or chemical properties may be obtained from, preferably derived from, the sample data. Specifically, biological and / or chemical properties may be obtained from the sample data based on characteristics of the sample data. In embodiments, one or more biological and / or chemical properties may be determinable and / or determineable based on the sample data. Where the sample data may be sample image data, a feature may be one or more portions of an image. Where the sample data may be sample numerical data, a feature may be a numerical value and / or a combination of two or more numerical values. Where the sample data may be sample tabular data, a feature may be at least a portion of a table and / or a combination of two or more portions of a table. Where the sample data may be sample text data, a feature may be a portion of a word and / or a combination of two or more portions of a word. Where the sample data may be sample audio data, a feature may be a sound and / or a combination of two or more sounds.
[0101] In this embodiment, sample data can be acquired using a sensor. Therefore, sample data can be sensor data. Sample data can be associated with a sample. The sample can have multiple biological and / or chemical properties. Sample data can indicate one or more biological and / or chemical properties associated with the sample. Specifically, sample data can be a representation of one or more biological and / or chemical properties.
[0102] In embodiments, the sample may be an object. Preferably, the sample may be a biological and / or chemical object. The sample may include biological and / or chemical materials. The sample, particularly one or more biological and / or chemical properties associated with the sample, may be monitored. The sample, particularly one or more biological and / or chemical properties associated with the sample, may be monitored to determine the quality associated with the sample and / or to determine the application of chemical and / or biological products to the sample.
[0103] In embodiments, training a classification model can refer to and / or the training process can be the process of building a classification model, particularly determining and / or updating the parameters of the classification model. During the training process, the classification model can be tuned to achieve a best fit with the training dataset, for example, best-fitting at least one input value with at least one target output value. For example, if the neural network is a feedforward neural network (e.g., a convolutional neural network (CNN)), the backpropagation algorithm can be applied to train the neural network. In the case of a recurrent neural network (RNN), the gradient descent algorithm can be used to achieve the training objective. The gradient descent algorithm uses gradients to update parameters. The gradient can indicate the degree of change of the parameters of the classification model. The gradient can be obtained through backpropagation. Therefore, the gradient descent algorithm can be based on backpropagation. The training process can terminate when the deviation between the output generated by the classification model and the target output specified by the training dataset falls within a predetermined range. When the training process can terminate, the determination and / or updating of the parameters of the classification model can be terminated. The output generated by the classification model can be additional sample data and / or biological and / or chemical properties. The target output specified by the training dataset can be biological and / or chemical properties and / or additional historical sample data. Training a classification model can be associated with the training process, and further training of the classification model can be associated with other training processes.
[0104] These and other objectives are addressed by the subject matter of the independent claims, and will become apparent upon reading the following description. The dependent claims relate to embodiments of the content of this disclosure.
[0105] In embodiments, biological and / or chemical properties can be first biological and / or chemical properties, and / or second biological and / or chemical properties, and / or third biological and / or chemical properties, and / or fourth biological and / or chemical properties. Biological and / or chemical properties can be related to biological characteristics and / or chemical characteristics and / or physical characteristics, can indicate biological characteristics and / or chemical characteristics, and can be a part of or a combination of biological objects and / or chemical objects. Biological objects can include living organisms, such as cells, animals, bacteria, viruses, etc. Chemical objects can include one or more compounds. Chemical characteristics can be characteristics defined by the structure of at least one chemical substance. Chemical characteristics can be characteristics that can be established by changing the structure of at least one chemical substance. Examples of chemical characteristics can be acidity, oxidation state, or reactivity. Physical characteristics can be one of the following: mechanical properties, electrical properties, optical properties, thermal properties, etc. For example, physical characteristics can include one or more of the following: density, scratch resistance, electrical conductivity, color, absorption, heat capacity, etc. In particular, end-of-use characteristics can refer to the characteristics of a product at the end of its use phase. Biological characteristics can refer to toxicity, bioactivity, biodegradability, growth rate, bioaccumulation, etc.
[0106] In embodiments, historical sample data can be historical sample image data, historical sample numerical data, particularly historical sample tabular data, historical sample text data, historical sample audio data, or a combination thereof. Historical sample data can be associated with historical biological and / or chemical properties. Specifically, historical sample data can represent historical biological and / or chemical properties. Additionally or alternatively, historical biological and / or chemical properties can be derived from historical sample data, preferably based on it. Classification models can be trained based on historical sample data indicating historical biological and / or chemical properties. Specifically, historical biological and / or chemical properties can be derived from historical sample data based on features of the historical sample data. Where historical sample data is historical sample image data, features can be one or more portions of an image. Where historical sample data is historical sample numerical data, features can be the value of a numerical value and / or a combination of two or more numerical values. Where historical sample data is historical sample tabular data, features can be at least a portion of a table and / or a combination of two or more portions of a table. Where historical sample data is historical sample text data, features can be a portion of a word and / or a combination of two or more portions of a word. In cases where historical sample data can be historical sample audio data, the features can be sounds and / or combinations of two or more sounds.
[0107] In an embodiment, the method may further include providing historical sample data and generating the additional sample data based on the historical sample data. The classification model may be trained based on the historical sample data. Generating additional sample data based on historical sample data may refer to generating the additional sample data by modifying at least one data point associated with the historical sample data to generate the additional sample data.
[0108] In an embodiment, the classification model can be further trained in response to a similarity score indicating the similarity between historical sample data and other sample data, within a first predefined range. In an embodiment, the predefined range can be a first predefined range and / or a second predefined range. The second predefined range can indicate a numerical range associated with a second biological and / or chemical attribute. Additionally or alternatively, the second predefined range can indicate a numerical range associated with a first biological and / or chemical attribute, and / or a third biological and / or chemical attribute, and / or a fourth biological and / or chemical attribute. The first predefined range can indicate a numerical range associated with the similarity score.
[0109] By doing so, the additional sample data and historical sample data are sufficiently difficult to identify from each other. This allows for significant modification of the decision boundary associated with the classification model. Therefore, this feature helps to tailor the classification model to use case specifications, resulting in robust and reliable identification and / or classification of sample data, enabling efficient resource utilization and a secure process. The similarity score can indicate the distance between historical sample data and other sample data within the data manifold associated with historical sample data. The data manifold can be a lower-dimensional representation of the historical sample data used to train the classification model. The data manifold can represent the distribution of historical sample data points within the historical sample data.
[0110] In an embodiment, generating additional sample data based on historical sample data can refer to providing historical sample data to a data generation model and receiving additional samples from the data generation model. The data generation model can be a generative classification model. The data generation model can be adapted and / or trained to generate additional sample data based on provided sample data, particularly historical sample data. The data generation model can be parameterized and / or trained based on sample data, particularly historical sample data, and / or additional sample data and / or additional historical sample data. Additional historical sample data can refer to additional sample data generated based on historical sample data. The data generation model can be parameterized and / or trained to generate additional sample data in response to receiving historical sample data by providing sample data. The data generation model can include an encoder and / or a decoder. The encoder can be adapted to generate a representation of the sample data by changing the dimensions of the sample data. The decoder can be adapted to project the representation of the sample data onto additional sample data. In an embodiment, the data generation model can be trained based on feedback from a classification model trained to distinguish between real and synthetic data. Examples of data generation models can include machine learning architectures associated with variational autoencoders, generative adversarial networks, normalized streams, and the like.
[0111] In an embodiment, generating additional sample data based on historical sample data can refer to providing historical sample data to a data generation model trained to provide sample data in response to receiving historical sample data, thereby generating additional sample data.
[0112] In an embodiment, the method may further include: providing additional sample data via a user interface to determine whether the additional sample data can be associated with a first biological and / or chemical attribute; and receiving the first biological and / or chemical attribute associated with the additional historical data point in response to providing the additional sample data via the user interface. The user interface allows the user to directly control the training of the classification model. By doing so, customized and robust identification and / or classification of chemical and biological processes becomes possible.
[0113] In one embodiment, in response to receiving historical sample data, a data generation model is trained based on the historical sample data. A classification model can be trained based on the historical sample data, and the data generation model can be trained based on the historical sample data to provide sample data in response to receiving the historical sample data. By doing so, training data from the classification model can be recycled to train another data generation model. This saves time and resources throughout the process of identifying and / or classifying chemical and biological processes.
[0114] In an embodiment, the classification model can be trained to be provided with sample data and preferably to provide biological and / or chemical properties associated with the sample data in response to receiving the sample data.
[0115] In an embodiment, the second biological and / or chemical attribute may be received from the classification model in response to the classification model determining that a confidence score associated with a second biological and / or chemical attribute associated with additional sample data is preferably within a second predefined range. Receiving the second biological and / or chemical attribute from the classification model may include determining a confidence score associated with the second biological and / or chemical attribute, which is associated with additional sample data. The classification model may be trained to determine a confidence score associated with a biological and / or chemical attribute associated with the same sample data. The confidence score may indicate the biological and / or chemical attribute. Therefore, receiving a biological and / or chemical attribute from the classification model may refer to receiving a confidence score from the classification model that indicates the biological and / or chemical attribute.
[0116] In this embodiment, the first predefined range and / or the second predefined range can be provided via a user interface. Doing so allows for intervention in the training of the classification model. This enables customized and robust identification and / or classification of chemical and biological processes.
[0117] In an embodiment, an additional trained classification model can be obtained during a separate training process. In an embodiment, training the classification model can refer to and / or the training process can be the process of establishing a classification model, particularly determining and / or updating the parameters of the classification model. During the training process, the classification model can be tuned to achieve a best fit with the training dataset, for example, best-fitting at least one input value with at least one target output value. For example, if the neural network is a feedforward neural network (e.g., a convolutional neural network (CNN)), the backpropagation algorithm can be applied to train the neural network. In the case of a recurrent neural network (RNN), the gradient descent algorithm can be used to achieve the training objective. The gradient descent algorithm uses gradients to update parameters. The gradient can indicate the degree of change in the parameters of the classification model. The gradient can be obtained through backpropagation. Therefore, the gradient descent algorithm can be based on backpropagation. The training process can terminate when the deviation between the output generated by the classification model and the target output specified by the training dataset falls within a predetermined range. When the training process can terminate, the determination and / or updating of the parameters of the classification model can be terminated. The output generated by the classification model can be additional sample data and / or biological and / or chemical properties. The target output specified by the training dataset can be biological and / or chemical attributes and / or other historical sample data. Training the classification model can be associated with the training process, and further training of the classification model can be associated with additional training processes.
[0118] The additional training process can be terminated by providing the classification model with second, additional sample data associated with a fourth biological and / or chemical attribute and receiving the fourth biological and / or chemical attribute from the classification model. The classification model can be trained to receive sample data and provide the biological and / or chemical attribute associated with the sample data during training. The training process can be an additional training process. The classification model can be further trained during this training process. The classification model can be trained before further training.
[0119] In an embodiment, the first biological and / or chemical property may be a third biological and / or chemical property and / or a fourth biological and / or chemical property. Alternatively, the second biological and / or chemical property may be a third biological and / or chemical property and / or a fourth biological and / or chemical property.
[0120] In this embodiment, the classification model can be a primary classification model or a secondary classification model. The primary and / or secondary classification models can be adapted to classify sample data and / or additional sample data based on biological and / or chemical properties. Determining whether sample data can be associated with biological and / or chemical properties can mean classifying the sample data based on the biological and / or chemical properties associated with it. Determining whether additional sample data can be associated with a first biological and / or chemical property can mean classifying the additional sample data based on the biological and / or chemical properties associated with it.
[0121] In an embodiment, the classification model may be further trained based on additional sample data associated with the first biological and / or chemical attribute in response to providing the classification model with such additional sample data, in order to determine whether the sample data can be associated with the first biological and / or chemical attribute.
[0122] In one embodiment, the method for determining one or more biological and / or chemical properties may be a method for adjusting a classification model to determine one or more biological and / or chemical properties. In another embodiment, the method for determining one or more biological and / or chemical properties may be a method for monitoring and / or controlling one or more biological and / or chemical properties. In yet another embodiment, the method for determining one or more biological and / or chemical properties may be a method for adjusting a classification model to monitor and / or control one or more biological and / or chemical properties.
[0123] In the embodiments, (multiple) biological and / or chemical properties can be obtained from sample data and / or synthetic sample data.
[0124] In embodiments, an adjusted classification model can be provided, particularly in response to providing a trigger signal for adjusting the classification model, and / or by adjusting the classification model using synthetic sample data and biological and / or chemical properties (particularly target biological and / or chemical properties) associated with the synthetic sample data. The biological and / or chemical properties associated with the synthetic sample data can be indicated by a synthetic quality metric. Providing a trigger signal for adjusting the classification model allows for adjustment of the classification model. Preferably, the classification model can be adjusted based on the generated synthetic sample data and synthetic quality metrics and / or biological and / or chemical properties associated with the synthetic sample data.
[0125] Preferably, the synthesis quality metric can be related to biological and / or chemical properties associated with the synthetic sample data. Specifically, the synthesis quality metric can be related to a comparison and / or matching between biological and / or chemical properties determined by providing the synthetic sample data to a data-driven classification model and a target classification of the synthetic sample data. The target classification of the synthetic sample data can be provided via a user interface. Providing a synthesis quality metric related to the generation of the synthesis data can include providing synthetic sample data for receiving the synthesis quality metric. Providing a synthesis quality metric can include receiving the synthesis quality metric.
[0126] In an embodiment, the device may be an apparatus for determining one or more biological and / or chemical properties associated with a biological and / or chemical product, wherein the (multiple) biological and / or chemical properties may be obtained from sample data associated with the biological and / or chemical product. Attached Figure Description
[0127] The disclosure will be further described below with reference to the accompanying drawings. In the drawings and the disclosure, the same reference numerals are intended to refer to the same or similar elements, components and / or portions.
[0128] Figure 1 Examples of apparatuses for controlling and / or monitoring chemical products based on one or more chemical properties associated with the input materials to be used in production are shown.
[0129] Figure 2 An example flowchart is shown for a method for generating one or more chemical and / or biological properties associated with chemical and / or biological products.
[0130] Figure 3 An example illustrating the effectiveness of the enhanced classification method is shown based on muscle tissue images.
[0131] Figure 4 An example user interface sequence for controlling and / or monitoring chemical and / or biological products based on an enhanced classification approach is presented.
[0132] Figure 5 Example model structures, including data-driven counterfactual models and data-driven classification models, are shown.
[0133] Figure 6 The potential spatial representation, including the boundary between two categories, is illustrated schematically.
[0134] Figure 7 A schematic flowchart for deploying the adapted classification model is shown.
[0135] Figure 8 A schematic flowchart is shown for the classification model to be used and adjusted by end users.
[0136] Figure 9 The system architecture used to tune the classification model is shown.
[0137] Figure 10 An example flowchart is shown for a method for generating one or more chemical and / or biological properties associated with chemical and / or biological products. Detailed Implementation
[0138] The following embodiments are merely examples for implementing the methods, systems, or application devices disclosed herein and should not be considered limiting.
[0139] Figure 1 Examples of apparatuses for controlling and / or monitoring chemical products based on one or more chemical properties associated with the input materials to be used in production are shown.
[0140] Chemicals are produced in large quantities through multiple processing steps. Small variations in production conditions can significantly affect the biological and / or chemical properties of chemicals. This can include minute changes in the composition of the input materials used to produce the chemical in one or more chemical reactions, such as changes in the composition of the input materials due to contamination or different suppliers with different production processes.
[0141] Due to the inherent nature of the aforementioned chemical reactions, chemical production may require reliable monitoring and / or control of the input materials supplied to the chemical production facility 108. One measure for monitoring and / or controlling the input materials supplied to the chemical production facility 108 may include verifying the quality of the input materials by analyzing the composition of the input material 106. This can typically be accomplished using non-invasive analytical tools, such as a spectrometer. Therefore, use cases applying the teachings of this disclosure can be found within the field of quality assurance. For example, sample 106 may be irradiated with infrared light, and the absorbance and / or reflectance of sample 106 to infrared light can provide an indication of the composition of sample 106. Data from an infrared spectrometer may be an example of sample data 104 relating to the characteristics of the sample of input material 106, such as its composition. Further examples of measuring the composition of input material 106 may include other spectral data, such as nuclear magnetic resonance spectroscopy, UVVis spectroscopy, electron spin resonance spectroscopy, etc.
[0142] Sample data 104 related to the properties of the input material 106 may include sensor data. Sensor data can be generated by recording signals associated with a sample of the input material 106. The sample of the input material 106 can be classified based on sample data 104, regarding various biological and / or chemical properties of the input material 106, such as, for example, in… Figures 2 to 5The context will describe this in more detail. In this example, (multiple) biological and / or chemical properties can include the composition of the input material.
[0143] A classification model can provide biological and / or chemical properties 120 to the chemical production facility 108. Specifically, (multiple) biological and / or chemical properties can be provided to the control and / or monitoring engine of the chemical production facility 108. The control and / or monitoring engine can be configured to receive biological and / or chemical properties to monitor and / or control the processing of multiple batches of input materials, obtain samples 106 from these multiple batches of input materials, and classify these samples according to their corresponding (multiple) biological and / or chemical properties. The chemical production facility 108 can process the input materials if the biological and / or chemical properties indicate that one or a subset of the input material batches 106 is suitable for production or meets quality constraints. The control and / or monitoring engine can provide instructions for providing the input materials to the chemical production facility to produce chemical product 140. Otherwise, the control and / or monitoring engine of the chemical production facility 108 can reject the input material 106, and different input materials may be required to produce chemical product 140. This ensures the reliable production of chemical product 140 by determining the biological and / or chemical properties associated with the input material 106.
[0144] The examples above should be considered non-limiting. Numerous other examples exist for using the enhanced classification methods disclosed herein to monitor and / or control chemical and / or biological operations. One example includes monitoring and / or controlling biological treatments on plants, where plant health can be classified based on images of the plant. Another example includes monitoring and / or controlling tissue analysis, where benign and non-benign tissues can be classified. Yet another example includes monitoring and / or controlling fermentation processes, where sugar and acid content can be classified based on electrochemical sensor measurements to categorize different stages of the fermentation process.
[0145] Figure 2 An example flowchart is shown for a method for generating one or more chemical and / or biological properties associated with chemical and / or biological products.
[0146] Sample data associated with biological and / or chemical products can be provided, for example, via a data provision interface associated with a measuring device configured to measure sample data. The sample data can be related to the characteristics of the chemical and / or biological product. The sample data can be related to sample quality, application characteristics, technical characteristics, physicochemical properties, tissue characteristics, etc., such as... Figure 1 As described in the example. Sample data may include measurements of the properties of chemical and / or biological products.
[0147] Sample data can be fed into a data-driven classification model, which can generate one or more biological and / or chemical attributes. The classification model can be parameterized based on historical sample datasets and corresponding biological and / or chemical attributes. The classification model can include a neural network architecture specifically designed for classification tasks related to sample data. Classification can include one or more categories associated with one or more biological and / or chemical attributes. For example, classification can be related to whether a biological and / or chemical attribute is true or false. Further, for example, classification can be related to chemical attributes associated with input materials used in chemical production. Further, for example, classification can be related to biological attributes associated with plant health. Further, for example, classification can be related to biological and / or chemical attributes associated with stages of the fermentation process. The classification model can be selected based on a specific sample data type or measurement type and the corresponding one or more biological and / or chemical attributes. The classification model can include a specialized classification model associated with chemical and / or biological products and corresponding biological and / or chemical attributes. The classification model can be customized to classify sample data associated with chemical and / or biological products according to one or more biological and / or chemical attributes. Classification attributes can be predefined regarding sample data and associated chemical and / or biological products. Based on the generation of one or more biological and / or chemical properties, one or more classification measures can be determined that are related to the quality of the classification generation. Such measures may include indicators based on statistical probability, such as accuracy or confidence level.
[0148] Synthetic sample data can be generated based on classified chemical and / or biological attributes. The identified sample data can be fed into a data-driven counterfactual model to generate counterfactual sample data associated with the sample data and one or more biological and / or chemical attributes. The data-driven counterfactual model can be parameterized to transform the sample data associated with the identified one or more biological and / or chemical attributes into counterfactual sample data associated with at least one distinct category of biological and / or chemical attribute. The counterfactual sample data can represent a classification of one or more biological and / or chemical attributes distinct from the sample data. In other embodiments, adversarial counterfactual sample data can be generated and provided. However, adversarial counterfactual sample data may hinder interpretability because noise or other artifacts may produce adversarial counterfactual sample data without increasing the interpretability of the classification model's decision strategy. The advantage of non-adversarial counterfactual sample data is that its generation can utilize the generative model as a regularizer and thus maintain a reasonably high density of sample data distribution, thereby enhancing the interpretability of the counterfactual sample data. Therefore, counterfactual sample data improves the credibility of evaluating and interpreting classification model results.
[0149] To generate counterfactual data, the counterfactual model can be based on a pre-trained or self-trained generative model. The generative model can include any model suitable for generating sample data based on a transformation from the sample data distribution space to the latent space (or vice versa). The generative model can be configured to ingest input data corresponding to the type of sample data. The generative model can be configured to transform sample data with relation to characteristics associated with the type of monitoring and / or control and / or the category of chemical and / or biological product.
[0150] Counterfactual models can be based on gradient mechanisms (such as gradient ascent), which perform gradient ascent steps until the classifier changes to another classifier, which is then associated with counterfactual sample data. Generating such counterfactual sample data can include a transformation from the sample data distribution space to a lower-dimensional latent space. Counterfactual models can be based on normalization flows, one or more variational autoencoders (VAEs), or generative adversarial networks (GANs). Normalization flows can be based on learning a reversible transformation of the data distribution to the latent space, sampling in the latent space, and generating data through an inverse transformation from the latent space to the data space. The latent space, latent feature space, or embedding space can encode the sample data into a transformed representation. The latent space, latent feature space, or embedding space can encode features associated with chemical and / or biological properties used to monitor and / or control biological and / or chemical products into a latent space representation. In the case of normalization flows, the model learns a mapping f: X -> Z, where X is the sample data distribution and Z is the distribution in the selected latent space. Based on this mapping, normalized flows can generate data by sampling z ~ pZ, where pZ is the probability density of z sampled from distribution Z; and applying the inverse transformation f⁻¹(z) = xgen to generate synthetic data. Normalized flows can operate on the sample data distribution and the corresponding latent space distribution with the same dimension. To reduce the data processing and storage footprint of variational autoencoders or GANs, transformations to latent spaces with dimensions lower than the sample data distribution space can be used. Counterfactual models can be based on transformations from sample data to lower-dimensional latent spaces. Lower-dimensional latent spaces can include latent spaces that encode features of the sample data into lower-dimensional representations, latent feature spaces, or embedding spaces. Counterfactual models can include autoencoder or GAN-based models for synthetic data generation. Examples of such methods are described in the following literature: Ann-Kathrin Dombrowski, Jan E. Gerken, Klaus-Robert Müller, and Pan Kessel, “Diffeomorphic Counterfactuals with Generative Models”, eprintarXiv:2206.05075, June 2022; 10.48550 / arXiv.2206.05075. For example, in Figures 2 to 6 The determination of counterfacts is described in more detail within the context of this study. Other options for generating counterfacts include models based on denoised diffusion models. A denoised diffusion model can refer to a model based on an invertible Markov chain or learned ordinary differential equations or stochastic differential equations.
[0151] At least one classification quality metric related to classification and / or at least one synthetic quality metric related to the generation of synthetic data can be provided. Based on determinations made by a data-driven classification model and / or a data-driven counterfactual model, one or more quality metrics can be generated and provided. One or more quality metrics can be related to statistical probability metrics of the data-driven model. One or more quality metrics can include accuracy or confidence intervals provided by the model. One or more quality metrics can include counterfactual sample data associated with the sample data and one or more biological and / or chemical attributes determined by the data-driven classification model. One or more quality metrics can include a counterfactual specification that associates the sample data and at least one associated biological and / or chemical attribute with the counterfactual sample data and at least one associated biological and / or chemical attribute different from the biological and / or chemical attribute associated with the sample data. This can also be referred to as a counterfactual interpretation or decision strategy of the classification model, which is related to the interpretability of the decision logic of one or more biological and / or chemical attributes determined by the data-driven classification model. Therefore, counterfactual sample data can be used as a quality metric for the data-driven classification model by providing insights into the internal decision logic of the classification model, such as regarding... Figures 2 to 6 As described in one of the examples.
[0152] Based on the provided classification quality metrics and / or synthetic quality metrics, adjustments to the classification model can be triggered, and an adjusted classification model can be provided. For example, if counterfactual sample data generated by a counterfactual model indicates that the model bases its decisions on features that are not considered robust and / or causal, the classification decision logic may not focus on the sample data features that would lead human experts to make classification decisions. In this case, the classification model can be considered non-robust and / or biased, thus triggering adjustments to improve the credibility of the classification model or the system implementing the classification model for users.
[0153] Based on a trigger signal, a sample dataset is generated that is associated with the classification quality metric and / or synthetic quality metric adjusted by the trigger. This sample dataset may include sample data that triggered the trigger signal, or sample data that indicates the classification model's classification is inaccurate via the classification quality metric, and / or indicates the classification model's classification is unrobust and / or biased via the synthetic quality metric. The sample dataset may include sample data provided to the classification model that leads to inaccurate and / or uninterpretable classification. The sample dataset may include one or more chemical and / or biological attributes provided by the classification model that lead to inaccurate and / or uninterpretable classification. The sample dataset may include pairings of sample data with corresponding one or more chemical and / or biological attributes provided by the classification model.
[0154] By training a classification model with synthetic data, the model can be tuned. Synthetic data can be generated for one or more tracking sample datasets. This synthetic data can be provided, for example, to human expert users for classification. The synthetic data can be labeled by human expert users, for example, with one or more chemical and / or biological attributes. Synthetic data and corresponding chemical and / or biological attributes can be provided. Synthetic data can be generated by providing sample data from the sample dataset to a counterfactual model. In this way, the training dataset of the synthetically generated data can focus on specific data subspaces where classification is not adequately performed. By using specific sample datasets and corresponding counterfactual data, classification boundaries between two or more categories can be established, and the model can be trained specifically on such subspaces(s). Synthetic data may or may not include historical sample data, depending on the historical data anchored and embedded in the model during training. Preferably, the synthetic data does not include the historical sample dataset used to train the classification model. The training of the classification model can be enhanced by using category labels with human feedback on new sample datasets. Additionally, the training dataset can be amplified by generating synthetic data based on this new data. Finally, the training quality can be improved by using labeled counterfactual training data.
[0155] Tuning a classification model by training it with labeled counterfactual sample data can involve iterative loops around the mapping of the tracked sample dataset until the labeling of the counterfactual sample data and the classification of the classification model converge. Thus, a classification model can be tuned by training it with synthetic sample data generated based on (multiple) tracked sample datasets.
[0156] Once a classification model is trained on synthetic data, a tuned classification model can be provided to classify sample data more robustly.
[0157] Figure 3 Examples of sample data and synthesized sample data are shown.
[0158] The examples shown are included for illustrative purposes only and are not intended to be limiting. This example shows an image of a human face. These images can be examples of sample data. The image on the left can be an example of the provided sample data. The image can show a human face, particularly a human facial expression. Images, particularly sample data, can be classified based on a human facial expression (i.e., smiling or not smiling). The sample data can be classified by a classification model. To represent the decision boundaries associated with the classification model, synthetic sample data can be generated. An example of synthetic sample data can be seen on the right. Synthetic sample data can be generated, for example, from sample data. Synthetic sample data can include, in particular, counterfactual data related to the sample data. Counterfactual data can be synthetic sample data generated from sample data and can be associated with biological and / or chemical properties different from those of the sample data. In a specific example, sample data can be associated with the label “not smiling.” Counterfactual data can be associated with smiling. As in Figure 3 As can be seen, two composite images can be generated from the sample data. The upper image may show meaningless deviations, i.e., those that can be associated with the same labels as those in the sample data. Meaningless deviations may refer to changes in the watermark in the lower right portion of the upper image. These changes are usually identifiable.
[0159] The lower image can reveal meaningful biases, i.e., labels that differ from those in the sample data. Therefore, the upper image can be a counterfactual of misclassification, and the lower image can be a counterfactual of correct classification.
[0160] This enables more reliable classification and more reliable use of the classification for further processing.
[0161] Similar to an example of human facial expressions, chemical and / or biological products can be classified via images based on the appearance of the product or based on numerical data characterizing the product (e.g., spectral or tabular data (not shown)).
[0162] Figure 4 A sample user interface for validating classifications by human experts is shown.
[0163] In the example, the misclassified counterfactual can be shown in contrast to the original image. The user interface can further display attributes associated with the original image and the synthesized counterfactual to allow human experts to review the labels to be assigned to the counterfactual and / or to review whether the synthesized sample data (i.e., the hypothetical counterfactual) is a true counterfactual. Human experts can be allowed to accept or reject the displayed classification via the user interface.
[0164] Figure 5 Example model structures, including data-driven counterfactual models and data-driven classification models, are shown.
[0165] In this example, the data-driven counterfactual model can be based on an autoencoder structure including an encoder and a decoder. In other examples, a diffusion model can be used. The data-driven classifier model can be based on any neural network architecture that can be used to monitor (multiple) specific chemical and / or biological products by determining one or more chemical and / or biological properties. The example shown illustrates a simple architecture including multi-layer classifiers, feature maps, softmax, and other layers. This should be considered non-limiting, and more complex architectures are feasible depending on the monitoring and / or control task. The classification model can be trained based on a historical sample dataset and the corresponding one or more biological and / or chemical properties. Therefore, the classification model can be configured to provide one or more biological and / or chemical properties based on sample data. The classification model can be trained separately from the data-driven counterfactual model.
[0166] Counterfactual models can include pre-trained generative models. Sample data that produce a specific classification of biological and / or chemical properties can be transformed into latent representations, such as lower-dimensional representations, by an autoencoder-based model (specifically, the encoder part of the model). In the latent space, a counterfactual latent representation corresponding to the minimum deformation x0 = x + δx that changes the classifier's prediction can be determined. In other words, in the latent space, a counterfactual latent representation corresponding to at least one distinct biological and / or chemical property and the minimum step to the decision boundary of at least one distinct biological and / or chemical property can be determined. Figure 6 The potential spatial representation, including the boundary between two categories, is illustrated schematically.
[0167] Counterfactual latent representations can be transformed into counterfactual sample data via an autoencoder-based model (specifically, the encoder portion of the model). This counterfactual sample data can then be fed into a data-driven classification model to determine one or more biological and / or chemical properties. In this way, synthetic data corresponding to the counterfactual aspects of the sample data can be generated, and this synthetic data can be used to retrain the classification model.
[0168] Figure 7 A schematic flowchart for deploying the adapted classification model is shown.
[0169] When deploying a classification model, the trained model can be examined for confounding factors, which can be eliminated by performing the methods disclosed herein. An initial classification model can be provided and trained based on initial object data, including historical datasets containing input-output pairs of sample data and corresponding biological and / or chemical properties. The classifier or classification model can be tested during training on a test dataset, and can be examined for confounding variables. If no confounding variables are found, the model can be deployed.
[0170] If confounding variables are found, synthetic data can be generated by a synthetic data generator that includes a generative model related to the classification model, such as, for example, in... Figures 2 to 6 This is described in the context of [the previous sentence]. Counterfactual sample data can be generated and provided to a user interface for labeling by human experts. The classification model can be tuned based on the training dataset of counterfactual sample data and the corresponding labels. The tuned classification model can be examined regarding its performance. Performance can be compared with, for example, [the following text is missing here]. Figures 2 to 6 The classification and / or synthetic data quality metrics described in the context are relevant. If the performance is accurate enough, the model can be deployed. If the performance is not accurate enough, the model can be further retrained based on the generated synthetic data.
[0171] Figure 8 A schematic flowchart is shown for the classification model to be used and adjusted by end users.
[0172] Similarly, for Figure 7 Initial sample data can be fed to a synthetic data generator to produce counterfactual sample data for labeling. The classification model's performance can be continuously or intermittently checked. If the performance is insufficient, the model can be further retrained based on the generated synthetic data until the performance is adequate. Figure 7 compared to, Figure 8 The flowcharts shown allow for complete end-user control. In particular, the methods, apparatus, and uses disclosed herein are based on counterfactual principles, which allows for easy and reliable use by end users. Furthermore, the intuitive nature of counterfactual principles and the fully automated training process based on synthetic sample data enable end users to ensure pattern robustness.
[0173] Figure 9 The system architecture used to tune the classification model is shown.
[0174] implement Figure 7 and Figure 8The system of the illustrated method may include at least one database storage device for storing initial or historical sample data and synthetic sample data. During initial training, a classification model provider can be configured to train a classification model based on an initial dataset of sample data and corresponding biological and / or chemical properties. Additionally, a generative model provider can be configured to further train a pre-trained generative model based on this data. Based on the sample data and corresponding biological and / or chemical properties, a synthetic data generator can be configured to generate counterfactual sample data. The counterfactual data can be provided to a user interface for labeling by human experts. The labeled counterfactual sample data can be provided to a storage device for storing and using data used to adjust the classification model through training.
[0175] Figure 10 An example of an enhancement method for counterfactual determination is shown.
[0176] Counterfactual sample data can be generated by transforming the sample data into a representation encoding characteristic features using a counterfactual generator trained on the sample data and corresponding biological and / or chemical properties, and an augmented data-driven classification model initialized with random weights. The augmented data-driven classification model initialized with random weights can include data-driven classification models trained on sample data with different weights. The augmented data-driven classification model can be trained and / or configured based on the sample data and the corresponding biological and / or chemical properties determined by the data-driven classification model, particularly after initializing the augmented data-driven classification model with random weights. This can be referred to as distillation. An augmented data-driven model can be obtained by distillation from a data-driven classification model.
[0177] The motivation behind obtaining an enhanced data-driven classification model from a data-driven classification model via distillation stems from the fact that the proof from Dombroski et al. assumes a perfectly trained generator where all directions deviating from the data manifold become 0, which is never precisely achieved for a real model. Therefore, there are always some adversarial directions in the gradients used to update the counterfactual predictor. Due to this adversarial component of the gradient, one can no longer trust the confidence of the predictor providing the gradient. However, adversarial attacks based on white-box access to the gradient do not transfer to other predictors but are specific to the weight configuration. Therefore, distilling the predictor and obtaining the gradient from the distilled predictor, while using the confidence of the original predictor as a stopping criterion, produces more reliable predictions. Furthermore, adversarial directions are often high-frequency noise orthogonal to the internally assumed data manifold of the predictor. If too much of this noise is added to the counterfactual over many steps to achieve the desired confidence value, this will produce a noisy image. However, various defenses against adversarial attacks are known to exist, making the predictor less vulnerable to attacks relative to adversarial directions. Applying such a technique produces less noisy gradients, thus producing less noisy counterfactuals.
[0178] Therefore, the architecture of an enhanced data-driven classification model can correspond to and / or be an architecture of a data-driven classification model. The enhanced data-driven classification model can be configured to determine changes in classified biological and / or chemical properties based on variations in sample data, particularly stepwise variations. Specifically, determining changes can preferably include determining the degree of change associated with the determination of biological and / or chemical properties associated with the sample data, preferably by the enhanced data-driven classification model. In this context, sample data can include provided sample data and / or synthetic sample data based on and / or obtained from the provided sample data. The enhanced data-driven classification model can be configured to receive sample data and / or synthetic sample data, particularly stepwise, and determine changes in classified biological and / or chemical properties associated with the sample data and / or synthetic sample data. Synthetic sample data can include counterfactual sample data and / or synthetic sample data other than counterfactual sample data, particularly synthetic sample data obtained by stepwise determination of changes in biological and / or chemical properties to preferably reach counterfactual sample data.
[0179] Based on changes in classified biological and / or chemical attributes, particularly gradual changes, data-driven classification models can be configured to determine the confidence level or class probability of the identified biological and / or chemical attribute. Once a predefined confidence level and / or target confidence level for the identified change can be reached, sample data associated with the change in the identified biological and / or chemical attribute or class change can be counterfactual sample data and / or can be provided. Counterfactuals can be determined with higher reliability by using a counterfactual generator trained on sample data for confidence or class probability determination and an enhanced data-driven classification model initialized with random weights for gradient determination. This achieves a quality metric that allows positive feature evaluation by assessing the correctness (positive) of the feature, rather than negative feature evaluation by assessing the incorrectness (non-negative). For the interpretability of AI, the enhanced reliability of counterfactual determination makes positive evaluations superior to non-negative evaluations, which are easier and faster for human operators.
[0180] The following illustrates the cutoff criteria for the search for counterfactual sample data on the logical and data manifolds for gradient determination. The logical line output should be considered for illustrative purposes only and not as restrictive. Determining the variation and / or degree of variation of biological and / or chemical properties can correspond to determining the gradient associated with the data-driven classification model and / or the augmented data-driven classification model. Reaching the target confidence level and / or a predefined confidence level can be a cutoff criterion for progressively determining the variation of biological and / or chemical properties. The data-driven classification model can be a predictor f. The augmented data-driven classification model can be a distilled predictor f_distilled. The counterfactual generator can be a generator g. The sample data can be the original sample x obtained by collecting data from the real world. Synthetic sample data can be obtained by progressively approaching the counterfactual sample data from the sample data.
[0181]
[0182]
[0183] The workflow for generating differential homeomorphism diffusion counterfactuals includes:
[0184] f: A classifier with confidence c_f and gradient G_f
[0185] f_distilled: A classifier with confidence c_f_d and gradient G_f_d distilled from f.
[0186] To generate `Generate f_distilled`, the following steps can be performed:
[0187] 1. Provide data with the following attributes: X and Y
[0188] 2. Provide a classifier f trained using X and Y.
[0189] 3. Initialize f_distilled using random weights but with the same topology / model architecture.
[0190] 4. Generate Y_pred: Y_pred = f(X)
[0191] 5. Train the classifier f_distilled using X and Y_pred.
[0192] Counterfactual generation is accomplished through optimizing loops:
[0193] This requires a stopping criterion (from the confidence level of f).
[0194] This requires direction (the gradient from f_distilled).
[0195] Different ways to implement workflows include:
[0196]
[0197]
[0198] The trusted AI methods outlined in this article enable the safe and reliable monitoring and / or control of chemical and / or biological products based on categorized chemical and / or biological properties. The safe use of artificial intelligence is particularly critical in chemical and pharmaceutical operations, as false positives or unexplained black-box functionality can lead to adverse effects.
[0199] This disclosure has also been described in conjunction with various preferred embodiments and examples. However, by studying the accompanying drawings, this disclosure, and the claims, those skilled in the art, as well as those who practice the claimed invention, will understand and implement other variations.
[0200] Any steps presented in this document can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps. Nor is it required that different steps be performed in a particular place or on a particular computing node in a distributed system; that is, each step can be performed on different computing nodes using different devices / data processing.
[0201] As used herein, "determine" also includes "initiating or causing determination," "generate" also includes "initiating and / or causing generation," and "provide" also includes "initiating or causing determination, generation, selection, sending, and / or receiving." "Initiating or causing an action" includes any processing signal that triggers a computing node or device to perform a corresponding action.
[0202] In the claims and specification, the word "comprising" or "including" or similar wording does not exclude other elements or steps and should not be construed as limiting oneself to the listed elements or steps. The indefinite article "a" or "an" does not exclude multiple. A single element or other unit may perform the function of several entities or items recited in the claims. The fact that certain measures are recited only in mutually different dependent claims does not indicate that a combination of these measures cannot be used in advantageous implementations or that additional elements may be included.
[0203] Within the scope of this disclosure, provision may include any interface configured to provide data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. Provision may include transmitting or submitting data to the interface, particularly displaying data to a user or having data used by a receiving entity.
[0204] Any disclosures and embodiments described herein relate to the methods, systems, apparatuses, devices, chemicals, materials, services, uses, and computer program elements listed above, and vice versa. Advantageously, the benefits provided by any embodiments and examples also apply to all other embodiments and examples, and vice versa.
[0205] All terms and definitions used in this article are to be understood in a broad sense and have their general meaning.
Claims
1. A method, particularly a computer-implemented method, for determining one or more biological and / or chemical properties associated with biological and / or chemical products, wherein, The biological and / or chemical properties can be obtained from sample data associated with the biological and / or chemical product, and the method includes: - Provide sample data associated with the biological and / or chemical product. - The sample data is provided to a data-driven classification model, which determines one or more biological and / or chemical properties. This classification model is parameterized based on historical sample datasets and the corresponding biological and / or chemical properties. - Synthetic sample data is generated by providing the sample data to a data-driven synthesis data generator, and the synthetic sample data is determined based on classified chemical and / or biological properties, wherein the data-driven synthesis data generator is configured to transform the sample data and generate synthetic sample data with respect to one or more biological and / or chemical properties. - Provide at least one synthetic quality metric related to the generation of the synthetic data and, optionally, at least one classification quality metric related to the classification determination, and / or - Provide trigger signals for adjusting the classification model based on the provided classification quality metric and / or synthetic quality metric.
2. The method as described in claim 1, wherein, The one or more biological and / or chemical attributes are associated with characteristic features embedded in the sample data, wherein these characteristic features can refer to the features that the classification model can generate or select as a classifier based on the one or more biological and / or chemical attributes.
3. The method as described in any of the preceding claims, wherein, Generating synthetic sample data includes providing sample data to a data-driven counterfactual generator, wherein the counterfactual generator is configured to transform the sample data into a representation encoding characteristic features embedded in the sample data, and to determine counterfactual sample data based on variations in the classified biological and / or chemical attributes, wherein the counterfactual sample data represents a different classification of one or more biological and / or chemical attributes from the sample data, wherein the counterfactual sample data is generated by the counterfactual generator transforming the sample data into the representation encoding characteristic features, the counterfactual generator being trained based on the sample data and the corresponding attributes and initialized with random weights.
4. The method as described in any of the preceding claims, wherein, Generating synthetic sample data includes generating and / or storing a set of synthetic sample data from a set of sample data, and providing each synthetic sample data in the set with one or more chemical and / or biological properties, at least one synthetic quality metric and / or at least one classification quality metric.
5. The method as described in any one of the preceding claims, wherein, Generating the synthetic sample data may further include determining the biological and / or chemical properties associated with the synthetic sample data, and wherein the at least one synthetic quality metric may include an indication of whether the biological and / or chemical properties associated with the synthetic sample data are target biological and / or chemical properties.
6. The method as described in any of the preceding claims, wherein, If the synthetic quality metric indicates a correlation that is not related to the mapping of sample data to the chemical and / or biological properties expected by human experts, then an adjustment to the classification model is triggered.
7. The method as described in any of the preceding claims, wherein, Trigger signals for adjusting the classification model are provided based on biological and / or chemical properties that are not equal to the target biological and / or chemical properties as determined by the classification model and associated with the synthetic data, wherein the target biological and / or chemical properties are baseline facts associated with the synthetic sample data.
8. The method as described in any of the preceding claims, wherein, Based on the trigger signal, sample datasets and / or synthetic sample datasets related to the classification quality metric and / or the synthetic quality metric adjusted by the trigger are provided.
9. The method as described in any of the preceding claims, wherein, Based on this trigger signal, synthetic data for one or more sample datasets is provided.
10. The method of claim 9, wherein, The generated synthetic data is provided to be associated with one or more biological and / or chemical properties and / or to adjust the classification model.
11. The method of claim 9 or 10, wherein, The synthetic data and associated biological and / or chemical properties are provided to the classification model, and the classification is adjusted by training the classification model with the synthetic data and associated biological and / or chemical properties.
12. The method as described in any of the preceding claims, wherein, The modified data-driven classification model is provided for classifying sample data into one or more biological and / or chemical attributes.
13. An apparatus for determining one or more biological and / or chemical properties associated with a biological and / or chemical product, wherein, The biological and / or chemical properties can be obtained from sample data associated with the biological and / or chemical product, and the device includes: - A sample data provision interface configured to provide sample data associated with the biological and / or chemical product. - A classification generator configured to feed the sample data to a data-driven classification model, which then determines one or more biological and / or chemical attributes, wherein the classification model is parameterized based on a historical sample dataset and the corresponding one or more biological and / or chemical attributes. - A synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator, and to determine the synthetic sample data based on classified chemical and / or biological properties, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to one or more biological and / or chemical properties to generate synthetic sample data. - A quality metric provider configured to provide at least one classification quality metric related to the classification determination and / or at least one synthetic quality metric related to the synthetic data generation. - A trigger signal provider configured to provide trigger signals for adjusting the classification model based on the provided classification quality metric and / or synthetic quality metric.
14. The use of one or more biological and / or chemical properties determined by the method according to claims 1 to 13 for monitoring and / or controlling the production and / or treatment of biological and / or chemical products.
15. The adjusted classification model generated by the method according to claims 1 to 13 is used to determine one or more biological and / or chemical properties and / or for monitoring and / or controlling the production and / or treatment of biological and / or chemical products.