Method for determining one or more biological and / or chemical attribute(s)

The method generates synthetic data to adapt classification models for chemical and biological products, addressing the lack of explainability and reliability in data-driven models, ensuring robust and scalable attribute determination.

WO2025228811A1PCT designated stage Publication Date: 2025-11-06BASF SE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/061344
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-04-25
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Data-driven models for monitoring and controlling chemical and biological products lack explainability and reliability, leading to untrustworthy predictions and potential adverse effects, especially in critical areas, and require human expert validation that is time-consuming and prone to errors.

Method used

A method and apparatus that utilize a data-driven classification model adapted based on historic data, generate synthetic sample data through a synthetic data generator, and provide quality measures to adapt the model, ensuring transparency and scalability, with triggers for model adjustment based on deviation scores from predefined ranges.

Benefits of technology

Enhances the robustness and reliability of classification models by providing transparent and scalable determination of biological and chemical attributes, reducing the need for human expert intervention and enabling large-scale, efficient resource use in complex production systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000057_0000
    Figure 00000057_0000
  • Figure 00000058_0000
    Figure 00000058_0000
  • Figure 00000059_0000
    Figure 00000059_0000
Patent Text Reader

Abstract

The disclosure relates to the technical field of trustworthy artificial intelligence for controlling and / or monitoring chemical and / or biological products. The disclosure relates to methods and apparatuses for determining a biological and / or chemical attribute related to sample data associated with the biological and / or chemical product.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR DETERMINING ONE OR MORE BIOLOGICAL AND / OR CHEMICAL ATTRIBUTE(S)

[0002] TECHNICAL FIELD

[0003] The disclosure relates to the technical field of trustworthy artificial intelligence for controlling and / or monitoring chemical and / or biological products. The disclosure relates to methods and apparatuses for determining a biological and / or chemical attribute related to sample data associated with the biological and / or chemical product, use of biological attributes as described herein, use of an adapted classification model as described herein.

[0004] TECHNICAL BACKGROUND

[0005] Chemical and / or biological products may be monitored in various ways relating to one or more chemical and / or biological attribute(s). Chemical and / or biological attribute(s) may include quality, application properties, technical properties, physiochemical properties, tissue properties or the like. For monitoring sample data measured from the chemical and / or biological products and data-driven models trained on sample data may be used.

[0006] However, data-driven models may not be explainable. Their output may not be reliable, and it may remain unclear to the user of the system what drives specific output predictions. This can cause obstacles for using data-driven models particularly in critical areas such as monitoring of chemical and / or biological products.

[0007] SUMMARY

[0008] In one aspect disclosed is a, in particular computer-implemented, method for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the method comprising: providing and / or obtaining, in particular receiving, preferably via an interface the sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data- driven classification model one or more biological and / or chemical attribute(s), wherein the adapting the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, providing at least one classification quality measure relating to the classification determination and / or at least one synthetic quality measure relating to the synthetic data generation, based on the provided classification quality measure and / or the synthetic quality measure providing a trigger for adapting the classification model, optionally provided the adapted classification model for classifying sample data to one or more biological and / or chemical attribute(s).

[0009] In another aspect disclosed is an apparatus for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the apparatus comprising: a sample data providing interface configured to provide the sample data associated with the biological and / or chemical product, a classification generator configured to provide the sample data to a data-driven classification model and to determine by the data-driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), a synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator and to determine synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, a quality measure provider configured to provide at least one classification quality measure relating to the classification determination and / or at least one synthetic quality measure relating to the synthetic data generation, a trigger provider configured to provide based on the provided classification quality measure and / or the synthetic quality measure a trigger for adapting the classification model, optionally a model provider configured to provide the adapted classification model to the classification generator for classifying sample data to one or more biological and / or chemical attribute(s).

[0010] In yet another aspect disclosed is a method for monitoring and / or controlling a biological and / or chemical product based on one or more biological and / or chemical attribute(s), the method comprising: providing and / or obtaining, in particular receiving, preferably via an interface, the sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data- driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, providing at least one classification quality measure relating to the classification determination and / or at least one synthetic quality measure relating to the synthetic data generation, based on the provided classification quality measure and / or the synthetic quality measure the generated synthetic data is provided to be associated with one or more biological and / or chemical attribute(s), wherein the classification is adapted by training the classification model with the synthetic data and associated biological and / or chemical attribute, wherein the adapted, data-driven classification model is provided for classifying sample data to one or more biological and / or chemical attribute(s).

[0011] In another aspect disclosed is an apparatus for monitoring and / or controlling a biological and / or chemical product based on one or more biological and / or chemical attribute(s), the apparatus comprising: a sample data providing interface configured to provide the sample data associated with the biological and / or chemical product, a classification generator configured to provide the sample data to a data-driven classification model and to determine by the data-driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), a synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator and to determine synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, a training module configured to provide based on the provided classification quality measure and / or the synthetic quality measure the generated synthetic data to be associated with one or more biological and / or chemical attribute(s), wherein the classification is adapted by training the classification model with the synthetic data and associated biological and / or chemical attribute, a model provider configured to provide the adapted, data-driven classification model to the classification generator for classifying sample data to one or more biological and / or chemical attribute(s).

[0012] In yet another aspect disclosed is the use of the one or more biological and / or chemical attribute(s) determined according to the methods disclosed herein for monitoring and / or controlling a biological and / or chemical product.

[0013] In yet another aspect disclosed is the use of the one or more biological and / or chemical attribute(s) determined according to the methods disclosed herein for monitoring and / or controlling producing and / or processing of a biological and / or chemical product.

[0014] In yet another aspect disclosed is the use of the synthetic data generated according to the methods disclosed herein for training the classification model. In yet another aspect disclosed is a method for generating synthetic sample data according to the methods disclosed herein for controlling and / or monitoring chemical and / or biological products.

[0015] In yet another aspect disclosed is the use of the synthetic data generated according to the methods disclosed herein for training the classification model. In yet another aspect disclosed is a method for generating synthetic sample data according to the methods disclosed herein for controlling and / or monitoring producing and / or processing of chemical and / or biological products.

[0016] In yet another aspect disclosed is the use of the trigger generated according to the methods disclosed herein for initiating training of the classification model. In yet another aspect disclosed is a method for training a classification model according to the methods disclosed herein, wherein the classification model is configured to classify sample data associated with a biological and / or chemical product according to a biological and / or chemical attribute.

[0017] In yet another aspect disclosed is the use of the adapted classification model generated according to the methods disclosed herein for determining one or more biological and / or chemical attribute(s) and / or for monitoring and / or controlling a biological and / or chemical product. In yet another aspect disclosed is a method for classifying sample data associated with a biological and / or chemical product according to a biological and / or chemical attribute by using the adapted classification model.

[0018] In yet another aspect disclosed is a computer element, such as a computer program product or a machine-readable medium, with instructions, which when executed on one or more computing node(s) or processor(s) is configured to carry out the steps of the method(s) disclosed herein or configured to be carried out by the apparatus(es) disclosed herein.

[0019] In an aspect, the disclosure relates to a, in particular computer-implemented, method for determining a biological and / or chemical attribute associated with sample data related to the biological and / or chemical attribute, the method comprising: providing and / or obtaining, in particular receiving, the sample data, in particular via a user interface, providing the sample data to a classification model for determining if the sample data is associated with a biological and / or chemical attribute, wherein the classification model is trained to being provided by sample data and provide a biological and / or chemical attribute associated with the sample data, wherein the classification model is further trained based on a further sample data associated with a first biological and / or chemical attribute in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute, providing the biological and / or chemical attribute in response to determining that the sample data is associated with the biological and / or chemical attribute. In another aspect, it relates to use of a classification model trained to being provided by sample data and provide a biological and / or chemical attribute associated with the sample data, and further trained based on a further sample data associated with a first biological and / or chemical attribute in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute for determining a biological and / or chemical attribute associated with sample data related to the biological and / or chemical attribute for determining a biological and / or chemical attribute.

[0020] In another aspect, it relates to use of a classification model as described herein for determining a biological and / or chemical attribute.

[0021] In another aspect, it relates to use of further sample data associated with a first biological and / or chemical attribute for further training a classification model as described herein for determining a biological and / or chemical attribute.

[0022] In another aspect, it relates to a, in particular computer-implemented, method for classifying sample data related to a biological and / or chemical attribute according to the biological and / or chemical attribute, the method comprising: providing and / or obtaining, in particular receiving, the sample data, in particular via a user interface, providing the sample data to a classification model for determining if the sample data is associated with a biological and / or chemical attribute, wherein the classification model is trained to being provided by sample data and provide a biological and / or chemical attribute associated with the sample data, wherein the classification model is further trained based on a further sample data associated with a first biological and / or chemical attribute in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute, providing the biological and / or chemical attribute in response to determining that the sample data is associated with the biological and / or chemical attribute.

[0023] In another aspect, it relates to a, in particular computer-implemented, method for generating further sample data associated with a first biological and / or chemical attribute to the classification model, the method comprising: providing and / or obtaining, in particular receiving, preferably via an interface historical sample data, providing the historical sample data to a data generating model trained to provide sample data in response to receiving the historical sample data for generating the further sample data, providing the further sample data for further training a classification model based on the further sample data in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute, wherein the classification model is trained to being provided by sample data and provide a biological and / or chemical attribute associated with the sample data.

[0024] In another aspect, it relates to a, in particular computer-implemented, method for training a classification model for determining a biological and / or chemical attribute associated with sample data related to the biological and / or chemical attribute, the method comprising: providing and / or obtaining, in particular receiving, preferably via an interface, further sample data associated with a first biological and / or chemical attribute to the classification model, wherein the further sample data is generated by providing historical sample data and generating the further sample data based on the historical sample data training the classification model based on the further sample data associated with a first biological and / or chemical attribute to the classification model in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute, optionally providing the classification model.

[0025] In another aspect, it relates to use of the one or more biological and / or chemical attribute(s) determined as described herein for monitoring and / or controlling producing and / or processing of a biological and / or chemical product.

[0026] In another aspect, it relates to a classification model generated adapted according to synthetic sample data as generated as described herein based on a trigger as described herein for monitoring and / or controlling producing and / or processing of a biological and / or chemical product.

[0027] In another aspect, it relates to a method, in particular a computer-implemented method, for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the method comprising: providing and / or obtaining, in particular receiving, preferably via an interface, the sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data- driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, determining if a distance score indicative of a distance measure between the synthetic sample data and the sample data, preferably based on the synthetic sample data and the sample data is within a predefined range, in response to determining that the distance score is within the predefined score, providing a trigger for adapting the classification model.

[0028] In another aspect, it relates to a method, in particular a computer-implemented method, for monitoring and / or controlling one or more properties of a biological and / or chemical product based on one or more current biological and / or chemical attribute(s) related to current sample data associated with the biological and / or chemical product, wherein current sample data may preferably relate to sample data generated for monitoring the biological and / or chemical product associated with the current sample data, method comprising: providing and / or obtaining, in particular receiving, sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data- driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, determining if a distance score indicative of a distance measure between the synthetic sample data and the sample data, preferably based on the synthetic sample data and the sample data is within a predefined range, in response to determining that the distance score is within the predefined score, providing a trigger for adapting the classification model, providing current sample data to the adapted classification model for providing the current biological and / or chemical attribute associated with the current sample data, generating the current biological and / or chemical attribute associated with the current sample data by the classification model, providing the generated current biological and / or chemical attribute.

[0029] In another aspect, it relates to a method, in particular a computer-implemented method, for monitoring and / or controlling one or more properties of a biological and / or chemical product based on one or more current biological and / or chemical attribute(s) related to current sample data associated with the biological and / or chemical product, wherein current sample data may preferably relate to sample data generated for monitoring the biological and / or chemical product associated with the current sample data, method comprising: providing and / or obtaining, in particular receiving, current sample data associated with the biological and / or chemical product, providing the current sample data to a data-driven classification model configured, obtained and / or trained as described herein for determining the current biological and / or chemical attribute, providing the determined current biological and / or chemical attribute.

[0030] In another aspect, it relates to an apparatus for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the apparatus comprising: a sample data providing interface configured to provide, in particular receive, the sample data associated with the biological and / or chemical product, a classification generator configured to provide the sample data to a data-driven classification model and to determine by the data-driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), a synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, a quality measure provider configured to determine if a distance score indicative of a distance measure between the synthetic sample data and the sample data based on the synthetic sample data and the sample data is within a predefined range, a trigger provider configured to provide in response to determining that the distance score is within the predefined score a trigger for adapting the classification model.

[0031] In another aspect, it relates to a, in particular computer-implemented, method for obtaining classification model for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the method comprising: providing and / or obtaining, in particular receiving, preferably via an interface, the sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data- driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, wherein the synthetic sample data is associated with a biological and / or chemical attribute different than the biological and / or chemical attribute associated with the sample data, determining a distance score indicative of a distance measure between the synthetic sample data and the sample data based on the synthetic sample data and the sample data, based on the distance score, providing the synthetic sample data and the biological and / or chemical attribute associated with the synthetic sample data to the classification model for updating the classification model optionally providing the classification model.

[0032] Any disclosure, embodiments and examples described herein relate to the methods, the systems, apparatuses, uses, and computer elements lined out above and below. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples. EMBODIMENTS

[0033] In the following, embodiments of the present disclosure will be outlined by ways of embodiments and / or examples. It is to be understood that the present disclosure is not limited to said embodiments and / or examples.

[0034] The disclosure relates to the field of trustworthy artificial intelligence in chemical and pharmaceutical industry. In this environment safe use of artificial intelligence is particularly relevant, since falsepositive or unexplainable black box functionalities can lead to adverse effects. In particular, transforming the sample data with respect to the one or more biological and / or chemical attribute(s) to generate synthetic sample data, yields to a second set of data that can be counterfactual or adversarial to the sample data and with that yields a second instance to check the result. Further by providing the at least one classification quality measure relating to the classification determination and / or at least one synthetic quality measure relating to the synthetic data generation, the quality measures can be used to assess the result of the classification and to ensure the trustworthiness of the result. This is particularly advantageous for classification models classifying chemical and / or biological attributes of chemical and / or biological products, since these operations are particularly sensitive and false results can lead to adverse effects when monitoring and / or controlling chemical and / or biological products. In some implementation, the ground truth is determined by human experts. This takes a lot of time and relies on the concentration of the human experts. Hence, such implementations are not scalable and prone to errors. Hence, it is desired to provide an even more robust, transparent and scalable determination of biological and / or chemical attributes of biological and / or chemical products. A solution is proposed by this disclosure. Determining if a deviation is within a predefined range allows to determine whether retraining of the classification model may be necessary independent of human experts. The deviation score may be a measure for a meaningful deviation of the synthetic sample data from the sample data. The meaningful deviation may be a deviation related to a difference of the biological and / or chemical attribute associated with the synthetic sample data in comparison to the biological and / or chemical attribute associated with the sample data. Based on the deviation, i.e. if a meaningful change is observed, the classification model may be obtained. Therefore, the classification model may be adapted based on meaningful deviations. This allows to improve the robustness of the classification model while taking an appropriate number of resources with respect to the improvement. Since the deviation score is a measure obtainable independent of human experts, determining the deviation score enables scale up of the herein presented solution to large datasets common in complex production systems such as Verbund systems.

[0035] The one or more biological and / or chemical attribute(s) related to sample data may be associated with and / or may comprise one or more characteristic(s) or properties of the chemical and / or biological product. The one or more biological and / or chemical attribute(s) may relate to and / or may comprise one or more characteristic(s) or properties of the production process and / or processing of chemical and / or biological products. The one or more biological and / or chemical attribute(s) may relate to and / or may comprise one or more characteristic(s) or properties of the treatment with or of chemical and / or biological products. The one or more biological and / or chemical attribute(s) may relate to characteristics or properties directly or indirectly associated with the chemical and / or biological product. Properties of the chemical and / or biological product may be chemical properties, physical properties, biological properties and / or environmental properties. Chemical property may be a property that can be established only by changing one or more chemical structures associated with the at least one chemical and / or biological product. Examples for chemical properties may be acidity, oxidation state or reactivity. Physical property may be one of the following: mechanical properties, electrical properties, optical properties, thermal properties or the like. For example, physical property may comprise one or more of the following density, scratch resistance, electrical conductivity, color, absorption, heat capacity or the like.

[0036] In an embodiment, environmental property of the chemical and / or biological product may comprise at least one of emission data of the chemical and / or biological product, recyclate content of the chemical and / or biological product, bio-based content of the chemical and / or biological product, renewable content of the chemical and / or biological product, chemical and / or biological product declaration data, chemical and / or biological product safety data, share of the chemical and / or biological product to be recycled, share of the chemical and / or biological product being biodegradable, share of the chemical and / or biological product being non-degradable or a combination thereof. Emission data may comprise any data related to environmental footprint. The environmental footprint may refer to an entity and its associated environmental footprint. The environmental footprint may be entity specific. For instance, the environmental footprint may relate to a chemical and / or biological product, a company, a process such as a manufacturing process, a raw material or basic substance, a chemical and / or biological product or material, a component, a component assembly, an end product, combinations thereof or additional entity-specific relations. Emission data may include data relating to carbon footprint of a chemical and / or biological product. Emission data may include data relating to greenhouse gas emissions e.g. released in production of the chemical and / or biological product. Emission data may include data related to greenhouse gas emissions. Greenhouse gas emissions may include emissions such as carbon dioxide (CO2) emission, methane (CH4) emission, nitrous oxide (N2O) emission, hydrofluorocarbons (HFCs) emission, perfluorocarbons (PFCs) emission, sulphurhexafluoride (SFe) emission, nitrogen trifluoride (NF3) emission, combinations thereof and additional emissions. Emission data may include data related to greenhouse gas emissions of an entities or companies own operations (production, power plants and waste incineration). Scope 2 comprise emissions from energy production which is sourced externally. Scope 3 comprise all other emissions along the value chain. Specifically, this includes the greenhouse gas emissions of raw materials obtained from suppliers. Product Carbon Footprint (PCF) sum up greenhouse gas emissions and removals from the consecutive and interlinked process steps related to a particular product. Cradle-to-gate PCF sum up greenhouse gas emissions based on selected process steps: from the extraction of resources up to the factory gate where the product leaves the company. Such PCFs are called partial PCFs. In order to achieve such summation, each company providing any products must be able to provide the scope 1 and scope 2 contributions to the PCF for each of its products as accurately as possible, and obtain reliable and consistent data for the PCFs of purchased energy (scope 2) and their raw materials (scope 3).

[0037] The sample data associated with a biological and / or chemical product may relate to measurement data signifying characteristics or properties of the production process of chemical and / or biological products. The sample data may relate to measurement data signifying characteristics or properties of the production process of chemical and / or biological products. The sample data may relate to measurement data signifying characteristics or properties of the treatment with or of chemical and / or biological products. The sample data may relate to measurement data signifying characteristics or properties directly or indirectly associated with the chemical and / or biological product.

[0038] The data-driven classification model and determining by the data-driven classification model one or more biological and / or chemical attribute(s). The one or more biological and / or chemical attribute(s) may serve as classifiers for the data-driven classification model. The data-driven classification model may be configured to ingest sample data at the input layer and to determine one or more classifier(s) relating one or more biological and / or chemical attribute(s). One biological and / or chemical attribute may serve as binary classifier, i.e. true or false classifier signifying the presence or absence of the biological and / or chemical attribute. At least two or more biological and / or chemical attribute(s) may serve as classifiers. The classification model may be parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s). The classification model may be trained with historic sets of sample data and corresponding one or more biological and / or chemical attribute(s). The historic sets may include multiple sets of input-output-pairs including sample data and corresponding one or more biological and / or chemical attribute(s). The classification model may be trained based on the input-output-pairs including sample data and corresponding one or more biological and / or chemical attribute(s). The classification model may include a neural network or a deep neural network architecture including multiple layer, such as input, output and hidden layers. The classification model may include a purpose specific model which may be trained for monitoring and / or controlling a specific type of chemical and / or biological product based on the one or more biological and / or chemical attribute(s). The classification model may be based on a neural network architecture that fits to the specific purpose. Since the gist of the disclosure is applicable to any classifier model suitable to classify one or more biological and / or chemical attribute(s) based on sample data associated with the chemical and / or biological product any suitable neural network architecture such as CNN, RNN, LSTM or the like may be applicable. It is to be understood that the person skilled in the art is capable of choosing a suitable architecture depending on the task to be fulfilled by the classification model.

[0039] The data-driven synthetic data generator may relate to a counterfactual or an adversarial synthetic data generator. The synthetic data generator may determine synthetic sample data based on the classified chemical and / or biological attribute by transforming the sample data with respect to the one or more biological and / or chemical attribute(s) and by generating synthetic sample data. The data- driven synthetic data generator may include a generative model that is suitable to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and generate synthetic sample data. The data-driven synthetic data generator may include a flow-based model, an autoencoder or a Generative Adversarial Network (GAN), in particular a generative model associated with a GAN. Further, the data-driven synthetic data generator may include a diffusion model. The diffusion model may be based on adding noise to training data. Diffusion models may include a latent variable model that maps to the latent space using a fixed Markov chain. The diffusion model may map the input to a latent space and via diffusion maps to an output space. The GAN, in particular the generative model associated with the GAN may be trained to generate synthetic sample data by mapping a numerical representation associated with the synthetic sample data, in particular noise associated with the synthetic sample data, to the synthetic sample data. The synthetic sample data may be provided to a detector associated with the GAN. The detector may be trained, in particular together with the generative model, to determine if sample data may be real sample data or synthetic sample data. The GAN, in particular the generative model associated with the GAN may be trained to generate the synthetic sample data determined by the detector to be real sample data. The flowbased model may be a normalizing flow.

[0040] Classification quality measure relating to the classification determination may signify a classifier performance or a statistical measure of the prediction accuracy. Classification quality measure may signify the confidence interval, classification error or classification accuracy. Classification quality measure may relate to the classification accuracy or classification error as a proportion or a ratio. Classification quality measure may signify the proportion of correct or incorrect predictions made by the classification model.

[0041] Synthetic quality measure relating to the synthetic data generation may relate to the synthetic sample data, in particular the counterfactual sample data. Adapting the classification model may include training the classification model with synthetic sample data and corresponding classifiers.

[0042] In one embodiment the one or more biological and / or chemical attribute(s) relate to characteristic features embedded in the sample data. Characteristic feature in this context may refer to features based on which the classification model may generate or select the one or more biological and / or chemical attributes as classifier. Characteristic feature in this context may refer to features based on which the human expert may select one or more biological and / or chemical attribute(s) as classifier.

[0043] In the context of explainable artificial intelligence, the characteristic features based on which the classification model may select the classifier may relate to or coincide with the characteristic features based on which the human experts may select the classifier.

[0044] In another embodiment the sample data is provided by a measurement interface configured to retrieve sample data as measured by a sensor. The sensor measuring the sample data may include any sensor suitable for measuring signals that include characteristics or properties directly or indirectly related to the chemical and / or biological product. The chemical and / or biological product may be monitored and / or controlled based on the classified chemical and / or biological attribute. Such monitoring and / or controlling may be directly or indirectly linked to the chemical and / or biological product. For example, a production process of the chemical and / or biological product, a property of the chemical and / or biological product and / or a treatment of or to the chemical and / or biological product may be controlled and / or monitored. In another embodiment generating synthetic sample data includes providing sample data to a data- driven counterfactual data generator. The counterfactual generator may be configured to transform the sample data to a representation, such as latent space representation, encoding characteristic features embedded in the sample data. The counterfactual generator may be configured to transform the sample data to a numerical representation associated with the sample data. The numerical representation and / or the representation encoding characteristic features may be associated with a different space than the sample data. Hence, a format of the numerical representation associated with the sample data and / or the representation encoding characteristic features may be different than a format of the sample data. In particular, the numerical representation and / or the representation encoding characteristic features may be associated with less dimensions and / or less numerical values than dimensions and / or numerical values associated with the sample data.

[0045] The counterfactual generator may be configured to determine counterfactual sample data based on a change in the classified biological and / or chemical attribute, preferably associated with generating synthetic sample data based on the sample data. The change in the classified biological and / or chemical attribute may be a change of the classified biological and / or chemical attribute. The change of the classified biological and / or chemical attribute may refer to a different chemical and / or biological attribute associated with the synthetic sample data than the chemical and / or biological attribute associated with the sample data. Hence, change of the classified biological and / or chemical attribute may refer to changing the biological and / or chemical attribute by generating the synthetic sample data based on the sample data, in particular from the sample data.

[0046] The counterfactual data generator may include a pre-trained generative model. For transformation, the counterfactual data generator may be configured to map sample data resulting in a specific classification of biological and / or chemical attribute to a latent representation, such as a lower dimensional representation. The counterfactual data generator may be configured to determine in the latent space the counterfactual latent representation that corresponds to the minimal distance such that the prediction of the classifier is changed. Change in classifier in this context includes the change from the original biological and / or chemical attribute to at least one different biological and / or chemical attribute. In other words, in the latent space the counterfactual latent representation corresponding to at least one different biological and / or chemical attribute and the minimum distance to the decision boundary between the original biological and / or chemical attribute and the at least one different biological and / or chemical attribute may be determined. The counterfactual data generator may be configured to map the counterfactual latent representation to the counterfactual sample data. This way counterfactual sample data may be generated which in contrast to adversial sample data resembles characteristics features embedded in the sample data rather than noise. This way the counterfactual sample data may be used to resemble the decision strategies of the classification model.

[0047] In another embodiment generating synthetic sample data includes generating and / or storing a set of synthetic sample data from a set of sample data and providing per synthetic sample data of the set one or more chemical and / or biological attribute(s), at least one synthetic quality measure and / or at least one classification quality measure. The generated and / or stored a set of sample data may be selected based on the synthetic quality measure and / or classification quality measure. The generated and / or stored a set of sample data may include the sample data that on classification and / or synthetic data generation lead to at least one insufficient synthetic quality measure and / or classification quality measure. In other words, the generated and / or stored a set of sample data may include the sample data that on classification and / or synthetic data generation is identified to trigger adaption of the classifier model. The synthetic sample data may be stored on classification, if the synthetic quality measure or classification quality measure signify a confounder or outlier. Such synthetic data may be provided to the adaption process to focus the adaption and the related training of the classification model to the data sub space the model performance is lacking. In other words, such synthetic data may be provided to the adaption process to more efficiently and targeted train the classification model.

[0048] In another embodiment the at least one synthetic quality measure includes the generated synthetic data and / or at least one counterfactual explanation relating to characteristic features embedded in the sample data that cause the classification model to determine one or more chemical and / or biological attribute(s). Generating the synthetic sample data may further comprise determining a biological and / or chemical attribute associated with the synthetic sample data. The at least one synthetic quality measure may include an indication whether a biological and / or chemical attribute associated with the synthetic sample data may a target biological and / or chemical attribute. The synthetic sample data may be associated with a target biological and / or chemical attribute. The target biological and / or chemical attribute may be a biological and / or chemical attribute associated correctly with the synthetic sample data. The biological and / or chemical attribute determined to be associated with the synthetic sample data, e.g. by the data-driven synthetic data generator, may be associated correctly or incorrectly with the synthetic sample data. The human expert may be qualified to judge whether the biological and / or chemical attribute may be associated correctly or incorrectly with the synthetic sample data. The at least one synthetic quality measure may include an indication whether the synthetic sample data may be associated with a different biological and / or chemical attribute than the sample data.

[0049] The co unterf actual explanation may relate to the difference between the original sample data and the counterfactual sample data. The counterfactual explanation may relate to human expert input signifying conformity or divergence of the characteristic features based on which the classification model classifies the sample data and characteristic features based on which the human expert would classify the sample data. Via the counterfactual explanation specific characteristic features of the decision strategies of the classification model may be identified and analyzed. This gives the human user or a system for monitoring and / or controlling a reliable and robust option to check the performance of the classification model and ensure reliability specifically for monitoring and / or controlling based on chemical and / or biological attribute(s)

[0050] In another embodiment adapting the classification model is triggered, if the classification synthetic quality measure is indicative of a correlation not related to the mapping of sample data to the chemical and / or biological attribute expected by human experts. Adapting the classification model may be triggered, if the classification quality measure is indicative of a confounder or outlier. The classification quality measure may lie outside a pre-defined range or cross a pre-defined threshold signifying low performance of the classification model. In other words, classification quality measure may signify incorrect classification by the classification model based on a statistical probability measure. The trigger for adapting the classification model may be provided in response to determining that a biological and / or chemical attribute unequal to a target biological and / or chemical attribute may be associated with the synthetic data by the classification model. The target biological and / or chemical attribute may be a ground truth associated with the synthetic sample data. The ground truth associated with the synthetic sample data and / or the sample data may be determined by a human expert and / or provided via a user interface.

[0051] In another embodiment adapting the classification model is triggered, if the counterfactual explanation relating to characteristic features embedded in the sample data, that cause the classification model to determine one or more chemical and / or biological attribute(s), is indicative of a correlation not related to the mapping of sample data to the chemical and / or biological attribute expected by human experts. The correlation may include non-causal correlation to the mapping of sample data to the chemical and / or biological attribute expected by human experts. The chemical and / or biological attribute expected by the human expert may be referred to as expected chemical and / or biological attribute. The expected chemical and / or biological attribute may be provided e.g. via a data providing interface such as a user interface. Adapting the classification model may be triggered if the counterfactual explanation relating to characteristic features embedded in the sample data, that cause the classification model to determine one or more chemical and / or biological attribute(s), is indicative of a confounder or outlier. In such cases the result of the classification model may be unexplainable with reference to a huma expert assessment.

[0052] In another embodiment based on the trigger sets of sample data and / or corresponding synthetic sample data related to the classification quality measure and / or the synthetic quality measure that triggered adaption are provided. The sets of sample data and / or corresponding synthetic sample data may be stored on classification, if the synthetic quality measure or classification quality measure signify a confounder or outlier. Such sets of sample data and / or corresponding synthetic sample data may be provided to the adaption process to focus the adaption and the related training of the classification model to the data sub space the model performance is lacking. In other words, such sets of sample data and / or corresponding synthetic sample data may be provided to the adaption process to more efficiently and targeted train the classification model.

[0053] In another embodiment based on the trigger synthetic data for one or more set(s) of sample data is generated by providing the one or more set(s) of sample data to the data-driven synthetic data generator. Based on the trigger synthetic data for one or more set(s) of sample data may be provided for adapting the classification model. Such generation may be based on the sets of sample data and / or corresponding synthetic sample data related to the classification quality measure and / or the synthetic quality measure that triggered adaption are provided. The of sample data and / or corresponding synthetic sample data may be used as anchor point for generating additional synthetic sample data in the proximity of the provided synthetic sample data. This way the sub data space relating to the confounder can be mapped and the classification model can be trained in a targeted manner.

[0054] In another embodiment the generated synthetic data is provided to be associated with one or more biological and / or chemical attribute(s), e.g. by providing the synthetic sample data via a user interface and / or receiving one or more biological and / or chemical attribute(s) via the user interface. Further, the synthetic sample data may be provided for adapting the classification model. In an embodiment, the synthetic sample data may be provided together with the one or more corresponding biological and / or chemical attribute(s), in particular the one or more corresponding target biological and / or chemical attribute(s). Since for trustworthy artificial intelligence the human expert is the benchmark, the one or more biological and / or chemical attribute(s) may be provided by the human expert. The process may be referred to as human expert labelling. The one or more biological and / or chemical attribute(s) may be provided via a user interface to label the synthetic data or associate the one or more biological and / or chemical attribute(s) with the generated synthetic data.

[0055] In an embodiment, synthetic data may be synthetic sample data.

[0056] In another embodiment the synthetic data and associated biological and / or chemical attribute is provided to the classification model and the classification is adapted by training the classification model with the synthetic data and associated biological and / or chemical attribute. The adaption may include changing the weights of the classification model by the training process.

[0057] In another embodiment the adapted, data-driven classification model is provided for classifying sample data to one or more biological and / or chemical attribute(s). Further, the synthetic sample data and associated biological and / or chemical attribute may be provided to the classification model for training the classification model with the synthetic data and associated biological and / or chemical attribute. This way a classification model can be provided allowing for more reliable and robust monitoring and / or control of the chemical and / or biological product.

[0058] Chemical and biological processes require high certainty and include high resource needs. Usually, these processes comprise a plurality of dependent steps wherefore errors in chemical and biological processes will be propagated along the value chain. As these products are the basis for end consumer products it is of utterly importance to ensure high product quality.

[0059] Further training a classification model based on further sample data associated with a first biological and / or chemical attribute in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute allows for robust and reliable recognition and / or classification of sample data. By doing so, efficient resource use and safe process are assured on a large scale. By further training the classification model in response to providing the further sample data to the classification model and receiving a second biological and / or chemical attribute the classification model is tailored according to use case specifications. Thereby, the performance of the classification model can be adapted to the needs of biological and / or chemical processes. Ultimately, this results in an improved product to waste ratio contributing to lowering the impact of chemical and biological processes on the environment.

[0060] In an embodiment, biological and / or chemical attribute may be a first biological and / or chemical attribute and / or a second biological and / or chemical attribute and / or a third biological and / or chemical attribute and / or a fourth biological and / or chemical attribute. Biological and / or chemical attribute may be related to a biological property and / or chemical property and / or a physical property, may be indicative of a biological property and / or chemical property and / or physical property, may be a part of a biological and / or chemical object or a combination thereof. A biological object may include living organisms such as cells, animals, bacteria, viruses or the like. A chemical object may include one or more chemical compounds. A chemical property may be a property defined by the structure of the at least one chemical substance. Chemical property may be a property that can be established by changing the structure of the at least one chemical substance. Examples for chemical properties may be acidity, oxidation state or reactivity. Physical property may be indicative of a physical parameter associated with a sample. Physical property may be a property that can be established by changing a physical state of a sample. Physical property may be one of the following: mechanical properties, electrical properties, optical properties, thermal properties or the like. For example, physical property may comprise one or more of the following density, scratch resistance, electrical conductivity, color, absorption, heat capacity or the like. In particular, end-of-use property may refer to a property of the use product at the end of use stage. Biological property may be indicative of a biological parameter associated with a sample. Biological property may be a property that can be established by performing a biological process. Biological property may refer to toxicity, biological activity, biodegradability, growth rate, bioaccumulation or the like.

[0061] In an embodiment, the classification model may be suitable for classifying the sample data, the further sample data, the historical sample data and / or the historical further sample data according to the biological and / or chemical attribute, in particular first biological and / or chemical attribute and / or second biological and / or chemical attribute associated with the sample data, the further sample data, the historical sample data and / or the historical further sample data. The classification model may determine a confidence score indicative of the biological and / or chemical attribute, in particular first biological and / or chemical attribute and / or second biological and / or chemical attribute. The confidence score may specify the biological and / or chemical attribute, in particular first biological and / or chemical attribute and / or second biological and / or chemical attribute associated with the sample data, the further sample data, the historical sample data and / or the historical further sample data.

[0062] In an embodiment, data generating model may be trained and / or parametrized to provide further sample data in response to receiving indications of further sample data. The indication of further sample data may include at least one of at least a part of the further sample data, a representation of the further sample data, a random tensor and / or a combination thereof. The representation of the further sample data may include a text representation of the further sample data to be generated and / or a visual representation of the further sample data. Text representation of the further sample data may include a text description of the further sample data. Visual representation of the further sample data may include a sketch of the further sample data.

[0063] In an embodiment, the further sample data may be further sample image data, further sample numerical data, in particular further sample tabular data, further sample text data, further sample audio data or the like. Further sample data may be associated with a first biological and / or chemical attribute. In particular, further sample data may be represent a first biological and / or chemical attribute. Additionally or alternatively, a first biological and / or chemical attribute may be derivable based on the further sample data, preferably from the further sample data. In particular, first biological and / or chemical attribute may be derivable from the further sample data based on a feature of the further sample data. Where the further sample data may be further sample image data, a feature may be one or more parts of an image. Where the further sample data may be further sample numerical data, the feature may be a numerical value and / or a combination of two or more numerical values. Where the further sample data may be further sample tabular data, the feature may be at least a part of a table and / or a combination of two or more parts of a table. Where the further sample data may be further sample text data, the feature may be a part of a word and / or a combination of two or more parts of a word. Where the further sample data may be further sample audio data, the feature may be a sound and / or a combination of two or more sounds.

[0064] Further sample data may include synthetic data. Further sample data may indicate a biological and / or chemical attribute associated with a sample e.g. a synthetic sample. The further sample data may be generated based on historical sample data. At least a part of the historical sample data may be associated with a first biological and / or chemical attribute. Providing at least the part of the historical sample data to the classification model may result in receiving a second biological and / or chemical attribute from the classification model. Hence, the classification model may provide a biological and / or chemical attribute other than the biological and / or chemical attribute associated with the historical sample data. The at least one part of the historical sample data associated with the biological and / or chemical attribute, preferably the first biological and / or chemical attribute may be part of a validation training data set and / or a test training data set. The further sample data may be generated from at least the part of the historical sample data associated with the biological and / or chemical attribute, preferably the first biological and / or chemical attribute, by manipulating a representation of the historical sample data in a model space to generate a representation of the further sample data in the model space associated with a biological and / or chemical attribute other than the biological and / or chemical attribute associated with the representation of the historical sample data, in particular the second biological and / or chemical attribute. The further sample data may be generated from at least the part of the historical sample data associated with the biological and / or chemical attribute, preferably the first biological and / or chemical attribute, by selecting a representation of the further sample data associated with a biological and / or chemical attribute other than the biological and / or chemical attribute associated with the historical sample data. A representation of the historical sample data may be a dimensionality-reduced representation of the historical sample data. The representation of the historical sample data may be obtained by providing the historical sample data to an encoder. The representation of the historical sample data may be a tensor. A representation of the further sample data may be a dimensionality-reduced representation of the further sample data. The representation of the further sample data may be obtained by providing the further sample data to an encoder. The representation of the further sample data may be a tensor.

[0065] In an embodiment, historical sample data may be historical sample image data, historical sample numerical data, in particular historical sample tabular data, historical sample text data, historical sample audio data or a combination thereof, historical sample data may be associated with a historical biological and / or chemical attribute. In particular, historical sample data may be represent a historical biological and / or chemical attribute. Additionally or alternatively, a historical biological and / or chemical attribute may be derivable based on the historical sample data, preferably from the historical sample data. The classification model may be trained based on historical sample data indicating the historical biological and / or chemical attribute. In particular, historical biological and / or chemical attributes may be derivable and / or may be derived from the historical sample data based on a feature of the historical sample data. Where the historical sample data may be historical sample image data, a feature may be one or more parts of an image. Where the historical sample data may be historical sample numerical data, the feature may be a numerical value and / or a combination of two or more numerical values. Where the historical sample data may be historical sample tabular data, the feature may be at least a part of a table and / or a combination of two or more parts of a table. Where the historical sample data may be historical sample text data, the feature may be a part of a word and / or a combination of two or more parts of a word. Where the historical sample data may be historical sample audio data, the feature may be a sound and / or a combination of two or more sounds.

[0066] The biological and / or chemical attribute of the historical sample associated with the historical sample data may be available and / or provided and / or received.

[0067] In an embodiment, historical sample data may be obtained with a sensor. Hence, historical sample data may be sensor data. Historical sample data may be associated with a historical sample. The historical sample may be a sample for which the biological and / or chemical attribute may be available and / or known and / or provided.

[0068] In an embodiment, predefined range may be a first predefined range and / or a second predefined range. Second predefined range may be indicative of a numerical range associated with the second biological and / or chemical attribute. Additionally or alternatively, second predefined range may be indicative of a numerical range associated with the first biological and / or chemical attribute and / or the third biological and / or chemical attribute and / or the fourth biological and / or chemical attribute. The first predefined range may be indicative of a numerical range associated with the similarity score. The predefined range may be indicative of one or more threshold value(s) and / or one or more range(s) of numerical values.

[0069] In an embodiment, sample data may be sample image data, sample numerical data, in particular sample tabular data, sample text data, sample audio data or a combination thereof. Sample data may be associated with and / or may be indicative of a biological and / or chemical attribute. In particular, sample data may represent a biological and / or chemical attribute. Additionally or alternatively, a biological and / or chemical attribute may be derivable from the sample data, preferably from the sample data. In particular, biological and / or chemical attribute may be derivable from the sample data based on a feature of the sample data. In an embodiment, the one or more biological and / or chemical attribute(s) may be determinable and / or may be determined from the sample data. Where the sample data may be sample image data, a feature may be one or more parts of an image. Where the sample data may be sample numerical data, the feature may be a numerical value and / or a combination of two or more numerical values. Where the sample data may be sample tabular data, the feature may be at least a part of a table and / or a combination of two or more parts of a table. Where the sample data may be sample text data, the feature may be a part of a word and / or a combination of two or more parts of a word. Where the sample data may be sample audio data, the feature may be a sound and / or a combination of two or more sounds.

[0070] In an embodiment, sample data may be obtained with a sensor. Hence, sample data may be sensor data. Sample data may be associated with a sample. The sample may have a plurality of biological and / or chemical attributes. The sample data may be indicative of one or more biological and / or chemical attributes associated with the sample. The sample data may be indicative of the one or more biological and / or chemical attributes associated with the sample. In particular, the sample data may be a representation of the one or more biological and / or chemical attributes.

[0071] In an embodiment, the sample may be an object. Preferably, the sample may be biological and / or chemical object. The sample may include a biological material and / or a chemical material. The sample, in particular the one or more biological and / or chemical attributes associated with the sample, may be monitored. The sample, in particular the one or more biological and / or chemical attributes associated with the sample, may be monitored for determining a quality associated with the sample and / or for determining an application of a chemical and / or biological product to the sample.

[0072] In an embodiment, training the classification model my refer to and / or the training process may be a process of building the classification model, in particular determining and / or updating parameters of the classification model. During the training process, the classification model may adjust to achieve best fit with the training data set, e.g. relating the at least on input value with best fit to the at least one target output value. For example, if the neural network is a feedforward neural network such as a convolutional neural network (CNN), a backpropagation-algorithm may be applied for training the neural network. In case of a recurrent neural network (RNN), a gradient descent algorithm may be employed for training purposes. Gradient descent algorithm may use gradient for updating parameters. Gradient may indicate the degree of change for a parameter of the classification model. The gradient may be obtained by backpropagation. Thus, gradient descent algorithm may be based on backpropagation. A training process may be terminated when a deviation of the output generated by the classification model in comparison to a target output specified by the training data set falls within a predetermined range. The determining and / or updating of parameters of the classification model may be terminated when the training process may be terminated. The output generated by the classification model may be further sample data and / or biological and / or chemical attribute. The target output specified by the training data set may be the biological and / or chemical attribute and / or historical further sample data. Training the classification model may be associated with a training process and further training the classification model may be associated with a further training process.

[0073] These and other objects, which become apparent upon reading the following description, are solved by the subject matters of the independent claims. The dependent claims refer to embodiments of the disclosure.

[0074] In an embodiment, biological and / or chemical attribute may be a first biological and / or chemical attribute and / or a second biological and / or chemical attribute and / or a third biological and / or chemical attribute and / or a fourth biological and / or chemical attribute. Biological and / or chemical attribute may be related to a biological property and / or chemical property and / or a physical property, may be indicative of a biological property and / or chemical property, may be a part of a biological and / or chemical object or a combination thereof. A biological object may include living organisms such as cells, animals, bacteria, viruses or the like. A chemical object may include one or more chemical compounds. A chemical property may be a property defined by the structure of the at least one chemical substance. Chemical property may be a property that can be established by changing the structure of the at least one chemical substance. Examples for chemical properties may be acidity, oxidation state or reactivity. Physical property may be one of the following: mechanical properties, electrical properties, optical properties, thermal properties or the like. For example, physical property may comprise one or more of the following density, scratch resistance, electrical conductivity, color, absorption, heat capacity or the like. In particular, end-of-use property may refer to a property of the use product at the end of use stage. Biological property may refer to toxicity, biological activity, biodegradability, growth rate, bioaccumulation or the like. Biological and / or chemical attribute may be a class label associated with the data-driven classification model.

[0075] In an embodiment, historical sample data may be historical sample image data, historical sample numerical data, in particular historical sample tabular data, historical sample text data, historical sample audio data or a combination thereof, historical sample data may be associated with a historical biological and / or chemical attribute. In particular, historical sample data may be represent a historical biological and / or chemical attribute. Additionally or alternatively, a historical biological and / or chemical attribute may be derivable based on the historical sample data, preferably from the historical sample data. The classification model may be trained based on historical sample data indicating the historical biological and / or chemical attribute. In particular, historical biological and / or chemical attribute may be derivable from the historical sample data based on a feature of the historical sample data. Where the historical sample data may be historical sample image data, a feature may be one or more parts of an image. Where the historical sample data may be historical sample numerical data, the feature may be a numerical value and / or a combination of two or more numerical values. Where the historical sample data may be historical sample tabular data, the feature may be at least a part of a table and / or a combination of two or more parts of a table. Where the historical sample data may be historical sample text data, the feature may be a part of a word and / or a combination of two or more parts of a word. Where the historical sample data may be historical sample audio data, the feature may be a sound and / or a combination of two or more sounds.

[0076] In an embodiment, the method may further comprise providing historical sample data and generating the further sample data based on the historical sample data. The classification model may be trained based on the historical sample data. Generating the further sample data based on the historical sample data may refer to generating the further sample data by modifying at least one datapoint associated with the historical sample data to generate the further sample data.

[0077] In an embodiment, the classification model may be further trained in response to further determining that a similarity score indicative of the similarity of the historical sample data and the further sample data is within a first predefined range. In an embodiment, predefined range may be a first predefined range and / or a second predefined range. Second predefined range may be indicative of a numerical range associated with the second biological and / or chemical attribute. Additionally or alternatively, second predefined range may be indicative of a numerical range associated with the first biological and / or chemical attribute and / or the third biological and / or chemical attribute and / or the fourth biological and / or chemical attribute. The first predefined range may be indicative of a numerical range associated with the similarity score.

[0078] By doing so, the further sample data and the historical sample data are difficult enough from each other. This enables a significant change of the decision boundary associated with the classification model. Thus, this feature contributes tailoring the classification model to the use case specifications resulting in robust and reliable recognition and / or classification of sample data for efficient use of resources and safe processes. The similarity score may be indicative of a distance between the historical sample data and the further sample data in a data manifold associated with the historical sample data. The data manifold may be a lower-dimensional representation of the historical sample data used for training the classification model. The data manifold may represent the distribution of on of the historical sample data points within the historical sample data.

[0079] In an embodiment, generating the further sample data based on historical sample data may refer to providing the historical sample data to a data generating model and receiving the further sample from the data generating model. The data generating model may be a generative classification model. The data generating model may be suitable and / or may be trained for generating further sample data based on being provided by sample data, in particular historical sample data. The data generating model may be parametrized and / or trained based on sample data, in particular historical sample data and / or further sample data and / or historical further sample data. Historical further sample data may refer to further sample data generated based on historical sample data. The data generating model may be parametrized and / or trained to provide sample data in response to receiving the historical sample data for generating the further sample data. The data generating model may comprise an encoder and / or a decoder. The encoder may be suitable for generating a representation of the sample data by changing the dimensions of the sample data. The decoder may be suitable for projecting the representation of the sample data to further sample data. In an embodiment, the data generating model may be trained based on feedback from a classification model trained for classifying real data from synthetic data. Examples for the data generating model may include a machine-learning architecture associated with a variational autoencoder, a generative adversarial network, a normalizing flow or the like.

[0080] In an embodiment, generating the further sample data based on the historical sample data may refer to providing the historical sample data to a data generating model trained to provide sample data in response to receiving the historical sample data for generating the further sample data.

[0081] In an embodiment, the method may further comprise providing the further sample data via a user interface for determining if the further sample data may be associated with the first biological and / or chemical attribute and receiving the first biological and / or chemical attribute associated with the further historical data point in response to providing the further sample data via the user interface. The user interface allows a user to directly control the training of the classification model. By doing so, tailored and robust recognition and / or classification of chemical and biological processes is enabled.

[0082] In an embodiment, the data generating model is trained based on the historical sample data in response to receiving the historical sample data. The classification model may be trained based on historical sample data and wherein the data generating model may be trained based on the historical sample data to provide sample data in response to receiving the historical sample data. By doing so, the training data from the classification model can be recycled for training another data generating model. This saves time and resources in the overall process of recognition and / or classification of chemical and biological processes.

[0083] In an embodiment, the classification model may be trained to being provided by sample data and provide a biological and / or chemical attribute associated with the sample data, preferably in response to receiving the sample data.

[0084] In an embodiment, the second biological and / or chemical attribute may be received from the classification model in response to determining that a confidence score associated with the second biological and / or chemical attribute being associated with the further sample data may be within a second predefined range, preferably by the classification model. Receiving the second biological and / or chemical attribute from the classification model may comprise determining the confidence score associated with the second biological and / or chemical attribute being associated with the further sample data. The classification model may be trained to determine the confidence score associated with a biological and / or chemical attribute being associated with the sample data. The confidence score may indicate the biological and / or chemical attribute. Hence, receiving biological and / or chemical attribute from the classification model may refer to receiving the confidence score indicating the biological and / or chemical attribute from the classification model.

[0085] In an embodiment, the first predefined range and / or the second predefined range may be provided via a user interface. By doing so, intervention in the training of the classification model is enabled. Thereby, tailored and robust recognition and / or classification of chemical and biological processes is achieved. In an embodiment, the further trained classification model may be obtained in a further training process. In an embodiment, training the classification model my refer to and / or the training process may be a process of building the classification model, in particular determining and / or updating parameters of the classification model. During the training process, the classification model may adjust to achieve best fit with the training data set, e.g. relating the at least on input value with best fit to the at least one target output value. For example, if the neural network is a feedforward neural network such as a convolutional neural network (CNN), a backpropagation-algorithm may be applied for training the neural network. In case of a recurrent neural network (RNN), a gradient descent algorithm may be employed for training purposes. Gradient descent algorithm may use gradient for updating parameters. Gradient may indicate the degree of change for a parameter of the classification model. The gradient may be obtained by backpropagation. Thus, gradient descent algorithm may be based on backpropagation. A training process may be terminated when a deviation of the output generated by the classification model in comparison to a target output specified by the training data set falls within a predetermined range. The determining and / or updating of parameters of the classification model may be terminated when the training process may be terminated. The output generated by the classification model may be further sample data and / or biological and / or chemical attribute. The target output specified by the training data set may be the biological and / or chemical attribute and / or historical further sample data. Training the classification model may be associated with a training process and further training the classification model may be associated with a further training process.

[0086] The further training process may be terminated based on providing the second further sample data associated with a fourth biological and / or chemical attribute to the classification model and receiving the fourth biological and / or chemical attribute from the classification model. The classification model may be trained to receive sample data and provide a biological and / or chemical attribute associated with the sample data in a training process. The training process may be the further training process. The classification model may be further trained in the training process. The classification model may be trained prior to be further trained.

[0087] In an embodiment, the first biological and / or chemical attribute may be a third biological and / or chemical attribute and / or a fourth biological and / or chemical attribute. Alternatively, the second biological and / or chemical attribute may be a third biological and / or chemical attribute and / or a fourth biological and / or chemical attribute. In an embodiment, the classification model may be a classification model may be a second classification model. The classification model and / or the second classification model may be suitable for classifying sample data and / or further sample data according to the biological and / or chemical attributes. Determining if the sample data may be associated with the biological and / or chemical attribute may refer to classifying the sample data according to the biological and / or chemical attribute associated with the sample data. Determining if the further sample data may be associated with the first biological and / or chemical attribute may refer to classifying the further sample data according to the biological and / or chemical attribute associated with the further sample data.

[0088] In an embodiment, the classification model may be further trained based on a further sample data associated with a first biological and / or chemical attribute in response to providing the further sample data to the classification model for determining the if the sample data may be associated with the first biological and / or chemical attribute.

[0089] In an embodiment, any one of the methods may further comprise determining if the distance score may be within a predefined score. The trigger may be provided in response to determining that the distance score may be within the predefined score. In particular, providing a trigger for adapting the classification model based on the distance score may comprise providing the trigger in response to determining that the distance score may be within the predefined score. The predefined range may be provided via a user interface and / or may be defined by a human expert. This allows to intervene and / or control evaluation of the generated synthetic sample data. Hence, biological and / or chemical attributes of chemical and / or biological products can be determined in a scalable yet expert-controlled and transparent manner.

[0090] In an embodiment, the distance score may be associated with and / or related to a distance measure associated with the synthetic sample data and the sample data, in particular based on the synthetic sample data and the sample data. The distance measure may be a difference between the synthetic sample data and the sample data. The distance measure may characterize a change occurred when generating the synthetic sample data based on the sample data, in particular from the sample data. The distance measure may refer to and / or may characterize a similarity between the sample data and the synthetic sample data. The distance score may be a similarity score, in particular associated with and / or related to a similarity between the sample data and the synthetic sample data. The distance score may be for example an Euclidean distance, a cosine distance, an earth mover’s distance, a Manhattan distance, a minkowski distance, a hamming distance, a Chebyshev distance, a jaccard distance, a haversine distance, a Sorensen-dice distance and / or the like.

[0091] In an embodiment, determining if a distance score indicative of a distance measure associated with the synthetic sample data and the sample data based on the synthetic sample data and the sample data may be within a predefined range may comprise determining the distance score indicative of a distance measure associated with the synthetic sample data and the sample data based on the synthetic sample data and determining if the distance score may be within a predefined range. Determining the distance score and / or the distance measure may comprise determining a difference between the synthetic sample data and the sample data. Determining the distance score and / or the distance measure may comprise comparing the synthetic sample data and the sample data. Determining the distance score and / or the distance measure may comprise calculating one or more difference(s) between one or more numerical value(s) associated with the synthetic sample data and one or more numerical value(s) associated with the sample data. Determining the distance score may comprise determining a change score indicative of a change from the sample data to the synthetic sample data and determining a distance score based on the change score and a distance factor indicative of a weighting of the change score, e.g. by multiplying the change score with the distance factor. The distance factor may be provided via a user interface. Determining if the distance score may be within the predefined range may comprise comparing the predefined range and the distance score. By doing so, meaningful deviations can be tailored to the provided sample data. For example, NMR data may require different assessment of a meaningful deviation than an image of a cell. The predefined range can be selected e.g. by human experts. Hence, determining one or more biological and / or chemical attribute(s) can be scaled to large amounts of data while human experts are enabled to intervene and / or control evaluation of the generated synthetic sample data.

[0092] In an embodiment, the distance score may be further determined based on a segmentation mask. The segmentation mask may be indicative of a significance of a deviation per part of the sample data and / or the synthetic sample data. The segmentation mask may be indicative of one or more weighting factor(s) associated with one or more parts of the synthetic sample data and the sample data, determining the distance score based on the segmentation mask, the synthetic sample data and the sample data comprises weighting the distance associated with the synthetic sample data and the sample according to the segmentation mask. Determining the distance score may comprise determining a change score indicative of a change from the sample data to the synthetic sample data. In particular, the change score may be indicative of a change of the sample data upon generating the synthetic sample data from the sample data. Determining the distance score based on the segmentation mask, the synthetic sample data and the sample data comprise weighting the change score, e.g. by multiplying the change score with one or more numerical value(s) associated with the segmentation mask. Determining the distance score based on the segmentation mask, the synthetic sample data and the sample data may comprise determining at least one change score indicative of a change of the sample data upon generating the synthetic sample data from the sample data per part of the synthetic sample data and / or the sample data. The synthetic sample data and / or the sample data may comprise two or more parts. At least one distance score may be determined by weighting the at least two change scores according to the segmentation mask.

[0093] In an embodiment, determining the distance score indicative of a distance measure between the synthetic sample data and the sample data may comprise determining a plurality of partial distance scores between parts of the synthetic sample data and the sample data and determining a distance score from the partial distance scores, e.g. by accumulating the partial distance scores. Additionally or alternatively, determining the distance score indicative of a distance measure associated with the synthetic sample data and the sample data may comprise determining a plurality of distance scores indicative of one or a plurality of distance measure associated with the synthetic sample data and the sample data in relation to two or more parts of the synthetic sample data and the sample data.

[0094] In a spectrum or an image, different parts may be changed, but only a subsection of the spectrum or the image may be associated with a meaningful change. For example, where a component may be added, a peak within a respective area or a cell wall of a cell in the image may be related to the added component. Hence, weighting the parts of the sample data allows for attributing different influence of different parts of the provided sample data on a meaningful change, i.e. the distance.

[0095] In an embodiment, the segmentation mask may be provided via a data providing interface. The data providing interface may be a user interface. The segmentation mask may be generated based on human-expert input relating to significance. The data providing interface may be a user interface. For detecting true and false counterfactuals, i.e. if the deviation may be a meaningful deviation or not, the human-expert may be the benchmark. Hence, to arrive at highly accurate classification as necessary in the chemical and / or biological production environment, human experts may be desired to be integrated in a scalable manner. This can be facilitated by the above-described feature. By doing so, human expert to specify relevant parts of the data associated with a meaningful change while providing a scalable determination of biological and / or chemical attributes of chemical and / or biological products.

[0096] In an embodiment, the segmentation mask may be generated by providing the synthetic sample data and / or the sample data to a segmentation data-driven model, wherein the segmentation data-driven model may be trained based on historic sets of sample data and corresponding segmentation masks. The segmentation mask may have a format and / or a size of the synthetic sample data and / or the sample data. The segmentation mask may assign a weighting factor to parts of the synthetic sample data and / or the sample data. Hence, the segmentation mask may comprise a number of weighting factors. The number of weighting factors may be equal to the number of datapoints associated with the synthetic sample data and / or the sample data. By doing so, suitable segmentation masks may be generated for a variety of different sample data and / or synthetic sample data. For example, first spectral data may show different rotations of peaks compared to a second spectral data. This may require different segmentation masks. Hence, generating segmentation masks by the segmentation data-driven model allows to treat a variety of different synthetic sample data and / or sample data. Hence, this enables to scale the herein proposed solution to a variety of use cases. Ultimately, chemical and / or biological products can be determined for a variety of use cases in a transparent, reliable and scalable manner. The segmentation mask provided by the segmentation data-driven model may be provided via a data-providing interface, in particular a user interface, preferably for validating the segmentation mask by the human expert. The segmentation mask may be provided via a data providing user interface for a validation of the segmentation mask based a human-expert input relating to a validation of the segmentation mask. The segmentation mask may be human- interpretable. Hence, determining biological and / or chemical attributes associated with biological and / or chemical products may be scaled towards larger scales by generating segmentation mask by the segmentation model while still allowing the human experts to control the classification of the synthetic sample data via validating the synthetic sample data.

[0097] In an embodiment, the synthetic sample data related to the synthetic quality measure that triggered adaption may be provided. Hence, providing the trigger may result in and / or may trigger providing the generated synthetic sample data, in particular together with a biological and / or chemical attribute associated with the synthetic sample data. The classification model may be adapted based on the synthetic sample data and the associated synthetic biological and / or chemical attribute. Providing the trigger for adapting may trigger adapting the classification model according to the synthetic sample data. This allows for controlling the synthetic data generation and enables use of the synthetic data, e.g. for retraining the classification model.

[0098] In an embodiment, generating the synthetic sample data may further comprise determining a biological and / or chemical attribute associated with the synthetic sample data by the classification model. Determining the biological and / or chemical attribute associated with the synthetic sample data may comprise providing the synthetic sample data to the classification model for determining the biological and / or chemical attribute associated with the synthetic sample data. This allows to evaluate the classification by the classification data-driven model directly. Hence, an already built model can be improved. This saves resources for developing a new model but takes advantage of already available models. Followingly, this contributes to determining reliably and transparently attributes of chemical and / or biological products based on already available classification models.

[0099] In an embodiment, synthetic sample data may be generated iteratively until the synthetic sample data generated may be associated with a different biological and / or chemical attribute than the biological and / or chemical attribute associated with the sample data. Hence, generating the synthetic sample data may comprise: generating the synthetic sample data based on the sample data, in particular by decreasing a confidence score associated with generating the synthetic sample data by the data-driven synthetic data generator, and determining a biological and / or chemical attribute associated with the generated synthetic data by the classification data-driven model until the synthetic sample data generated may be associated with a different biological and / or chemical attribute than the biological and / or chemical attribute associated with the sample data and / or until the determined biological and / or chemical attribute differs from the biological and / or chemical attribute associated with the sample data. By doing so, synthetic sample data close to the decision boundary of the classification model may be identified. This enables shifting of a potentially falsely learned decision boundary. Hence, this enables adjustment of the classification model towards robust determination of attributes. Followingly, this contributes to determining reliably, transparently and in a scalable manner attributes of chemical and / or biological products.

[0100] In an embodiment, any one of the methods may further comprise adapting the classification model in response to providing the trigger for adapting the classification model. Providing the trigger for adapting may comprise adapting the classification model. Further, a trigger for determining one or more biological and / or chemical attribute(s) by the adapted classification model may be provided in particular in response to providing the trigger for adapting the classification model. In an embodiment, the trigger for determining one or more biological and / or chemical attribute(s) by the adapted classification model may be provided upon adapting the classification model. Providing the trigger for determining one or more biological and / or chemical attribute(s) by the adapted classification model may comprise and / or may result in determining one or more biological and / or chemical attribute(s) by the adapted classification model. Adapting of the classification model may be necessary upon detecting a misclassification by the classification model. A misclassification may refer to determining a deviation of the classification by the classification model from the classification by the humanexpert. The classification by the human expert may be referred to as expected classification. The expected classification may be provided e.g. via a data providing interface such as a user interface. Hence, this feature allows to improve a decision boundary of the classification model via a concrete and interpretable misclassification.

[0101] In an embodiment, any one of the methods may further include displaying an indication of the classification model being adapted, in particular in response to providing the trigger for adapting and providing the trigger for determining the one or more biological and / or chemical attribute(s) by the adapted classification model. Preferably the sample data and / or the synthetic sample data the trigger for adapting was provided and / or the classification model was adapted may be further provided, in particular together with the indication of the classification model being adapted.

[0102] In an embodiment, a distance between a numerical representation of the synthetic sample data may be within a predefined distance from a plurality of numerical representations associated with generator training data used for configuring and / or training the synthetic data generator. By doing so, meaningful synthetic sample data that may be similar to the sample data may be generated. This allows to create realistic sample data to prove the classification model. Hence, this feature allows to test the classification model under realistic conditions. As a consequence, the reliability of the classification model can be significantly improved while providing scalable classification. Numerical representation may be a projection of the synthetic sample data, the sample data and / or the generator training data into a representation space. The representation space may be associated with a different dimensionality than the sample data space associated with sample data, the generator training data and / or the synthetic sample data. In an embodiment, the method for determining one or more biological and / or chemical attribute(s) may be a method for adapting a classification model for determining one or more biological and / or chemical attribute(s). In an embodiment, the method for determining one or more biological and / or chemical attribute(s) may be a method for monitoring and / or controlling one or more biological and / or chemical attribute(s). In an embodiment, the method for determining one or more biological and / or chemical attribute(s) may be a method for adapting a classification model for monitoring and / or controlling one or more biological and / or chemical attribute(s). In an embodiment, monitoring and / or controlling a chemical and / or biological product may refer to monitoring and / or controlling production and / or processing of the biological and / or chemical product.

[0103] In an example embodiment, monitoring and / or controlling the biological and / or chemical product, in particular the one or more biological and / or chemical attribute(s) and / or the production and / or processing of the biological and / or chemical product may comprise determining if the biological and / or chemical product, in particular the biological and / or chemical product to be produced and / or processed and / or currently produced and / or processed and / or as previously produced and / or processed, corresponds to a target specification associated with the biological and / or chemical product. The target specification may be associated with a target characteristic and / or property of the biological and / or chemical product. In an example embodiment, monitoring and / or controlling the biological and / or chemical product, in particular the one or more biological and / or chemical attribute(s) and / or the production and / or processing of the biological and / or chemical product may comprise monitoring and / or controlling production and / or processing of the biological and / or chemical product according to the one or more biological and / or chemical attribute(s).

[0104] BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0105] In the following, the present disclosure is further described with reference to the enclosed figures.

[0106] The same reference numbers in the drawings and this disclosure are intended to refer to the same or like elements, components, and / or parts.

[0107] Fig. 1 illustrates an example of an apparatus for controlling and / or monitoring a chemical product based on one or more chemical attribute(s) associated with the input material to be used for production. Fig. 2 illustrates an example flowchart of a method for generating one or more chemical and / or biological attribute(s) associated with the chemical and / or biological product.

[0108] Fig. 3 illustrates an example of sample data and synthetic sample data.

[0109] Fig. 4 illustrates an example user interface for validating a classification by a human-expert.

[0110] Fig. 5 illustrates an example model structure including data-driven counterfactual model and data-driven classification model.

[0111] Fig. 6 illustrates schematically the latent space representation including the boundary between two classes.

[0112] Fig. 7 illustrate a schematic flow chart for deploying an adapted classification model.

[0113] Fig. 8 illustrate a schematic flow chart for using and adapting the classification model by an end-user.

[0114] Fig. 9 illustrates a system architecture for adapting the classification model.

[0115] Fig. 10 illustrates an embodiment of determining a distance score.

[0116] DETAILED DESCRIPTION

[0117] The following embodiments are mere examples for implementing the method, the system or application device disclosed herein and shall not be considered limiting.

[0118] Fig. 1 illustrates an example of an apparatus for controlling and / or monitoring a chemical product based on one or more chemical attribute(s) associated with the input material to be used for production. Chemicals are produced in large quantities via a plurality of processing steps. Small variations in production conditions can largely influence the biological and / or chemical attributes of the chemicals. This may include small variations of input material compositions used for producing the chemicals in one or more chemical reactions e.g. due to contamination or different suppliers with different production processes resulting in variations in the composition for the input material.

[0119] Because of the above-described sensible nature of chemical reactions, chemical production may require reliable monitoring and / or controlling of input materials provided to the chemical production facilities 108. One of the measures for monitoring and / or controlling input materials provided to chemical production facilities 108 may include verifying the quality of the input material by analyzing the composition of the input material 106. This may be typically done by non-invasive analysis tools such as spectroscopy. For example, the sample 106 may be illuminated by infrared light and the absorbance and / or reflectance of the infrared light by the sample 106 may provide an indication on the composition of the sample 106. Data from infrared spectroscopy may be an example for sample data 104 relating to properties of a sample of the input material 106 such as its composition. Further examples for measuring the composition of the input material 106 may include other spectroscopical data such as nuclear magnetic resonance spectroscopy, UVVis spectroscopy, electron spin resonance spectroscopy or the like.

[0120] Sample data 104 relating to properties of the input material may include sensor data. The sensor data may be generated by recording a signal associated with the sample of input material 106. The sample of input material 106 may be classified based on the sample data 104 with respect to the biological and / or chemical attribute(s) of the input material 106 as will be described in more detail e.g. in the context of Figs. 2 - 5. The biological and / or chemical attribute(s) may in this example include the composition of the input material.

[0121] The classification model may provide the biological and / or chemical attribute 120 to the chemical production facility 108. In particular, the biological and / or chemical attribute(s) may be provided to a control and / or monitoring engine of the chemical production facility 108. The control and / or monitoring engine may be configured to receive the biological and / or chemical attribute for monitoring and / or control the processing of multiple input material batches, from which samples 106 are taken and classified according to their respective biological and / or chemical attribute(s). Where the biological and / or chemical attribute indicates that one or a subset of input material batch(es) 106 are suitable for production or comply with quality constraints, the input material may be processed by the chemical production facility 108. The control and / or monitoring engine may provide instructions for providing the input material to the chemical production for producing a chemical product 140.

[0122] Otherwise, the control and / or monitoring engine of the chemical production facility 108 may reject the input material 106 and may require a different input material for producing the chemical product 140. This ensures a reliable production of chemical products 140 by determining the biological and / or chemical attribute associated with the input material 106.

[0123] The example described above shall be considered non limiting. Multiple other examples for using the enhanced classification approach disclosed herein to monitor and / or control chemical and / or biological operations exist. One example includes monitoring and / or controlling of biological treatment of plants, where the plant health may be classified based on images of the plant. Another example includes monitoring and / or controlling tissue analysis, where begin and non-begin tissue may be classified. Another example includes the monitoring and / or controlling of fermentation processes, where sugar and acid content may be classified based on electrochemical sensors measurements to classify different stages of the fermentation process.

[0124] Fig. 2 illustrates an example flowchart of a method for generating one or more chemical and / or biological attribute(s) associated with the chemical and / or biological product.

[0125] The sample data associated with the biological and / or chemical product may be provided e.g. via a data providing interface, preferably associated with a measurement apparatus configured to measure sample data. The sample data may relate to properties of the chemical and / or biological product. The sample data may relate to sample quality, application properties, technical properties, physiochemical properties, tissue properties or the like as described in the examples of Fig. 1 . The sample data may include property measurements of the chemical and / or biological product. The sample data may comprise analytical data associated with the chemical and / or biological product.

[0126] The sample data may be provided to a data-driven classification model for generating one or more biological and / or chemical attribute(s). The classification model may be parameterized and / or trained based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s). The classification model may include a neural network architecture specifically designed for the classification task with respect to the sample data. The classification may include two or more class(es) related to two or more different biological and / or chemical attribute(s). For example, the classification may relate to the biological and / or chemical attribute being true or false. Further for example, the classification may relate to the chemical attributes relating to input materials for chemical production. Further for example, the classification may relate to the biological attributes relating to crop health. Further for example, the classification may relate to the biological and / or chemical attributes relating to fermentation process stages. The classification model may be chosen based on the specific sample data type or measurement type and corresponding one or more biological and / or chemical attribute(s). The classification model may include a purpose specific classification model associated with the chemical and / or biological product and the corresponding one or more biological and / or chemical attribute(s). The classification model may by tailored to the classification of sample data associated with the chemical and / or biological product according to one or more biological and / or chemical attribute(s). The classification attributes may be pre-defined with respect to the sample data and associated chemical and / or biological product. Based on the generation of the one or more biological and / or chemical attribute(s), one or more classification measures relating to the quality of the classification generation may be determined. Such measures may include indicators based on the statistical probability such as accuracy or confidence level.

[0127] Synthetic sample data may be generated based on the chemical and / or biological attribute and the sample data. The sample data may be provided to a data-driven synthetic data generator comprising a data-driven counterfactual model to generate synthetic sample data. The synthetic sample data may be counterfactual data if the biological and / or chemical attribute may be different to the biological and / or chemical attribute associated with the sample data. The data-driven counterfactual model may be parameterized to transform the sample data related to the determined one or more biological and / or chemical attribute(s) to counterfactual sample data related to at least one different biological and / or chemical attribute class. The counterfactual sample data may hence represent a different classification of the one or more biological and / or chemical attribute(s) than the sample data. In other embodiments adversarial counterfactual sample data may be generated and provided. Adversarial counterfactual sample data, however, may impede interpretability, since noise or other artefacts may lead to the adversarial counterfactual sample data and not add to the interpretability of the classification models decision strategy. The advantage of non-adversarial counterfactual sample data is that the generation may utilize a generative model as regularizer and hence remains reasonable high density of the sample data distribution thus enhancing interpretability of the counterfactual sample data. Counterfactual sample data, hence, increases trustworthiness in assessing and interpreting the results of the classification model.

[0128] To generate the counterfactual data, the counterfactual model may be based on a pre-trained or selftrained generative model. The generative model may include any model adapted for sample data generation based on a transformation from the sample data distribution space to a latent space and vice versa. The generative model may be configured to ingest input data corresponding to sample data type. The generative model may be configured to transform sample data relating to the features relevant to the monitoring and / or controlling type and / or the chemical and / or biological product class.

[0129] The counterfactual model may be based on a gradient mechanism, such as a gradient ascent, which performs a gradient ascent step until the classifier changes to another classifier which is then related to the counterfactual sample data. The generation of such counterfactual sample data may include a transformation from the sample data distribution space to a lower dimensional latent space. The counterfactual model may be based on a normalizing flow, one or more Variational Autoencoder(s) (VAEs) or Generative Adversarial Network(s) (GANs). Normalizing flows may be based on learning an invertible transformation of a data distribution to a latent space, sampling in the latent space and generating data through inversion from the latent space to the data space. The latent space, latent feature space or embedding space may encode the sample data to a transformed representation. The latent space, latent feature space or embedding space may encode features associated with the chemical and / or biological attribute for monitoring and / or controlling the biological and / or chemical product into a latent space representation. In the case of normalizing flows models learn the mapping f: X -> Z, where X is the sample data distribution and Z is the distribution in a chosen latent space. Based on such mapping normalizing flows may generate data by sampling z~pZ, with pZ being the probability density of sampling z under distribution Z, and applying the inverse transformation f- 1 (z)=xgen to generate synthetic data. Normalizing flows can operate on sample data distribution and the corresponding latent space distribution with same dimensionality. For reducing the data processing and storage footprint with variational autoencoders or GANs, transformations to latent spaces with lower dimensionality than the sample data distribution space may be used. The counterfactual model may be based on a transformation of the sample data to a lower dimensional latent space. The lower dimensional latent space may include a latent space, latent feature space or embedding space, that encodes features of the sample data to a lower dimensional representation. The counterfactual model may include an Autoencoder or GAN-based model for synthetic data generation. Examples of such approaches are described in Ann-Kathrin Dombrowski, Jan E. Gerken, Klaus-Robert Muller, and Pan Kessel, “Diffeomorphic Counterfactuals with Generative Models”, eprint arXiv:2206.05075, June 2022; 10.48550 / arXiv.2206.05075. The determination of the counterfactual is described in more detail e.g. in the context of Figs. 2-6. Other options for generating counterfactuals include models based on denoising diffusion models. Denoising diffusion models may refer to models based on invertible Markov chains or learned ordinary or stochastic differential equations. The generated synthetic data, in particular the counterfactual data, may be classified by the classification data-driven model. The classification data-driven model may determine if the biological and / or chemical attribute associated with the synthetic sample data may be different from the biological and / or chemical attribute associated with the sample data. Hence, synthetic sample data may be generated iteratively, in particular until the synthetic sample data generated at a time step t may be associated with a different biological and / or chemical attribute than the biological and / or chemical attribute associated with the sample data and / or the synthetic sample data at time step t-1 , ie at a previous generation of synthetic sample data. Hence, synthetic sample data may be generated until the biological and / or chemical attribute associated with the synthetic sample data changes from the biological and / or chemical attribute associated with the sample data to another biological and / or chemical attribute. By doing so, synthetic sample data close to the decision boundary of the classification model may be generated. Hence, the true decision boundary can be found in relation to the decision boundary of the classification model. This allows for correcting the decision boundary efficiently and thus, explaining why a certain decision might be taken correctly and / or incorrectly. Hence, biological and / or chemical attributes are determined efficiently and robustly while being explainable.

[0130] If the decision boundary of the classification model is correct, the determined counterfactual data will be correct counterfactual data. If the decision boundary of the classification model is at least partially incorrect, the determined counterfactual data may be incorrect counterfactual data. The ground truth of the biological and / or chemical attribute associated with the incorrect counterfactual data may be the same biological and / or chemical attribute associated with the sample data. Hence, for determining if the decision boundary of the classification model may be correct, it may be assessed whether the generated synthetic data, in particular the counterfactual data may be a correct or an incorrect counterfactual. For this purpose, at least one synthetic quality measure may be determined. The synthetic quality measure may be related to associating the synthetic biological and / or chemical attribute with the synthetic sample data, i.e. indicative if the determined counterfactual data may be incorrect or correct counterfactual data.

[0131] The at least one synthetic quality measure may be an indication whether synthetic biological and / or chemical attribute determined by the classification model may correspond to a target biological and / or chemical attribute associated with the synthetic sample data, i.e. if the synthetic biological and / or chemical attribute may be assigned correctly or incorrectly. Said synthetic quality measure can be provided by domain experts. This is cumbersome and relies on humans providing always the correct synthetic quality measure. Followingly, synthetic quality measures determined by humans may underlay error rates due to humans losing concentration. Further, said evaluation is not scalable towards large amounts of data. Hence, it is desired to improve the robustness of determining biological and / or chemical attributes while allowing for scalability of determining biological and / or chemical attributes to large amounts of data.

[0132] This can be achieved by determining a distance measure associated with the synthetic sample data and the sample data. Further, an indication whether a distance measure associated with parts of the synthetic sample data and the sample data may be target distance measures or independent distance measures. A target distance measure may be a distance measure associated with at least parts of the synthetic sample data and the sample data associated with the distance measure of the biological and / or chemical attribute of the synthetic sample data and the sample data. For example, where the data comprises images, the target distance measure may be a distance measure associated with a predefined part of the images. For example, where the attribute may indicative of a presence of a chemical and / or biological substance, a predefined area of a spectrum may indicate the presence of the chemical and / or biological substance. Hence, correct counterfactual data may be images where at least a part of the predefined image area may be changed in comparison to the sample data. Incorrect counterfactual data or adversarial data may be images where at least a part of the predefined image area may be unchanged in comparison to the sample data. An example for determining a distance measure between the synthetic sample data, in particular the counterfactual data and the sample data may be seen in Fig. 10.

[0133] Where the synthetic sample data may be incorrectly classified by the classification model, the classification model may be adapted to correct the decision boundary of the classification model. Hence, based on the provided classification quality measure and / or the synthetic quality measure adaption of the classification model may be triggered, and an adapted classification model may be provided. For example, if the counterfactual sample data generated by the counterfactual model indicate that the model bases its decision on features, which are not considered robust and / or causal, the decision logic of the classification may not focus on the features of the sample data that would cause a human expert to take classification decisions. In such instances, the classification model may be view non-robust and / or biased and hence adaption may be triggered to increase trustworthiness for a user of the classification model or a system implementing the classification model. If the classification model classified the biological and / or chemical attribute in relation to the synthetic sample data correctly, the classification model may not be adapted and / or may be provided as is. In an embodiment, more synthetic sample data may be generated and synthetic quality measure may be determined. In particular, one set of synthetic sample data may be generated per set of sample data. The synthetic sample data may be generated until sets of synthetic sample data may be generated for each set of sample data.

[0134] Based on the trigger, the synthetic sample data, in particular the counterfactual data, may be provided, e.g. for training the classification model. By training the classification model with synthetic data the classification model may be adapted. The synthetic data may be provided for classification e.g. to a human expert user. The synthetic data may be labelled with the one or more chemical and / or biological attribute(s) e.g. by a human expert user. The synthetic data and the corresponding one or more chemical and / or biological attributes may be provided. The synthetic data may be generated by providing the sample data of the sample data set to the counterfactual model. This way the training data set of synthetically generated data can be focused on the specific data sub-spaces the classification does not perform adequately. By using specific sample data sets and the corresponding counterfactuals, the classification boundaries between two or more classifications may be mapped out and the model may be trained specifically focused on such sub-space(s). The synthetic data may or may not take historical sample data depending on the anchoring such historical data embedded into the model on training. Preferably the synthetic data does not include historic sample data sets that were used for training the classification model. By using the class label with human feedback on the new sample data set, the training of the classification model can be enhanced. In addition, by generating synthetic data based on such new data, the training data set can be enlarged. Lastly, by using labelled counterfactual training data, the training quality can be enhanced.

[0135] Adaption of the classification model by training with labelled counterfactual sample data may include an iterative loop mapping around tracked sample data sets until the labelling of the counterfactual sample data and the classification by the classification model converge. This way the classification model may be adapted by training the classification model with synthetic sample data generated based on the tracked sample data set(s).

[0136] Once the classification model is trained based on the synthetic data, the adapted classification model may be provided for more robust classification of sample data. Fig. 3 illustrates an example of sample data and synthetic sample data.

[0137] The example shown is included for illustrating purposes only and is not considered to be limiting. The example shows facial images of a person. These images may be examples of sample data. The image on the left side may be an example of provided sample data. The image may show the face of the person, in particular the facial expression of the person. The images, in particular the sample data, may be classified according to the facial expression of the person, i.e. smiling or non-smiling. The sample data may be classified by the classification model. For signifying a decision boundary associated with the classification model, synthetic sample data may be generated. Examples of synthetic sample data may be seen on the right side. Synthetic sample data may be for example generated from the sample data. Synthetic sample data may comprise a counterfactual, in particular in relation to the sample data. The counterfactual may be synthetic sample data generated from the sample data and may be associated with a different biological and / or chemical attribute than the sample data. In the concrete example, the sample data may be associated with the label “nonsmiling”. A counterfactual may be associated with smiling. Two synthetic images may be generated from the sample data as it can be seen in Fig. 3. The upper image may show a non-meaningful deviation, i.e. may be associated with the same label as the sample data. The non-meaningful deviation may refer to a change of a watermark in the lower right part of the upper image. Said changes may often be recognized.

[0138] The lower image may show a meaningful deviation, i.e. may be associated with a different label than the sample data. Hence, the upper image may be a falsely classified counterfactual. The lower image may be a correctly classified counterfactual.

[0139] This allows for more reliable classification and more reliable use of the classification for further processing.

[0140] Analogous to the example with the facial expression of the person, chemical and / or biological products may be classified according to an appearance of the products via images or according to numerical data characterizing the products, e.g. spectra or tabular data (not shown).

[0141] Fig. 4 illustrates an example user interface for validating a classification by a human-expert. In the example, the falsely classified counterfactual may be shown in comparison to the original image. The user interface may further display the attributes associated with the original image and the synthetically generated counterfactual to allow a human-expert to review assigning the displayed label to the counterfactual and / or to review whether the synthetic sample data, i.e. the assumed counterfactual, may be a true counterfactual. The human-expert may be allowed to accept or decline the displayed classification via the user interface.

[0142] Fig. 5 illustrates an example model structure including data-driven counterfactual model and data- driven classification model.

[0143] In this example the data-driven counterfactual model may be based on an autoencoder structure including encoder and decoder. In other examples diffusion models may be used. The data-driven classifier model may be based on any neural network architecture usable for monitoring specific chemical and / or biological product(s) by determination of the one or more chemical and / or biological attribute(s). In the example shown a simple architecture including multi-layer classifier, feature map, softmax and other layers is illustrated. This shall be considered non limiting and more complex architectures are feasible depending on the monitoring and / or controlling task. The classification model may be trained based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s). The classification model may hence be configured to provide one or more biological and / or chemical attribute(s) based on sample data. The classification model may be trained separately from the data-driven counterfactual model.

[0144] The counterfactual model may include a pre-trained generative model. The sample data resulting in a specific classification of biological and / or chemical attribute may be transformed to a latent representation, such as a lower dimensional representation, by the autoencoder-based model, specifically the encoder part of the model. In the latent space the counterfactual latent representation may be determined that corresponds to the minimal deformations xO = x+5x such that the prediction of the classifier is changed. In other words, in the latent space the counterfactual latent representation corresponding to at least one different biological and / or chemical attribute and the minimum crossing of the decision boundary to the at least one different biological and / or chemical attribute may be determined. Fig. 6 illustrates schematically the latent space representation including the boundary between two classes. Via the autoencoder-based model, specifically the encoder part of the model the counterfactual latent representation may be transformed to the counterfactual sample data. The counterfactual sample data may be provided to the data-driven classification model to determine the one or more biological and / or chemical attribute(s). This way the synthetic data may be generated that corresponds to counterfactuals of sample data and may be used for re-training the classification model.

[0145] Fig. 7 illustrate a schematic flow chart for deploying an adapted classification model.

[0146] On deployment of the classification model the trained classification model may be checked for confounders, which may be eliminated by performing the methods disclosed herein. An initial classification model may be provided and trained based on an initial object data comprising a historic set of data including input-output pairs of sample data and corresponding biological and / or chemical attribute. The classifier or classification model may be tested during training based on a test dataset and may be checked for confounding variables. If none are found the model may be deployed.

[0147] If confounding variables are found, synthetic sample data may be generated by the synthetic data generator including a generative model in connection the classification model as described e.g. in the context of Figs. 2-6. The counterfactual sample data may be generated and classified based on a distance score indicative of a distance measure associated with the synthetic sample data and the sample data. Optionally, the classification may be provided for verification by a human expert to a user interface. The classification model may be adapted based on the training data set of the synthetic sample data and corresponding labels, i.e. attributes. The adapted classification model may be checked with regard to model performance. Performance may relate to the classification and / or synthetic data quality measure as described e.g. in the context of Figs. 2 to 6. If the performance is sufficiently accurate, the model may be deployed. If the performance is not sufficiently accurate, the model may be further retrained based on synthetic sample data generated. In the workflow shown, the synthetic sample data may be labeled by a human-in-the-loop. As described throughout this disclosure, this may be facilitated by a distance score and predefined range specified by the humanexpert. The human-expert may still be in control while speeding up data processing as the humanexpert does not have to assess every data point. Furthermore, the human-expert may interfere at, e.g., preselected, points in time where predefined criteria may be met.

[0148] Fig. 8 illustrate a schematic flow chart for using and adapting the classification model by an end-user. Similarly, to Fig. 7 the initial sample data may be provided to the synthetic data generator to generate counterfactual sample data for labelling. The classification model may be constantly or intermittently checked with regard to model performance. If the performance is not sufficiently accurate, the model may be further retrained based on synthetic data generated until performance is sufficient. In contrast to Fig. 7 the flow chart shown in Fig, 8 allows for full end user control. In particular the methods, apparatuses and uses disclosed herein are based on counterfactuals which allow for an easy and reliable use by end users. Moreover, the intuitive nature of counterfactuals and the fully automatable training procedure based on synthetic sample data enables end users to ensure the mode robustness.

[0149] Fig. 9 illustrates a system architecture for adapting the classification model.

[0150] The system executing the methods illustrated in Fig. 7and 8 may include at least one data base storage storing the initial or historic sample data and the synthetic sample data. On first training a classification model provider may be configured to train the classification model based on the initial data set of sample data and corresponding biological and / or chemical attributes. In addition, the a generative model provider may be configured to train the pr-trained generative model further on such data. Based on the sample data and the corresponding biological and / or chemical attribute the synthetic data generator may be configured to generate counterfactual sample data. The counterfactual data may be provided to a user interface for labelling by a human expert. As described earlier, labelling every data point solely by a human-expert may take a lot of time and may rely on the concentration of the human-expert throughout a high amount of data points. Thus, the system may comprise a distance determination engine configured to determine the distance measure associated with the synthetic sample data and the sample data. The synthetic sample data may be classified according to the distance score. Optionally, the classification may be assessed as a random sample or when predefined criteria may be met. Hence, the system may optionally comprise a user interface for validating the classification by human-expert input. The labeled counterfactual sample data may be provided to a storage for storing and using the data for adapting the classification model by training.

[0151] The trustworthy Ai approach line out herein enables safe and reliable monitoring and / or controlling of chemical and / or biological products based on the classified chemical and / or biological attribute. Particularly in chemical and pharmaceutical operations safe use of artificial intelligence is important, since false-positive or unexplainable black box functionalities can lead to adverse effects. Fig. 10 illustrates an embodiment of determining a distance score.

[0152] Facial images are chosen to allow for a clear distinction between changed features. The attribute corresponds to a smiling or non-smiling person in the image. This example is not intended to limit the scope but for illustrative purposes only. Spectrums may show minor and less prominent changes. Thus, the example in Fig. 10 should allow for better understanding.

[0153] From the sample data in the bottom part of the Figure, two sets of synthetic sample data may be generated. The classification model may determine the attribute associated with the two set of synthetic sample data to be smiling. The classification model may determine the attribute associated with the sample data to be non-smiling. Hence, the classification model indicated that the two sets of synthetic sample data may be sets of counterfactual data. A difference between the sets of synthetic sample data and the sample data may be determined per set of synthetic sample data. Hence, pixel values associated with a first set of the synthetic sample data (left side) may be subtracted from the sample data. This may result in the absolute difference on the left side. Further, pixel values associated with a second set of the synthetic sample data (right side) may be subtracted from the sample data. This may result in the absolute difference on the right side. A segmentation mask may be provided. The segmentation mask may be indicative of a positive contribution to a distance measure, i.e. a target distance measure, or a negative contribution to the distance measure, i.e. an independent deviation. However, the segmentation mask may indicate a weighting of parts of the absolute difference. The absolute difference may be an example of a change score. For example, the segmentation mask may indicate +1 for a region associated with the face, 0 for a region associated with a background and -1 associated with a region associated with the number. Further, gradations are possible. The segmentation mask may be provided e.g. via an interface or by a user. Further, the segmentation mask may be learned by a segmentation data-driven model. The segmentation data- driven model may be trained based on sample data and corresponding segmentations. The segmentation data-driven model may be configured to provide a segmentation of sample data in response to receiving sample data. The segmentation may be indicative of a relation between at least a part of the sample data and the one or more biological and / or chemical attributes. Hence, segmentation data-driven models may be trained specifically for one or more biological and / or chemical attributes. Thus, the segmentation data-driven model may be selected based on the one or more biological and / or chemical attribute to be determined. Further, the segmentation may be provided by a human expert. As the segmentation is provided once by an expert or by an algorithm, determining the distance score may be scaled to large amounts of data. For example, where cells may be evaluated as described in the context of Fig. 3, the segmentation mask may be indicative of features associated with benign or non-benign tissue.

[0154] From the segmentation and the absolute difference, a distance score may be determined. For example, the distance score may be determined by multiplying the absolute difference, eg in pixel values or contrasts, with the segmentation mask, e.g. with +1 or -1. A distance score may be indicative of a degree of the change of predefined parts of the synthetic sample data and the sample data. The distance score may relate the change of predefined parts of the synthetic sample data and the sample data to a change of other parts of the synthetic sample data and the sample data. For example, where the degree of change of predefined parts of the synthetic sample data in relation to the sample data prevail, the distance score may be positive. Where the degree of change of nonpredefined parts of the synthetic sample data in relation to the sample data prevail, the distance score may be negative. A positive distance score may indicate correct counterfactual data. A negative distance score may indicate incorrect counterfactual data. Other implementations of the distance score may be possible. Predefined ranges associated with the distance score may be provided e.g. a via a data input interface such as a user interface. This allows for scalable control of evaluating the synthetic sample data by experts. Hence, Biological and / or chemical attributes can be determined in a scalable yet controlled manner.

[0155] Additionally or alternatively, the distance score may be determined based on an earth mover’s distance. In particular, the distance score may comprise the earth mover’s distance associated with the synthetic sample data and the sample data.

[0156] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims.

[0157] Any steps presented herein can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps. It is also not required that the different steps are performed at a certain place or in a certain computing node of a distributed system, i.e. each of the steps may be performed at different computing nodes using different equipment / data processing. As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.

[0158] In the claims as well as in the description the word “comprising” or “including” or similar wording does not exclude other elements or steps and shall not be construed limiting to the elements or steps lined out. The indefinite article “a” or “an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or further elements may be included.

[0159] Providing in the scope of this disclosure may include any interface configured to provide data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Providing may include communication of data or submission of data to the interface, in particular display to a user or use of the data by the receiving entity.

[0160] Any disclosure and embodiments described herein relate to methods, systems, apparatuses, devices, chemicals, materials, services, uses, computer program elements lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.

[0161] All terms and definitions used herein are understood broadly and have their general meaning.

Claims

Claims1 . A method for determining one or more biological and / or chemical attribute(s) associated with a biological and / or chemical product, wherein the one or more biological and / or chemical attribute(s) are determinable from the sample data the method comprising: providing the sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data-driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, determining a distance score indicative of a distance measure between the synthetic sample data and the sample data, based on the distance score, providing a trigger for adapting the classification model and optionally a trigger for determining one or more biological and / or chemical attribute(s) by the adapted classification model.

2. The method according to claim 1 , wherein the one or more biological and / or chemical attribute(s) may be associated with one or more characteristic(s) or properties of the chemical and / or biological product.

3. The method of claim 1 or 2, wherein the synthetic sample data generated is associated with a different biological and / or chemical attribute than a biological and / or chemical attribute associated with the sample data.

4. The method of any one of the preceding claims, further comprising determining if the distance score is within a predefined range, and wherein the trigger is provided in response todetermining that the distance score is within the predefined range, and wherein the distance score is further determined based on a segmentation mask indicative of a significance of a deviation per part of the sample data and / or the synthetic sample data.

5. The method of any one of the preceding claims, further comprising determining if the distance score is within a predefined range, and wherein the trigger is provided in response to determining that the distance score is within the predefined range, and wherein the distance score is further determined based on a segmentation mask indicative of one or more weighting factor(s) associated with one or more parts of the synthetic sample data and the sample data, wherein the distance score is associated with a significance of a deviation per part of the sample data and / or the synthetic sample data.

6. The method of any one of the preceding claims, wherein the distance score is determined based on a segmentation mask indicative of a significance of a deviation per part of the sample data and / or the synthetic sample data, and wherein the segmentation mask is provided via a data providing interface, wherein the segmentation mask is generated based on provided human-expert input relating to a significance of a deviation per part of the sample data and / or the synthetic sample data.

7. The method of any one of the preceding claims, wherein the distance score is determined based on a segmentation mask indicative of a significance of a deviation per part of the sample data and / or the synthetic sample data, and wherein the segmentation mask is generated by providing the synthetic sample data and / or the sample data to a segmentation data-driven model, wherein the segmentation data-driven model is trained based on historic sets of sample data and corresponding segmentation masks.

8. The method of any one of the preceding claims, wherein the distance score is determined based on a segmentation mask indicative of a significance of a deviation per part of the sample data and / or the synthetic sample data, and wherein the segmentation mask is provided via a data providing user interface for a validation of the segmentation mask based on provided human-expert input relating to a validation of the segmentation mask.

9. The method of any one of the preceding claims, wherein generating the synthetic sample data may further comprise determining a biological and / or chemical attribute associated with the synthetic sample data by the classification model.

10. The method of any one of the preceding claims, further comprising adapting the classification model in response to providing the trigger for adapting the classification model.11 . The method of any one of the preceding claims, wherein based on the trigger, the synthetic sample data that triggered adaption are provided, in particular via a data providing interface.

12. A method for obtaining classification model for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the method comprising: providing the sample data associated with the biological and / or chemical product, providing the sample data to a data-driven classification model and determining by the data-driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), generating synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, wherein the synthetic sample data is associated with a biological and / or chemical attribute different than the biological and / or chemical attribute associated with the sample data, determining a distance score indicative of a distance measure between the synthetic sample data and the sample data based on the synthetic sample data and the sample data, based on the distance score, providing the synthetic sample data and the biological and / or chemical attribute associated with the synthetic sample data to the classification model for updating the classification modeloptionally providing the classification model.

13. An apparatus for determining one or more biological and / or chemical attribute(s) related to sample data associated with a biological and / or chemical product, the apparatus comprising: a sample data providing interface configured to provide the sample data associated with the biological and / or chemical product, a classification generator configured to provide the sample data to a data-driven classification model and to determine by the data-driven classification model one or more biological and / or chemical attribute(s), wherein the classification model is parameterized based on historic sets of sample data and corresponding one or more biological and / or chemical attribute(s), a synthetic data provider configured to generate synthetic sample data by providing sample data to a data-driven synthetic data generator and determining synthetic sample data based on the classified chemical and / or biological attribute, wherein the data-driven synthetic data generator is configured to transform the sample data with respect to the one or more biological and / or chemical attribute(s) and to generate synthetic sample data, a quality measure provider configured to determine if a distance score indicative of a distance measure between the synthetic sample data and the sample data based on the synthetic sample data and the sample data is within a predefined range, a trigger provider configured to provide in response to determining that the distance score is within the predefined score a trigger for adapting the classification model.

14. Use of the one or more biological and / or chemical attribute(s) determined according to the methods of claims 1 to 13 for monitoring and / or controlling producing and / or processing of a biological and / or chemical product.

15. Use of a classification model generated adapted according to synthetic sample data as generated according to any one of claims 1 to 13 based on a trigger according to any one of claims 1 to 13 for monitoring and / or controlling producing and / or processing of a biological and / or chemical product.

Citation Information

Patent Citations

  • Generating machine learning data in salient regions of a feature space

    US11509674B1

  • Systems and methods for algorithmically estimating protein concentrations

    WO2022246224A1

  • Synthetic generation of training data

    WO2023242236A1

  • Processes, machines, and articles of manufacture related to predicting effects of combinations of items

    WO2024015798A1