Determining the stability of a substance or substance mixture
The method leverages near-infrared spectroscopy and chemometric analysis to efficiently assess the stability of substance mixtures, overcoming resource-intensive traditional methods by quantifying changes using mathematical distance measures and providing intuitive data output.
Patent Information
- Application Number
- PCT/EP2023/081994
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-21
- Filing Date
- 2023-11-16
- Publication Date
- 2025-07-10
AI Technical Summary
Traditional stability studies in pharmaceutical, cosmetic, and food sectors require extensive personnel and equipment resources due to the need for numerous analytical methods, leading to inefficiencies and potential oversight of critical stability parameters.
A computer-implemented method using near-infrared spectroscopy and chemometric analysis to determine the stability of substance mixtures by analyzing holistic chemical profiles, incorporating various data sets, and employing mathematical distance measures to quantify changes over time.
Enables rapid, comprehensive evaluation of complex mixtures with reduced resources, identifying the most stable candidates for further investigation, and providing intuitive data output without requiring extensive programming knowledge.
Smart Images

Figure EP2023081994_10072025_PF_FP_ABST
Abstract
Description
[0001] DETERMINATION OF THE STABILITY OF A SUBSTANCE OR MIXTURE
[0002] TECHNICAL FIELD
[0003] The present invention relates generally to the technical field of analytics, in particular pharmaceutical analytics.
[0004] BACKGROUND OF THE INVENTION
[0005] The development of new products, for example in the pharmaceutical, cosmetic, and food sectors, is subject to an increasingly complex list of requirements and strict quality criteria. Stability studies play a particularly important role in such developments, covering a wide range of potential influencing factors (e.g., chemical or physical incompatibilities between molecules in a product, resistance to temperature and / or humidity, resistance to UV / VIS radiation, oxidation, interactions between different reactants, microbial contamination, biodegradation, etc.).
[0006] According to established approaches, testing each parameter in a product individually for its response to the factors mentioned above and others can only be characterized using a large number of different technical methods. This typically requires several different HPLC (high-performance liquid chromatography) methods for different analytes, wet-chemical colorimetric group reactions, moisture determination, and testing of physical parameters (e.g., breaking strength, friability, and disintegration time of tablets; flowability, bulk and tapped density of powders; turbidity measurements of liquids; and many more).
[0007] In order to map as many of these potential interactions as possible, it is common practice to plan projects using a "Design of Experiments" (DoE, synonym: statistical experimental design). This usually takes into account not only all possible qualitative compositions, but also the quantitative ratios of the substances contained. In addition, it can also be considered, for example, which packaging (primary and / or secondary or additional packaging) best protects a product from alteration. Packaging refers to products used to package products. If all of these aspects are to be integrated into a test plan, the number of product variants is multiplied again by the packaging options to be tested.
[0008] A general distinction is made between primary packaging and secondary packaging. Primary packaging for pharmaceuticals and medical devices includes containers or components made of glass, rubber, plastic, aluminum, composite materials, and films. These materials come into direct contact with the medical devices and must therefore meet certain requirements regarding safety, efficacy, and reliability. Manufacturers of primary packaging must meet the expectations of pharmaceutical manufacturers and be able to demonstrate that their production processes are subject to an integrated quality management system (QMS) and the rules of Good Manufacturing Practice (GMP), thus meeting the required quality standards. Secondary packaging is understood to be outer packaging that is not in direct contact with the items to be packaged, i.e., pharmaceuticals or other substances, and which usually serves a protective and control function.
[0009] This results in a project design with a large number of samples coupled with numerous analytical methods, which ideally would have to be performed seamlessly for all samples. However, since such experimental designs are expected to yield enormous information gains, their implementation is desirable in order to ultimately determine the best possible product with the longest product lifespan. This is offset by the very large personnel and equipment expenditure required for implementation using traditional analytical methods. With the established personnel base and equipment fleet, such a comprehensive study could only be carried out with great difficulty, as the required working time and financial investment would be hardly acceptable, and timelines for such product developments would be massively extended.The alternative would be a drastic reduction in measurement effort, which could, however, lead to unacceptable product quality.
[0010] In addition to the enormous effort required to collect a comprehensive dataset for a stability study, interpreting such a massive dataset also presents a significant challenge. This is due to the fact that the "traditional" analysis of the original data, which often only evaluates a small subset of the total data (e.g., examining only the three strongest signals in an HPLC chromatogram instead of the entire "fingerprint range"), can mask a weak trend or anomaly or even fail to detect it at all. Modern chemometric methods can generally provide a solution here. These include, for example, principal component analysis, correlation analysis, distance measures, partial least squares regression, support vector machines, and neural networks.Such methods are capable of converting a complex original data set with numerous variables (here: test parameters) into a few latent variables that are still capable of describing the data set with sufficient accuracy. Therefore, the implementation of chemometric methods for the simple and efficient evaluation of a data set is essential to solving the problem. Due to the aforementioned personnel, financial, and time-consuming aspects, products are currently only tested for stability with respect to a few parameters. This leads to certain other unmeasured parameters that can influence stability either not being detected, thus resulting in a substandard product.It is therefore of immense advantage to record the stability properties of a product as a whole and not on the basis of individually selected parameters which may not correctly reflect the stability of a substance or mixture of substances from which the product or, for example, its packaging is made. Important parameters could thus be ignored and lead to suboptimal products. The present invention solves these problems by recording the properties of a product or substance or mixture of substances (hereinafter "mixture of substances") as a whole and enabling a concrete statement regarding the stability of the entire product. The analysis of a large number of parameters is therefore no longer necessary thanks to the present invention.
[0011] It is therefore an object of the present invention to provide a method for determining the stability of a substance mixture that overcomes the aforementioned disadvantages of the prior art. Furthermore, the calculation method can also be used to monitor processes in which a large number of monitoring parameters are recorded in parallel.
[0012] DESCRIPTION
[0013] The present invention comprises a computer-implemented method for determining the stability of a substance mixture, as well as for monitoring processes in which a plurality of monitoring parameters are recorded in parallel. The method comprises at least one data acquisition step in which at least one measurement data set is received, as well as at least one subsequent data acquisition step in which at least one further measurement data set is received. The first data acquisition step comprises the at least one initial value measurement data set. Each measurement data set can represent a chemical, in particular phytochemical, profile of a respective substance mixture. The method can further comprise a data evaluation step. Evaluation is the processing of (raw) data from an experiment with the (raw) data from the at least one initial value measurement data set and the at least one further measurement data set to generate concrete knowledge gain.The data evaluation step comprises determining a starting value of the respective substance mixture for each measurement data set based on the measurement data set. Furthermore, the data evaluation step comprises quantifying the change in the respective substance mixture with respect to the starting value using a mathematical distance measure by recording at least one further, subsequent data acquisition step, with which at least one further measurement data set is recorded. The method can further comprise a data output step in which the change in the respective substance mixture is graphically displayed for each measurement data set.
[0014] A mixture of substances according to the present invention includes both individual substances and combinations or mixtures of substances, as well as chemical and / or biological products. The terms "substance or mixture of substances," "combination," and "product" are understood as equivalent terms within the scope of the present invention.
[0015] The stability of a substance mixture within the meaning of the present invention is understood as a measure of the period in which the change in the respective substance mixture remains within previously defined limits. This includes, in particular, chemical and / or physical stability. Within the framework of a stability study, the stability can be investigated according to the method according to the invention. A stability study is therefore understood as an experiment that investigates whether different substance mixtures are stable at all and how quickly stability decreases. Alternatively or additionally, it can be provided that different substance mixtures, in particular formulations, are compared with one another within the framework of such a study. Stability comprises the change in a substance or mixture of substances over a period of time that is measured between at least two points in time.The first time point comprises the first data acquisition step with the at least one initial value measurement data set; the second time point comprises the second and all subsequent data acquisition steps with the at least one further measurement data set. Furthermore, this method can also be used to monitor production processes or reactions (see Application Example 3). These include chemical or biological production processes and reactions, for example, extractions, tabletting, capsule filling, and mixing processes. For example, the establishment of equilibrium in an extraction process can be monitored using the novel method without prior calibration of a method, and the time point of equilibrium can be determined. In the present method, equilibrium is indicated by an asymptotic approximation of the calculated distance values to a plateau, as shown in Figs. 8 and 9.
[0016] A chemical profile of a mixture of substances is understood as the totality of the chemical compounds in the mixture and / or the resulting chemical properties or reactivity of the mixture. A phytochemical profile of a mixture of substances is understood as the totality of the chemical compounds in the mixture derived from plant components and / or the resulting properties of the mixture.
[0017] The challenges and requirements described above define the criteria that a novel analytical concept must meet. The method according to the invention enables rapid and comprehensive investigation of mixtures of substances prepared in parallel and their comparison with each other, even over time. This advantageously enables a relative comparison of product variants to select a selection of promising candidates, for example, in the context of drug research.
[0018] For the purposes of the present disclosure, a medicinal product is understood to mean: a) all substances or combinations of substancesZ-mixtures intended as agents having properties for curing or preventing human or animal diseases, or b) all substances or combinations of substancesZ-mixtures that can be used in or on the human or animal body or administered to a human or animal in order to either restore, correct or influence human or animal physiological functions through a pharmacological, immunological or metabolic action or to make a medical diagnosis.
[0019] A particular advantage underlying the method according to the invention is that a significantly reduced selection of possible product candidates can be generated with very little effort during product development, so that, for example, only these product candidates subsequently need to be further investigated using other, specific, classical analytical methods. By concentrating on certain promising product candidates, further detailed knowledge can be efficiently obtained to enable a final selection of the best product variant. In particular, the proposed method according to the invention enables the analytical processing of large-scale product comparisons with very limited personnel and equipment capacities in a reasonable time.The process according to the invention is particularly advantageous for complex mixtures of substances, but can also be used for products with few individual components and individual compounds.
[0020] The method according to the invention makes it possible, through a comprehensive evaluation of complex measurement data sets that represent the most comprehensive (phyto-)chemical profile possible of all samples of a stability study, to identify those substance mixtures that show the smallest change in relation to their starting value and the smallest variance and can therefore be assumed to be the most stable.
[0021] Preferably, the method, in particular the data evaluation step, comprises the use of a machine learning model, which was preferably generated or trained by unsupervised and / or supervised machine learning. Unsupervised machine learning is a mathematical or information technology method in which a data processing device / computer processes data without additional external information and attempts to independently find a solution. One example of this is principal component analysis (PCA). Supervised machine learning is a mathematical or information technology method in which a data processing device / computer processes data with the aid of external information (e.g. training data, in particular information about concentrations used in an experiment).An independent solution search typically takes place after calibration on already known data.
[0022] It is particularly preferred that the at least one measurement data set comprises data obtained by means of near-infrared spectroscopy.
[0023] Near-infrared spectroscopy (NIR) allows samples of mixtures to be measured non-destructively without complicated sample preparation to generate measurement data sets. Near-infrared spectroscopy, abbreviated to NIR spectroscopy or NIRS, is a physical analysis technique based on spectroscopy in the short-wave infrared range. In NIRS, detection takes place in the near infrared, preferably from approximately 760 to 2500 nm (equivalent to approximately 13160 cm²). 1 up to 4000 cm' 1). This technology records combination and overtone vibrations of all molecules in a sample and is therefore able to record them holistically. Mathematical methods (chemometrics) are used in particular to evaluate this data, since the information contained in these spectra (e.g. presence / absence of certain chemical compounds, concentrations, etc.) would otherwise generally remain hidden from the observer. Near-infrared spectroscopy can be used advantageously for liquid samples (e.g. measurement of transmission, i.e. the measuring beam passes through the sample in a defined layer thickness and the non-absorbed light is detected by a detector), or transflectance, i.e. the measuring beam hits the sample. The reflected light is detected by a detector.The absorption indicates the difference between incident and reflected light) as well as for solid samples (measurement of diffuse reflection or transmission, for example).
[0024] One aspect of the method according to the invention is to provide a machine-based and machine-reproducible data analysis of measurement data sets. The advantage of data sets obtained, for example, using near-infrared spectroscopy is that these data sets—compared to data sets obtained from other analytical techniques—contain significantly more information in a single observational datum, because conventional analytical techniques only observe a few isolated signals simultaneously. Near-infrared spectroscopy enables a broader analysis of a multitude of signals.
[0025] Preferably, it is provided that the at least one measurement data set additionally or alternatively comprises: data obtained by means of UV / VIS spectroscopy, data obtained by means of Raman spectroscopy, a (U)HPLC fingerprint, a GC fingerprint, a peak table from a chromatographic process, and / or at least one physical, biological or chemical parameter, in particular sugar content, disintegration rate, color, breaking strength, disintegration time, friability, density, viscosity, refractive index and / or optical rotation angle.
[0026] UV / VIS photometry is a spectroscopic technique belonging to optical molecular spectroscopy that utilizes electromagnetic waves of ultraviolet (UV) and visible (VIS) light. A light source emits ultraviolet and visible light in the wavelength range from approximately 200 nm to approximately 800 nm. In Raman spectroscopy, the material to be examined is irradiated with monochromatic light, preferably with a laser. In the spectrum of the light scattered by the sample, additional frequencies are observed in addition to the incident frequency (Rayleigh scattering). The frequency differences from the incident light correspond to the energies of rotational, vibrational, phonon, or spin-flip processes characteristic of the material. Conclusions about the substance under investigation can be drawn from the resulting spectrum.
[0027] (U)HPLC ((ultra) high performance liquid chromatography) is a standard analytical method (liquid chromatography technique) that not only separates substances but also identifies and quantifies them using standards. (U)HPLC can also be used to analyze non-volatile substances. This can be connected to a sensitive detector system such as mass spectrometry (MS, or collectively UHPLC MS / MS).
[0028] During the analysis, it is possible to record the complete results as a sample (“fingerprint”). These “fingerprints” can be compared with those in a database or with results from other analyses using the same method to identify the analyzed substance.
[0029] The method according to the invention advantageously enables various analytical data sets to be combined and incorporated into an evaluation. For example, the at least one measurement data set for a substance mixture can comprise data sets obtained using UV / VIS spectroscopy, Raman spectroscopy, (U)HPLC fingerprints (in particular from a coupling with different detectors, for example, diode array detector (DAD), refractive index detector, electrochemical detector, UV / VIS detector, mass spectrometry detector (MS), evaporative light scattering detector (ELSD), GC fingerprints (in particular from a coupling with various detectors, for example, flame ionization detector (FID), MS), or, quite generally, peak tables from all conceivable chromatographic methods.In other words, the method according to the invention can be applied to a variety of measurement techniques that assign one or more X values (e.g. wavelengths, m / z ratios (mass / charge ratio), retention times, measurement points) to one or more Y values (e.g. intensities, absorptions, voltages, measured values). The method is therefore particularly versatile and flexible. Furthermore, with appropriate analytical methods (e.g. HPLC-DAD, HPLC-MS), multidimensional data sets can also be processed. Furthermore, large data matrices from individually (i.e., also using different techniques) collected test parameters (e.g., physical, biological, or chemical analyses (without being limited thereto)) can be combined for evaluation and subjected to the new method.For example, the measurement of sugar content, disintegration rate, tablet color, and breaking strength could be compiled into a data matrix and evaluated from a data set measured over several points in time.
[0030] It may also be provided that the quantification of the change in the respective mixture of substances is based on the totality of the ingredients of the respective mixture of substances.
[0031] By considering all components of a mixture under investigation, it is particularly easy to generate reliable results. In particular, all spectral changes of all components relative to a starting value are taken into account in a stability test. Unlike traditional marker analysis (e.g., using HPLC), this allows for a holistic view of all components of complex mixtures (e.g., using near-infrared spectroscopy), thus mapping the fundamental changes in the entire sample.
[0032] The change in the mixture of substances can be calculated using a mathematical distance measure.
[0033] Preferably, the mathematical distance measure is selected from the following group: Euclidean distance, Mahalanobis distance, Manhattan distance, Pearson distance and / or Gower distance.
[0034] Selecting a mathematical distance measure from the group mentioned above ensures that the change in a substance mixture under investigation is quantified over time, allowing the complexity of the underlying data sets to be processed holistically. Selecting a suitable mathematical strategy to quantify the degree of change in a sample over time thus enables efficient quantification of the change in a particular substance mixture. In particular, a very good comparative overview of a data set from a stability study is obtained without having to individually examine numerous separate parameters. It is particularly advantageous that the mathematical distance measure is user-selectable.
[0035] The ability for a user to select a mathematical distance measure makes the method according to the invention more flexible and adaptable. In particular, a user can incorporate their experience to select a suitable mathematical distance measure for determining the stability of a substance mixture using the method according to the invention.
[0036] In particular, it can be provided that the user can iteratively select several mathematical distance measures one after the other to quantify the change in the respective substance mixture, with the quantification being performed on the basis of each selected mathematical distance measure. This allows a user to compare the results of various selectable mathematical distance measures with each other. This makes the method for determining the stability of a substance mixture particularly flexible and versatile. Furthermore, by providing this option, an unintentional selection of a mathematical distance measure for quantifying a change in the respective substance mixture can be prevented.
[0037] Preferably, the method further comprises a data preprocessing step. This may include performing a stray light correction on the at least one measurement data set, particularly if it comprises near-infrared measurement data. Furthermore, the data preprocessing step may include performing a centering, normalization, and / or scaling of the at least one measurement data set. Furthermore, the data preprocessing step may also include performing a principal component analysis.
[0038] Data preprocessing is understood here as the mathematical processing of raw data with the aim of preparing the actual evaluation. This can include, for example, stray light correction, an increase in the signal-to-noise ratio, or simply reformatting acquired measurement data sets or raw data so that data fed in can be correctly processed by the algorithm during a subsequent evaluation. Suitable data preprocessing can be used in particular to increase the quality of the results of the method according to the invention. Data preprocessing preferably takes place after the data acquisition step of the method according to the invention. Through appropriate data preprocessing, which can in particular include a reduction in the amount of data, possible parameter interactions within a complex overall data set are to be automatically recorded and taken into account.In other words, the accuracy of the method according to the invention or the evaluation quality can be increased by carrying out a stray light correction (for example in the case of near-infrared spectroscopy data), in particular before the actual data evaluation.
[0039] Normalization or scaling of the data, as well as a prior principal component analysis, can also separate unwanted noise from the information in a data set, thus further increasing the quality, validity, and robustness of the data. Noise is defined as a disturbance with a broad, non-specific frequency spectrum that can potentially overlay or mask desired, information-bearing signals. An example of noise is unwanted scattered light that, along with desired excitation light, accidentally falls on the detector of a near-infrared spectroscopy device and is also measured. The application of special mathematical techniques can help separate noise from usable information, thus increasing the signal-to-noise ratio and making the analysis more robust and reliable.
[0040] It can further be provided that the data output step comprises: displaying the change in the respective substance mixture as a box plot; and / or displaying mean values, medians, 0.25 / 0.75 quantiles, highlighting possible outlier candidates; and / or performing at least one statistical test, in particular t-test, Wilcoxon rank sum test, one-way ANOVA and / or Kruskall-Wallis test.
[0041] This design of the data output step enables the analysis of very large measurement data sets, especially with multiple samples, with greater significance and at the same time, with less time expenditure. Furthermore, intuitive evaluation is advantageously provided. This means that this technology can be used by users without in-depth mathematical knowledge, and no individual programming is required for each study. For example, in a data output step that outputs the change in the substance mixture under investigation as a box plot, a user can gain insights at a glance and in a simple, intuitive manner. Highlighting means, medians, 0.25 / 0.75 quantiles, and outliers further increases user-friendliness.Measurement data sets collected by a user can, for example, be fed into the process directly from a measuring device or after a data preprocessing step as described above, and an easily interpretable evaluation is immediately generated.
[0042] It can further be provided that the determination of the starting value of the respective substance mixture is based on metadata of the respective measurement data set.
[0043] Using metadata to determine the seed value can advantageously save examination and processing time. The seed value is determined quickly based on metadata.
[0044] Preferably, the data evaluation step further comprises: performing an additional principal component analysis for the measurement data set, which is not included in the quantification of the change by means of the mathematical distance measure, wherein the result of the additional principal component analysis is presented in the data output step.
[0045] The additional principal component analysis, which is primarily a qualitative analysis that is not included in the distance calculation, can support the user's evaluation of output results after the data output step. The data output step thus provides the user with both a quantitative analysis, which particularly includes the distance measure, and a qualitative analysis, which allows for further conclusions to be drawn at a glance and in an intuitive manner.
[0046] In particular, it can be provided that the mixture of substances comprises solid and / or liquid and / or gaseous mixtures of substances.
[0047] It can further be provided that the mixture of substances contains biological, chemical, plant, animal, human substances or mixtures of substances, pharmaceutical compositions, plant medicinal products, chemical and / or biological medicinal products, cells, cell therapeutics (for example gene therapeutics, e.g. CAR T cells (Chimeric Antigen Receptor T cells), NK cells (Natural Killer cells), somatic cell therapeutics, biotechnologically processed tissue products / tissue engineered products, tissue, stem cells, stem cell products or preparations, e.g.CD34+ cells, CD19+ cells, CD20+ cells, HEK295 cells, TCR alpha / beta cells, TCR gamma / delta cells, CD3+, CD4+, CD8+, CD133+ cells), blood, blood products, organs, medicinal teas, extracts, in particular verbena extract, drops, tablets, coated tablets, capsules, powders, granules, solutions, suspensions, juices, foodstuffs, in particular meat or minced meat, fruit juice, in particular orange juice, food supplements, cosmetics, emulsions, ointments and / or creams, as well as packaging, packaging materials, films, in particular polyethylene, polyvinyl chloride, etc.
[0048] The method for determining the stability of a mixture of substances is very versatile. It can be used to advantageously examine mixtures of substances in any state of matter.
[0049] The computer program according to the invention comprises instructions which, when the program is executed by a computer, cause the computer to carry out the method as described above.
[0050] The device according to the invention can, in particular, be a measuring instrument or a server computer and comprises means for executing the method as described above. A mobile electronic device, such as a smartphone, a programmable logic controller, or a so-called "edge device," is also conceivable as a device according to the invention. In addition, the method can also be provided as a cloud solution within the framework of "software-as-a-service" or generally as a "serverless" application.
[0051] The technical advantages and embodiments described with regard to the method according to the invention apply equally to the computer program according to the invention and to the device according to the invention.
[0052] BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Preferred embodiments of the present disclosure are described below with reference to the following drawings:
[0054] Fig. 1 shows schematically a rough overview of possible steps of the method according to embodiments of the invention.
[0055] Fig. 2 shows a schematic overview of possible steps of the method according to embodiments of the invention.
[0056] Fig. 3 shows a spectral overlay of measured data sets from a near-infrared spectroscopy instrument with detector saturation removed. Fig. 4 shows a spectral overlay of measured data sets from a near-infrared spectroscopy instrument without data preprocessing (top) and with exemplary data preprocessing (scattered light correction) (bottom).
[0057] Fig. 5 shows the boxplot of a comparison of the determined Euclidean distances, calculated based on NIR spectra of orange juice that was stored for several days at room temperature (=“RT”) or 40°C (=“40C”) (,,SW“=start value).
[0058] Fig. 6 shows the mean + / - standard error of the determined Euclidean distances, calculated based on NIR spectra of minced meat samples that were stored for several days at room temperature (=“RT”) or 40°C (=“40C”) (,,SW“=start value).
[0059] Fig. 7 shows the photographic change in the minced meat samples over the observed period at 40 °C and room temperature (RT). A clear change in the sample at 40 °C is evident (1). At RT, the marbling of the minced meat is clearly preserved (2 and 3). However, with the method according to the invention, this qualitative deterioration can be quantified and monitored much more precisely.
[0060] Fig. 8 shows the boxplot of a comparison of the determined Euclidean distances, calculated based on NIR spectra from a verbena extraction experiment carried out over a defined period of 180 minutes.
[0061] Fig. 9 shows a more detailed kinetic evaluation of the data from Fig. 8, in which in particular the rate constants (distance max and k) of the test, and derived therefrom the estimated times required for 50, 90 and 95 percent extraction.
[0062] Fig. 10 shows a more detailed kinetic analysis of data from a cell culture experiment, here an untreated HEK295 cell culture without additional stress agent (=control). The times to 50% (K m ), 90% and 95% reaching a plateau value.
[0063] Fig. 11 shows the detailed kinetic analysis of data from a cell culture experiment under the influence of a 3% ethanol solution, which was added as a stress agent. The times to 50% (K m), 90% and 95% plateau values. The different viabilities and kinetics (see also Fig. 12) correlate very well with the preliminary experiment (Example 4), in which significantly more concentrations were tested.
[0064] Fig. 12 shows a detailed kinetic analysis of data from a cell culture experiment under the influence of a 6% ethanol solution, which was added as a stress agent. The times to 50% (K m ), 90% and 95% reaching a plateau value.
[0065] Fig. 13 shows a summary of the plateau values determined from the cell culture experiment. The distance value determined for the 6% ethanol mixture is the largest, indicating a strong influence of this high concentration on the cell culture. This measurement mixture is therefore the most "unstable." The value for the 3% and control mixtures is correspondingly lower. These results are also reflected in the previously conducted cell viability assay.
[0066] Fig. 14 shows the correlation between cell viability and applied ethanol concentration (in %) in the preliminary experiment (A), observed maximum Euclidean distance and applied ethanol concentration (in %) in the main experiment (B), as well as the correlation between cell viability and observed maximum Euclidean distance (C), which relates both analytical techniques. For each subgraph (A), (B), and (C), a linear trend line was added, each representing an R 2of 0.955, 0.999, and 0.999, respectively.
[0067] Fig. 15 Boxplot of the Euclidean distances of a NIR measurement of a medicinal tea preparation under different storage conditions (refrigerator (KS) / room temperature (RT) / 40°C climate chamber).
[0068] Fig. 16 Comparison of stability measurements of a medicinal tea preparation. Measurements were performed using a validated reference method (photometric assay in the UV / VIS range, referred to as GPP in the legend) and using NIR using the novel evaluation method according to the invention.
[0069] Fig. 17 Example graphical representation of the HPLC / MS raw data set for a 20% ethanolic extract of a medicinal herb mixture of thyme, rosemary, and chamomile. These raw data were processed directly using the novel evaluation method.
[0070] Fig. 18 Boxplot representation of the calculated Euclidean distances for the HPLC / MS data set for the 20% ethanolic extract shown in Fig. 17. The Euclidean distances were plotted against the weekly values (observation period: 24 weeks) or, on the far right, after irradiation with light (SunT). Key: Kll = climate zone II; AC = climate zone AC; B = amber glass bottle; W = white glass bottle.
[0071] Fig. 19 Fit of the determined kinetic function for the Euclidean distances calculated from the HPLC / MS raw data for the 20% ethanolic extract from Fig. 17. Key: Kll = climate zone II; AC = climate zone AC.
[0072] Fig. 20 Stability curve of the 20% extract of a medicinal drug mixture of thyme, rosemary and chamomile over a representative, single, selected marker compound (m / z 329.17; retention time = 8.16 min) under different climatic conditions.
[0073] Fig. 21 Boxplot representation of the Euclidean distances of a medicinal powder mixture of rosemary, thyme, and chamomile before (time 0) and after irradiation with light for 20 hours (time 20). Key: P1 = powder mixture packaged in a paper bag; P2 = powder mixture stored openly.
[0074] Fig. 22 Bar chart of reference data using wet-chemical photometric determination of total polyphenols based on Monograph 2.8.14 of the European Pharmacopoeia. A medicinal powder mixture was irradiated with light for 20 hours, both packaged in a paper bag (P1) and stored openly (P2). The recovery is to be interpreted as the percentage agreement of the absorption after irradiation with the initial value of the unexposed sample.
[0075] Fig. 23 Fit of the determined kinetic function for the Euclidean distances calculated on latent variables based on the NIR spectra (SNV-corrected) for a medicinal drug powder mixture over an observation period of 24 weeks in different climate zones. Key: Kll = climate zone II; AC = climate zone AC. Fig. 24 Stability curve of a medicinal drug powder mixture using a representative, single, selected marker compound (m / z 329.17; retention time = 8.16 min) under different climatic conditions.
[0076] Fig. 25 Visual change of a medicinal powder mixture after storage in different climate zones after 24 weeks. The sample of the mixture stored in climate zone AC is significantly darker in color.
[0077] Fig. 26 Boxplot representation of the calculated Euclidean distances from a NIR data set (SNV-corrected) for hawthorn film-coated tablets. The Euclidean distances were plotted against the weekly values (observation period: 24 weeks). Key: FT = film-coated tablets in polyethylene bag; PP = film-coated tablets in blister and folding box.
[0078] Fig. 27 Stability curve of hawthorn film-coated tablets over a representative dimeric procyanidin (m / z 577.14; retention time = 2.69 min) under different climatic conditions and different packaging materials. Key: FT = film-coated tablets in polyethylene bag; PP = film-coated tablets in blister and folding box.
[0079] DESCRIPTION OF PREFERRED EMBODIMENTS
[0080] As shown in Fig. 1, data acquisition 102 initially takes place. At least one measurement data set is received, which can be generated by a measurement using one or more measuring devices. Each measurement data set comprises a chemical, in particular phytochemical, profile of the substance mixture to be examined. The measurement data set or the multiple measurement data sets can, in particular, comprise data obtained using near-infrared spectroscopy. Using near-infrared spectroscopy (NIR), samples of substances or mixtures of substances can be measured without complicated sample processing in order to generate measurement data sets. This technology captures combination and overtone vibrations of all molecules in a sample and is therefore capable of capturing them holistically.The advantage of data sets obtained using near-infrared spectroscopy (NIR), for example, is that these data sets—compared to data sets obtained using other analytical techniques—contain significantly more information in a single observation, because classical analytical techniques only observe a few isolated signals simultaneously. In other words, near-infrared spectroscopy enables a broader evaluation of a multitude of signals.
[0081] Alternatively or additionally, the measurement data set or multiple measurement data sets can comprise data acquired using UV / VIS spectroscopy, Raman spectroscopy, (U)HPLC fingerprints (particularly from a coupling with different detectors, e.g., DAD, MS, ELSD), GC fingerprints (particularly from a coupling with various detectors, e.g., FID, MS), or, more generally, peak tables from all conceivable chromatographic methods. In other words, the method can be applied to a variety of measurement techniques that assign one or more X values (e.g., wavelengths, m / z ratios, retention times, measurement points) to one or more Y values (e.g., intensities, absorptions, voltages, measured values). The method is therefore particularly versatile and flexible.
[0082] In a further, optional step, data preprocessing 104 can take place. This can include, in particular: performing a stray light correction of the at least one measurement data set, in particular if this includes near-infrared measurement data; performing a centering, normalization and / or scaling of the at least one measurement data set; and / or performing a principal component analysis. Suitable data preprocessing can be used, in particular, to increase the quality of the results of the method according to the invention. The data preprocessing, which can, in particular, include a reduction in the amount of data, is intended to automatically capture and take into account possible parameter interactions within a complex overall data set. In other words: the accuracy of the method according to the invention orThe evaluation quality can be increased by performing a stray light correction (e.g. in the case of near-infrared spectroscopy data), especially before the actual data evaluation.
[0083] In a further step, data evaluation 106 is performed for each measurement data set. This involves a starting value calculation 112 and a distance measure calculation 114, each for each measurement data set. During the starting value calculation 112, a starting value is determined based on the measured data, starting from which a change in the substance mixture to be examined is observed and quantified within the framework of the distance measure calculation 114. The distance measure calculation 114 can, in particular, be performed on the entirety of the ingredients of the substance mixture to be examined in order to increase the robustness and validity of the result. The distance measure in the distance measure calculation 114 can, in particular, be selected from the following group: Euclidean distance, Mahalanobis distance, Manhattan distance, Pearson distance, and / or Gower distance.Selecting a mathematical distance measure from the aforementioned group ensures that the change in a substance or mixture of substances under investigation is quantified over time, allowing the complexity of the underlying data sets to be processed holistically. Selecting a suitable mathematical strategy to quantify the degree of change in a sample over time thus enables efficient quantification of the change in a specific substance or mixture of substances. In particular, a very good comparative overview of a data set from a stability study is obtained without having to individually examine numerous separate parameters.
[0084] It is particularly advantageous that the mathematical distance measure can be selected by a user.
[0085] The ability for a user to select a mathematical distance measure makes the method according to the invention more flexible and adaptable. In particular, a user can incorporate their experience to select a suitable mathematical distance measure for determining the stability of a substance or mixture of substances using the method according to the invention.
[0086] In particular, it can be provided that the user selects several mathematical distance measures to quantify the change in the respective substance or mixture of substances, with the quantification being performed based on each selected mathematical distance measure. This allows the user to compare the results of different selectable mathematical distance measures.
[0087] Alternatively or additionally, the user can select various mathematical distance measures one after the other to quantify the change in the respective substances or mixtures of substances. This makes the method for determining the stability of a substance or mixture of substances particularly flexible and versatile. Furthermore, providing this option can prevent an inadvertent selection of a mathematical distance measure for quantifying a change in the respective substances or mixtures of substances.
[0088] Optionally, a PCA calculation 110 can be performed as part of the data evaluation 106, which precedes the initial value calculation 112 and the distance measure calculation 114. The abbreviation PCA stands for Principal Component Analysis. The data evaluation 106 can thus be performed both with the original data and after data scaling and / or data reduction to latent variables, for example, using the PCA calculation 110.
[0089] In a final step, shown in Fig. 1, data output 108 occurs. The change in the respective substance mixture is graphically displayed for each measurement data set. The data output 108 can, in particular, comprise: displaying the change in the respective substance mixture as a box plot; and / or displaying mean values, medians, 0.25 / 0.75 quantiles, highlighting possible outlier candidates; and / or performing at least one statistical test, in particular a t-test, Wilcoxon rank sum test, one-way ANOVA, and / or Kruskall-Wallis test. This configuration of the data output 108 enables the examination of very large measurement data sets, particularly with a plurality of samples, with greater significance while simultaneously reducing the time required. Furthermore, an intuitive evaluation is advantageously provided.This means that this technology can be used by users without in-depth mathematical knowledge, and no individual programming is required for each study. For example, a user can gain insights at a glance and in a simple, intuitive way, for example, in a data output step that displays the change in the substance mixture under investigation as a box plot. Highlighting means, medians, 0.25 / 0.75 quantiles, and outliers further increases user-friendliness. Measurement data sets collected by a user can, for example, be fed into the process directly from a measuring device or after a data preprocessing step, as described above, and an easily interpretable evaluation is immediately generated.
[0090] Fig. 2 shows a schematic overview of possible steps of the method according to embodiments of the invention. The present invention is explained below with reference to Fig. 2.
[0091] In a step 202, raw data from one or more measuring devices are compiled. As described above, various measurements of the same substance mixture or different samples of a substance mixture can be performed to acquire the raw data (data acquisition step).
[0092] In a step 204, a decision must be made as to whether preprocessing is desired or not. If no preprocessing is desired, the distance measure is selected in a next step 211. The distance measure can be selected from the distance measures mentioned above. Once the distance measure has been selected, a quantitative and / or a qualitative evaluation is carried out - depending on the design of the method. As part of the quantitative evaluation, start values or zero values are first identified for each measurement data set in a step 212. In a further step 214, the change in a substance mixture to be examined is determined, for example over time, using the selected distance measure. The result is grouped and displayed in a step 216, for example in a box plot. In addition, accompanying statistical tests can be carried out as part of the quantitative evaluation in a step 217.Within the scope of a possible qualitative evaluation 316, for example, through an additional principal component analysis, a user can obtain a useful statement in addition to the quantitative evaluation. In a step 400, the results are graphically displayed and output.
[0093] If it is decided in step 204 that data preprocessing is desired, a decision is first made in step 2041 as to whether cropping of the measurement data sets is desired. If cropping of the measurement data sets is desired, this is carried out in step 2042. A decision is then made in step 2043 as to whether saturation correction is desired. If saturation correction is desired, this is carried out in step 2044. Furthermore, a decision is made in step 2045 as to whether stray light correction is desired. If stray light correction is desired in step 2045, this is carried out in step 2046. A decision is then made in step 2047 as to whether derivation with smoothing of the measurement data sets is desired. If such smoothing is desired, this is carried out in step 2048.Furthermore, in step 2049, a decision is made as to whether a calculation based on latent variables is desired, and a scaling is selected in step 2050. The scaling can be selected, for example, from "none," "UV" (univariance, not to be confused with "ultraviolet"), or "Pareto." It is understood that during the optional preprocessing, only a selection of the measures described above can be performed.
[0094] A PCA calculation can then be performed in step 210. The distance measure is then selected as described above. IMPLEMENTATION EXAMPLE
[0095] The experiments described below were conducted using a reference implementation in the programming language R. The reference implementation includes the following steps:
[0096] Calculating the distances of individual observations / samples to their starting value ("zero-month value") using suitable distance measures (e.g., Euclid, Mahalanobis, Manhattan, Pearson, Gower, etc.). This distance calculation can be performed on both raw data and latent variables from preprocessing (e.g., PCA, data scaling, etc.). The calculated distance measures also make it possible to determine the kinetics of a product's degradation. In this way, half-lives can also be used to identify the most stable variant(s).
[0097] Easy-to-interpret representation of the calculated distances in box plots including automatic execution of descriptive statistics (calculation and display of means, medians, display of the 0.25 / 0.75 quantiles, plotting of possible outlier candidates) and defining / executing the most suitable statistical test (t-test / Wilcoxon rank sum test / one-way ANOVA / Kruskall-Wallis test) for comparing the means.
[0098] Possibility to track changes in samples over time, including estimation of statistical significance (e.g., whether further degradation occurs in the samples after n months or whether a plateau has been reached)
[0099] Converting the spectral data into latent variables using principal component analysis (PCA) allows for additional estimation of different types of degradation or rearrangement mechanisms via score plots. Two cases can be distinguished here:
[0100] A) A PCA can be calculated from the raw data, the scores of which are then used to calculate the distance and subsequently generate boxplots and statistics. These scores are then also displayed graphically.
[0101] B) The actual distance calculation (and boxplot display, as well as the associated statistics) can also be carried out without a preceding PCA - this is then still calculated and displayed, but is not included in the actually critical distance calculation.
[0102] The created app enables intuitive evaluation for end users without any programming or in-depth mathematical knowledge. The corresponding calculation method implemented in one embodiment of the invention is described below:
[0103] The starting point in this embodiment is a raw data table prepared by the user, which contains the samples to be analyzed and meta information (for example, the climate zone or month in which the sample was measured). Each row represents a sample, each column represents a "property," such as intensity at a measured wavelength or a numerical value for breaking strength, the concentration of a contained substance, or other recorded parameters. The number of columns can vary depending on the measurement method and the availability / relevance of additional meta information. An excerpt from an example raw data table is shown here: Table 1: Example excerpt from the raw data set for application example 1. This includes both the meta information and the NIR spectra at different times under different climatic conditions for a stability study on orange juice.
[0104] Especially with NIR spectra, it may be desirable to use only certain wavenumber ranges for the evaluation, so that these can be cut out and reassembled if desired.
[0105] In NIR spectra, it can happen during the measurement process that the sensitivity of the detector is exceeded in certain wavenumber ranges (for example, in the case of aqueous samples in the range of approx. 5400 cm' 1 up to 4900 cm' 1 , which is very sensitive to water content in the sample). Since detector saturation does not contain any relevant information and could compromise the evaluation quality, such areas can be deliberately cut out and excluded from further analysis if desired.
[0106] In NIR spectra, a baseline shift may occur in individual spectra if unwanted stray light occurred during the measurement. This can be corrected by stray light correction. One possibility is the so-called vector normalization (SNV normalization), which works as follows: A sample matrix P is first transposed, then for each value in each column, the The corrected value is calculated using the formula x' = . Where: x: value in a column, x: mean value of the column, o: standard deviation of the column values. The resulting matrix is then transposed back so that a sample is again listed in each row. Furthermore, for stray light correction in the NIR range, multiplicative stray light correction (MSC) can also be used in embodiments of the invention. Both procedures are well-known standard methods. Fig. 4 shows an exemplary NIR spectra overlay before and after data pretreatment. Furthermore, it may be advantageous to differentiate NIR spectra (or other raw data matrices) before the actual analysis and / or to subject them to a smoothing algorithm. The Savitzky-Golay algorithm is used for these requirements (cf. citation
[0016] ), which performs both data treatments automatically.
[0107] If desired, a Principal Component Analysis (PCA) of the prepared raw data can be performed prior to the actual distance calculation. A subsequent distance calculation would then be performed on the determined score values. This step can be useful when processing highly noisy raw data. Prior to PCA, the data can also be centered, UV (univariance) or Pareto scaled (centering: x' = x - x, UV scaling: x' = Pareto scaling: x' = different from the stray light correction does not require transposition of the data matrix before calculation).
[0108] The user can then select the desired distance measure for the actual calculation. Typically, this will be Euclidean distance, but other distance measures are also conceivable, in particular Manhattan distance, Pearson distance, Gower distance, or Mahalanobis distance.
[0109] Once this pre-processing has been completed, the application uses the metadata of the raw data to determine the start value or “zero value” (in other words: the individual data points at time 0) for each sample contained in the raw data. This means the chemical / physical / biological state of the sample, in particular - but not limited to - sugar content, disintegration rate, color, breaking strength, disintegration time, friability, density, viscosity, refractive index and / or optical rotation angle prior to storage, e.g. under certain climatic conditions or targeted stress tests such as irradiation with UV / VIS light, forcing redox reactions, or similar. The start value or zero value is defined or determined by the user. The sample metadata precisely indicates which samples or batches of samples were recorded at time 0.If there are multiple replicates of time 0 from a batch, the mean spectrum is calculated and stored in memory for later use. This so-called starting date for each corresponding sample then serves as a reference to which a distance calculation is performed.
[0110] The distance is then calculated for each time point, which can be determined from the metadata, for each batch, or for each sample batch, using the method selected by the user. The distance values obtained in this way can then be sorted, aggregated, and graphically displayed, as well as subjected to further statistical analysis.
[0111] In the described embodiment, an additional PCA is also calculated from the pre-processed raw data, which is not included in the distance calculation but supports a further, albeit purely qualitative, interpretation of the raw data.
[0112] APPLICATION EXAMPLE 1 - NIR measurements of a liquid food: Orange juice
[0113] The subject of the investigation was:
[0114] Orange juice “Beste Wahl” from REWE (100% juice from direct juice, with pulp, 8.5g sugar per 100ml (according to the label))
[0115] Opening the package and storing an aliquot at room temperature (22°C)
[0116] Storage of another aliquot at 40 °C, relative humidity = 75 %
[0117] A near-infrared spectroscopy device MPA II (Bruker Optik GmbH) was used as the measuring instrument:
[0118] Measurement of orange juice in transmission;
[0119] Wavenumber range: 12500 cm' 1 up to 3950 cm' 1 (spectral ranges with an absorption > 4.5 were not considered for the evaluation)
[0120] Data pretreatment of the NIR spectra:
[0121] Formation of the 1st derivative with 9 smoothing points
[0122] The following calculation parameters were chosen:
[0123] Distance measure: Euclidean distance
[0124] Fig. 5 shows the boxplot of the calculated distances over 5 days, with day 0 being the "starting value" (= "SW"). In the first two days (day 1 and day 2), the change is practically identical under both storage conditions and there is no statistically significant difference. From day 3 onwards, however, it becomes clear that the determined Euclidean distance of the samples becomes significantly larger at 40°C and has also increased further on day 4 (Kruskal-Wallis test results in p<0.05, pairwise Wilcox test results in statistical significance on days 3 and 4 with p<0.05, all other days are not significant). This is consistent with the generally expected assumption that storage of food at elevated temperatures leads to chemical or microbial changes (e.g., decomposition of individual components or bacterial growth).
[0125] APPLICATION EXAMPLE 2 - NIR measurements of a solid food: minced meat
[0126] The subject of the investigation was:
[0127] Freshly prepared minced pork from REWE
[0128] Store an aliquot immediately after purchase at room temperature (22°C)
[0129] Storage of another aliquot at 40°C, relative humidity = 75%
[0130] A near-infrared spectroscopy device MPA II (Bruker Optik GmbH) was used as the measuring instrument:
[0131] Measuring NIR spectra in diffuse reflection as absorption spectra
[0132] Wavenumber range: 12500 cm' 1 up to 3950 cm' 1
[0133] Data pretreatment of the NIR spectra: performing SNV scattered light correction
[0134] The following calculation parameters were chosen: Euclidean distance
[0135] Fig. 6 shows the mean values (error bars: standard error) of the calculated Euclidean distances. The largest change in all cases occurs from day 0 to day 1. In general, it can be seen that, as expected, the minced meat changes more significantly at higher temperatures than at room temperature (all time points differ statistically significantly with p<0.05 (Kruskal-Wallis or Wilcoxon rank sum test)), and the plateau formation in the 40°C samples is not yet complete.
[0136] In parallel with the measurements mentioned above, photo documentation (Fig. 7) was conducted (start and end of the observation period). In the sample stored at elevated temperature, a distinctly discolored "ring" (marked with (1)) is visible. Furthermore, a liquid had formed on top of the meat sample, which was not present in the room-temperature sample. Furthermore, the texture of the sample (clear separation of meat portion (marked with (2)) and fat (marked with (3))) had changed significantly, which was significantly less pronounced in the control sample. These changes suggest food spoilage. APPLICATION EXAMPLE 3 - Online NIR measurements of a verbena extraction
[0137] The subject of the investigation was:
[0138] 150 g chopped verbena (includes all dried above-ground parts of the plant)
[0139] Extraction in a beaker with stirring at 20 °C in 1800 ml of demineralized water
[0140] An MPA II near-infrared spectroscopy device (Bruker Optik GmbH) was used as the measuring instrument. The samples to be measured were filtered prior to measurement.
[0141] Wavenumber range: 12500 cm' 1 up to 3950 cm' 1 (spectral ranges with an absorption > 4.5 were not considered for the evaluation)
[0142] Data pretreatment of the NIR spectra: performing SNV scattered light correction
[0143] The following calculation parameters were chosen: Euclidean distance
[0144] Fig. 8 shows the boxplot of the distances calculated using the new method over a period of 0 to 180 minutes of the extraction process. It is clear that the change in the solution is greater as the extraction process progresses at the beginning of the experiment and slowly reaches a plateau towards the end. The crucial and new feature of the calculation method according to the invention, however, is that in this experiment, not a single lead substance is tracked as a representative of the entire extraction process, but rather the entire dissolved substances are tracked across the entire NIR spectrum. Different chemical plant components can exhibit different extraction rates / solubility kinetics. If only a single lead substance is used to assess an "exhaustive" extraction, it is possible that this particular substance's dissolution kinetics are not representative of all the substances contained.This could potentially lead to an extraction being terminated prematurely because the lead substance no longer indicates any change in the solution after a certain period of time, even though other, undetected components have not yet completely dissolved into the solution. The present method eliminates this problem, as it captures all substances in the solution, thus allowing the "overall kinetics" of all dissolution processes to be recorded.
[0145] Determining an identical extraction from production batch to production batch is particularly important in the manufacture of pharmaceutical extracts, as the active ingredient content in the manufactured products must always be identical. Fluctuating active ingredient contents can lead to "out-of-specification" events during subsequent quality control, which may result in the rejection of a production batch. The method according to the invention can contribute to significantly simplifying this requirement with minimal effort, or even making it possible in the first place.
[0146] Fig. 9 shows the overall kinetics calculated from the raw data of Fig. 8. The points represent the mean values with standard errors from the individual measurements of the box plot (from Fig. 8), as well as a subsequently calculated, non-linear kinetic curve. For this, the formula [ ] was used. Similarly, another, adequate mathematical description. Where Distance(t) = distance at time t; distance max = distance value in the plateau; k = speed constant.
[0147] From the kinetics presented in this example for the experiment conducted, it can also be deduced that the plant sample used was 50% extracted after approximately 4 minutes, 90% extraction was achieved after approximately 35 minutes, and a nearly exhaustive extraction of 95% was completed after approximately 75 minutes. This also means that in this case, the extraction can be terminated after 75 minutes, eliminating the need to continue for several more hours, which would be associated with considerable costs on an industrial scale.
[0148] APPLICATION EXAMPLE 4 - Tracking the influence of stress factors on cell cultures
[0149] Before the main experiment described below, a preliminary experiment was conducted to determine whether HEK295 cells could a) grow confluently in NIR measuring vessels and b) be stressed by ethanol, and if so, at what concentrations, resulting in a corresponding decrease in viability / confluence. This was verified, in particular, using the complex but proven trypan blue staining technique. For this purpose, concentrations of 0–10% ethanol were tested.
[0150] Based on the results of this preliminary experiment, it was decided to use two different concentrations of ethanol in the main experiment, as a clear influence on cell viability was seen in a period that could be easily covered later.
[0151] Subsequently, the main experiment was carried out, which confirmed the decrease in viability / confluence by an NIR measurement using the calculation method according to the invention.
[0152] The subject of the investigation was:
[0153] HEK295 cells, provided by Microbify, at a concentration of 2.2 * 10 5 Cells / ml in medium, confluently grown, ie adherently grown cells that have formed a continuous monolayer / cell layer, in the NIR measuring vessels with transflectance plunger made of stainless steel from BrukerOptics o Use of nutrient medium DMEM (Dulbecco's Modified Eagle's Medium), with high glucose content, HEPES buffer, without phenol red indicator.
[0154] Medium was purchased ready-to-use from ThermoFisher, catalog number 21063045, and supplemented with 10% fetal calf serum (v / v) and 1% penicillin-strepomycin solution (v / v). Cell culture for 48 hours before measurement directly in the measuring vessel.
[0155] Addition of several concentrations of ethanol: 0% (control), 3%, and 6% (each v / v). These EtOH concentrations were selected based on the preliminary experiment, which elicited a corresponding death / detachment response in a cell viability assay. Before the experiment, the cells were confluent, i.e., 100% viable.
[0156] Measurement of cell changes via NIR over a period of 380 min
[0157] The measuring instrument used was a near-infrared spectroscopy device microPHAZIR GP (Thermo Fisher Scientific Inc.) with a wavenumber range of 6266 cm' 1 up to 4172 cm' 1 Data pretreatment of the NIR spectra: No special pretreatment was performed, spectra were processed directly.
[0158] The following calculation parameters were chosen: Euclidean distance
[0159] The cells were removed from the incubator and treated with a chemical stress factor (ethanol) at the specified concentrations. Three independent replicates were prepared with 0%, 3%, and 6% ethanol and measured via NIR at close intervals (every 2 minutes) over a total period of 6 hours.
[0160] From the measured spectra, Euclidean distances were calculated according to the method of the invention, averaged per measurement time point, and kinetic curves were fitted (see Application Example 3; identical procedure). These are shown in Figs. 10, 11, and 12, respectively for control, treatment with 3% ethanol, and 6% ethanol. The data points are the mean values of three independent replicates, each measured three times, thus resulting in nine measurement points per point. The error bars represent the standard error of the mean. Also shown are the determination points for K m (time at which 50% of the plateau value was reached), Kgo (time at which 90% of the plateau value was reached), and K95 (time at which 95% of the plateau value was reached).
[0161] These values are reached very quickly in the control (K m~ 2.4 min, Kg5~43.7 min), whereby the plateau value itself also represents the lowest value within the experiments. Obviously, removing the cells from the incubator and using them on the lab bench already represents a minor stress factor. The K m - Values of the ethanol-stressed cells are statistically identical (at 3% EtOH ~ 13 min, at 6% EtOH ~ 11 min), which could be due to a uniform, concentration-independent initial reaction of the cells to the addition of ethanol.
[0162] However, with the addition of ethanol at different concentrations, a successively higher plateau is reached, indicating a significantly higher stress load on the cells. This is also consistent with the fact that the cells in the measuring vessels (as already observed in the preliminary experiment) detached from the surface over time and lost confluence. This is also clearly evident from the bar chart (Fig. 13) (statistical significance, one-way ANOVA with p<0.001; post-hoc test: paired t-test, all p-values <0.001), which was generated from the parameters of the curve fits from Figs. 10, 11, and 12.Here, too, the key innovation is that detecting cell stress does not require first identifying, isolating, and quantifying a representative individual substance from the cell suspension, nor does it require establishing and conducting complex cell staining and counting assays, which, moreover, cannot be designed as online measurements. Instead, a very large sample volume can be directly measured spectroscopically without complex external sample preparation and evaluated with minimal effort using the method according to the invention. Furthermore, it has been demonstrated that the present method can also be successfully used to monitor the stability / viability of living cells.
[0163] The observed cell viabilities at the end of the preliminary experiment (which were linearly interpolated due to slightly different EtOH concentrations used) correlate with the observed maximum Euclidean distances (with R 2 > 95%) (see Fig. 14 and the table shown below).
[0164] Table 2: Comparison of distance measures calculated as Euclidean distance with the results of a colorimetric cell viability assay (staining with trypan blue).
[0165] APPLICATION EXAMPLE 5 - Comparison of a medicinal product (medicinal tea) with a photometric reference method
[0166] The subject of the investigation was:
[0167] Medicinal tea: Bad Heilbrunner Liver and Gallbladder Tea (1 filter bag with 1.75 g contains, according to the manufacturer: 0.61 g peppermint leaves, 0.35 g dandelion, 0.26 g Javanese turmeric, 0.18 g yarrow herb; without quantity specified: fennel, chamomile flowers, caraway, licorice root)
[0168] Extraction of 3 tea bags in a beaker under reflux (boiling temperature of water, approx. 100 °C) in 150 ml of demineralized water
[0169] An MPA II near-infrared spectroscopy device (Bruker Optik GmbH) was used as the measuring instrument. The samples to be measured were filtered prior to measurement.
[0170] Wavenumber range: 12500 cm' 1 up to 3950 cm' 1 (spectral ranges with an absorption > 4.5 were not considered for the evaluation)
[0171] Data pretreatment of the NIR spectra: performing SNV scattered light correction
[0172] The following calculation parameters were chosen: Euclidean distance
[0173] A validated photometric method for the detection of total polyphenols (measurement of absorption at 760 nm after reaction of the solution with a molybdate tungstate reagent and 29% sodium carbonate solution, based on Monograph 2.8.14 of the European Pharmacopoeia) was used as the reference method.
[0174] The aim of the study was to determine whether a spectroscopic analysis of tea samples at time points 0 and 10 days when the sample was stored at 4°C in a refrigerator (CS), at 20°C (room temperature, RT) and at 40°C could demonstrate a similar recovery rate / stability as a validated photometric reference method that detects total polyphenols (GPP) in a solution.
[0175] Fig. 15 shows the Euclidean distances determined from the NIR measurements on day 0 and day 10. These increase least for the samples in the refrigerator (CS), moderately when stored at room temperature (RT) and most strongly at 40°C, which is intuitively expected.
[0176] The comparison with the recovery rates and stability from the photometric method (labeled "GPP" in the legend) in Fig. 16 also shows that the new calculation using distance measures agrees well with this reference. The fact that it is somewhat lower than in the reference method is understandable, since the new method detects all substances in the mixture, whereas the reference method only detects a subset. To calculate a percentage stability rate, a hypothetical blank spectrum without analytes was assumed for the NIR data, and its Euclidean distance to the averaged initial spectrum was calculated. This distance was then used to normalize all other distances within this interval. For the reference method, the total polyphenols were determined in mg / 100g, and their degradation after 10 days was also normalized to a percentage value.
[0177] APPLICATION EXAMPLE 6 - COMPARISON OF AN EXTRACT MIXTURE IN DIFFERENT CLIMATIC ZONES USING HPLC / MS
[0178] The subjects of the investigation were:
[0179] Ethanolic-aqueous extract of a medicinal plant mixture of thyme, rosemary and chamomile (1 :1 :1) using a 20% ethanolic aqueous mixture
[0180] Storage conditions:
[0181] AC (Accelerated conditions): Temperature = 40 °C, relative humidity = 75 %
[0182] Kll (climate zone II): temperature = 25°C, relative humidity = 60%
[0183] Storage for 24 weeks each
[0184] An additional, independent test: Irradiation for 20 hours at 150 klx in the Sun Tester in amber glass / white glass vials (labeled “B” or “W” in the legend of Fig. 18)
[0185] Device used: Suntest CPS+ (Atlas Material Testing Technology GmbH)
[0186] Wavelength: 320 nm to 800 nm
[0187] Irradiation dose: 3000 klx*h
[0188] The following devices were used as measuring instruments:
[0189] - Chromatographic system: 1290 Infinity II (Agilent Technologies Germany GmbH & Co. KG)
[0190] - Detector: TripleTOF (Sciex)
[0191] The following calculation parameters were chosen:
[0192] Distance measure: Euclidean distances
[0193] Calculation on LC / MS peakable (2602 individual signals)
[0194] Preprocessing: Calculation on latent variables, variance coverage > 95% For the kinetic calculations (Fig. 19) the equation Distance = - — used and distance max , and k are determined mathematically.
[0195] The results of the study are shown in Figure 19. It is clearly evident that in the more severe climate condition ("AC"), the changes occur much more rapidly, as a greater distance is reached earlier. However, the samples in both climate conditions reach virtually the same distance plateau towards the end of the study.
[0196] By irradiating the samples in the Suntester (see Fig. 18), the same level of change is achieved within approximately 20 hours, which would otherwise only be achieved after approximately 3.5 weeks of storage under climate zone II conditions. There is no significant difference between amber and white glass vials.
[0197] A classical approach using a representative marker substance (m / z 329.17; retention time = 8.16 min, Fig. 20) reveals, as with the NIR analysis, an expected degradation of the compound over time, which reaches the endpoint significantly faster under condition AC than under condition II. In contrast to processing a data set containing information from many compounds (such as NIR spectra here), this analytical approach using only a single marker substance provides only a very limited picture of this product and cannot truly represent the entire multicomponent mixture. Furthermore, the evaluation of data sets with regard to individual compounds is significantly more complex and time-consuming.
[0198] APPLICATION EXAMPLE 7 - COMPARISON OF PACKAGING MATERIALS USING NIR UNDER THE INFLUENCE OF IRRADIATION WITH INTENSIVE LIGHT AND ELEVATED TEMPERATURE
[0199] The subjects of the investigation were:
[0200] A medicinal powder mixture of rosemary, thyme and chamomile in a ratio of 1:1:1
[0201] Storage conditions: stored closed in a paper bag (=“P1”) or open (=“P2”). Irradiation with light was carried out under the following conditions:
[0202] Device: Suntest CPS+ (Atlas Material Testing Technology GmbH)
[0203] Wavelength: 320 nm to 800 nm
[0204] Irradiation dose: 3000 klx*h The measuring instrument used was a near-infrared spectroscopy device MPA II (Bruker Optik GmbH):
[0205] Measurement in diffuse reflection
[0206] The following calculation parameters were chosen:
[0207] Distance measure: Euclidean distance
[0208] Calculation on SNV-corrected NIR spectra without any restriction of wavenumber ranges
[0209] Distance calculation on latent variables with total variance covered > 95%
[0210] The results of the investigation were as follows:
[0211] The result is shown in Fig. 21 in the form of a box plot. At time 0 (before the start of irradiation in the Sun Tester), the distance values in both samples are identical. However, the values at time 20 hours not only differ from the starting value, but P1 (powder mixture in a paper bag) and P2 (openly stored powder mixture) are also significantly different from each other. This is due, on the one hand, to the strong light exposure (hence the strong increase in sample P2), but also to the elevated temperature reached here in the Sun Tester, which also affects sample P1 (in the paper bag). These results were also confirmed by a conventional analytical method. For this purpose, the total polyphenols were determined photometrically in accordance with general monograph 2.8.14 of the European Pharmacopoeia. This is a very complex wet-chemical analytical method.Here too, sample P2 showed a significant deviation from the initial value (difference from initial value = 26%), whereas sample P1 (difference from initial value = 1%) remained almost unchanged in the paper bag (see Fig. 22).
[0212] The method according to the invention clearly demonstrates that temperature and light exposure together have a measurable influence compared to thermal exposure alone, or neither of these factors being present. The very time-consuming wet chemical analysis of the polyphenols confirms the occurrence of changes in the sample, but can only depict a portion of these changes, since polyphenols represent only a portion of the overall chemical profile. The method according to the invention, in contrast, captures the complete profile and therefore provides a much more comprehensive picture. APPLICATION EXAMPLE 8 - COMPARISON OF DIFFERENT CLIMATIC ZONES USING NIR
[0213] The subjects of the investigation were:
[0214] A medicinal powder mixture of rosemary, thyme and chamomile in a ratio of 1:1:1
[0215] Storage conditions:
[0216] AC (Accelerated conditions): Temperature = 40 °C, relative humidity = 75 %
[0217] Kll (climate zone II): temperature = 25°C, relative humidity = 60%
[0218] Storage for 24 weeks each
[0219] A near-infrared spectroscopy device MPA II (Bruker Optik GmbH) was used as the measuring instrument:
[0220] Measurement in diffuse reflection
[0221] The following calculation parameters were chosen:
[0222] Distance measure: Euclidean distance
[0223] Calculation on SNV-corrected NIR spectra without any restriction of wavenumber ranges
[0224] Distance calculation on latent variables with total variance covered > 95%
[0225] The following devices were used as measuring instruments for the parallel, classical analytical approach:
[0226] Chromatographic system: 1290 Infinity II (Agilent Technologies Germany GmbH & Co. KG)
[0227] Detector: TripleTOF (Sciex)
[0228] The results of the investigation were as follows:
[0229] It is immediately apparent from Figure 23 that the more severe climate zone AC leads to significantly greater distances even in the first week compared to climate zone II. This is of course in line with expectations, which are also depicted in parallel using a classic analytical approach via HPLC / MS using a single representative compound. Fig. 24 shows the signal intensity curve of a main compound (m / z = 329.17; retention time = 8.16 min) from the powder mixture. Here, a significantly faster degradation is observed in climate AC than in climate zone II. However, this classic approach, which considers only a single substance, does not do justice to the complexity of a plant product. The novel approach considers a large data set for a complex product in its entirety. Analogous to Example 1, kinetics were calculated for the two different climate zones using the distances determined from the NIR data set.
[0230] The change after 24 weeks is also visually evident (Fig. 25), where the powder in the sample vial from climate zone AC is significantly darker than its counterpart from climate zone II. However, this purely visual assessment is highly subjective and generally difficult to assess. Our calculation method, however, makes this change measurable not only qualitatively but also quantitatively.
[0231] APPLICATION EXAMPLE 9 - COMPARISON OF PACKAGING MATERIALS OF A SOLID
[0232] DOSAGE FORM USING NIR
[0233] The subjects of the investigation were:
[0234] Hawthorn film-coated tablets:
[0235] Each film-coated tablet contains 450 mg dry extract of hawthorn leaves with flowers
[0236] The film-coated tablets were stored in a blister pack in a folding box and in polyethylene bags.
[0237] Storage conditions:
[0238] AC (Accelerated conditions): Temperature = 40 °C, relative humidity = 75 %
[0239] Kll (climate zone II): temperature = 25°C, relative humidity = 60%
[0240] Storage for 24 weeks each
[0241] A near-infrared spectroscopy device MPA II (Bruker Optik GmbH) was used as the measuring instrument:
[0242] Measurement in diffuse reflection
[0243] The following calculation parameters were chosen:
[0244] Distance measure: Euclidean distance
[0245] Calculation on SNV-corrected NIR spectra without any restriction of wavenumber ranges
[0246] The following devices were used as measuring instruments for the parallel, classical, analytical approach:
[0247] Chromatographic system: 1290 Infinity II (Agilent Technologies Germany GmbH & Co. KG) Detector: TripleTOF (Sciex)
[0248] The analysis of the NIR dataset shows an expected result. Firstly, the distance increases significantly faster under climate conditions AC and also reaches larger maximum values over a period of 24 weeks. Furthermore, a stabilizing effect of the higher-quality packaging material in the form of blisters in combination with a folding box (labeled "PP" in the figure legend) can be clearly deduced, in contrast to the polyethylene bags (labeled "FT" in the figure legend) (see Fig. 26). The same result can also be derived using a representative marker compound (dimeric procyanidin, m / z 577.14; retention time = 2.69 min) in a conventional evaluation approach (see Fig. 27). Here, too, a largely constant degradation can be seen for the samples from climate zone II and a clear, rapid degradation in climate zone AC.However, the traditional approach involves significantly greater measurement and evaluation effort and is unable to represent the product in its full material complexity. The novel method overcomes both limitations, as no specific selection and processing of individual raw data points by a user is required. Instead, the complete spectrum of constituents is captured via NIR spectra, and the new method is guaranteed to reflect the behavior of the sample under investigation.
[0249] It should be understood that aspects of the embodiments described herein that have been described in the context of an apparatus also represent a description of a corresponding method. Some or all of the method steps may be performed by (or using) a hardware device, such as a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the key method steps may be performed by such a device.
[0250] Embodiments of the invention can be implemented in hardware and / or software. The implementation can be carried out using a non-volatile storage medium, such as a digital storage medium, such as a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM and EPROM, an EEPROM, or a FLASH memory, on which electronically readable control signals are stored that can interact with a programmable computer system such that the respective method is carried out. Therefore, the digital storage medium can be computer-readable. Some embodiments according to the invention comprise a data carrier with electronically readable control signals that can interact with a programmable computer system such that one of the methods described herein is carried out.
[0251] In general, embodiments of the present invention can be implemented as a computer program product with a program code, wherein the program code is effective for executing one of the methods when the computer program product is running on a computer. The program code can, for example, be stored on a machine-readable medium. Further embodiments include the computer program for performing one of the methods described herein, which is stored on a machine-readable medium.
[0252] Another embodiment of the present invention is a storage medium (or a data carrier or a computer-readable medium) comprising a computer program stored thereon for performing one of the methods described herein when executed by a processor. The data carrier, digital storage medium, or recorded medium is typically tangible and / or non-transitory. Another embodiment of the present invention is an apparatus as described herein, comprising a processor and the storage medium.
[0253] A further embodiment of the invention is a data stream or signal sequence representing the computer program for performing one of the methods described herein. The data stream or signal sequence can, for example, be configured to be transmitted via a data communication connection, for example, via the Internet.
[0254] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured or adapted to carry out any of the methods described herein.
[0255] A further embodiment comprises a computer on which the computer program for carrying out one of the methods described herein is installed.
[0256] A further embodiment according to the invention comprises a device or system configured to transmit a computer program for executing one of the methods described herein to a recipient. The recipient may, for example, be a computer, a mobile device, a storage device, or the like. The device or system may, for example, comprise a file server for transmitting the computer program to the recipient.
Claims
CLAIMS 1. A computer-implemented method for determining the stability of a substance or mixture of substances, the method comprising the following steps: - a data acquisition step (102) in which (raw) and / or metadata are received with a starting value measurement data set and at least one further measurement data set, wherein the starting value measurement data set and each further measurement data set each represent a chemical, in particular phytochemical, profile of a respective substance or mixture of substances; - a data evaluation step (106) comprising for each further measurement data set: - determining (112) the initial value measurement data set of the respective substance or mixture of substances on the basis of (raw) and / or meta-data of the initial value measurement data set; and - quantifying (114) the change in the respective substance or mixture of substances over time with respect to the initial value measurement data set and the at least one further measurement data set by means of a mathematical distance measure; and - a data output step (108) in which the change in the respective substance or mixture of substances is graphically displayed for each further measurement data set.
2. The method according to claim 1, wherein the data evaluation step (106) comprises using a machine learning model which was preferably generated and / or trained by unsupervised and / or supervised machine learning.
3. The method according to claim 1 or 2, wherein the at least one measurement data set comprises data obtained by means of near-infrared spectroscopy, NIR.
4. The method according to any one of the preceding claims, wherein the at least one measurement data set comprises: - data obtained by UV / VIS spectroscopy; - data obtained by Raman spectroscopy; - a (U)HPLC fingerprint; - a GC fingerprint; - a peak table from a chromatographic procedure; and / or - at least one physical, biological or chemical parameter, in particular sugar content, disintegration rate, colour, breaking strength, Disintegration time, friability, density, viscosity, refractive index and / or optical rotation angle.
5. The method according to any one of the preceding claims, wherein the quantification of the change in the respective substance or mixture of substances is based on the totality of the ingredients of the respective substance or mixture of substances.
6. The method according to any one of the preceding claims, wherein the mathematical distance measure is selected from the following group: Euclidean distance, Mahalanobis distance, Manhattan distance, Pearson distance and / or Gower distance.
7. The method of any one of the preceding claims, wherein the mathematical distance measure is selectable by a user.
8. The method according to any one of the preceding claims, further comprising: a data preprocessing step (104) comprising: - performing a stray light correction of the at least one measurement data set, in particular if it comprises near-infrared measurement data; - performing a centering, normalization and / or scaling of the at least one measurement data set; and / or - Conduct a principal component analysis.
9. The method according to any one of the preceding claims, wherein the data output step (108) comprises: - Displaying the change in the respective substance or mixture of substances as a box plot; and / or - Displaying means, medians, 0.25 / 0.75 quantiles, dispersion measures, highlighting possible outlier candidates; and / or - Conduct at least one statistical test, in particular t-test, Wilcoxon rank sum test, one-way ANOVA and / or Kruskall-Wallis test.
10. The method according to any one of the preceding claims, wherein the data evaluation step (106) further comprises: - performing (110) an additional principal component analysis for the measurement data set, which is not included in the quantification of the change by means of the mathematical distance measure; wherein the result of the additional principal component analysis is displayed in the data output step (108).
11. The method according to any one of the preceding claims, wherein the substance or mixture of substances comprises solid and / or liquid and / or gaseous substances.
12. The method according to any one of the preceding claims, wherein the substance or mixture of substances comprises biological, chemical, plant, animal, human substances or mixtures of substances, pharmaceutical compositions, plant medicaments, chemical and / or biological medicaments, cells, cell therapeutics (for example gene therapeutics, e.g. CAR T cells (Chimeric Antigen Receptor T cells), NK cells (Natural Killer cells), somatic cell therapeutics, biotechnologically processed tissue products / tissue engineered products, tissue, stem cells, stem cell products or preparations, e.g. CD34+ cells, CD19+ cells, CD20+ cells, HEK295 cells, TCR alpha / beta cells, TCR gamma / delta cells, CD3+, CD4+, CD8+, CD133+ cells), blood, blood products, organs, medicinal teas, extracts, in particular verbena extract, thyme, rosemary and chamomile. Medicinal drug mixtures, e.g.as an ethanolic extract or in powder form, drops, tablets, dragees, capsules, powders, granules, solutions, suspensions, juices, foodstuffs, in particular meat or minced meat, fruit juice, in particular orange juice, food supplements, cosmetics, emulsions, ointments and / or creams, as well as packaging, packaging materials, films, in particular polyethylene, polyvinyl chloride.
13. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1-12.
14. A device, in particular a measuring instrument or a server computer, comprising means for carrying out the method according to any one of claims 1-12.