System and method for generating predictive model based on input data set

By combining evolutionary processing technology with multiple prediction models, the overfitting problem of blood biomarker prediction models was solved, the prediction accuracy of spectral measurement data was improved, and more reliable biomarker predictions were achieved.

CN120604242APending Publication Date: 2025-09-05ESCUEL VERDAD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380092404.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-12-12
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the prior art, prediction models for blood biomarkers are susceptible to overfitting problems, resulting in limited accuracy during operation and making it difficult to effectively utilize spectral measurement data for accurate predictions.

Method used

Evolutionary processing technology is used to explore the parameter space using multiple prediction models. Through evolutionary algorithms such as genetic algorithms, mixed prediction models are generated to generate multiple generations of prediction models to reduce the risk of overfitting and improve prediction accuracy.

Benefits of technology

Through multi-generation evolutionary processing, the accuracy of the prediction model is improved, the dependence on specific training data is reduced, the prediction ability in spectrometric data is enhanced, and more reliable output data is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604242A_ABST
    Figure CN120604242A_ABST
Patent Text Reader

Abstract

A method, system, and computer program are presented. The method includes providing a training data set and training a selected number of prediction models using a first portion of the training data set, thereby providing a first generation prediction model. The first generation prediction model is processed using evolutionary processing, and a next generation prediction model is generated. The next generation prediction model is trained using the first portion of training data, and a plurality of current generation prediction models are generated. The selected number of current generation prediction models is processed using evolutionary processing, and a selected number of next generation prediction models is generated. The selected number of predictive models is repeatedly trained and evolved to obtain accuracy metrics within a selected accuracy threshold, and output data indicative of the selected number of predictive models is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to techniques for generating one or more prediction models for predicting selected output data based on input data. The techniques of the present disclosure are particularly useful for predicting the concentration of blood biomarkers based on spectrometric measurements. Background Art

[0002] Analysis of blood biomarkers is one of the most widely used medical tests. Blood typically contains a variety of biomarkers that provide valuable information about an individual's general health. To obtain data about blood biomarkers, a clinician / physician typically draws a blood sample from an individual's vein, and the blood sample is transported to a laboratory for testing to obtain quantitative data about the individual's blood biomarkers.

[0003] Predictive models and general machine learning processing techniques typically require specific training of processing modules based on input data. This training typically involves processing the input data and adjusting specific processing elements to minimize a cost function. This cost function can indicate the difference between the processed output and the expected output given the input data.

[0004] Various techniques for streamlining the blood testing process have been described. Such techniques generally include various techniques related to chemical testing of a blood sample and the use of one or more predictive models to obtain data regarding one or more biomarkers.

[0005] US Pat. No. 10,815,518 describes a sampler and a method for parameterizing it by calibrating a digital circuit and determining the concentrations of several biomarkers simultaneously and non-invasively in real time. The method utilizes an apparatus that applies a digital filter to a set of luminescence characteristics (spectra) provided by a spectrophotometer (E5) (E6), decomposing the spectrum into sub-spectra that exhibit digital characteristics of the relevant markers, and obtaining the concentrations of the set of several biomarkers simultaneously and in real time through a digital decoder.

[0006] US2006 / 281982 discloses an apparatus for non-invasively sensing a biological analyte in a sample, the apparatus comprising an optical system having at least one radiation source and at least one radiation detector; a measurement system operably coupled to the optical system; a control / processing system operably coupled to the measurement system and having an embedded software system; a user interface / peripheral system operably coupled to the control / processing system for providing user interaction with the control / processing system; and a power supply system operably coupled to the measurement system, the control / processing system, and the user interface system for providing power to each of the systems. The embedded software system of the control / processing system processes a signal obtained from the measurement system to determine the concentration of the biological analyte in the sample. Summary of the Invention

[0007] The term overfitting is associated with data analysis and relates to the problem that the calibration of a selected data set is too closely correlated with the selected data set. Therefore, the overfit calibration may not correspond to the fitting additional data provided as input during operation. The present disclosure provides a technology and corresponding system for generating one or more prediction models using data processing, and the one or more prediction models are suitable for predicting one or more parameters, and typically a set of parameters, based on input data. The technology typically utilizes processing of a selected number of prediction models, each of which is determined based on different initial parameters and can be associated with different topologies, and the selected number of prediction models evolves to generate other generations of prediction models. This evolutionary process is used to allow the prediction model to explore a wide area of ​​the parameter space, and thereby can limit the overfitting problem of the input data that may limit the prediction accuracy to the training data set. In this regard, the term evolutionary process, evolutionary process or evolutionary algorithm can refer to one or more techniques for optimization using a population-based metaheuristic optimization (PBMOPT) algorithm.

[0008] Essentially, using evolutionary processing techniques increases the probability of finding a better local optimum or even a global optimum. In conventional optimization processes, and within a very wide solution space, the probability of finding a better solution or even a global optimum can decrease as the set of gradient vectors converges towards a local optimum. Evolutionary processing operates by taking a population of current solutions and mixing and / or applying a perturbation process to them to relocate some local optima within a range of values ​​within the wide solution space that have not yet been evaluated. With each generation of the predictive model, a new population of solutions is generated.

[0009] In this regard, the present disclosure provides a method, typically implemented by one or more computers or processor and memory circuits (PMCs). The method includes providing a training data set, such as obtained from a storage unit or transmitted via network communication, and training a selected number of prediction models using at least a first portion of the training data set to provide a selected number of first-generation prediction models. The method also includes processing the selected number of first-generation prediction models using one or more evolutionary processing techniques to generate a plurality of next-generation prediction models. Each of the next-generation prediction models is further trained using the first portion of the training data to generate a plurality of current-generation prediction models. Furthermore, each current-generation prediction model is processed using evolutionary processing to generate a next-generation model, which is trained starting with initial parameters determined by the evolutionary processing technique. After a selected number of cycles, or when a selected accuracy metric reaches a selected threshold or stabilizes, the processing terminates, thereby providing output data for the current-generation prediction model, which serves as a final-generation prediction model. In this regard, the number of prediction models can be any selected number, such as 3, 5, 8, 12, 18, 23, 34, or any other selected number of two or more prediction models.

[0010] Therefore, the present disclosure provides a technique that utilizes a selected number (typically two or more) of different prediction models to predict output data based on input data. The disclosed technique also utilizes evolutionary processing (e.g., genetic algorithms) for hybrid prediction models. This is to explore the parameter space of the prediction model while enhancing higher accuracy models. Multiple evolutionary generations of the prediction model can optimize the prediction model and reduce the risk of overfitting to specific training data.

[0011] Generally speaking, the techniques of the present disclosure can be directed to predicting a selected number of blood biomarkers based on input data comprising spectrometric data obtained from an individual. In this regard, a training dataset can typically be in the form of spectrometric data obtained from the skin of a plurality of individuals, along with corresponding blood biomarker data for the individuals. Such a training dataset typically includes blood biomarker data obtained by performing laboratory analysis on blood samples from the plurality of individuals, for blood samples collected at a time close to the time the spectrometric data was collected.

[0012] The present disclosure also utilizes, during an operational phase, multiple prediction models to predict a selected set of output data segments in response to selected input data, and provides reliable data regarding the accuracy of the predictions. More specifically, the present disclosure also utilizes receiving input data and operating a selected number of trained prediction models to predict a set of output data segments based on the input data. The present disclosure still further includes processing the predicted output data segments and determining one or more statistical parameters associated with the predicted outputs from different prediction models, and determining at least a variation of the predicted outputs from the different prediction models relative to a first threshold limit and a second threshold limit. When the variation of the output data segments is within the first threshold limit, the technique provides output data comprising an average prediction of the predicted outputs. When the variation exceeds the first threshold but is within the limits of the second threshold, the technique provides output data comprising a set of predicted output data segments. When the variation exceeds the second threshold limit, the technique provides an output indicating the unreliability of the predictions.

[0013] This may be advantageous in the prediction of biomedical parameters, where regulatory requirements often require an indication of the accuracy of the analysis and measurements.The present technology may be used for a selected panel of predictive biomarkers based on spectrometric data obtained from a user's biopsy tissue.

[0014] Thus, according to a broad aspect, the present disclosure provides a method implemented by a processor and memory circuit (PMC), the method comprising:

[0015] a. Provide training data set;

[0016] b. training a selected number of prediction models using at least a first portion of the training data set to provide a selected number of first-generation prediction models;

[0017] c. using one or more evolutionary algorithm processes to process the selected number of first generation prediction models and generate a selected number of next generation prediction models;

[0018] d. using the first portion of the training data to train the selected number of next-generation prediction models to generate multiple current-generation prediction models;

[0019] e. using one or more evolutionary algorithm processing techniques to process the selected number of current generation prediction models and generate a selected number of next generation prediction models;

[0020] f. Repeat actions (d) and (e) for a selected number of generations until the accuracy metric of the current generation prediction model reaches a pre-selected training accuracy threshold;

[0021] g. Provide output data including the selected number of current generation prediction models.

[0022] According to some embodiments, training the selected number of prediction models (b) includes determining random initial parameters for each prediction model, and training the prediction model starting from the random initial parameters.

[0023] According to some embodiments, training the selected number of next generation prediction models (d) includes: using parameters of the next generation prediction models as initial parameters for training.

[0024] According to some embodiments, processing the selected number of current generation predictive models using one or more evolutionary algorithm processing techniques includes introducing a selected rate of mutations in the evolutionary algorithm processing.

[0025] According to some embodiments, the method may further include: using at least a second portion of the training data set to validate a training state of the selected current generation prediction model.

[0026] According to some embodiments, the method may further include determining an accuracy metric for the validation training of the selected current generation prediction model, and repeating the training (d) when the accuracy metric is below a selected threshold.

[0027] According to some embodiments, the method may further comprise: mixing the first portion and the second portion of the training dataset for repeated training.

[0028] According to some embodiments, the method may further include: testing the selected number of current-generation prediction models using at least a third portion of the training data set, and determining a test accuracy metric for the selected number of current-generation prediction models. According to some embodiments, the method may further include: determining the test accuracy metric relative to a preselected accuracy threshold, and generating a request for additional training data when the test accuracy metric is below the preselected threshold.

[0029] According to some embodiments, the training dataset includes spectral profile data segments obtained from a plurality of individuals and corresponding data on a selected set of blood biomarkers for the individuals, and the training of the prediction model is directed to predicting the selected set of biomarkers based on the input spectral profile data of the individuals.

[0030] According to some embodiments, the selected panel of biomarkers comprises biomarkers selected based on biological correlations between the biomarkers.

[0031] According to some embodiments, the selected panel of biomarkers includes two or more biomarkers characterized by a typical spectral effect above a first threshold, and one or more biomarkers characterized by a typical spectral effect below a second threshold.

[0032] According to some embodiments, selecting one or more panels of biomarkers comprises pairing two or more biomarkers characterized by a typical spectral effect above a first threshold with one or more biomarkers characterized by a typical spectral effect below a second threshold.

[0033] According to some embodiments, the spectrogram data obtained from the plurality of individuals includes a plurality of spectrogram readings collected within a selected time range associated with blood circulation of a selected portion of the individual's blood volume.

[0034] According to some embodiments, the spectrogram data indicates spectral absorption in a range between 600 nm and 2700 nm.

[0035] According to some embodiments, the selected number of prediction models includes prediction models having different topologies between the prediction models.

[0036] According to some embodiments, the selected number of prediction models includes prediction models selected from the group consisting of principal component analysis, principal component regression, partial least squares, parallel factor analysis, N-way partial least squares, multiple linear regression, spectral matching values, moving blocks, hierarchical cluster analysis, K-nearest neighbors, support vector machines, naive Bayes, linear or normal discriminant analysis, soft independent modeling of class analogies, feedforward neural networks, recurrent neural networks, Bayesian regularization, convolutional neural networks and / or generative adversarial networks.

[0037] According to some embodiments, the method may further include storing the output data comprising a selected number of current generation prediction models in a computer-readable medium in the form of a collection of pre-trained prediction models for predicting blood biomarkers based on input data comprising spectrometric data obtained by performing non-invasive spectrometric readings of living tissue of an individual.

[0038] According to some embodiments, the blood biomarkers used for prediction include:

[0039] using a spectrometer to obtain one or more spectra from living tissue of the individual;

[0040] processing the one or more spectrograms using the set of pre-trained predictive models to obtain a set of predictions for one or more biomarkers;

[0041] processing the set of predictions for the one or more biomarkers and determining at least a statistical variation between predictions of the set of pre-trained prediction models;

[0042] Processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that:

[0043] determining output data comprising an average value of the output prediction data segments when the statistical variation is within a first variation limit;

[0044] When the statistical variation is within a second variation limit, determining output data including the output prediction data segment and the statistical variation; and

[0045] determining the output data as undetermined when the statistical variation data is outside the second variation limit; and

[0046] An output signal is generated that includes the output data indicative of one or more biomarkers of the individual.

[0047] According to some embodiments, the collection of pre-trained prediction models is configured to predict a common set of biomarkers.

[0048] According to some embodiments, at least determining a statistical change comprises determining a percentage change.

[0049] According to some embodiments, the first threshold is between 5% change and 15% change.

[0050] According to some embodiments, the second threshold is between 20% change and 30% change.

[0051] According to another broad aspect, the present disclosure provides a program storage device readable by a machine, the program storage device tangibly embodying a program of instructions executable by the machine to perform a method comprising:

[0052] Provide training dataset;

[0053] training a selected number of predictive models using at least a first portion of the training dataset to provide a selected number of first generation predictive models;

[0054] processing the selected number of first generation predictive models using one or more evolutionary algorithm processing techniques and generating a selected number of next generation predictive models;

[0055] using the first portion of the training data to train the selected number of next-generation prediction models to generate a plurality of current-generation prediction models;

[0056] processing the selected number of current generation predictive models using one or more evolutionary algorithm processing techniques and generating a selected number of next generation predictive models;

[0057] repeating said training of said next generation prediction models to generate a plurality of current generation prediction models, and processing said selected number of current generation prediction models using one or more evolutionary algorithm processing techniques, and generating a selected number of next generation prediction models for a selected number of generations, until an accuracy metric of the current generation prediction models reaches a preselected training accuracy threshold; and

[0058] Provides output data including the selected number of current generation prediction models.

[0059] According to another broad aspect, the present disclosure provides a system including a processor and memory circuit (PMC), wherein the PMC is configured to:

[0060] a. Obtain training data set;

[0061] b. using at least a first portion of the training data set to train a selected number of prediction models, and providing a selected number of first-generation prediction models;

[0062] c. processing the selected number of first-generation prediction models by one or more evolutionary algorithm processing techniques, and generating a selected number of next-generation prediction models;

[0063] d. using the first portion of the training data to train the selected number of next-generation prediction models to generate multiple current-generation prediction models;

[0064] e. using one or more evolutionary algorithm processing techniques to process the selected number of current generation prediction models and generate a selected number of next generation prediction models;

[0065] f. Repeating actions (d) and (e) for a selected number of generations until the accuracy metric of the current generation prediction model reaches a pre-selected training accuracy threshold; and

[0066] g. Provide output data including the selected number of current generation prediction models.

[0067] According to another broad aspect, the present disclosure provides a method for predicting output data in response to input data, the method comprising:

[0068] Provides a collection of pre-trained prediction models;

[0069] Processing the input data through the set of prediction models and obtaining output prediction data segments;

[0070] processing the output prediction data segments and determining at least statistical variations between the output prediction data segments;

[0071] Processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that:

[0072] determining output data comprising an average value of the output prediction data segments when the statistical variation is within a first variation limit;

[0073] When the statistical variation is within a second variation limit, determining output data including the output prediction data segment and the statistical variation; and

[0074] determining the output message as undetermined when the statistical variation data is outside the second variation limit; and

[0075] An output signal is generated including the output data.

[0076] According to some embodiments, the input data comprises one or more spectrograms obtained from living tissue of the individual.

[0077] According to some embodiments, the one or more spectral graphs include data indicative of spectral absorption in a range between 600 nm and 2700 nm.

[0078] According to some embodiments, the output data includes data regarding the concentration of one or more biomarkers in the individual's blood.

[0079] According to some embodiments, at least determining a statistical change comprises determining a percentage change.

[0080] According to some embodiments, the first threshold is between 5% change and 15% change.

[0081] According to some embodiments, the second threshold is between 20% change and 30% change.

[0082] According to yet another broad aspect, the present disclosure provides a program storage device readable by a machine, the program storage device tangibly embodying a program of instructions executable by the machine to perform a method comprising:

[0083] Provides a collection of pre-trained prediction models;

[0084] Processing the input data through the set of prediction models and obtaining output prediction data segments;

[0085] processing the output prediction data segments and determining at least statistical variations between the output prediction data segments;

[0086] Processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that:

[0087] determining output data comprising an average value of the output prediction data segments when the statistical variation is within a first variation limit;

[0088] When the statistical variation is within a second variation limit, determining output data including the output prediction data segment and the statistical variation; and

[0089] determining the output message as undetermined when the statistical variation data is outside the second variation limit; and

[0090] An output signal is generated including the output data.

[0091] According to another broad aspect, the present disclosure provides a system comprising a processor and memory circuit (PMC), wherein the memory comprises a set of pre-trained prediction models, wherein the PMC is configured to:

[0092] obtaining input data comprising one or more spectroscopic data obtained from living tissue of an individual;

[0093] processing the input data through each predictive model in the set of predictive models and obtaining output predictive data segments indicative of one or more biomarkers in the blood of the individual;

[0094] processing the output prediction data segments and determining at least statistical variations between the output prediction data segments;

[0095] Processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that:

[0096] determining output data comprising an average value of the output prediction data segments when the statistical variation is within a first variation limit;

[0097] When the statistical variation is within a second variation limit, determining output data including the output prediction data segment and the statistical variation; and

[0098] determining the output message as undetermined when the statistical variation data is outside the second variation limit; and

[0099] An output signal is generated including the output data.

[0100] According to some embodiments, the system may further comprise a spectrometer unit.According to some embodiments, the spectrometer unit may be configured to provide spectrogram data indicative of spectral absorption in a range between 600 nm and 2700 nm.

[0101] According to yet another broad aspect, the present disclosure provides a computer program product comprising a computer usable medium having computer readable program code embodied therein for predicting output data in response to input data, the computer program product comprising:

[0102] computer readable program code for causing a computer to provide a collection of pre-trained predictive models;

[0103] computer readable program code for causing a computer to process said input data through said set of prediction models and obtain output prediction data segments;

[0104] computer readable program code for causing a computer to process said output prediction data segments and to at least determine statistical variations between said output prediction data segments;

[0105] Computer readable program code for causing a computer to process the statistical variation and determine correlation of the statistical variation with at least a first variation limit and a second variation limit to:

[0106] computer readable program code for causing a computer to determine whether the statistical variation is within first variation limits and accordingly determine output data comprising an average value of said output prediction data segments;

[0107] Computer readable program code for causing a computer to determine whether the statistical variation is within a second variation limit and accordingly determine output data comprising the output prediction data segment and the statistical variation; and

[0108] Computer readable program code for causing a computer to determine whether the statistical variation data is outside of said second variation limit and accordingly determining an output message as undetermined; and

[0109] Computer readable program code for causing a computer to generate an output signal comprising said output data. BRIEF DESCRIPTION OF THE DRAWINGS

[0110] In order to better understand the subject matter disclosed herein and to illustrate how it may be implemented in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:

[0111] Figure 1 A method for training a predictive model system according to some embodiments of the present disclosure is illustrated;

[0112] Figure 2 illustrates a computer system configured to perform operations according to some embodiments of the present disclosure;

[0113] Figure 3illustrates a method for training a predictive model system with additional detail according to some embodiments of the present disclosure;

[0114] Figure 4 A method for predicting one or more data segments using a predictive model system according to some embodiments of the present disclosure is illustrated;

[0115] Figure 5 A system for non-invasively determining blood biomarkers according to some embodiments of the present disclosure is illustrated;

[0116] Figure 6A and Figure 6B Illustrate spectral data according to some embodiments of the present disclosure ( Figure 6A ) and spectral data labeled by biomarker data ( Figure 6B );

[0117] Figure 7 Cholesterol prediction based on spectrogram data is illustrated, and the accuracy metric R according to some embodiments of the present disclosure is shown. 2 ;

[0118] Figure 8 The operation of an evolutionary algorithm in processing an ensemble of predictive models is illustrated according to some embodiments of the present disclosure;

[0119] Figure 9 illustrates a flow chart illustrating the use of evolutionary processing to train an ensemble of predictive models according to some embodiments of the present disclosure; and

[0120] 10A to 10D Predictive optimization according to some embodiments of the present disclosure relative to conventional techniques is illustrated. Figure 10A Show possible arrangements of solutions in a hypothetical solution domain, Figure 10B Shown with Figure 10A The expected divergence trend associated with the optimization steps after the state in , Figure 10C shows possible arrangements of solutions for each group in the solution domain according to some embodiments of the present disclosure, and Figure 10D Examples of some embodiments of the present disclosure are given. Figure 10C The expected convergence trend associated with the optimization steps following the state in . DETAILED DESCRIPTION

[0121] As indicated above, the present disclosure provides a technique for generating a selected number of prediction models. The multiple prediction models can be configured to predict any selected parameter and can be formed by various selected topologies, including various machine learning configurations, artificial neural networks, PLS, ANN, CNN, PCA, PCR, NPLS, PARAFAC, etc. The techniques of the present disclosure can be directed to determining biomarkers based on spectrometric readings from patient tissue and can be used to determine the patient's blood biomarker levels.

[0122] In this regard, the term "blood biomarker" as used herein refers to any molecule, macromolecule, or cluster of molecules that is or may be present in the blood of an individual and that can be a target for blood test analysis. The term blood biomarker may be used herein in combination with additional terms such as marker, analyte, biological analyte, and molecule, as used below.

[0123] Unless otherwise specifically stated, as will be apparent from the following discussion, throughout this specification, discussions utilizing terms such as "obtain," "use," "feed," "determine," "estimate," "generate," and the like refer to the action and / or process of a computer manipulating data and / or converting data into other data, where the data is represented as physical (such as electronic) quantities and / or where the data represents physical objects.

[0124] The term "computer" or "computerized system" should be interpreted broadly to include any kind of hardware-based electronic device having at least one data processing circuit (e.g., a digital signal processor (DSP), CPU, GPU, TPU, field programmable gate array (FPGA), application specific integrated circuit (ASIC), microcontroller, microprocessor, etc.). The processing circuitry may include, for example, one or more processors operably connected to a computer memory loaded with executable instructions for performing operations, as further described below. The processing circuitry encompasses a single processor or multiple processors, which may be located in the same geographic area, or may be located at least partially in different areas, and may be able to communicate together.

[0125] The present technology can utilize the input data in the form of spectrogram data collected from individual living tissue by non-invasive testing means. Such individuals can be patients in clinics, homes or other places, or users in the case of collecting training data sets, wherein at least some of the data are collected from healthy users. For example, a spectrometer can be used to collect spectrograms from individual skin (e.g., wrist, fingertips, forehead, arms, neck, chest, cheeks, legs, etc.). Generally, body areas with high capillary blood flow or high capillary density may be preferred. Spectrogram data can be collected in the infrared part of the electromagnetic spectrum (including wavelengths in the range between 600nm to 2700nm or between 800nm ​​to 2700nm). Generally speaking, the present technology can be relevant to any wavelength range that obtains spectrogram data, and is used to train prediction models as described below. In addition, a spectrometer can be configured to obtain spectrogram data in a wider range. Such a spectrometer can also be suitable, and in the analysis, the scope of the prediction model prepared for it is used.

[0126] As used herein, the term infrared or near infrared (NIR) refers to the portion of the electromagnetic spectrum commonly referred to as infrared or near infrared. However, it should be noted that the technology of the present invention can generally be used with other spectral ranges, including, for example, the visible spectrum, mid-infrared and / or far-infrared wavelength ranges. Typically, due to the absorbance of functional groups that characterize biochemical compounds, near infrared spectra in the range of 600nm to 2700nm, or 600nm to 2000nm, or 700nm to 2000nm, or 700nm to 2700nm, or 800nm ​​to 2000nm, or 800nm ​​to 2700nm are used to analyze biomarkers.

[0127] For example, the spectrogram data may include data regarding spectral absorbance / reflectance within a selected range using a spectral resolution between 2 nm and 60 nm. As indicated herein, the spectrogram data may be collected over a selected time period, wherein multiple spectrograms are collected over a time period between 1 second and 300 seconds. Thus, a collection of spectrogram data from an individual may include multiple spectrograms, each of which is assigned an acquisition time. The collection of spectrograms may be averaged to determine an average spectrogram reading for the individual.

[0128] Spectral data collected from an individual typically includes a list of absorption levels (typically presented as a graph) that indicates the absorption levels of the individual's tissue for different wavelengths within a selected range. The present technology preferably utilizes spectrogram data that includes multiple spectral measurements obtained from the user's living tissue (e.g., skin) within a selected measurement time (e.g., 1 second to 300 seconds, or 2 seconds to 180 seconds, or approximately 120 seconds). Typically, the collection of spectrogram data may take about one second, so that multiple spectrogram measurements can be collected within the selected measurement time. The multiple spectrogram measurements can vary based on the blood circulation through the individual's blood vessels, so that within a measurement time of approximately 90 seconds, the spectrogram measurements indicate the entire blood volume circulating through the individual's body. In addition, in some embodiments, the collection of spectrogram data for training a predictive model as described herein below can be performed within the selected measurement time. The consistency of the multiple spectrograms thus collected can be further analyzed, for example, by determining the standard variation of the spectrogram data between different spectral measurements; spectrogram data segments with a standard deviation exceeding a preselected threshold can be determined to be inconsistent and, therefore, can be omitted from the training data set.

[0129] refer to Figure 1 , which illustrates a method for training a selected number of predictive models according to some embodiments of the present disclosure. As shown, the method includes providing or obtaining a training dataset 1010. The training dataset can be any type of training data and typically includes labeled data segments. For example, the training dataset can include a collection of spectral data segments from various individuals and corresponding data regarding the concentration of one or more biomarkers in the blood of the respective individuals. As indicated, the training dataset can be obtained from a computer storage device, a network storage device, a remote data source, etc.

[0130] Generally speaking, in some embodiments, the present technology can utilize the following: grouping selected biomarkers according to one or more parameters, and generating corresponding prediction models for different groups of biomarkers. In this regard, Figure 1 The examples and additional examples provided herein relate to a collection of prediction models that are directed to predicting a selected set of biomarkers. Various collections of prediction models can be used in parallel or in series to predict different sets of biomarkers.

[0131] After obtaining the training data set, the present technology utilizes one or more computer systems, such as processors and memory circuits, to train a selected set of prediction models 1020. The selected set of prediction models can include multiple prediction models with selected one or more topologies, and utilize random different initial parameters for each prediction model in the prediction model. Initial training can be used as calibration to the prediction model, and can utilize a portion of the training data set that is usually assigned as the calibration data set. Use evolutionary processing to further process the calibrated prediction model 1030 for mixing and merging the trained prediction models, and generate multiple next generation prediction models 1040. Typically, multiple prediction models can be selected as two or more, or three or more prediction models. The actual number can vary between generations during the evolutionary processing, and can, for example, be in the range of between 2 and 50 or 3 and 25 prediction models.

[0132] In this regard, the term evolutionary processing or evolutionary algorithm as used herein refers to one or more population-based processing techniques that use a selected combination of processing mechanisms such as replication, mutation, and recombination. The technology of the present disclosure can utilize one or more specific evolutionary processing techniques, such as genetic algorithms, genetic programming, evolutionary programming, differential evolution, neuroevolution, and / or one or more learning classifiers.

[0133] In some examples, the evolutionary process can operate on trained predictive models by evaluating the most successful models, selecting successful models (optionally including selected less successful models) for replication, multiplying / merging selected predictive models, and introducing a specific level of stochastic / random mutations to generate the next generation of predictive models.

[0134] The next generation prediction model is typically formed based on the parameters of the trained prediction model of the previous generation. To optimize predictions, the present technique utilizes additional training / calibration 1050 of the next generation prediction model using a calibration portion of the training dataset. While the initial calibration typically utilizes random initial parameters, the calibration / training of the next generation prediction model utilizes initial parameters 1060 obtained from the evolution process, which are typically a specific combination of the trained previous generation models. For simplicity, the calibrated next generation prediction model is referred to herein as the current generation prediction model.

[0135] In general, the present technology can be operated to evolve the current generation prediction model to the next generation prediction model, and further calibrate the next generation prediction model for a selected number of generations. In some embodiments, the present technology can be operated to determine an accuracy metric (e.g., R 2parameter, or other selected accuracy metric) 1070, and when the accuracy metric is determined to be below a selected threshold, the evolution and calibration process is repeated. For example, the threshold may be selected as R 2 ≥0.6.2, or R 2 ≥0.7, or R 2 ≥0.72, so that for R 2 The evolution process is repeated for lower values ​​of .

[0136] After completing a selected number of evolution cycles, or after determining that the accuracy metric of the current generation prediction model is within a threshold limit, the present technology can provide output data about the training prediction model. In some embodiments, the present technology can further proceed to verify the current generation prediction model 1080 using a second portion of the training data set, which is not typically used for calibration. Generally speaking, in some embodiments, the present technology can further determine the accuracy metric of the prediction model after verification. When it is determined that the accuracy metric is below a desired threshold, the two can be shuffled between the calibration portion and the verification portion of the training data set, and the prediction model can be further evolved and calibrated as indicated by action 1030. When it is determined that the accuracy metric is within the desired threshold limit, the prediction model can be tested 1090 using a third portion of the training data set, and when the test is successful, output data 1110 in the form of a selected number of prediction models is generated.

[0137] In general, the selected set of prediction models can utilize various machine learning or artificial intelligence techniques, including, for example, principal component analysis, principal component regression, partial least squares, parallel factor analysis, N-way partial least squares, multiple linear regression, hierarchical cluster analysis, K-nearest neighbors, support vector machines, naive Bayes, linear or normal discriminant analysis, soft independent modeling of class analogies, feedforward neural networks, recurrent neural networks, Bayesian regularization, convolutional neural networks, generative adversarial networks, or any other suitable prediction model configuration. In addition, the prediction models can utilize parallel or sequential chains of artificial intelligence techniques, chemometric techniques, and / or statistical modeling. The selected set of prediction models can utilize prediction models of different topologies. For example, the prediction models can differ in the number of nodes in each layer, the number of processing layers, initial parameters, or their other characteristics.

[0138] As indicated above, methods according to some embodiments of the present disclosure may be implemented using one or more processors and memory circuits. Figure 2A computer-based system 500 for generating one or more predictive models according to some embodiments of the present disclosure is shown. The system 500 includes at least one processor 510 and memory 520, which are generally defined as a processor and memory circuit (PMC). The system 500 also includes suitable input and output communication modules 530 and may include a user interface 540, such as a display, a keyboard, etc. The PMC is operable to implement one or more algorithms that are suitable for generating data indicative of a selected number of predictive models 551 to 55N according to the techniques described herein. In some additional embodiments described below, the PMC can be configured to implement a selected number of predictive models 551 to 55N based on input data, such as one or more spectrograms collected from an individual's tissue, to provide output data indicative thereof (e.g., a selected list of biomarkers). In particular, the processor 510 can execute a number of computer-readable instructions implemented on a computer-readable memory stored or included in the PMC, wherein execution of the computer-readable instructions enables data processing of training input data (e.g., spectrogram data labeled by biomarker data) to generate predictive models and determine predictive model parameters.

[0139] Further references Figure 3, which provides an additional example of a block diagram of a method for training a prediction model according to some embodiments of the present technology. As shown, specific training data (e.g., a collection of spectrograms S obtained from multiple individuals, and corresponding blood biomarker data B obtained from the same collection of individuals and determined in the laboratory based on blood samples) typically collected within a short period of time of obtaining the spectrogram data is used as a training data set 3010. The collection of SB pairs can be divided into a calibration set (approximately 30% to 60%, typically 40%), a validation set (approximately 30% to 60%, typically 40%), and a test set (approximately 10% to 30%, typically 20%) 3020. The different subsets can be stored in a storage device (memory, HDD), and the calibration set is used to calibrate a selected set 3030 of prediction models. This initial calibration typically provides a selected number of trained prediction models. One or more selected evolutionary processing techniques of the trained prediction models are applied to the evolved prediction models to form a set 3035 of next-generation prediction models. Such an evolutionary process may typically include selecting successful prediction models based on their accuracy metrics, merging the parameters of the prediction models, and introducing specific random changes to form the next generation of prediction models. Each of the next generation of prediction models is calibrated 3040 starting from the initial parameters obtained in the evolutionary process to obtain a set of trained current generation prediction models. Generally speaking, the present technology can repeat the evolutionary process and calibration of the next generation prediction models for a selected number of generations. In addition, in some embodiments of the present disclosure, the technology can be operated to determine an accuracy metric 3045 of the trained current generation prediction model, and when the accuracy metric is below a selected threshold (e.g., R 2 ≤0.7, or 0.8, or 0.9), the technique may repeat the evolution process of the predictive model for additional generations 3048. In some embodiments, the disclosed techniques may be operable to repeat the evolution process for a selected number of generations and determine whether the accuracy metric is within or below a threshold limit. When it is determined that the accuracy metric is below a desired limit, the technique may be operable to evolve the predictive model for a selected additional number of generations.

[0140] In this regard, calibration of the prediction model can include changes in prediction parameters to minimize a selected loss function based on a comparison between the prediction and the corresponding known target to provide the smallest possible gap. The loss function can be optimized (i.e., the gap is made smaller) by any method known in the art, such as, but not limited to, gradient descent applied to the loss function and algorithms derived therefrom. Optimizing the loss function causes the parameters of the algorithm to be changed so that any value for which the loss function is optimal can be achieved. For cases with a small number of targets, it is easier to achieve the optimal value of the loss function than for cases with a large number of targets. In these cases, the loss function exhibits multiple optima, where the best optimal value is called the global optimal value and all other optimal values ​​are called local optimal values. The more targets there are, the more difficult it is to achieve the global optimal value. Therefore, the calibration phase can utilize one or more prediction model training techniques, such as stochastic gradient descent, or any other suitable training technique for the prediction model, to make the predicted output (Y) of each input spectrogram (S) closer to the corresponding blood biomarker data (B).

[0141] In some embodiments, calibration may be sufficient. However, to ensure the accuracy of the prediction model, the present technology can further validate and / or test its accuracy.

[0142] A second validation phase can be performed using the second portion of the data set. For each prediction route, the accuracy of the predicted output (Y) is validated using a data segment from the second portion of the data set relative to the actual blood biomarker data (B). The validation phase can include further optimizing the prediction route to minimize the corresponding loss model. After the validation phase, the data indicating the loss function can be processed according to corresponding thresholds associated with the expected results. When the prediction accuracy is insufficient, the first and second portions of the data set can be reshuffled and re-split, and the calibration phase and validation phase can be repeated. In some configurations, additional training data sets may be required, particularly after insufficient accuracy is determined after testing the prediction model. In this case, the system, for example using a computer processor of the PMC (as defined below), can generate an output signal requesting an additional data set.

[0143] The third testing phase involves testing the prediction model's accuracy on the third portion of the dataset. The testing phase typically avoids further optimization of the prediction model and directly tests its accuracy using the input data (S) and target output (B) not used in the calibration and validation phases. If the prediction accuracy of the output data (Y) is insufficient, additional reference datasets in the form of spectrogram data (S) and blood biomarker data (B) may be requested, depending on selected thresholds, and the calibration and validation phases may be repeated using the existing and additional data.

[0144] Generally speaking, in some embodiments, the present technology can provide a fully trained prediction model after a selected number of evolution generations. However, in some embodiments, after a selected number of generations have been prepared and the accuracy metric is within desired limits, the present technology can proceed to validate the predictions of the prediction model 3050. To this end, the present technology can utilize a selected validation dataset selected from the initial dataset used for training. Similarly, the accuracy metric after calibration can be determined 3055, and when it is below a desired threshold, the training datasets for calibration and validation can be shuffled and the training process can be repeated 3058.

[0145] At this stage, the unused third portion of the data set can be used to further test the predictive model 3060. Generally speaking, calibration and testing are used to estimate and avoid overfitting of the predictive model to the initial training data set. To this end, additional elements of training data can be used to verify that the predictive model can infer the additional data to ensure that it is directly trained. In some cases, if the accuracy metric after testing is insufficient (i.e., below a desired threshold), the training process cannot be continued with the same training data set, and the present technology can be operated to generate a request for additional training data 3070.

[0146] However, when the accuracy metric is determined to be within the desired limits, the training of the predictive model is complete, and the present technology can generate output data indicative of a set of predictive models 3080. The output data can be in the form of an indication to a user and can include one or more data tables and / or computer-readable instructions to enable one or more processors to operate the predictive models so trained.

[0147] In this regard, it should be noted that the use of evolutionary processing within the training operation of one or more predictive models provides an effective technique for exploring various regions of the solution space that would not necessarily be explored during conventional training. This is particularly advantageous in the prediction of multidimensional data, such as predicting a set of blood biomarkers based on spectrogram data obtained from an individual.

[0148] Furthermore, according to some embodiments of the present disclosure, the use of multiple prediction models enables prediction of one or more biomedical data segments based on input data collected by non-invasive techniques, and provides output data regarding the accuracy and error level of the predictions. Thus, the use of a selected number of prediction models (typically differing between them in at least one of topology and initial parameters) provides increased accuracy, mitigates overfitting, provides enhanced prediction accuracy, and enables the generation of output data indicative of prediction accuracy. In this regard, Figure 4The method for predicting output data in response to input according to some embodiments of the present disclosure is illustrated. As shown in the figure, Figure 4 The illustrated process involves operating a predictive data set, typically in response to one or more biomedical parameters of input data. Here, the input data is provided to a computer system 4010 capable of executing a selected set of trained predictive models. For example, the input data may be in the form of spectral data obtained from an individual's skin. The technique also includes operating a selected set of predictive models for prediction based on the input data 4020. The set of predictive models is typically pre-stored in a local or remote memory unit of the computer system. The set of predictive models may include two or more, or three or more, predictive models having two or more different topologies. Generally speaking, the set of predictive models may be trained according to the aforementioned techniques and an evolutionary process may be employed across two or more generations to further enhance exploration of the solution space. Processing may be performed in parallel or serially. Each predictive model in the set of predictive models is used to process the input data to determine output data 4030. For example, a predictive model may be trained to predict a selected set of blood biomarkers, grouped according to their biological correlation or concentration levels. Therefore, each prediction model in the prediction model can generate output data in the form of a numbered list, each number indicating the concentration of a blood biomarker from the group of biomarkers. At this stage, the present technology utilizes a selected number of output data fragments for further estimating the prediction accuracy. The technology typically operates one or more processors to determine the statistical behavior 4040 of different prediction outputs. The statistical behavior can be determined based on the mean and variation measure (e.g., using mean and standard deviation) of the prediction distribution. In some other embodiments, the statistical data can be an average value, and the percentage change is defined by DIF% (i) (k) = abs (b (k) ib (k)) / b (k), where i traverses the number of the prediction model, k indicates each fragment of the output data (e.g., each biomarker in the group), b (k) is the average output, and b (k) i is the output of each prediction model i.

[0149] The variation between the predicted outputs is compared to a selected first threshold 4050. For example, for some operations, the desired variation is less than 10%. When the variation is within the desired threshold (yes), the average data is used to determine output data 4060; the output data may also indicate that it is the average of the number of predictions and is within the variation threshold. When the variation exceeds the desired first threshold (no), the variation may be compared to a selected second threshold (e.g., a 25% variation) 4070. In this case, when the variation exceeds the second threshold, the technique may generate an indication that no meaningful output data can be determined 4090. When the variation is between the desired thresholds, the output data may be provided as a collection of predictions as obtained by multiple prediction models 4080.

[0150] The technique provides reliable output data, as well as providing an indication of the reliability of the output data (thereby allowing for prediction of medical and biomedical parameters with sufficient confidence) and an indication of whether the output data is insufficient.

[0151] In this regard, Figure 5 A system 70 for non-invasively determining one or more biomarkers (e.g., blood biomarkers) in a patient is shown. System 70 generally includes a computer-based system 700 comprising at least one processor 710 and memory 720, generally defined as a processor and memory circuit (PMC). System 70 may include or be associated with a spectrometer 770 configured to measure absorption levels from a user's tissue (such as skin) within a selected spectral range, typically comprising, for example, 600 nm to 2700 nm, as described above. Computer-based system 700 also includes a suitable input and output communication module 740 for receiving input spectrogram data from the spectrometer when in use, and may include a user interface 750, such as a display, keyboard, etc. The PMC is pre-stored with computer-readable data indicating a plurality of prediction models 730 trained to predict selected biomarkers based on spectrometric data. The plurality of prediction models may include a set of three or more prediction models trained to predict a common set of biomarkers and to exploit specific variations in model topology. In addition, the system may include operating instructions for implementing the prediction model 730 and for processing the output prediction data, as described above. Figure 4In particular, the processor may execute a plurality of computer-readable instructions implemented on a computer-readable memory stored or included in the PMC, wherein execution of the computer-readable instructions enables data processing of input data in the form of spectral profile data for determining one or more or a group of biomarkers based on the input data. In some additional embodiments, the PMC may implement computer-readable instructions for processing input data indicating spectral profiles for operating one or more predictive models and generating output data indicating predicted biomarker concentrations in the blood of the corresponding individual.

[0152] Generally, it should be noted that, for clarity, system 70 is shown herein in conjunction with spectrometer 770. Generally, system 70 can operate as a computer-based system and obtain patient spectrogram data from a selected storage unit via a communication link for non-invasively determining data regarding blood biomarkers at a remote location.

[0153] As indicated above, the processor 710 can be operated to receive input data in the form of one or more spectrograms from the spectrometer 770 and operate the plurality of prediction models 730 in series or in parallel for predicting a set of biomarkers based on the input spectrogram data. The processor can also obtain output prediction data comprising predictions of quantified biomarkers from a plurality of prediction outputs from three or more (e.g., 3, 5, 8, 13, 17, 25, or 30) prediction outputs. The process can be operated to process the output prediction data and determine statistical parameters associated with at least the variation of the different prediction outputs. The processor also determines whether the variation measure of the prediction output is within a first threshold limit, exceeds the first threshold limit and is within a second threshold limit, or exceeds the second threshold limit, and determines the output data accordingly.

[0154] Generally speaking, when the statistical variation of the predicted output is within a first threshold limit (e.g., 8% variation, or 10% variation, or 15% variation), the prediction is considered relatively accurate, and the processor 710 can generate output data indicating the amount of the selected one or more biomarkers as an average of the predicted output.

[0155] In the event that the statistical variation of the predicted output exceeds a first threshold but is within a second threshold limit (e.g., a 20% variation, or a 25% variation, or a 30% variation), the prediction is considered to have a lower reliability. Accordingly, the processor 710 generates corresponding output data, for example, providing an average prediction and a collection of different predicted outputs, or an average prediction marked as having a lower reliability.

[0156] If the statistical variation exceeds the second threshold, the prediction output has low reliability, and the processor 710 operates to generate an output indication of a low quality prediction. The output indication may indicate that the biomarker prediction is uncertain, that the input data is incorrect, or other suitable message, and may request improved quality spectrometric data.

[0157] In general, the selected set of prediction models can utilize various machine learning or artificial intelligence techniques, including, for example, principal component analysis, principal component regression, partial least squares, parallel factor analysis, N-way partial least squares, multiple linear regression, hierarchical cluster analysis, K-nearest neighbors, support vector machines, naive Bayes, linear or normal discriminant analysis, soft independent modeling of class analogies, feedforward neural networks, recurrent neural networks, Bayesian regularization, convolutional neural networks, generative adversarial networks, or any other suitable prediction model configuration. In addition, the prediction models can utilize parallel or sequential chains of artificial intelligence techniques, chemometric techniques, and / or statistical modeling. The selected set of prediction models can utilize prediction models of different topologies. For example, the prediction models can differ in the number of nodes in each layer, the number of processing layers, initial parameters, or their other characteristics.

[0158] Therefore, the present disclosure also provides a method and system for making data predictions using a set of three or more prediction models that have been trained to predict selected data segments in response to common input data. The method includes providing input data, processing the input data through a set of three or more prediction models (e.g., prediction models trained as described herein), and obtaining three or more output data segments. The method also includes processing the output data segments and determining statistical parameters of the output data segments, and determining based on the statistical parameters between one or more output data. According to the present technology, when the statistical variation of the output data segments is within a first selected limit, the output data is determined by the average value of the output data segments; when the statistical variation is within a second selected limit, the output data is determined by providing a complete set of output data segments and an indication of the statistical variation; when the statistical variation is outside the selected limit, output data is generated indicating that a prediction cannot be made.

[0159] The above method can be implemented by one or more processors and memory circuits (PMC), for example, as shown above. Figure 5 As illustrated, the PMC includes pre-stored data regarding a set of prediction models that are pre-trained to predict output data in response to common input data, as described herein.

[0160] In some configurations, a set of prediction models can be trained to predict a selected set of blood biomarkers in response to input data in the form of one or more spectrograms obtained non-invasively from a living tissue of an individual. The selected set of biomarkers can be a group of biomarkers grouped together according to one or more parameters, such as biological relevance, concentration level, etc.

[0161] Thus, the present technology with improved training and possibly also improved data output reliability can be used for accuracy and quantification of blood biomarkers.The present disclosure provides non-invasive and simple prediction of blood biomarkers and eliminates or at least significantly reduces the need to draw blood from patients.

[0162] In this regard, it will be clear to those skilled in the art that the present disclosure can utilize various types of input data depending on the objectives of the technology, training, and functional requirements. As indicated herein, in some embodiments, the present disclosure can utilize spectral data for predicting blood biomarkers. Such spectral data can be in any selected spectral range, including but not limited to infrared, near infrared, mid-infrared, far infrared, visible light, UVA, UVB, UVC, microwave, Raman, RF frequencies, etc.

[0163] In this regard, reference Figure 6A and Figure 6B , which illustrate spectral data and spectrogram-blood biomarker pairs, respectively. Figure 6A Illustrated are spectrogram data showing absorption levels at different wavelengths collected over a selected number of sampling instances. Figure 6B A data segment including a spectrogram and corresponding biomarker data is shown, wherein the biomarker data is presented in the form of biomarkers obtained by analyzing a blood sample (b n). Generally speaking, each spectrogram data segment can be collected from an individual using spectrometric detection based on reflection, transmission and / or absorption, and can be collected from the individual's skin (e.g., at the wrist or elsewhere). Such spectrometric data can be collected by non-invasive means that do not cause pain or discomfort to the patient, other than requiring a period of relative stillness for spectrometric data collection. Generally speaking, a complete near-infrared (NIR) spectrogram may require an acquisition time of approximately 1 second. However, in order to provide reliable output data, the present technology can utilize the collection of a selected number of spectrogram scans collected over a period of time between 10 seconds and 2 minutes. This acquisition time is based on the blood circulation time through the patient's blood system. In order to provide labeled data, each individual also provides a blood sample, which is taken to a laboratory to obtain a blood analysis indicating the desired blood biomarkers. Blood biomarker results (B) are collected and used to label the spectrogram data, thereby providing multiple data segments indicating spectrogram data obtained from multiple individuals and corresponding blood biomarker data.

[0164] Generally speaking, in some embodiments, the input spectrogram data may be preprocessed. Preprocessing may include selected operations on the collected spectrograms, such as, but not limited to, baseline normalization, spectral truncation, signal-to-noise optimization, signal cleaning, spectral derivation, band deletion, and spectral averaging. Preprocessing may be performed while the spectrograms are being collected, or directly on the input data.

[0165] As is known in the art, spectrogram data can be noisy, and incorrect spectrum acquisition will present more noise. For at least one channel of the spectrogram, when the current measurement value is outside a predetermined range (e.g., outside the standard deviation value), the entire spectrogram is discarded. Otherwise, the spectrogram is retained and quantization is performed using this set as input. In a preferred embodiment, the predetermined range is anywhere from 0.01 to 0.1.

[0166] In order to provide correct labeling of input training data, the training dataset may include the actual concentrations b of n blood markers of the user. n This data can be obtained by any known conventional means, such as by means of a conventional blood test for any commonly tested blood marker. This data is then used by the method of the invention as a vector of values ​​B, i.e. as a table with one column and n number of rows, each row comprising the actual concentration b of one blood marker n This vector B will be used as target data for the method of the present invention for the purpose of calibration or building a prediction model.

[0167] Thus, a data set for the methods of the present disclosure may include a plurality of "SB pairs." An SB pair is added to a database, which may include any number i of SB pairs, each individually named S i -B i Yes. Figure 6B As seen, S i can refer to the matrix of spectra, and B i It can refer to vector B, which are input and target respectively. Each S i -B i The SB pairs are all from one user and are added to the database. For the purpose of calibration, verification and testing, the database is used by the method of the present invention. As will be understood by those skilled in the art, the SB pairs refer to all S pairs in the database. i -B i For, where S and B refer individually to all S i The matrix and refers to all B i vector.

[0168] The spectrogram data S can be processed using one or more techniques such as baseline normalization, spectral truncation, signal-to-noise optimization, signal cleaning, spectral derivation, band deletion, spectral averaging, etc. i The segments are pre-processed or pre-treated. Pre-processing of the spectrogram data segments can be directed to reducing noise associated with the measurement device and movement, as well as defining the spectral range across the entire dataset. Pre-processing can be performed at the time of spectrogram collection and / or after the collection of the entire dataset. In addition, the collected spectrograms can be evaluated for noise level and consistency. For example, spectrogram data segments that are excessively noisy or inconsistent (e.g., the standard deviation between instances of spectrograms collected from a single individual is above a selected threshold) can be ignored and discarded along with the corresponding blood biomarker data.

[0169] In order to improve the predictive accuracy of the model, the present disclosure generally utilizes a group of selected numbers of biomarkers to determine the biomarkers. The list of biomarkers for which the predictive model is targeted can be any list of materials that may be present in a person's blood and may be the target for analysis. Such biomarkers can include, for example: red blood cell count (RBC), hemoglobin, hematocrit, mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCH C), platelets, mean platelet volume (MPV), red blood cell distribution width (RDW), absolute neutrophil count, absolute lymphocyte count, absolute monocyte count, absolute eosinophil count, absolute basophil count, total cholesterol, triglycerides, low-density lipoprotein (LDL-C), high-density lipoprotein (HDL-C), glucose, blood urea nitrogen, creatinine, sodium, potassium, chloride, carbon dioxide, uric acid, albumin, globulin, calcium, phosphorus, alkaline phosphatase, alanine aminotransferase (ALT or SGPT), aspartate aminotransferase (AST or SGOT), LD H, total bilirubin, GGT, iron, TIBC, C-reactive protein, cortisol, DHEA-sulfate, estimated glomerular filtration rate (eGFR), estradiol, ferritin, folate, hemoglobin A1c, homocysteine, progesterone, prostate-specific Ag (PSA), testosterone, thyroid-stimulating hormone, vitamin D, or any other biomarker that may be the target of a biomedical analysis.

[0170] From this list, biomarkers can be divided into groups of two or more biomarkers selected according to one or more parameters (such as typical concentration levels and / or the level of influence of biomarkers on spectral graph data). Each biomarker can be characterized by the typical amount present in the selected blood volume, and this amount can be weighted to indicate the spectral absorption effect of the biomarker. Biomarkers can be arranged in groups, including two or more biomarkers characterized by a high amount (concentration) and one or more biomarkers characterized by a low amount (concentration). Various biomarkers are arranged in a selected number of groups so that each group includes two or more (preferably three or more) biomarkers. Generally speaking, some or all of the biomarkers included in the model are placed in two or more groups. In some embodiments, various biomarkers can be arranged in a selected number of groups so that each group includes two or more (preferably three or more) biomarkers. Generally speaking, some or all of the biomarkers included in the model are placed in two or more groups.

[0171] The biomarkers of these groups can be set up based on biochemical correlation, concentration level or other selected features known in physiology and medicine.As an example, it is known in the field of physiology that the metabolic changes in kidney can affect the concentration changes of some blood markers (such as uric acid, sodium and hemoglobin etc.), and this blood marker is directly affected by the physiology of kidney.All blood markers with the metabolic relationships of these kinds are considered to be " biochemically relevant ", and are grouped accordingly.The blood markers mentioned in this example are grouped and can include the group consisting of uric acid, sodium and hemoglobin.Groups can be set up so that identical blood markers are present in at least one different group, preferentially in at least two different groups, each group preferentially includes at least two blood markers, and can include the blood biomarkers between 2 and 20 kinds.In a preferred embodiment, each group optimizes a kind of blood marker, but more blood markers can also be optimized from a group.

[0172] Given that the concentrations of dozens of blood markers can be calculated by the present technology, there can be many different groups. Therefore, the present technology provides a prediction route that includes a selected number of prediction models that are generally trained as indicated above to predict the parameters of each group. Generating the prediction output can utilize the above Figure 4 and Figure 5 The illustrated statistics provide output prediction data for a set of biomarkers and data on prediction accuracy. Thus, each set of biomarkers can be predicted by a selected set of prediction models, the selected set of prediction models comprising, for example, 2 to 35 prediction models that differ in some prediction parameters, topology, etc.

[0173] Typically, in some embodiments, an accuracy metric can be associated with a correlation value for each specific biomarker. In some embodiments, the correlation value can be a statistical coefficient of determination, also known as R 2 In this regard, reference Figure 7 , which illustrates the correlation factor R for cholesterol determined based on the input data set 2 Generally speaking, in order to provide accurate predictions, the accuracy measure is preferably high enough, i.e., closer to unity. However, since the prediction models of the present disclosure are generally aimed at predicting multiple biomarkers that are usually not directly related, the accuracy measure R 2 It is preferably below unity to allow for some flexibility in inter-individual variability and in particular to avoid overfitting to the training dataset. Thus, in some embodiments, an R between 0.70 and 0.99 may be used. 2In addition, in the testing phase, the accuracy can be determined to be sufficient based on the R in the range between 0.65 and 0.99. 2 To determine sufficient accuracy. In general, the selected range can vary depending on the size of the dataset and the clinical relevance of the blood biomarkers. As indicated above, when the correlation value of at least one biomarker in the panel is outside the desired range, the corresponding prediction path can be retrained by repeating the calibration and validation phases after shuffling the dataset.

[0174] Figure 8 and Figure 9 Specific processes for techniques according to some embodiments of the present disclosure are illustrated. Figure 8 The operation of performing evolutionary processing on a set of current generation prediction models to generate a set of next generation prediction models is illustrated by way of example. Figure 8 FIGURE 1 is a flow chart illustrating the application of the evolutionary process in a specific and non-limiting example of the present technology. During the first instance of the calibration phase, the chemometric algorithms are optimized starting with randomly selected parameters. These parameters are then adjusted for all algorithms in order to achieve the optimal These algorithms constitute the initial population that will be submitted to the EA (evolutionary algorithm processing technology / module). After the results of the mixing and / or interference process performed by the EA, a new algorithm is obtained, which presents initial parameters derived from the parameters of the selected algorithm, which initial parameters are different from the random parameters from which the first chemometric algorithm starts. For this reason, the new algorithm is considered a new generation, which is Figure 8 The above is called the population P. Each new generation is recalibrated but uses a mixed set of parameters from the previous generation as a starting point. Since this process is repeated infinitely, a new population P is always obtained from the algorithm of the previous generation. i In this optimization process using EA, if Figure 8 As shown, some chemometric algorithms may no longer be selected and may not even participate in future groups, i.e., EA additionally filters the type of chemometric algorithms that obtain the most consistent results from those with poorer results. This filtering occurs for each group G of biochemically relevant blood markers, thereby obtaining the best method for each group, or even obtaining a method that provides a global optimal solution for each group. EA optimization additionally mitigates the problem of over-optimization in a local optimum by not over-optimizing in one local optimum. Overfitting of predictions.

[0175] In this regard, Figure 9A specific and non-limiting example of a set of training prediction models according to some embodiments of the present disclosure is shown. As shown, the technique includes obtaining / providing input training data 9010. The training data can typically include multiple pairs of spectral data S and blood biomarker data B, thereby forming SB pairs. The input data set is divided into three parts 9015, including a calibration set, a validation set, and a test set. In some cases, the calibration set can include 40% to 60% of the data, the validation set can include 30% to 50% of the data, and the test set can include 10% to 30% of the data. Generally speaking, the data segments may not overlap between sets. The biomarker data B can be divided into multiple groups of biomarkers 9020 based on biological relevance, concentration level, or other factors, and predictions are performed independently for each group of biomarkers. Therefore, the technique includes initializing a prediction model for each group of biomarkers 9025.

[0176] A random initial population 9035 is used to calibrate the prediction model (including a collection of different prediction topologies and algorithms) 9030 for a selected set of biomarkers. After calibrating the first generation of prediction models, the technology utilizes one or more evolutionary processing techniques (evolutionary algorithms) to determine the next generation of prediction models 9040 and calibrate the next generation of prediction models 9050 to provide the current generation of prediction models. This process can be performed for a selected number of generations or by determining the optimization level of the current generation of prediction models 9060. When the optimization level is insufficient, the technology repeats the evolutionary process 9040 and calibration 9050 until the optimization level is sufficient.

[0177] After sufficient optimization has been performed in calibration, the technique further operates to validate the predictions using a validation set of input data 9070. Validation may also typically include optimization of the prediction operation, as in training. The optimization level after validation is determined 9080, and when the optimization level is insufficient, the method may operate to reshuffle the calibration and validation datasets and repeat calibration 9030. When the optimization level is verified to be sufficient, the prediction model may be further tested using a test dataset 9090. Testing of the prediction model may typically not include any optimization adjustments and may simply verify the accuracy of the predictions for use with a new dataset. When the optimization level is sufficient, determining the optimization level after testing 9100 may result in generating an output set 9110 for the prediction model. When the optimization level is insufficient, it may be associated with insufficient training data, and the method may generate a request for additional training data 9150.

[0178] In general, as indicated above, the present disclosure utilizes evolutionary processing to better explore the solution space. In addition, the present disclosure can also utilize grouping of target biomarkers according to pre-selected features. This enables the present technology to provide improved prediction accuracy. Reference 10A to 10D, which shows the optimization process for a predictive model and illustrates the advantages of the technology of the present invention. Figure 10A and Figure 10B The use of multiple prediction solutions is shown respectively. to specific prediction solutions, each associated with a single biomarker and expected convergence after further optimization steps. Figure 10C and 10D The prediction solutions of multiple groups of biomarkers are shown separately to and expected convergence after the optimization step.

[0179] Generally speaking, Figure 10A The dispersion of the prediction solutions for predicting biomarkers individually, shown in , has similar characteristics to the dispersion associated with the predictions for the entire set of biomarkers. In both cases, the level of data variability is high enough to cause the predictions to overfit to the data provided for training, so that the final solution does not converge with new inputs.

[0180] Alternatively, the technology of the present invention utilizes optimization of a selected set of biomarkers for each panel. The technology maintains a small number of local optima, and the similarity of the characteristics of the different biomarkers causes several panels to converge together.

[0181] The technology disclosed in this invention is tested against the commonly used methods in the prior art. The concentration of blood markers is predicted from SB to one at a time. Predict all at once or by dividing B into different B G Group to make predictions The calibration phase uses a calibration set that has 60% of the total number of SB pairs present in the database. The validation phase and the test phase use a validation set and a test set, respectively, that have 20% of all SB pairs in the database.

[0182] The results are summarized in the table below. For each method the correlation values ​​(R 2 ) are all selected for each blood marker or All predicted R 2 As can be seen, the algorithm optimization was initially more efficient for the method that quantifies one blood marker at a time. However, the final results were consistently better for the method with grouping (i.e., the method of the present invention).

[0183] Table 1

[0184] method <![CDATA[Calibration phase (R 2 ) - S C - B C > <![CDATA[Verification phase (R 2 )-S V -B V > <![CDATA[Test phase (R 2 )S T -B T > one by one 0,965 0,932 0,613±0.227 everything at once 0.867 0.860 0.400±0.196 With grouping 0,942 0,914 0,861±0.181 With grouping + EA optimization 0,970 0,941 0,891±0,176

[0185] Even though the results of the method of quantifying one blood marker at a time are better in the calibration and validation phases, the results from the algorithm are not necessarily clinically relevant. Since there is no grouping, the algorithm will not observe the relationship between the concentrations of the blood markers. Therefore, even if one or more concentration values ​​are physiologically impossible, the results obtained will be any optimal value obtained by optimizing the loss function. When the concentrations of all blood markers are calculated simultaneously, it is extremely difficult to obtain a global optimum, and the relationship between blood markers is uncertain. Therefore, the results are not only poor, but also most likely not clinically relevant, which can be observed in the results of the testing phase. When the blood markers are grouped based on biochemical correlations, the algorithm will observe the relationship between the concentration values, which helps to obtain a global optimum and also helps to obtain coherent concentration values. The application of the EA optimization step has been shown to further improve the results obtained, and is therefore a good optimization step for the method of the present disclosure.

[0186] As indicated above, the present technology can be implemented by one or more computer systems using corresponding one or more processors and memory circuits. The system can be directly connected to the spectrometer for providing on-site blood biomarker data, or be located at a selected location to provide network processing and / or offline biomarker prediction processing.

[0187] It should be noted that the various features described in the various embodiments can be combined according to all possible technical combinations.

[0188] It should be understood that the present invention is not limited in its application to the details set forth in the description contained herein or shown in the accompanying drawings. The present invention is capable of other embodiments and can be practiced and carried out in various ways. Therefore, it should be understood that the words and terms used herein are for descriptive purposes and should not be construed as limiting. Therefore, those skilled in the art will appreciate that the concepts upon which this disclosure is based can readily be used as a basis for designing other structures, methods, and systems for achieving the several purposes of the presently disclosed subject matter.

[0189] It will be readily appreciated by those skilled in the art that various modifications and changes may be applied to the embodiments of the invention as described above without departing from the scope of the invention as defined by the accompanying claims.

Claims

1. A method implemented by a processor and memory circuit (PMC), the method comprising: (a) Provide a training dataset; (b) training a selected number of predictive models using at least a first portion of the training data set to provide a selected number of first-generation predictive models; (c) processing the selected number of first-generation prediction models using one or more evolutionary algorithm processing techniques and generating a selected number of next-generation prediction models; (d) using the first portion of the training data to train the selected number of next-generation prediction models to generate a plurality of current-generation prediction models; (e) processing the selected number of current generation prediction models using one or more evolutionary algorithm processing techniques and generating a selected number of next generation prediction models; (f) repeating actions (d) and (e) for a selected number of generations until the accuracy metric of the current generation prediction model reaches a pre-selected training accuracy threshold; as well as (g) providing output data comprising a selected number of current generation prediction models.

2. The method of claim 1 , wherein said training a selected number of prediction models (b) comprises: Random initial parameters are determined for each prediction model, and the prediction model is trained starting from the random initial parameters.

3. The method according to claim 1 or 2, wherein said training said selected number of next generation prediction models (d) comprises: The parameters of the next generation prediction model are used as initial parameters for training.

4. The method of any one of claims 1 to 3, wherein processing the selected number of current generation predictive models using one or more evolutionary algorithm processing techniques comprises: Introducing mutations at a selected rate in the evolutionary algorithm process.

5. The method according to any one of claims 1 to 4, further comprising: At least a second portion of the training data set is used to validate a training state of the selected current generation prediction model.

6. The method according to claim 5, further comprising: An accuracy metric is determined for the validation training of the selected current generation prediction model, and when the accuracy metric is below a selected threshold, the training is repeated (d).

7. The method according to claim 6, further comprising: The first portion and the second portion of the training data set are mixed for repeated training.

8. The method according to any one of claims 1 to 7, further comprising: The selected number of current generation prediction models are tested using at least a third portion of the training data set, and a test accuracy metric is determined for the selected number of current generation prediction models.

9. The method of claim 8, wherein a request for additional training data is generated when the test accuracy metric is below a preselected threshold.

10. A method according to any one of claims 1 to 9, wherein the training dataset comprises spectral data segments obtained from a plurality of individuals and corresponding data on a selected set of blood biomarkers of the individuals, and the training of the prediction model is directed to predicting the selected set of biomarkers based on the input spectral data of the individuals.

11. The method of claim 10, wherein the selected panel of biomarkers comprises biomarkers selected based on biological correlations between the biomarkers.

12. The method of claim 10, wherein the selected set of biomarkers comprises: Two or more biomarkers with typical spectral effects above a first threshold, and one or more biomarkers characterized by a typical spectral effect below a second threshold.

13. The method of claim 10, wherein selecting one or more panels of biomarkers comprises: Two or more biomarkers characterized by a typical spectral effect above a first threshold are paired with one or more biomarkers characterized by a typical spectral effect below a second threshold.

14. The method of any one of claims 10 to 13, wherein the spectrogram data obtained from a plurality of individuals comprises a plurality of spectrogram readings collected within a selected time range associated with blood circulation of a selected portion of an individual's blood volume.

15. The method of any one of claims 10 to 14, wherein the spectrogram data indicates spectral absorption in the range between 600 nm and 2700 nm.

16. The method of any one of claims 1 to 15, wherein the selected number of prediction models includes prediction models having different topologies between the prediction models.

17. The method of any one of claims 1 to 16, wherein the selected number of prediction models comprises a prediction model selected from the group consisting of principal component analysis, principal component regression, partial least squares, parallel factor analysis, N-way partial least squares, multiple linear regression, spectral matching values, moving blocks, hierarchical cluster analysis, K-nearest neighbors, support vector machines, naive Bayes, linear or normal discriminant analysis, soft independent modeling of class analogies, feedforward neural networks, recurrent neural networks, Bayesian regularization, convolutional neural networks, and generative adversarial networks.

18. The method according to any one of claims 1 to 17, further comprising: The output data, including a selected number of current generation prediction models, is stored in a computer-readable medium in the form of a collection of pre-trained prediction models for predicting blood biomarkers based on input data, wherein the input data includes spectrometric data obtained by non-invasive spectrometric readings of living tissue of an individual.

19. The method of claim 18, wherein the method for predicting blood biomarkers comprises: (a) obtaining one or more spectra from living tissue of an individual using a spectrometer; (b) processing the one or more spectrograms using the set of pre-trained predictive models to obtain a set of predictions for one or more biomarkers; (c) processing the set of predictions for the one or more biomarkers and determining at least the statistical variation between the predictions of the set of pre-trained prediction models; (d) processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that: i) determining output data comprising an average value of said output prediction data segments when the statistical variation is within said first variation limit; ii) when the statistical variation is within the second variation limit, determining output data including the output prediction data segment and the statistical variation; as well as iii) determining the output data as undetermined when the statistical variation data is outside the second variation limit; (e) generating an output signal comprising said output data indicative of one or more biomarkers of said individual.

20. The method of claim 19, wherein the collection of pre-trained prediction models is configured to predict a common set of biomarkers.

21. The method of claim 19, wherein said at least determining a statistical change comprises determining a percentage change.

22. The method of claim 21, wherein the first threshold is between a 5% change and a 15% change.

23. The method of claim 21, wherein the second threshold is between a 20% change and a 30% change.

24. A method for predicting output data in response to input data, the method comprising: (a) Provide a collection of pre-trained prediction models; (b) processing the input data through each set of prediction models and obtaining output prediction data segments; (c) processing the output prediction data segments and determining at least statistical variations between the output prediction data segments; (d) processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that: i) determining output data comprising an average value of said output prediction data segments when the statistical variation is within said first variation limit; ii) when the statistical variation is within the second variation limit, determining output data including the output prediction data segment and the statistical variation; as well as iii) determining the output message as undetermined when the statistical variation data is outside the second variation limit; (e) generating an output signal including the output data.

25. The method of claim 24, wherein the input data comprises one or more spectrograms obtained from living tissue of an individual.

26. The method of claim 25, wherein the one or more spectral graphs include data indicative of spectral absorption in a range between 600 nm and 2700 nm.

27. The method of claim 24, wherein the output data comprises data regarding the concentration of one or more biomarkers in the individual's blood.

28. The method of claim 24, wherein said at least determining a statistical change comprises determining a percentage change.

29. The method of claim 24, wherein the first threshold is between a 5% change and a 15% change.

30. The method of claim 24, wherein the second threshold is between a 20% change and a 30% change.

31. A program storage device readable by a machine, the program storage device tangibly embodying a program of instructions executable by the machine to perform a method comprising: (a) Provide a training dataset; (b) training a selected number of predictive models using at least a first portion of the training data set to provide a selected number of first-generation predictive models; (c) processing the selected number of first-generation prediction models using one or more evolutionary algorithm processing techniques and generating a selected number of next-generation prediction models; (d) using the first portion of the training data to train the selected number of next-generation prediction models to generate a plurality of current-generation prediction models; (e) processing the selected number of current generation prediction models using one or more evolutionary algorithm processing techniques and generating a selected number of next generation prediction models; (f) repeating actions (d) and (e) for a selected number of generations until the accuracy metric of the current generation prediction model reaches a pre-selected training accuracy threshold; (g) providing output data comprising a selected number of current generation prediction models.

32. A program storage device readable by a machine, the program storage device tangibly embodying a program of instructions executable by the machine to perform a method comprising: (a) Provide a collection of pre-trained prediction models; (b) processing the input data through the set of prediction models and obtaining output prediction data segments; (c) processing the output prediction data segments and determining at least statistical variations between the output prediction data segments; (d) processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that: i) determining output data comprising an average value of said output prediction data segments when the statistical variation is within said first variation limit; ii) when the statistical variation is within the second variation limit, determining output data including the output prediction data segment and the statistical variation; as well as iii) determining the output message as undetermined when the statistical variation data is outside the second variation limit; (e) generating an output signal including the output data.

33. A system comprising a processor and memory circuit (PMC), wherein the PMC is configured to: (a) Obtain a training dataset; (b) using at least a first portion of the training data set to train a selected number of prediction models and provide a selected number of first generation prediction models; (c) processing the selected number of first-generation prediction models by one or more evolutionary algorithm processing techniques and generating a selected number of next-generation prediction models; (d) using the first portion of the training data to train the selected number of next-generation prediction models to generate a plurality of current-generation prediction models; (e) processing the selected number of current generation prediction models using one or more evolutionary algorithm processing techniques and generating a selected number of next generation prediction models; (f) repeating actions (d) and (e) for a selected number of generations until the accuracy metric of the current generation prediction model reaches a pre-selected training accuracy threshold; (g) providing output data comprising a selected number of current generation prediction models.

34. A system comprising a processor and memory circuit (PMC), wherein the memory comprises a set of pre-trained prediction models, wherein the PMC is configured to: (a) obtaining input data, the input data comprising one or more spectrogram data obtained from living tissue of an individual; (b) processing the input data through each predictive model in the set of predictive models and obtaining output predictive data segments indicative of one or more biomarkers in the blood of the individual; (c) processing the output prediction data segments and determining at least statistical variations between the output prediction data segments; (d) processing the statistical variation and determining a correlation of the statistical variation with at least a first variation limit and a second variation limit such that: i) determining output data comprising an average value of said output prediction data segments when the statistical variation is within said first variation limit; ii) when the statistical variation is within the second variation limit, determining output data including the output prediction data segment and the statistical variation; and iii) determining the output message as undetermined when the statistical variation data is outside the second variation limit; (e) generating an output signal including the output data.

35. The system of claim 34, further comprising a spectrometer unit.

36. The system of claim 35, wherein the spectrometer unit is configured to provide spectrogram data indicative of spectral absorption in a range between 600 nm and 2700 nm.

37. A computer program product comprising a computer usable medium having computer readable program code embodied therein for predicting output data in response to input data, the computer program product comprising: computer readable program code for causing a computer to provide a collection of pre-trained predictive models; computer readable program code for causing the computer to process the input data through each set of prediction models and obtain output prediction data segments; computer readable program code for causing the computer to process the output prediction data segments and at least determine statistical variations between the output prediction data segments; Computer readable program code for causing the computer to process the statistical variation and determine correlation of the statistical variation with at least a first variation limit and a second variation limit to: computer readable program code for causing the computer to determine whether the statistical variation is within the first variation limit and accordingly determine output data comprising an average value of the output prediction data segments; computer readable program code for causing the computer to determine whether the statistical variation is within the second variation limit and accordingly determine output data comprising the output prediction data segment and the statistical variation; and computer readable program code for causing the computer to determine whether the statistical variation data is outside the second variation limit and accordingly determine an output message as undetermined; and Computer readable program code for causing the computer to generate an output signal comprising the output data.

Citation Information

Patent Citations

  • Sampler and method of parameterizing of digital circuits and of non-invasive determination of the concentration of several biomarkers simultaneously and in real time

    US10815518B2

  • Method and apparatus for the non-invasive sensing of glucose in a human subject

    US20060281982A1