An ai-based method and device for biofluid classification using drying droplet patterns
The AI-based method and device analyze spatio-temporal drying patterns of biofluids using textural features to address scalability and resource constraints in healthcare diagnostics, achieving high-accuracy disease detection with minimal samples.
Patent Information
- Application Number
- PCT/JP2025/006180
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2025-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
Current healthcare diagnostics for diseases like cancer and kidney disorders are labor-intensive, resource-intensive, and not suitable for point-of-care applications due to reliance on complex laboratory techniques and limited accessibility, especially in low-resource settings, and the scalability of droplet-based AI pattern recognition is hindered by the need for numerous droplets and sufficient bio-fluid volumes.
An AI-based method and device that analyzes spatio-temporal drying patterns of biofluids using textural features, integrating traditional machine learning and deep learning to classify biofluids with minimal sample requirements, employing image acquisition and computational tools like ImageJ and Python libraries for feature extraction and model training.
Achieves high-accuracy classification of biofluid patterns exceeding 90% prediction accuracy, enabling efficient, low-resource disease detection and diagnosis with reduced sample volumes, and adaptable to diverse biofluids and environments.
Smart Images

Figure JP2025006180_28082025_PF_FP_ABST
Abstract
Description
AN AI-BASED METHOD AND DEVICE FOR BIOFLUID CLASSIFICATION USING DRYING DROPLET PATTERNS
[0001] This invention relates a computer-implemented technology for classifying distinct patterns formed by the deposition and drying of multi-component complex fluids on a substrate.
[0002] The late detection of diseases, coupled with the shortage of skilled personnel, sophisticated equipment, and limited access to essential resources, poses a significant challenge to healthcare systems globally. Human bio-fluids, such as urine, saliva, and blood, are conventionally analyzed in pathological laboratories to diagnose diseases, metabolic disorders, and other abnormalities in our bodies. Blood, for example, is a naturally occurring multi-component bio-fluid, where abnormalities arise from changes in the shape, size, or composition of its components, often leading to various diseases. For instance, elevated blood sugar levels in diabetes cause compositional changes that impair blood viscosity and flow properties, while anemia reduces red blood cell (RBC) counts and deforms their characteristic bi-concave shape. Kidney-related diseases are associated with irregular filtration and excretion of proteins such as albumin and waste products like creatinine. The long-term effects of many such diseases often culminate in severe complications, including heart disease, stroke, kidney failure, vision loss, liver impairment, and hearing loss. Furthermore, the severity of abnormalities in bio-fluid compositions is often correlated with the progression and intensity of disease conditions. For example, significant deviations in blood glucose levels, RBC counts, or protein concentrations in urine can indicate the advancement of diseases such as bladder cancer.
[0003] Despite the recent development, current healthcare diagnostics heavily rely on biochemical assays, which require specialist knowledge, sophisticated and expensive machines, and specialized reagents, thereby limiting their adaptability for point-of-care (POC) applications in home settings. For instance, diagnosing diseases such as cancer and kidney disorders often involves complex and time-intensive laboratory techniques, including electrophoresis, chromatography, immunoassays, mass spectrometry, flow cytometry, and spectroscopy. Although Raman and other spectroscopy methods are the closest to implementation due to their widespread availability, offering advantages in terms of reduced sample preparation, these are destructive. Moreover, transportation and accessibility issues restrict the feasibility of these solutions in low-resource settings, such as rural areas or regions with inadequate infrastructure, further increasing costs and limiting their suitability for POC applications in home environments. There are some recent advancements in paper-based microfluidics, and they have led to significant progress in the detection methods. Commercially available lateral flow assays, which utilize colorimetric detection, are simple and rapid, making them effective diagnostic tools. However, these methods face significant challenges related to reproducibility, stability, and reliability.
[0004] Recent advancements have demonstrated that healthy biofluids and those associated with diseases, abnormalities, or disorders can be effectively distinguished through the qualitative analysis of residue patterns formed when a droplet of biofluid is deposited onto a surface and allowed to dry. These residue patterns exhibit unique morphological features that reflect the composition and condition of the biofluid, given that these are undisturbed and sit under minimal fluctuations in the environment, such as temperature and relative humidity.
[0005] While qualitative analysis provides valuable insights into potential pathologies, the integration of artificial intelligence (AI) with dried patterns has enabled advanced pattern recognition and classification. AI has demonstrated significant potential in analyzing the final dried patterns of various biofluids, including blood, urine, and cerebrospinal fluids. For example, dried blood patterns have been correlated with different physiological states, while cerebrospinal fluid patterns have been associated with conditions such as Alzheimer’s disease. Furthermore, the combination of AI and dried droplet analysis has been applied to classify patterns of biomolecules such as peptides and DNA, further emphasizing the potential of this approach in disease diagnostics. Importantly, this method is not limited to biological systems; it extends to non-biological systems such as tap water, salts, and milk. For instance, it can differentiate water quality, detect contamination in milk, and classify various types of salts, thereby broadening its applicability across diverse domains.
[0006] Pal, A., Gope, A., & Sengupta, A. (2023). Drying of bio-colloidal sessile droplets: Advances, applications, and perspectives. Advances in Colloid and Interface Science, 314, 102870.Sefiane, K., Duursma, G., & Arif, A. (2021). Patterns from dried drops as a characterisation and healthcare diagnosis technique, potential and challenges: A review. Advances in Colloid and Interface Science, 298, 102546.
[0007] The current approach of applying AI to analyze spatial features or patterns solely from dried droplets presents several practical challenges. While effective for small-scale studies, scaling this methodology for large-scale validation needs the deposition of at least 50 to 100 droplets to generate a comprehensive image corpus for various classes or sample types. This process is not only time-consuming and labor-intensive but also requires handling to ensure uniformity and accuracy across datasets. Addressing these limitations calls for advancements in automated droplet deposition systems, data augmentation techniques, and optimized image acquisition workflows. Alternatively, new methods should be explored to achieve similar diagnostic results without needing to produce so many droplets. Additionally, a major obstacle arises in obtaining sufficient volumes of bio-fluid from each patient for creating the image corpus. This challenge is particularly pronounced in cases involving complex (serious) diseases, such as cancer, where the availability of bio-fluid is limited, making the collection process costly and logistically demanding. These constraints point out the need for innovative solutions to minimize sample requirements while maintaining the integrity and scalability of the diagnostic process using drying droplets.
[0008] One aspect of this invention that solves the problems of the conventional arts is a classification processing method using a classification device equipped with a processor, wherein the processor acquires a microscopy image of biofluids dropped on a predetermined surface, executes predetermined classification processing using the captured image and outputs the classification results from said classification processing.
[0009] Figure 1 is a diagram comparing conventional diagnostic methods with diagnostic methods according to the embodiments of the present invention.Figure 2 is a flowchart showing an example of the classification processing method of this embodiment.Figure 3 is an explanatory diagram showing examples of image data obtained during the drying process and examples of temporal changes in their associated textural features.Figure 4 is an explanatory diagram showing examples of temporal changes in textural features.Figure 5 is an explanatory diagram showing other examples of temporal changes in textural features.Figure 6 is an explanatory diagram showing an example of a normalized confusion matrix and F1-score for traditional machine learning.Figure 7 is an explanatory diagram showing an example of image data obtained during droplet drying.Figure 8 is an explanatory diagram showing another example of image data obtained during droplet drying.Figure 9 is an explanatory diagram showing examples of temporal changes in textural features.Figure 10 is an explanatory diagram showing other examples of temporal changes in textural features.Figure 11 is an explanatory diagram showing an example of a normalized confusion matrix and F1score for traditional machine learning.Figure 12 is an explanatory diagram showing an example of a deep learning model used for classification.Figure 13 is an explanatory diagram showing an example of a normalized confusion matrix for random forest.Figure 14 is an explanatory diagram showing an example of a normalized confusion matrix for a deep learning model.Figure 15 is an explanatory diagram showing an example of machine learning status for the VGG16 model.Figure 16 is an explanatory diagram of image data from the first healthy serum sample during droplet drying.Figure 17 is an explanatory diagram of image data from the second healthy serum sample during droplet drying.Figure 18 is an explanatory diagram of image data from the third healthy serum sample during droplet drying.Figure 19 is an explanatory diagram of image data from the first diabetic serum sample during droplet drying.Figure 20 is an explanatory diagram of image data from the second diabetic serum sample during droplet drying.Figure 21 is an explanatory diagram of image data from the third diabetic serum sample during droplet drying.Figure 22 is an explanatory diagram of image data from the fourth diabetic serum sample during droplet drying.Figure 23 is an explanatory diagram of image data from the fifth diabetic serum sample during droplet drying.Figure 24 is an explanatory diagram of image data from the sixth diabetic serum sample during droplet drying.Figure 25 is an explanatory diagram showing another example of a deep learning model used for classification.Figure 26 is an explanatory diagram showing an example of a normalized confusion matrix and F1-score for traditional machine learning.Figure 27 is an explanatory diagram of image data from the healthy serum samples during droplet drying.Figure 28 is an explanatory diagram of image data from the first cancerous serum sample during droplet drying.Figure 29 is an explanatory diagram of image data from the second cancerous serum sample during droplet drying.Figure 30 is an explanatory diagram of image data from the third cancerous serum sample during droplet drying.Figure 31 is an explanatory diagram showing an example of a normalized confusion matrix and F1-score for traditional machine learning.Figure 32 is an explanatory diagram of image data from the bio-mimetic samples during droplet drying.Figure 33 is an explanatory diagram of image data from another bio-mimetic sample during droplet drying.Figure 34 is an explanatory diagram of image data from different bio-fluid samples during droplet drying.Figure 35 is an explanatory diagram of image data from human and mouse serum samples during droplet drying.Figure 36 is an explanatory diagram comparing image data from serum samples of patients with different diseases during droplet drying.Figure 37 is an explanatory diagram comparing image data from serum samples of patients with different diseases after droplet drying.Figure 38 is a block diagram showing an example configuration of a system for using the method according to an embodiment of the present invention.Figure 39 is a flowchart showing an example process performed by an information processing apparatus using the method of an embodiment of the present invention.
[0010] Disease Diagnosis Technology Based on Drying Pattern Classification of Biological Fluids Using Machine Learning Definition of Terms: 1) A sessile droplet refers to a small liquid droplet resting on a solid surface. When such a droplet contains non-volatile particles and undergoes drying, it transitions from a liquid to a solid or semi-solid state due to the loss of the solvent (evaporation), leaving behind residue patterns that reflect the bio-fluid’s initial states, such as composition and concentrations. 2) Colloids, or colloidal systems, are defined as minute particles dispersed in a medium, which in this context is an aqueous solution. Bio-colloids, which have biological relevance, are a specific subset of colloids and can be classified into two categories: active bio-colloids and passive bio-colloids. Active bio-colloids include living entities such as microbes or artificially created entities such as Janus particles. Passive bio-colloids, on the other hand, are further divided into biomolecules and bio-fluids. Biomolecules refer to bio-colloids essential for biological processes and include DNA, peptides, proteins, etc. Bio-fluids are bio-colloids excreted (e.g., urine, sweat, saliva), secreted (e.g., breast milk, bile, tears), obtained through invasive means (e.g., blood, cerebrospinal fluid), or formed due to pathological circumstances (e.g., blister or cyst fluid). The drying of a colloidal solution enables sedimentation and the formation of residue patterns, which provide crucial insights into its initial properties.
[0011] 3) Importantly, when no external fields, such as magnetic or mechanical agitation, are applied, and there is minimal influence from environmental factors like temperature and relative humidity during the drying process, the constituent particles in the solution naturally assemble, leading to unique pattern formation. Each pattern serves as a unique fingerprint or blueprint of its respective case. To preserve consistency in the signatures of these patterns, the droplets' size and shape need to remain uniform.
[0012] 4) The patterns can be differentiated into two categories in a crude sense. One is the dried patterns, which represent the final morphology left after the liquid has completely evaporated. Unlike dried patterns, which provide static information spatially, drying patterns encode temporal and spatial data, offering richer insights into the dynamics occurring during evaporation.
[0013] 5) This spatio-temporal evolution of the drying patterns is used as an input for implementing AI, as different flow behaviors, phase changes, mechanical instabilities, multi-component aggregation, etc., occur during drying. Due to these events, the texture of the images changes and is directly correlated with the different initial compositions and the drying patterns.
[0014] 6) Two approaches are employed -- first, quantitative textural analysis is conducted on the images to derive numerical data, which is subsequently used in traditional machine learning algorithms such as random forests, decision trees, etc. In the second approach, images obtained during the drying process are directly input into deep learning (DL) neural networks, where feature extraction is automated by the algorithm, in contrast to the first approach, where features are manually extracted based on textural changes. This dual, parallel approach has the potential to uncover new insights and achieve high-accuracy classification of drying patterns, even with smaller fluid volumes, compared to the current methods that focus solely on final dried-state images.
[0015] 7) Instead of using physics-informed data in AI, we have used the data formed due to drying-induced textural changes. Although these physics-informed parameters (for example, the normalized ring width and crack spacing) provide valuable insights into the drying dynamics, they come with certain limitations. These parameters are highly sensitive to external conditions and do not generalize well across various samples with different compositions and disease conditions. For instance, crack spacing in some cases does not show a consistent correlation with concentration due to the irregularities in crack formation. The advantage of using these textural parameters over any physics-informed parameters is that these capture the intricate variations in the drying patterns of droplets, offering a data-driven approach to quantify the underlying structural heterogeneity within any drying droplets. One major advantage is that these textural parameters provide a more generalized description of the drying process by capturing subtle pixel intensity variations across different regions of the droplet. This is particularly useful when complex factors, such as the composition of blood components, ions, external environmental conditions, etc., influence the droplet’s drying behavior. While important, physics-informed parameters typically focus on a limited set of measurable variables, which might not always be easy to measure consistently across all droplets. This makes the use of such parameters less robust when compared to texture-based quantitative measures. Furthermore, textural analysis is inherently flexible and adaptable. It can be applied to any image, regardless of the underlying physical properties, making it an excellent candidate for analyzing complex, real-world biological systems where the physical models might be incomplete or insufficient. The extraction of textural features also allows for scaling up analyses in a high-throughput manner, particularly when combined with AI that can process and interpret large volumes of data efficiently.
[0016] 8) Integrating AI with the drying process achieves a prediction accuracy of above 90% in classifying different sessile droplets. Image textures were quantified using first-order statistics (FOS) and gray-level co-occurrence matrix (GLCM) to capture changes during drying. Other second-order statistics can also be applied, such as Gray Level Size Zone Matrix (GLSZM), Gray Level Run Length Matrix (GLRLM), Gray Level Dependence Matrix (GLDM), etc., or the combinations of any of these statistical features. These techniques, individually or in combination, provide deeper insight into textural variations by capturing spatial relationships, structural complexity, and intensity distribution of the images captured at every drying pattern. FOS and GLCM features can be computed using ImageJ, while a comprehensive set of textural statistical parameters can be extracted using PyRadiomics in Python. These computational tools enable efficient and automated quantification of image textures, facilitating robust feature extraction for AI-driven pattern classification.
[0017] 9) The embodiments described are not exhaustive, and additional algorithms can be explored by considering different traditional machine learning. Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), K-nearest neighbors (KNN), and Naive Bayes (NB) are implemented as traditional machine learning using Scikit-Learn in Python in this disclosure.
[0018] 10) To implement DL, we utilized Keras in conjunction with the TensorFlow library or Pytorch in Python. Different pre-trained models, such as VGG16, ResNet-18, etc., can be employed. This base model had not previously been exposed to these droplet images. The hyperparameters are tuned while training for different numbers of epochs and batch sizes. Categorical cross-entropy was used when there was a multi-class identification as the loss function, and the model was optimized using the Adam optimizer.
[0019] 11) The Gradient-weighted Class Activation Mapping (Grad-CAM) technique is applied to DL models to generate visual explanations that highlight the specific regions of the drying patterns contributing to the model’s predictions. This enhances the interpretability of the DL models by enabling users to understand how the AI associates particular features of the patterns with specific abnormalities or diseased conditions. By providing these insights, Grad-CAM not only improves the transparency of the classification process but also builds confidence in the system’s predictive accuracy.
[0020] 12) The various metrics were employed using Scikit-Learn in Python in this disclosure to assess the efficacy of AI models in classifying the samples. For example, a confusion matrix provides a quantitative measure that contrasts actual and predicted values based on True Positives (TP) --- indicating correctly predicted positive instances, False Positives (FP) --- representing incorrect positive predictions, True Negatives (TN) --- signifying accurate negative predictions, and False Negatives (FN) --- corresponding to inaccurate negative predictions. Based on these, other metrics are also computed, such as Accuracy, Precision, Recall, and F1-Score, and the model’s performance is compared.
[0021] 13) The described embodiments of the whole pipeline are not exhaustive, and additional variations can be explored by considering different combinations of the explicitly detailed embodiments. Unless otherwise specified in this disclosure, it will be apparent to those skilled in the field that the described embodiments can be integrated or adapted in various ways. For instance, unless explicitly stated otherwise, all features outlined in the embodiments --- whether related to the method or system for analyzing drying patterns --- can be combined with or substituted by other features from different embodiments to optimize performance and applicability.
[0022] 14) According to the present disclosure, the minimally required components for implementing the system are provided. For instance, at least 100 - 300 images of a sessile droplet on a substrate are captured throughout the drying process, with at least one image processing software performing computational analysis on the captured images. Such image processing software may be installed on a computing device and may include ImageJ, MATLAB, and / or Python scripts for feature extraction and quantification. According to one or more embodiments, the system further comprises a database that stores sets of numerical or image-based datasets collected under different conditions. Each dataset includes extracted textural statistical parameters, such as first-order statistics (FOS), Gray Level Co-occurrence Matrix (GLCM), and other second-order statistical features. The system also accounts for variations in environmental and initial conditions that impact the drying process, ensuring robust and reproducible analysis. In certain embodiments, the system further includes a camera for real-time image acquisition of the drying droplets. AI-based analysis is implemented using Python scripts, facilitating automated pattern recognition and classification. While deep learning (DL) models require high GPU capabilities and advanced computational hardware, the disclosure also provides an alternative for low-resource scenarios by utilizing traditional machine learning (ML) approaches based on extracted textural features. Thus, this disclosure presents an integrated pipeline that spans from data collection to feature extraction, analysis, and predictive modeling, offering a scalable and adaptable solution for biofluid pattern classification.
[0023] 15) This disclosure shows some diverse systems where this classification pipeline has been applied, demonstrating a versatile classification protocol and highlighting its adaptability and robustness. While the primary focus is on disease detection, we have also applied this approach to other bio-mimic systems, including blood mixtures, protein-protein, protein-salt, and liquid crystal mixtures. This adaptability is only possible due to textural quantification and correlation to the subtle compositional and concentration changes, making it a powerful tool for classifying any pattern dynamics. It highlights the significance of using drying patterns with AI for precise classification, presenting a proof-of-concept for developing sustainable, low-volume, and highly accurate screening tools.
[0024] Figure 1 compares the current healthcare diagnostic system with the proposed drying droplet-based AI method.
[0025] Figure 1 provides a comparison between the conventional healthcare diagnostic pathway and the method of this disclosure involving drying droplets combined with AI. In Figure 1, the conventional healthcare system is illustrated. In the conventional healthcare system, the diagnostic process begins with the collection of bio-fluids such as blood, urine, or saliva. These samples are analyzed through advanced pathological tests, including biochemical assays and microscopy techniques, to detect diseases or abnormalities. Once the test results are available, the doctor interprets the findings, diagnoses the condition, and prescribes medication to the patient. While effective, this system relies heavily on skilled personnel, sophisticated equipment, and multiple stages of analysis, making it time-consuming and resource-intensive. This dependence limits its accessibility, especially in low-resource settings for preliminary diagnosis. In contrast, this disclosure highlights the transformation of the diagnostic process by integrating the dynamics of drying droplets, where these droplets undergo evaporation, forming unique patterns. The integration of these drying pattern dynamics with AI in this disclosure opens new possibilities for disease classification while addressing global healthcare challenges effectively.
[0026] Figure 2 illustrates a detailed methodology for classifying biological samples using drying droplets and AI.
[0027] Figure 2 illustrates a flowchart representing the classification pipeline, where the entire drying process is utilized as input rather than relying solely on the final dried morphologies. The system captures evolving patterns throughout the drying phase. It integrates them with quantitative image textural analysis and AI-based classification techniques, including traditional machine learning and deep learning. This approach enables the identification and classification of various samples, including biofluids with abnormalities, disease markers, or bio-mimetic systems.
[0028] (A) SAMPLES(S21): Samples include both bio-mimetic and real bio-fluid samples. The bio-mimetic samples could be the system containing (i) globular proteins (lysozyme protein derived from chicken), optically active liquid crystals (thermotropic 5CB), and different amounts of phosphate saline buffer (PBS), and (ii) globular protein (lysozyme protein derived from chicken) and different amounts of PBS. The real biofluids are derived from humans, for example, the whole blood consisting of cellular components (red blood cells (RBCs), white blood cells (WBCs), and platelets), proteins, ions, and coagulant agents. When the whole blood fluid excludes the described cellular components, it is referred to as plasma. The plasma fluid, excluding these coagulant agents, is referred to as serum. These bio-fluids can serve as bio-mimetic samples by diluting blood with water or buffer solutions to simulate various abnormalities. Additionally, they can be collected directly from patients diagnosed with different diseases, such as cancer or diabetes at various stages, to analyze and classify disease-specific patterns.
[0029] (B) DROPLET PREPARATION(S22): A simple bio-physical methodology is proposed where the circular droplets of a volume of approximately 1 μL and approximately 2 mm diameter are deposited on the solid surface, and the liquid (water) evaporates with time within the droplets, referred to as drying droplets. It takes only 5-10 mins (rapid) in which drying occurs under a fixed environmental condition (temperature [T] of 20-30 °C and relative humidity [RH] of 35-55%), and the solid surface has an average contact angle of 40-60 degrees with the droplet.
[0030] (C) IMAGE ACQUISITION(S23): Microscopy comprises a magnification lens ranging from 2x (minimum) to 5x (maximum) for capturing the entire 2 mm diameter droplet; a camera is attached to the microscopy to facilitate capturing images at a given frame per second, and a computer with camera software is installed, wherein the camera is configured to save captured images on a computer. To ensure high-resolution images, the images should not be compressed or filtered during the camera’s initial click; however, the minimum requirement is to click one frame in two seconds repeatedly. Depending on the nature of the bio-fluids, microscopy under different configurations, such as bright-field and crossed-polarizing configurations, can be used. Each sample’s images were captured throughout the drying process, and the light remained constant to minimize background fluctuations.
[0031] (D) QUANTIFICATION OF IMAGES: ImageJ software was utilized to select a circular region of interest (ROI) using the oval tool. Gray values in the 8-bit images ranged from 0 to 255. All images were transformed into 8-bit images for improved visualization. Textural features, including first-order statistics (FOS) and gray-level co-occurrence matrix (GLCM) attributes, were extracted. FOS parameters include mean, standard deviation, Kurtosis, and skewness of the image. GLCM parameters include angular second moment, contrast, correlation, inverse difference moment, and entropy. GLCM parameters were computed using the Texture Analyzer plugin in ImageJ. The Pyradiomics tool in Python can also be used to extract the quantitative features during drying. The images are re-sized, and all images are converted into 8-bit grayscale images. Following this, a mask is selected on the time series of the droplet to highlight the region of interest (ROI) for the feature extraction. A systematic approach was employed to select these variables. Features related to shape were disregarded due to the limitation of top-view images, which capture the droplet in 2D. The droplet shape remains consistent during drying as its edge is pinned to the substrate. Subsequently, all 89 quantitative features from the Pyradiomics package can be used as the features, or the visual inspection can be used to identify the features that exhibit distinct behavior. These features were then selected for further analysis. However, these are the standard features that can also be generated using any other image processing software. This includes First Order Statistics (FOS), Gray Level Co-occurrence Matrix (GLCM), Gray Level Run Length Matrix (GLRLM), Gray Level Size Zone Matrix (GLSZM), and Gray Level Dependence Matrix (GLDM). The respective classes are RadiomicsFirstOrder, RadiomicsGLCM, RadiomicsGLRLM, RadiomicsGLSZM, and RadiomicsGLDM, which are used to get quantitative textures for all the images.
[0032] (E) MACHINE LEARNING(S24): Two analytical approaches are employed. Deep Learning is employed as approach 1, where entire images are used directly as input for neural networks, allowing the model to learn complex patterns and relationships in the data. This approach requires high computational resources, including GPUs, and excels at handling large datasets. Traditional Machine Learning is employed as approach 2, where features are extracted from the images through texture analysis from different libraries (Python or ImageJ) as shown in (D) QUANTIFICATION OF IMAGES and fed into ML models. This method is more suitable for low-resource environments as it demands fewer computational resources.
[0033] The common step for implementing traditional machine learning using Python is to load the CSV file using pandas. Separate the features (all columns except the last) and the target variable (the last column). Handle missing data if required, using techniques like imputation or removing irrelevant features. This dataset (S) can be denoted as S=(x1,y1),(x2,y2),...(xn,yn) with n representing the total dataset size. Each x comprises a feature vector. Correspondingly, y corresponds to the respective class, encompassing distinct types of samples. All ML implementations are supervised learning. The dataset was divided randomly into two mutually exclusive groups: Strainand StestDifferent ML algorithms were assessed using Strainand subsequently tested for performance on DtestThe dataset was split into Strainand Stestusing the train_test_split() function from the scikit-learn library. A test size of 0.2 or 0.3 was employed, indicating that 80% or 70% of S was designated as the training set, while the remaining 20% or 30% served as the testing set.
[0034] (F) TRADITIONAL MACHINE LEARNING: Five distinct traditional ML algorithms were employed --- Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), and Naive Bayes (NB). Implementing these algorithms was facilitated through the Scikit-Learn library, which offers a Python interface for traditional ML methods. Prior to applying any ML technique, the data is scaled. Random Forest (RF) is built upon Decision Trees (DT). Each individual tree in the ensemble predicts outcomes, and a majority vote determines the final class prediction. RF employs bagging and feature randomness during tree construction, resulting in an uncorrelated forest of multiple trees. The RF classifier was instantiated using RandomForestClassifier from the sklearn.ensemble module. Decision Trees (DT) follow a branching approach to illustrate possible decision outcomes. We utilized the DecisionTreeClassifier with criterion = 'entropy'. This classifier was imported from sklearn.tree. K-Nearest Neighbors (KNN) uses a distance metric to classify data points. The KNeighborsClassifier was employed with parameters such as n_neighbors set to 5, weights as 'distance', and metric as 'minkowski'. These settings were imported from sklearn.neighbors. Support Vector Machine (SVM) is based on identifying a hyperplane and support vectors closest to the hyperplane to separate different classes. We used the SVC() function from sklearn.svm for the analysis. Naive Bayes (NB) leverages Bayes’ Probability Theorem to compute the probability of data belonging to a specific class. We imported GaussianNB() from sklearn.naive_bayes for classification purposes.
[0035] (G) DEEP LEARNING: To implement a neural network, we utilized Keras and Pytorch in conjunction with the TensorFlow library. Neural networks operate on the concept of artificial neurons, which loosely mimic the neurons in the human brain. These networks establish probabilistic-weighted connections between inputs and output classes. The architecture of these networks is constructed and manipulated to predict output classes based on the inputs. Throughout the training process, the network’s weighted connections are adjusted, aiming to achieve increasingly accurate output predictions. Training concludes after a sufficient number of iterations. For this, the images are resized to 224 x 224 pixels. We can implement different architectures, such as VGG16, ResNet18, etc., from either Keras or Pytorch and pre-trained on the ImageNet dataset. This base model had not been previously exposed to these droplet images. However, the fully connected layers were retrained using our images. The base VGG16 model, with frozen weights, was integrated into a sequential model, followed by a GlobalAveragePooling2D (GAP) layer and two different dense layers. The VGG16 or ResNet18 base model can be optimized to address overfitting by reducing its complexity, such as selecting only a subset of the convolutional layers or removing the pooling layer. Additionally, the dense layers can be modified to simplify the architecture further and make it resource-efficient, minimizing computational overhead. The final layer, equipped with softmax activation, generated class probabilities for multiple categories. For binary classification, it is replaced with a sigmoid activation function, which outputs a single probability score for one of the two classes, simplifying the decision-making process and making it more suitable for binary outcomes. The model was trained for different epochs with early stopping applied to halt training if validation performance plateaued. A batch size was selected to balance computational efficiency with stable training. Categorical cross-entropy and binary cross-entropy were used as the loss functions for multi-class identification and binary classification, respectively. The model is optimized using the Adam optimizer with a learning rate of 0.01-0.0001.
[0036] (H) VISUALIZATION OF DEEP LEARNING USING GRAD-CAM: Grad-CAM (Gradient-weighted Class Activation Mapping) is a powerful visualization technique used to interpret deep learning models by highlighting regions in an input image that contribute most significantly to the model’s predictions. The implementation of Grad-CAM begins with specifying essential parameters, such as the image size (e.g., 224 x 224 pixels, as used in many pre-trained models), the image path for loading the data, and identifying the last convolutional layer of the model. Identifying this layer requires a careful examination of the model’s architecture, as the final convolutional layer generates the feature maps used for heatmap computation. The process starts by iterating over the images in the specified directory, where each image is loaded and converted into a numerical array suitable for the deep learning model. This is followed by defining a function, make_gradcam_heatmap(), which leverages tools from the Keras library to compute the Grad-CAM heatmap. The function calculates the gradients of the predicted class score concerning the feature maps of the last convolutional layer. These gradients are globally averaged (pooled) to obtain the importance weights, which are then multiplied with feature maps to generate the heatmap. The computed Grad-CAM heatmap is then superimposed on the original image, providing a visually intuitive overlay that highlights the most relevant regions to the model’s prediction.
[0037] (I) EVALUATION METRICS: Various metrics were employed to assess the efficacy of ML models in classifying blood samples. These metrics encompass: (i) 10-fold Cross-Validation: A 10-fold cross-validation procedure was applied to access the overfitting and underfitting issues in the traditional MLs, involving dataset shuffling and division into ten groups. Each group was utilized as a test set, while the remaining nine served as the training set. The model was fitted on the training data, the test data was evaluated, and an evaluation score was derived. This process was repeated for all ten groups, enabling a comprehensive comparison. The KFold function from sklearn.model_selection was employed for this purpose. (ii) Confusion Matrix: The confusion matrix provides a quantitative measure that contrasts actual and predicted values. This m x m matrix, where m denotes the number of classes (in this case, 11), showcases the relationship between true and predicted class labels. Diagonal elements correspond to correct predictions, while off-diagonal elements denote misclassifications. The confusion_matrix function from sklearn.metrics was employed. Typically, these matrix values are normalized and presented as a color map within the range of 0 to 1. (iii) Accuracy, Precision, Recall, and F1-Score: A series of classification metrics were computed, where accuracy represents the proportion of correctly predicted instances out of the total occurrences. Precision, Recall, and F1-score were also calculated. Mathematically,
[0038]
[0039] Here, True Positives (TP) indicate correctly predicted positive instances, False Positives (FP) represent incorrect positive predictions, True Negatives (TN) signify accurate negative predictions, and False Negatives (FN) correspond to inaccurate negative predictions. The accuracy_score and classification_report functions from sklearn.metrics were utilized to compute these metrics. Of note is that the F1 score, which balances precision and recall, was employed to compare the performance of each algorithm.
[0040] For deep learning, in addition to these metrics, the training and validation loss and accuracy curves provide critical insights into model performance. These curves help assess whether the model is underfitting, overfitting, or generalizing well to unseen data. A smooth and gradually decreasing training loss, along with a validation loss that plateaus or decreases similarly, indicates a well-trained model. However, if the validation loss starts increasing while the training loss continues to decrease, it suggests overfitting, meaning the model is memorizing training data instead of learning generalizable patterns.
[0041] This method of disclosure is implemented by a computer system illustrated in Figure 38. This computer system comprises a processor 11, a memory 12, an operation unit 13, a display unit 14, and an input / output unit 15, and is connected to a microscope device 2. This microscope device 2 captures images at predetermined intervals of biological fluids (taken from subjects including humans and animals) deposited on a solid surface, and outputs the image data to the information processing device 1 each time image data is obtained.
[0042] The processor 11 operates according to programs stored in the memory 12 and receives the image data repeatedly output by the microscope device 2. Additionally, the processor 11 uses at least one of: (1) A deep learning model that has been machine-learned to take the received image data as input and output disease types and progression states as classification results, or (2) A model that has been machine-learned to take features extracted from the received image data as input and output disease types and progression states as classification results (classifier) to obtain classification results of the subject's disease type or disease state and so forth. The processor 11 then outputs these obtained classification results. The operation of the processor 11 will be described in detail later.
[0043] The memory 12 stores the programs executed by the processor 11. These programs may be provided stored on computer-readable non-transitory storage media and copied to the memory 12. The memory 12 also works as a working memory for the processor 11. The operation unit 13 consists of devices such as a mouse and keyboard, receives user operations, and outputs information representing the content of these operations to the processor 11. The display unit 14 consists of devices such as a display monitor and displays information according to the processor 11's instructions. The input / output unit 15 is an interface such as USB, and in this disclosure's example, is connected to the microscope device 2.
[0044] The microscope device 2 repeatedly captures magnified images of the target object at predetermined intervals and outputs the obtained image data to the computer system. In this embodiment, the target object is biological fluid deposited on a solid surface, and the microscope device 2 repeatedly captures images of this biological fluid over a predetermined time period (for example, every 5 to 10 minutes).
[0045] Furthermore, in this embodiment, the microscope device 2 is placed in an environment where there are no external field effects such as magnetic fields or mechanical agitation on the imaging target, and where environmental factors such as temperature and relative humidity have minimal impact during the drying process by maintaining constant temperature and humidity, without applying external fields. As one example, microscope device 2 captures images of the target object in a room environment with a temperature of 20 to 30 degrees Celsius and relative humidity of 35 to 55%. To place the droplets in the drying process within an environment with constant temperature and humidity selected from this range, well-known temperature and humidity control methods are used to control temperature and humidity while allowing fluctuations within the range of ±5-10°C and ±5-10% relative humidity without affecting the processing method.
[0046] In this embodiment, the processor 11 executes programs stored in the memory 12 and performs machine learning and inference processing. In the machine learning processing, it trains different models, such as decision trees (DT), random forests (RF), support vector machines (SVM), k-nearest neighbors (KNN), naive Bayes (NB), etc., as well as deep learning models, such as VGG-16, ResNet-18.
[0047] In the machine learning processing, the operator prepares multiple biological fluid samples obtained from patients whose classification results (disease types and progression states) are unknown in advance, and deposits each of these samples on a solid surface. The microscope device 2 then captures time-series image data (multiple image data capturing the drying stages) of each deposited biological fluid. The processor 11 performing the machine learning processing receives multiple time-series image data associated with each of multiple samples captured by the microscope device 2, and preprocesses the multiple time-series image data to provide input image data corresponding to each sample based on these multiple image data. This input image data is used for machine learning of the deep learning model.
[0048] Specifically, since image data must be input as vector information to the deep learning model, the processor 11 arranges the pixels of each image data in a predetermined order to vectorize it, and then concatenates the vector data obtained from each image data in a predetermined order (for example, in chronological order of image capture) to generate a single vector data. The processor 11 then uses this generated single vector data as input image data for the deep learning model. However, this is just one example, and any method that can input a series of image data is acceptable. In this embodiment's example, when using the deep learning model, the same method is used to generate a single input image data from a series of image data in both machine learning and inference processing.
[0049] The processor 11 inputs the input image data corresponding to each sample into a predetermined deep learning model (specific examples of this model will be described later), and performs machine learning by updating the deep learning model using its output and the classification results (training data) corresponding to that sample. This updating method can adopt the widely known backpropagation method.
[0050] Additionally, the processor 11 extracts predetermined feature information from each of the multiple image data obtained for each sample. It is also preferable to use texture features as the feature information. Available texture features include first-order statistics (FOS), gray-level co-occurrence matrix (GLCM), gray-level run length matrix (GLRLM), gray-level size zone matrix (GLSZM), gray-level dependence matrix (GLDM), etc. The processor 11 then preprocesses the extracted feature information to provide training data by pairing information representing the temporal changes in feature information obtained from multiple image data for each sample with the classification results (training data) corresponding to that sample, and performs machine learning on a predetermined classifier. This classifier is machine-learned to receive input of information representing temporal changes in feature information and output classification results such as disease types. Such classifiers can be implemented using widely known models such as decision trees (DT), random forests (RF), support vector machines (SVM), k-nearest neighbors (KNN), naive Bayes (NB), etc.
[0051] Through the above machine learning processing, the deep learning model reaches a state where it has been machine-learned to output corresponding disease types and other classification results when input image data is input, and the classifier reaches a state where it has been machine-learned to output corresponding disease types and other classification results when the above feature information is input.
[0052] When performing inference processing, the operator deposits the biological fluid to be classified on a solid surface, and the microscope device 2 captures time-series image data (multiple image data capturing the drying stages) of that deposited biological fluid. The processor 11 executes the processing shown in Figure 39 according to operator instructions, first receiving multiple image data captured by microscope device 2 (S21). Then the processor 11 preprocesses the multiple image data to provide input image data for the deep learning model based on these multiple image data (S22). The method for generating this input image data is kept the same as during machine learning processing. Additionally, the processor 11 extracts predetermined feature information from each of the image data sequentially received in step S21 and preprocesses the feature information to provide input data (S23). The feature information extracted here is also kept the same as that used during machine learning processing. Processor 11 may perform the processing of Steps S22 and S23 in parallel.
[0053] The processor 11 inputs the input image data provided in step S22 into the deep learning model that was machine-learned during machine learning processing and obtains its inference output (S24). This output represents classification results such as disease types. Additionally, processor 11 inputs the input data into the classifier, which was machine-learned during the machine learning process. The input data is provided through preprocessing of feature information extracted in Step S23, wherein the preprocessing is performed in the same manner as preprocessing during the machine learning process. Processor 11 then obtains its inference output (S25). This output also represents classification results such as disease types. The processor 11 presents the outputs obtained in steps S24 and S25 to the operator (S26: result output).
[0054] Note that while both deep learning model and classifier are used here, when using only the deep learning model, steps S23 and S25 processing can be skipped. Similarly, when using only the classifier, steps S22 and S24 processing can be skipped.
[0055] Now, RESULTS AND DISCUSSIONS ON THE BIO-MIMETIC SYSTEMS are described. (A) The droplets contain lysozyme, liquid crystals (LC), and different amounts of phosphate saline buffer (PBS) AIM: The objective is to classify five distinct types of droplets with varying PBS concentrations --- 0x, 0.25x, 0.5x, 0.75x, and 1x --- by analyzing their drying patterns and employing traditional ML techniques.
[0056] Figure 3 (A) shows the drying pattern dynamics of the droplets containing lysozyme, liquid crystals (LC), and different amounts of phosphate saline buffer (PBS). The scale bar is 0.2 mm in length.(B) shows the texture analysis that includes First-Order Statistics (FOS) encompassing a range of features: (I) Mean, (II) Variance, (III) Skewness, (IV) Kurtosis, (V) Root mean square, (VI) Uniformity, (VII) Entropy, and (VIII) Energy. Dynamic changes in FOS parameters are presented over normalized time and calculated as the instantaneous time divided by the total time. Noteworthy stages -- initial, middle, and final -- are marked with the yellow, pink, and blue colors, respectively.
[0057] Optical images depict the progressive drying stages of liquid crystal (LC)-protein droplets with varied initial PBS concentrations (0x, 0.25x, 0.5x, 0.75x, and 1x). These images are captured under a crossed polarizing configuration and are represented by crossed arrows in Figure 3(A). The experiments were performed under ambient conditions (the room temperature of 25°C and the relative humidity of 50%). At the beginning of the experiment, the droplet’s contact angle was approximately 40°.
[0058] The changes in the dynamics of the patterns are found in the initial, middle, and final stages of the drying process, as shown in Figure 3(A). The initial stage reveals gradual changes in these parameters. However, limited variations in these parameters are observed for specific PBS concentrations. The middle stage shows a rapid rise and is more vibrant than the initial and final stages. In contrast, the final stage stabilizes the parameters, where their values become relatively constant. This three-stage FOS evolution provides a comprehensive picture of the dynamic and settling behaviors exhibited by the droplet throughout drying (see Figure 3(B)).
[0059] Figure 4(A) shows the features from the Gray Level Co-occurrence Matrix (GLCM) encompassing a range of variables: (I) Contrast, (II) Correlation, (III) Inverse Difference Moment, (IV) Maximum Probability, (V) Difference Average, (VI) Difference Variance, (VII) Sum Entropy, and (VIII) Difference Entropy. (B) shows the features from the Gray Level Run Length Matrix (GLRLM) encompassing a range of variables: (I) Variance, (II) Run Variance, (III) Nonuniformity, and (IV) Run Entropy. Dynamic changes in GLCM and GLRLM parameters are presented over normalized time and calculated as the instantaneous time divided by the total time. The liquid crystal (LC)-protein drying droplets with varied initial buffered concentrations include 0x, 0.25x, 0.5x, 0.75x, and 1x. Noteworthy stages -- initial, middle, and final -- are marked with the yellow, pink, and blue colors, respectively.
[0060] Figure 4A(I-VIII) shows the quantitative drying evolution of the LC-protein droplets with varied initial PBS concentrations using the Gray Level Co-occurrence Matrix (GLCM). In the context of image analysis, various texture features derived from a GLCM provide valuable insights into local intensity patterns. The variables include Contrast, Correlation, Inverse Difference Moment (IDM), Maximum Probability, Difference Average, Difference Variance, Sum Entropy, and Difference Entropy [Figure 4A(I-VIII)]. The middle stage is the most dynamic phase, characterized by peaks and troughs in the parameter values of GLCM.
[0061] The Gray Level Run Length Matrix (GLRLM) computes Variance, Run Variance (RV), Non-Uniformity, and Run Entropy (RE) [Figure 4B(I-IV)]. Variance quantifies the variation in gray level intensity for the runs, while RV assesses the variance in runs concerning run lengths. The variance exhibits a trend starting with lower values and increasing toward the end of the process [Figure 4B(I-II)]. Non-Uniformity quantifies the similarity of gray-level intensity values in the image, with a lower value indicating greater similarity in intensity values. On the other hand, RE measures uncertainty and randomness in the distribution of run lengths and gray levels. A higher RE value suggests increased heterogeneity in the texture patterns. The expected relationship between Non-Uniformity and RE is that they should exhibit opposite trends, given that the former measures homogeneity while the latter measures heterogeneity. This anticipated Contrast is observed in Figure 4B(III-IV), where the initial drying stage is characterized by uniform texture, while the resulting patterns in the final drying stage are heterogeneous.
[0062] Figure 5 (A) shows the features from the Gray Level Size Zone Matrix (GLSZM) encompassing a range of variables: (I) Non-Uniformity, (II) Zone Non-Uniformity, (III) Variance, (IV) Zone Variance, and (V) Zone Entropy. (B) shows the features from the Gray Level Dependence Matrix (GLDM) encompassing a range of variables: (I) Non-Uniformity, (II) Dependence Non-Uniformity, (III) Variance, (IV) Dependence Variance, and (V) Dependence Entropy. Dynamic changes in GLSZM and GLDM parameters are presented over normalized time and calculated as the instantaneous time divided by the total time. The liquid crystal (LC)-protein drying droplets with varied initial buffered concentrations include 0x, 0.25x, 0.5x, 0.75x, and 1x. Noteworthy stages -- initial, middle, and final -- are marked with the yellow, pink, and blue colors, respectively.
[0063] Figure 5A-B(I-V) shows GLSZM features, including nonuniformity, zone nonuniformity, variance, zone variation (ZV), and zone entropy (ZE). In contrast, the GLDM feature consists of Non-Uniformity, Dependence Non-Uniformity (DN), Variance, Dependence Variance (DV), and Dependence Entropy. Variance quantifies the variance in gray level intensities for the zones, while ZV measures the variance in zone size volumes for the zones. ZE evaluates the uncertainty and randomness in the distribution of zone sizes and gray levels, with a higher value indicating increased heterogeneity in the texture patterns. Non-Uniformity assesses the variability of gray-level intensity values in the image, and a lower value signifies more homogeneity in intensity values. In contrast, Zone Non-Uniformity gauges the variability of size zone volumes in the image, with a lower value indicating more homogeneity in size zone volumes.
[0064] The observed trend between Non-Uniformity and Zone Non-Uniformity follows a similar pattern, exhibiting lower values at the initial stage and higher values at the final drying stage. However, the trend in Variance and Zone Variance is not closely aligned, suggesting that the zone’s impact may vary depending on the specific parameter being considered. Not only this, but the Gray Level Variance under the GLDM measures the variance in gray levels in the image, while DV assesses the variance in dependence size in the image (Figure 5A-B(I-IV)).
[0065] Figure 6 displays the normalized confusion matrix illustrating the performance of various traditional ML: (A) Naive Bayes (NB), (B) Support Vector Machine (SVM), (C) Decision Tree (DT), (D) K-Nearest Neighbors (KNN), and (E) Random Forest (RF). These MLs are employed to classify different drying patterns of the bio-mimetic system of protein, liquid crystals, and buffer. (F) The radial plot depicting the average F1-score in percentage provides a comparative assessment of each ML’s performance.
[0066] The normalized confusion matrix in Figure 6(A-F) provides a comprehensive view of how well each ML performs in classifying the samples into different categories (see Figure 6(A-E)). NB performs relatively poorly compared to other MLs. SVM’s performance improves when it replaces NB. DT performs better than both NB and SVM. RF shows the best performance, with minimal off-diagonal values in the normalized confusion matrix. The average F1-score, a balanced metric considering precision and recall, is calculated for each ML. The hierarchy of performance based on F1-scores is NB (75%) < SVM (93%) < DT (95%) < KNN (97%) < RF (98%) (see Figure 6(F)). The main conclusion is that the features (FOS, GLCM, GLRLM, GLSZM, and GLDM) derived from texture analysis using PyRadiomics in Python exhibit significant potential as critical parameters for characterizing different drying patterns of these droplets. When these features are fed into traditional MLs, we observe a robust classification performance, at least with more than a predictive F1-score of above 90% in all MLs except for NB. The F1-scores indicate that all ML can effectively differentiate between the samples using the extracted features, demonstrating their ability to distinguish between different categories in the drying patterns effectively. This highlights the reliability and scalability of texture-based analysis combined with ML for applications in screening the samples.
[0067] (B) The droplets contain blood mixtures with water and buffer mimicking the blood abnormalities AIM: The objective is to classify eleven distinct types of blood droplets with varying abnormalities induced by different volumes of added water and buffer ranging from 75% to 12.5% (v / v) using the drying patterns and employing AI methods.
[0068] To modify the concentration of whole blood (initially at 100% by volume), we employed a dilution strategy involving the addition of 1x phosphate buffer saline (PBS) and de-ionized water. Notably, 1x PBS contained a composition of 0.137 M NaCl, 0.0027 M KCl, and 0.119 M phosphates while maintaining a consistent pH within the range of 7.3-7.5. Diverse blood samples were prepared to cover a spectrum of concentrations, ranging from 12.5 to 75% (by volume). Importantly, all experimental procedures were executed soon after sample preparation. This rapid processing ensured that the lysis of various blood components did not occur before the initiation of the drying process. A total of eleven blood samples was prepared, encompassing various compositions: healthy (100 (v / v)%), blood mixed with PBS (12.5, 25, 50, 62, and 75 (v / v)%), and blood mixed with water (12.5, 25, 50, 62, and 75 (v / v)%). Optical images depict the progressive drying stages of these blood mixture droplets. These images are captured under bright-field configuration and are represented in Figures 7 and 8.The experiments were performed under ambient conditions (the room temperature of 25°C and the relative humidity of 50%). At the beginning of the experiment, the droplet’s contact angle was approximately 40°.
[0069] For blood+water sample at 75% (v / v), an inhomogeneous texture first appears in the central region (visible at normalized time of 0.7), eventually spreading across the entire droplet by the end of the drying process. Comparisons between 75% (v / v) and 12.5% of the blood texture samples reveal notable differences in the inhomogeneity, which is the differing component counts and interactions within the droplet drive. The 75% (v / v) sample shows dark and light gray patches due to higher RBC counts, while the 12.5% (v / v) sample exhibits an overall lighter gray shade, indicating a more dilute composition. The central region expands as the blood concentration decreases from 75% to 12.5% (v / v), reflecting changes in fluid dynamics and pattern formation. The crack patterns also differ significantly: the 75% (v / v) blood+water sample shows large radial and chaotic cracks in the central region, while the 12.5% (v / v) sample displays fewer radial cracks, concentrated around the peripheral ring, with no cracks in the central region (see Figure 7).
[0070] Figure 7 illustrates time-lapse images of blood droplets with varying abnormalities induced by different volumes of added water ranging from 75% to 12.5% (v / v). Healthy blood corresponds to 100% (v / v). The final drying time represents the point at which no further visible changes occur. The timestamps indicate the normalized time (time divided by the drying time), with a scale bar of 0.2 mm.
[0071] Blood+buffer samples follow a similar initial drying behavior. However, during the middle-to-final stages of drying, distinct patterns emerge, influenced by the blood’s composition and ionic balance. At 75% (v / v) concentration, dendritic structures form in the central region of the droplet. This behavior, influenced by the ionic balance and reduced cellular counts, highlights the importance of composition in driving pattern formation. In blood droplets with PBS concentrations ranging from 75% to 50% (v / v), radial and orthoradial cracks develop, further emphasizing how the unique component interactions at different concentrations impact the stress distribution and pattern formation. Surprisingly, no radial cracks are observed at 25% (v / v). Instead, a distinct ring-like pattern appears in the central region, demonstrating how drastically lower component counts and altered blood composition affect the final structure. At 12.5% (v / v), the central region exhibits a heterogeneous and grainy texture with no significant cracks, again reflecting the influence of reduced blood composition and component interactions on the drying dynamics (see Figure 8).
[0072] Figure 8 illustrates time-lapse images of blood droplets with varying abnormalities induced by different volumes of added buffer ranging from 75% to 12.5% (v / v). Healthy blood corresponds to 100% (v / v). The final drying time represents the point at which no further visible changes occur. The timestamps indicate the normalized time (time divided by the drying time), with a scale bar of 0.2 mm.
[0073] Figure 9 shows the textural statistics, including FOS and GLCM features, which are analyzed for these different concentrations and plotted as a function of normalized time for added water. FOS parameters include mean, standard deviation, skewness, and kurtosis, while GLCM features encompass angular second moment, contrast, correlation, inverse difference moment, and entropy.
[0074] Figure 10 displays the textural statistics, including FOS and GLCM features, which are analyzed for these different concentrations and plotted as a function of normalized time for added buffer. FOS parameters include mean, standard deviation, skewness, and kurtosis, while GLCM features encompass angular second moment, contrast, correlation, inverse difference moment, and entropy.
[0075] Figures 9 and 10 provide insight into the textural variation in blood droplets with and without abnormalities. FOS includes parameters like mean, standard deviation, skewness, and kurtosis of the image. In contrast, GLCM incorporates features like angular second moment, contrast, correlation, inverse difference moment, and entropy. These textural statistics are extracted from time-lapse images captured during drying for different blood abnormalities (+water and +buffer) and healthy blood (without any abnormalities). A unique trendline is evident in all these textural parameters when plotted as a function of the normalized time. Notably, during the middle stage of the drying process, most of these statistical features exhibit dynamic behaviors characterized by peaks and dips, which are absent in the initial and final drying stages. The textural statistics presented in Figures 9 and 10 depict variations on both spatial and temporal scales. For instance, the temporal variation of the mean statistical feature of the droplet shows a slow initial increase, followed by a rapid rise, a gradual decrease, and eventual saturation. On a spatial scale, during the initial drying stage, the mean values of blood droplets with 100 to 62%(v / v) concentration exhibit a range between 15-25 a.u., whereas it is 50-80 a.u. for droplets with 50 to 12.5%(v / v) concentration. In the final drying stage, the mean values range between 40-60 a.u. for 100 to 62%(v / v), 60-80 a.u. for 50 to 25%(v / v), and surpass 80 a.u. for droplets with 12.5% (v / v) concentration.
[0076] The FOS and GLCM statistics observed for blood+PBS (Figure 10) exhibit similar patterns capturing both temporal and spatial variations, comparable to those observed in blood+water droplets (Figure 9). Regardless of whether water or PBS as a diluent is added, the drying process follows a consistent pattern in terms of its stages: an initial stage (for example, normalized time of ~0.06-0.5), a middle stage (~0.5-0.8), and a final stage (~0.8-1.0). Interestingly, blood+PBS droplets maintain their dynamic behavior from the middle stage to the final stage (normalized time of ~0.8-1.0). In contrast, blood+water droplets exhibit their most dynamic behavior during the middle stage (normalized time of ~0.5-0.8) of the drying process, as illustrated in Figure 9. This suggests that the presence of salts in the PBS solution has a retarding effect on the process of pattern formation compared to the behavior observed when only water is present.
[0077] Figure 11 shows the normalized confusion matrix illustrating the performance of various traditional ML: (I) Naive Bayes (NB), (II) Support Vector Machine (SVM), (III) K-Nearest Neighbors (KNN), (IV) Decision Tree (DT), and (V) Random Forest (RF). These MLs are employed to classify healthy blood droplets and blood droplets with added water or phosphate buffer saline (PBS). (VI) The radial plot depicting the average F1-score in percentage provides a comparative assessment of each ML’s performance.
[0078] The normalized confusion matrix provides a comprehensive view of how well each ML performs in classifying the blood samples into healthy, blood+water, and blood+PBS categories (see Figure 11(I-V)). The performance evaluation includes (i) NB: The accuracy is around 58% for healthy blood, 39% for blood+water, and 87% for blood+PBS. The misclassifications are mainly between blood+PBS and healthy blood or blood+water. NB performs relatively poorly compared to other ML. (ii) SVM: After SVM replaces NB, the performance improves. The accuracy is around 74% for healthy blood, 84% for blood+water, and 100% for blood+PBS. SVM performs better than NB but still has room for improvement, especially in classifying blood+water. (iii) KNN: It performs better than both NB and SVM. It achieves accuracy ranging from 90% to 99% across all classes, demonstrating relatively minimal misclassification. (iv) DT: The accuracy is between 97% and 100% for all classes. DT performs well in accurately classifying the blood samples into their respective categories. (v) RF: It achieves similar accuracy to DT, ranging from 97% to 100%. RF shows the best performance, with minimal off-diagonal values in the normalized confusion matrix.
[0079] The average F1-score, a balanced metric considering precision and recall, is calculated for each ML. The hierarchy of performance based on F1-scores is NB (63.7%) < SVM (90.8%) < KNN (97.8%) < DT (99.3%) < RF (99.7%) (see Figure 11(VI)). The F1-scores indicate that all ML can effectively differentiate between the blood samples with and without abnormalities using the extracted drying features. In fact, the different types of blood abnormalities, i.e., +water and +PBS, could also be identified. Both DT and RF demonstrate excellent performance, with accuracy above 99%, making them equally strong candidates for predicting the different blood sample categories. Therefore, while all the evaluated ML can differentiate between blood samples with and without abnormalities, DT and RF stand out as the best performers, offering accurate classification across the tested categories.
[0080] The main conclusion is we do not need many features, even just FOS and GLCM, derived from texture analysis using the Texture Analysis plugin in ImageJ, exhibiting significant potential as critical parameters for characterizing different drying patterns of these droplets. When these features are fed into traditional MLs, we observe a robust classification performance, at least with more than a predictive F1-score of above 90% in all MLs except for NB. The F1-scores indicate that all ML can effectively differentiate between the samples using the extracted features, demonstrating their ability to distinguish between different categories in the drying patterns effectively. This highlights the reliability and scalability of texture-based analysis combined with ML for applications in screening the samples. The manual extraction of quantitative features, such as textural statistics, followed by input into the traditional RF model, has achieved a prediction accuracy of 99%.
[0081] To explore if neural networks (NN) can offer comparable performance, we utilized a convolutional neural network (CNN), which directly processes the images as input and autonomously extracts features for classifying the blood samples. The CNN architecture and the implementation of GRAD-CAM to classify the different blood samples are shown in Figure 12(I-III).
[0082] Figure 12(I) shows a visual representation of a convolutional neural network (CNN) utilizing the VGG-16 architecture, enhanced by Gradient-weighted Class Activation Mapping (Grad-CAM), to provide insights into the model’s feature extraction process. (II) Bar charts comparing the class-wise and overall performance of CNN and RF models evaluated using F1-scores. (III) Grad-CAM snapshots overlaid on the original images, capturing key regions of interest throughout the drying process.
[0083] To classify images directly without relying on extracted numerical data, we implemented a convolutional neural network (CNN) using transfer learning. Specifically, we utilized the VGG16 architecture pre-trained on the ImageNet data. We configured the base VGG16 model without its top layers, freezing its pre-trained weights to retain them during training. The overall model was built sequentially, with the frozen base model at its core. A GlobalAveragePooling2D (GAP) layer was added to reduce the spatial dimensions of the output from the base model, followed by two fully connected dense layers, both activated using ReLU. The final layer was a dense layer with 11 units and softmax activation, producing class probabilities for the eleven blood samples. We also applied Grad-CAM to the final convolutional layer of the pre-trained model to generate heatmaps that highlight the key areas influencing the model's predictions. This helped us interpret the network's focus during classification and ensured that the model was attending to relevant image regions.
[0084] The performance of CNN and RF is evaluated based on the F1-score depicted in Figure 12(II). This pattern is consistent when the F1-scores for each class are visualized in a bar plot in Figure 12(II). The average F1-score in the radial plot demonstrates that the RF and CNN classify all the blood samples with an F1-score of 99% and 96%, respectively (see Figure 12(II)). Both the RF and CNN demonstrated strong performance in classifying the blood samples. RF achieved a marginally higher F1 score of 99%, outperforming CNN's 96%, indicating that RF was slightly more effective at capturing the distinguishing features of the dataset. However, both models are highly reliable and capable of accurate classification with minimal misclassifications. Since the texture of these samples appears to be drastically unique for each blood sample, the RF model’s ability to handle simpler patterns more effectively gave it a slight edge, but CNN still provided valuable insights through its automated feature extraction.
[0085] Figure 12(III) presents Grad-CAM heatmaps overlaid on original images of blood droplets during the drying process, illustrating the most influential model for CNN in classifying different blood samples. Three distinct conditions are shown: healthy blood at 100% (v / v), blood mixed with water at 75% (v / v), and blood mixed with buffer at 62% (v / v). These heatmaps offer visual insights into how the model identifies key features in each sample as the drying progresses, providing a deeper understanding of the physical changes driving classification. Across all samples, the Grad-CAM heatmaps reveal that the model's focus shifts dynamically during the drying process, aligning with the fact that the drying process is highly dynamic and information-rich. Initially, as the droplet is deposited, the heatmaps indicate that the model does not focus on the droplets themselves, evidenced by the red-highlighted regions outside the droplets in Figure 12(III). This suggests that the model largely ignores the early fluid flow. However, as the drying progresses into the middle stage when the fluid front begins to move, the texture changes start from the droplet edge, and the model shifts its attention to this change. As drying continues, the model's focus gradually moves toward the center of the droplet, particularly as cracks and central patterns begin to form. The increasing intensity near the center reflects the model's attention to these internal morphological features, which likely play a significant role in the final classification. In the case of blood mixed with water at 75% (v / v), the heatmaps display a similar but more fragmented activation pattern. While the focus is still on the fluid front, the attention becomes more dispersed across the droplet as drying continues. This behavior could indicate the irregular drying patterns caused by the evaporation of water, which alters the formation of cracks and structures. The model seems to rely on both peripheral and internal features to classify the water-mixed sample, though with less central focus compared to the healthy blood sample. A similar pattern is observed for the blood sample mixed with buffer at 62% (v / v). Thus, the Grad-CAM heatmaps provide valuable insights into how the CNN interprets drying patterns in different blood samples.
[0086] Figure 13 displays the normalized confusion matrix of random forest (RF) using a testing dataset to classify 11 blood samples, i.e., 0 (healthy), +water (12.5, 25. 50, 62, and 75 (v / v)%, and +PBS (12.5, 25. 50, 62, and 75 (v / v)%.
[0087] Figure 14 exhibits the normalized confusion matrix of the convolution neural network (CNN) using a testing dataset to classify 11 blood samples as 0 (healthy), +water (12.5, 25, 50, 62, and 75 (v / v)%), and +PBS (12.5, 25, 50, 62, and 75 (v / v)%).
[0088] The normalized confusion matrices for RF and CNN are shown in Figures 13 and 14, respectively. In the CNN model, certain misclassifications occur, such as the healthy blood being confused with blood diluted with water at 75% (v / v), resulting in a classification accuracy of 0.91. Additionally, there are classification errors between blood samples with varying concentrations of abnormalities. For instance, while 95% of the blood sample diluted with water at 62% (v / v) is correctly classified, 5% is misclassified as the 50% (v / v) water-diluted sample. Moreover, only 89% of the blood sample mixed with water at 75% (v / v) is accurately identified, with 10% being mistakenly classified as the 50% (v / v) blood-buffer mixture. While the RF also struggles with some minor confusion, it appears to outperform CNN in terms of overall accuracy. The RF misclassifies the blood diluted with buffer at 62 (v / v)% and 12.5 (v / v)%. It shows a stronger distinction between the various blood abnormality levels compared to CNN, with most off-diagonal elements being very close to zero.
[0089] Figure 15 presents the training and validation performance curves for a VGG-16 model over epochs used for classifying different blood samples. The left plot depicts the accuracy progression, while the right plot illustrates the loss curves. In both plots, training performance, and validation performance are shown in gray dots.
[0090] Figure 15 illustrates the loss and accuracy curves as functions of epochs. Both the training and validation curves show consistent improvements, with near convergence in accuracy and loss. The close alignment between the training and validation curves for both accuracy and loss suggest that the model is well-regularized and capable of generalizing unseen data without overfitting. These results suggest that the model is robust and performs well across the data, making it suitable for implementation in similar classification tasks.
[0091] The main conclusion is that both traditional ML and deep learning approaches demonstrate robust classification performance in the three-class identification task (healthy, abnormalities with water, and abnormalities with buffer). Furthermore, when classifying all eleven sample categories, the models achieve a predictive F1-score ranging from 96% to 99%, highlighting their reliability and effectiveness in distinguishing between diverse sample types by using the patterns that emerged during the drying process.
[0092] RESULTS AND DISCUSSIONS ON REAL DISEASED SYSTEMS (A) Blood serum of different diabetic stages AIM: The objective is to classify three distinct types of droplets with varying diabetic stages---healthy, low, and high diabetes by analyzing their drying patterns and employing AI methods.
[0093] Diabetes mellitus (DM) is related to the concentration of glucose in our bodies. There are two types---Type I and Type II. The difference between these two types lies in their causes, onset, insulin dependence, risk factors, prevention, and progression. Type I diabetes is an autoimmune condition where the immune system mistakenly attacks insulin-producing beta cells in the pancreas, resulting in little to no insulin production. In contrast, Type II diabetes occurs due to insulin resistance, where the body’s cells do not respond properly to insulin, leading to impaired glucose regulation over time. The onset of Type I diabetes is usually in childhood or adolescence. In contrast, Type II diabetes typically develops in adulthood, though it is becoming more prevalent in younger individuals due to lifestyle factors. People with Type I diabetes require lifelong insulin therapy because their bodies do not produce insulin. In contrast, Type II diabetes can often be managed through lifestyle changes, oral medications, and, in some cases, insulin therapy at later stages. The risk factors also differ---Genetic and autoimmune factors primarily influence Type I diabetes. In contrast, Type II diabetes is linked to genetics, obesity, poor diet, lack of physical activity, and other metabolic conditions. When it comes to prevention, Type I diabetes cannot be prevented as it results from an autoimmune response. In contrast, Type II diabetes can often be prevented or delayed through a healthy diet, regular physical activity, and weight management.
[0094] The study primarily focuses on Type II Diabetes mellitus (T2DM). Although diabetes does not have strict pathological stages, pathologically, the HbA1C levels are measured from blood. HbA1C, also known as glycated hemoglobin, is a key biomarker used to assess long-term blood glucose levels by measuring the percentage of hemoglobin bound to glucose over approximately three months. Higher HbA1C levels indicate poor glucose control, with values below 6.5% considered healthy, 7.5-9.5% indicating moderate diabetes risk, and levels above 14.5% suggesting severe diabetes, often associated with additional health complications.
[0095] Table 1 presents a structured dataset of patients categorized into "Healthy or Control," "Stage 1," and "Stage 2" groups based on their diabetes diagnosis and HbA1C levels. It includes various columns such as Patient ID, storage conditions (-80°C for all samples), age, height, weight, gender, and race. The diagnosis section specifies whether a patient is healthy or has diabetes, along with their HbA1C values.
[0096]
[0097] The table 1 displays patient data categorized by health stages, including demographics, diagnosis, HbA1C levels, medications, comorbidities, and viral test results of healthy and Type II Diabetes mellitus (T2DM).
[0098] Medications are listed, with healthy individuals having no medications, while diabetic patients in Stage 1 take oral anti-diabetic drugs like Metformin and Glimepiride, and those in Stage 2 rely on insulin therapy. The "Other diseases" column highlights additional conditions such as obesity, hypertension, and chronic gastritis in diabetic patients. Lastly, the table confirms that all patients tested negative for HIV -- Human Immunodeficiency Virus, HCV -- Hepatitis C Virus, HBV -- Hepatitis B Virus, and RPR -- Rapid Plasma Reagin (a test for syphilis).
[0099] The study utilizes serum samples from patients, with three samples per stage. These real serum samples are initially frozen and stored at -80°C before being at room temperature for droplet deposition. Optical imaging is employed to capture the progressive drying stages of these serum droplets. All experiments were conducted under ambient conditions, with a room temperature of ~25°C and a RH of ~35%. At the start of each experiment, the droplet’s contact angle measured approximately 50°. Each drying process was repeated twice to ensure consistency, while the dried droplet morphology was examined 15-20 times, with images captured at each instance. The fps used for these samples is 1 frame per second (fps), giving rise to ~300 images per sample.
[0100] Figure 16 shows the time-sequence images depicting the drying dynamics and final morphological structures of healthy serum droplets of Patient ID 66 (with HbA1C < 6.5%) at various stages of evaporation. The timestamps are indicated in black throughout the drying process.
[0101] Figure 17 shows the time-sequence images depicting the drying dynamics and final morphological structures of healthy serum droplets of Patient ID 79 (with HbA1C < 6.5%) at various stages of evaporation. The timestamps are indicated in black throughout the drying process.
[0102] Figure 18 shows the time-sequence images depicting the drying dynamics and final morphological structures of healthy serum droplets of Patient ID 29 (with HbA1C < 6.5%) at various stages of evaporation. The timestamps are indicated in black throughout the drying process.
[0103] Figures 16-18 show a sequential visualization of the drying process of healthy serum droplets, highlighting changes in morphology over time. The top row displays a time-lapse progression of a single serum droplet, starting from an initial spherical shape to the fluid front movement from the droplet edge to the central region and the propagation of cracks. The subsequent rows showcase multiple repetitions of the dried droplet morphology, highlighting the reproducibility of the pattern formation. These images exhibit well-defined peripheral cracks and a central region with distinct textural differences, characteristics of drying-induced stresses in biofluid droplets. The consistent formation of these structures across multiple healthy samples (Patient IDs 66, 79, and 29) highlights the robustness of the drying patterns, which can be analyzed for potential correlations with disease conditions or compositional variations in the serum. The study was conducted under controlled ambient conditions, ensuring that the observations were representative of the natural drying behavior of healthy serum droplets.
[0104] Figure 19 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from Stage 1 Type II Diabetes Mellitus (T2DM) for Patient ID 5 (with HbA1C = 7.7%), captured at various stages of evaporation. The timestamps are indicated in black throughout the drying process.
[0105] Figure 20 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from Stage 1 Type II Diabetes Mellitus (T2DM) for Patient ID 6 (with HbA1C = 7.7%), captured at various stages of evaporation. The timestamps are indicated in black throughout the drying process.
[0106] Figure 21 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from Stage 1 Type II Diabetes Mellitus (T2DM) for Patient ID 7 (with HbA1C = 9.5%), captured at various stages of evaporation. The timestamps are indicated in black throughout the drying process.
[0107] Figures 19-21 provide a sequential visualization of the drying process of serum droplets from Stage 1 Type II Diabetes Mellitus (T2DM) patients (Patient IDs 5, 6, and 7) with HbA1C levels ranging from 7.7% to 9.5%, captured at different evaporation stages. The top row illustrates the time-lapse progression of a single serum droplet, transitioning from its initial spherical shape to the inward movement of the fluid front from the droplet edge toward the center, followed by the formation and propagation of cracks. The subsequent rows display multiple repetitions of the dried droplet morphology, confirming the reproducibility of the observed patterns. A comparative analysis with healthy serum droplets (Figures 16-18) reveals a higher number of peripheral cracks in the samples. In contrast, the dried droplets in Figures 19-21 (Stage 1 T2DM) exhibit a more intact morphology with fewer fractures, suggesting lower mechanical stress accumulation during the drying process. Despite these differences, the characteristic "coffee-ring effect" remains consistently present in both healthy and diabetic samples, with solute accumulation at the periphery. However, the overall crack density is notably lower in Stage 1 T2DM droplets compared to their healthy counterparts.
[0108] Figure 22 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from Stage 2 Type II Diabetes Mellitus (T2DM) for Patient ID 4S (with HbA1C > 14.5%), captured at various stages of evaporation. The timestamps are indicated in red throughout the drying process.
[0109] Figure 23 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from Stage 2 Type II Diabetes Mellitus (T2DM) for Patient ID 7S (with HbA1C > 14.5%), captured at various stages of evaporation. The timestamps are indicated in red throughout the drying process.
[0110] Figure 24 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from Stage 2 Type II Diabetes Mellitus (T2DM) for Patient ID 8S (with HbA1C > 14.5%), captured at various stages of evaporation. The timestamps are indicated in red throughout the drying process.
[0111] Figures 22-24 present a sequential visualization of the drying process of serum droplets from Stage 2 Type II Diabetes Mellitus (T2DM) patients (Patient IDs 4S, 7S, and 8S) with HbA1C levels exceeding 14.5%, captured at different stages of evaporation. The top row illustrates the time-lapse progression of a single serum droplet, transitioning from its initial spherical shape to the inward movement of the fluid front from the droplet edge toward the center. Unlike the previous samples from healthy individuals and Stage 1 T2DM patients, the drying process in Stage 2 T2DM droplets exhibits no visible cracks. The subsequent rows display multiple repetitions of the dried droplet morphology, confirming the reproducibility of the observed patterns. Notably, these droplets maintain a uniform structure with almost no peripheral cracks or fractures, indicating a different drying behavior compared to earlier stages.
[0112] A comparative analysis of crack formation across different conditions reveals a clear trend. Healthy serum droplets (Figures 16-18) exhibited a high density of well-defined radial cracks, originating from the periphery and propagating inward. This extensive cracking suggests significant mechanical stress buildup during drying, likely due to solute aggregation at the droplet edges. In contrast, Stage 1 T2DM droplets (Figures 19-21) showed a lower number of cracks, with fractures appearing primarily at the edges rather than extending toward the center. While the coffee-ring effect was still present, these droplets demonstrated greater structural integrity than healthy samples, indicating a reduction in drying-induced stress. Stage 2 of T2DM droplets in Figures 22-24 stand in stark contrast to both healthy and Stage 1 samples, as they exhibit no crack formation and a distinctly different drying pattern. The absence of cracks and the coffee-ring effect suggest a significant shift in the serum properties, influencing the drying dynamics. The progressive reduction in crack density from healthy samples to Stage 1 of T2DM to Stage 2 of T2DM highlights a clear correlation between disease progression and the mechanical response of drying droplets. This comparison underscores the impact of serum composition on the drying process, affecting droplet central-edge height variations and leading to distinct final morphologies. These findings emphasize the potential diagnostic value of droplet analysis in detecting diabetes-related changes in biofluids and offer new insights into the physicochemical behavior of serum at different stages of the disease.
[0113] Figure 24 provides the AI result on classifying different stages of diabetes. Figure 24(I) showcases the ResNet-18 architecture used for feature extraction, highlighting its key layers, including convolutional layers (Conv2D), batch normalization, ReLU activations, pooling layers, and fully connected layers. The model is pre-trained on ImageNet and fine-tuned for classifying serum droplet images. Droplet images undergo preprocessing before being passed through the network, where feature maps are generated for classification. Figure 24(II) shows the feature extraction block; the training accuracy and loss curves provide insights into model performance over multiple epochs. The left graph displays the accuracy progression for both training and testing datasets, showing a consistent improvement as training progresses. The right graph presents the loss curves, indicating a steady decline in both training and validation loss, suggesting proper model convergence without significant overfitting. Figure 24(III) visualizes Grad-CAM heatmaps overlaid on the images of serum droplets, highlighting the most informative regions that contribute to classification decisions. The heatmaps for healthy droplets display distinct peripheral patterns with scattered activation zones, suggesting a high concentration of textural features in these regions. Stage 1 droplets exhibit reduced crack formations, and the heat maps indicate localized activation at the edges; however, they show significant leakage in this class. In contrast, Stage 2 droplets show more uniform and widespread activations, correlating with a distinct morphology characterized by minimal cracking and central region variations. Figure 24(IV) exhibits the normalized confusion matrix that summarizes the model’s classification performance across the three classes: Healthy, Stage 1, and Stage 2 of T2DM. The high diagonal values (close to 1) indicate strong classification accuracy, while the off-diagonal values remain minimal, demonstrating that the model effectively differentiates between the different serum droplet categories, with an average F1-score of 99%.
[0114] The comparison of healthy, Stage 1, and Stage 2 droplets through both classification accuracy and heatmap visualizations emphasizes the progressive alteration in drying patterns associated with disease progression. The study underscores the potential of deep learning-assisted droplet analysis as a non-invasive diagnostic tool for assessing diabetes-related changes in the serum, highlighting image-based feature extraction and classification to distinguish between different diseased stages.
[0115] Figure 25(I) shows a visual representation of a convolutional neural network (CNN) utilizing the ResNet-18 architecture, enhanced by Gradient-weighted Class Activation Mapping (Grad-CAM), to provide insights into the model’s feature extraction process. (II) training and testing curves for loss and accuracy. (III) Grad-CAM snapshots overlaid on the original images, capturing key regions of interest throughout the drying process. (IV) Normalized classification matrix of predicting different groups of Type II Diabetes Mellitus(T2DM)--- Healthy, Stage 1, and Stage 2.
[0116] Figure 26 shows the normalized confusion matrix illustrating the performance of various traditional ML: (I) Naive Bayes (NB), (II) Support Vector Machine (SVM), (III) K-Nearest Neighbors (KNN), (IV) Decision Tree (DT), and (V) Random Forest (RF). These MLs are employed to classify healthy Stage 1 and Stage 2 of Type II Diabetes Mellitus (T2DM). (VI) The radial plot depicting the average F1-score in percentage provides a comparative assessment of each ML’s performance.
[0117] Figure 25(I-V) showcases the performance evaluation of various traditional ML models employed to classify health conditions and different stages of T2DM, specifically distinguishing between healthy individuals, Stage 1, and Stage 2 patients. The evaluation is depicted through normalized confusion matrices for five distinct ML classifiers: (I) Naive Bayes (NB), (II) Support Vector Machine (SVM), (III) K-Nearest Neighbors (KNN), (IV) Decision Tree (DT), and (V) Random Forest (RF). Additionally, a radial plot in Figure 25(VI) provides a summary of the average F1-scores for these models, offering a comparative visualization of their overall performance.
[0118] The confusion matrix in the NB classifier reveals a relatively lower performance compared to the other models. The values along the diagonal, representing correct classifications, are not consistently close to 1, indicating notable misclassifications. This is reflected in the off-diagonal values, suggesting that instances from different classes are often confused with one another. The normalized values indicate that NB struggles to distinguish between the health stages effectively. The SVM classifier displays a notable improvement, with higher values along the diagonal, particularly for Stage 1 and Stage 2 classes. The healthy class, however, exhibits some degree of misclassification. Interestingly, KNN, DT, and RF classifiers reveal minor misclassification errors, with the diagonal values equal to 1 in most cases, signifying perfect classification for the majority of samples. The off-diagonal values are near zero, indicating minimal confusion between classes. The F1-scores in Figure 25(VI) give the hierarchy of the performance as RF= DT= KNN ~ 99%, SVM with 97%, while NB lags significantly behind with an F1-score of 58%.
[0119] The study highlights the potential of traditional ML methods as an alternative to deep learning-assisted droplet analysis, with both approaches proving equally effective as non-invasive diagnostic tools for detecting diabetes-related changes in serum, achieving an accuracy of 99%.
[0120] (B) Blood serum of cancer AIM: The objective is to classify types of droplets with healthy serum and lung cancer samples by analyzing their drying patterns and employing AI methods.
[0121] Lung cancer is one of the most common and life-threatening malignancies worldwide, characterized by the abnormal and uncontrolled growth of cells in the lungs. It is broadly classified into two main types: non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC). NSCLC is the more prevalent type, accounting for approximately 85% of lung cancer cases, while SCLC is more aggressive and tends to spread rapidly. The causes of lung cancer are multifaceted, with smoking being the primary risk factor; however, these samples are taken from non-smokers.
[0122] The progression of lung cancer is typically described using the Tumor, Node, Metastasis (TNM) staging system. This system evaluates the size and extent of the primary tumor (T), lymph node involvement (N), and the presence of metastasis (M). Early-stage lung cancer (Stage I) is often confined to the lungs and has a better prognosis when treated promptly. As the disease advances to Stages II, III, and IV, the tumor spreads to nearby lymph nodes and other organs, significantly reducing survival rates. Histological grading also plays a vital role in lung cancer diagnosis, indicating the degree of tumor differentiation. Low-grade tumors (G1) are well-differentiated and tend to grow slowly, whereas high-grade tumors (G3) are poorly differentiated and exhibit aggressive behavior with a higher likelihood of metastasis.
[0123] Table 2 displays patient data categorized by health stages, including diagnosis, stages, tumor levels, medications, and comorbidities of lung cancer patients with Caucasian ethnicity and who are non-smokers. Table 2 shows the details of the patients. The cases described with G3 T1bN0M0 IA2, G3 T2aN0M0 IB, and G1 T1bN0M0 IA2 all fall under the category of early-stage lung cancer (Stage I), which is typically characterized by a localized tumor without lymph node involvement or distant metastasis. Specifically, Stage IA2 (T1bN0M0) represents a small tumor (1 - 2 cm) confined to the lung, while Stage IB (T2aN0M0) indicates a slightly larger tumor (3- 4 cm), still localized and without nodal spread. Grade G1 suggests well-differentiated, slower-growing cancer cells, whereas G3 represents poorly differentiated, more aggressive cells. Despite the differences in grade and tumor size, all these cases are considered early-stage lung cancer.
[0124]
[0125] The study utilizes serum samples from early-stage lung cancer patients. These real serum samples are initially frozen and stored at -80°C before being at room temperature for droplet deposition. Optical imaging is employed to capture the progressive drying stages of these serum droplets. All experiments were conducted under ambient conditions, with a room temperature of ~18°C and a RH of ~25%. At the start of each experiment, the droplet’s contact angle measured approximately 55°. Each drying process was repeated twice to ensure consistency, while the dried droplet morphology was examined 8-10 times, with images captured at each instance. The fps used for these samples is 1 frame per second (fps), giving rise to ~300 images per sample.
[0126] It is crucial to evaluate whether the serum drying pattern changes significantly when room temperature and relative humidity. Even subtle variations in these parameters can influence the dynamics of droplet evaporation, fluid flow, and final deposition patterns. Specifically, a decrease in room temperature by 10°C from the usual 25°C, coupled with a 10% variation in RH, may alter the solvent evaporation rate, surface tension gradients, and the formation of cracks patterns during serum droplet drying. Understanding these effects is vital to ensure the reproducibility and robustness of the diagnostic patterns, particularly when transitioning from controlled laboratory settings to real-world applications where temperature and humidity may fluctuate.
[0127] Figure 27 shows the time evolution images of a drying droplet of healthy human serum (Patient ID 29) captured under constant experimental conditions (droplet size, surface, volume). The top and bottom rows represent the droplet drying under room temperatures of 18°C and 25°C with relative humidity (RH) of 25% and 35%, respectively. The timestamps in red indicate the drying time in seconds.
[0128] Figure 27 presents the time evolution images of a drying healthy serum droplet from Patient ID 29 under two distinct environmental conditions. While variations in temperature and relative humidity can influence the onset and progression of crack formation, the overall drying behavior and final pattern remain largely consistent. In both conditions, the radial cracks predominantly emerge along the droplet’s periphery, following a comparable sequence of morphological stages. Although slight shifts in the timing of crack initiation and propagation may occur, these fluctuations do not significantly alter the characteristic pattern for serum samples. This suggests that while maintaining consistent environmental conditions is generally important, deviations within a range of 5-10 units (°C for temperature or % for RH) appear to have minimal impact on the reproducibility of serum droplet patterns.
[0129] Figure 28 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from lung cancer for Patient ID (01LC), captured at various stages of evaporation. The timestamps are indicated in red throughout the drying process.
[0130] Figure 29 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from lung cancer for Patient ID (10LC), captured at various stages of evaporation. The timestamps are indicated in red throughout the drying process.
[0131] Figure 30 shows the time-sequence images illustrating the drying dynamics and final morphological structures of serum droplets from lung cancer for Patient ID (15LC), captured at various stages of evaporation. The timestamps are indicated in red throughout the drying process.
[0132] Figures 28-30 provide a sequential visualization of the drying process of serum droplets from early-stage cancer patients (Patient IDs 01LC, 10LC, and 15LC) captured at different evaporation stages. The top row illustrates the time-lapse progression of a single serum droplet, transitioning from its initial spherical shape to the inward movement of the fluid front from the droplet edge toward the center, followed by the formation and propagation of cracks. The central regions of these cancer-affected serum droplets exhibit some dense, textured networks of irregular, web-like structures dispersed with small granular formations. The characteristic "coffee-ring effect" remains consistently present in all these samples, with solute accumulation at the periphery.
[0133] Figure 31 shows the normalized confusion matrix illustrating the performance of various traditional ML: (I) Naive Bayes (NB), (II) Support Vector Machine (SVM), (III) Decision Tree (DT), (IV) K-Nearest Neighbors (KNN), and (V) Random Forest (RF). These MLs are employed to classify healthy people and patients suffering from the early stages of lung cancer. (VI) The radial plot depicting the average F1-score in percentage provides a comparative assessment of each ML’s performance.
[0134] Figure 31 evaluates the performance of traditional machine learning models---Naive Bayes (NB), Support Vector Machine (SVM), Decision Tree (DT), K-Nearest Neighbors (KNN), and Random Forest (RF)---in classifying serum drying patterns from healthy individuals and early-stage lung cancer patients. The confusion matrices reveal that NB performs the weakest (F1-score: 61.2%), while SVM (99.3%), DT (99.8%), KNN, and RF (all ~100%) demonstrate near-perfect classification accuracy. The radial plot summarizes this comparison, emphasizing that SVM, DT, KNN, and RF are the most robust models, highlighting the high reliability of serum drying patterns analyzed through machine learning as a non-invasive diagnostic tool for early lung cancer detection.
[0135] RESULTS AND DISCUSSIONS ON THE VISUAL DIFFERENCES IN DRYING PATTERNS Figure 32 shows the dried droplet patterns containing different initial concentrations of lysozyme protein derived from chicken (in wt% along the x-axis) and different amounts of PBS (in dilution factor---0x, 0.25x, 0.5x, and 1x along the y-axis).
[0136] Figure 32 displays a morphological grid of the samples varying the initial concentrations of lysozyme in wt% (along the Y-axis) and the initial concentrations of the buffer from 1x to 0.25x (along the X-axis). The 0x embodies the lysozyme solution prepared in the de-ionized water. Though all these deposits show the “coffee-ring” effect, diverse patterns are observed for each concentration. The lysozyme films show a mound-like structure when the solution is prepared without external salts. A dimple (or depression) is also noticed within this mound. The mound area gets wider as the protein concentration increases. The random cracks are only observed in the peripheral ring; however, these cracks spread throughout the film as the protein concentration increases. The radial and orthoradial cracks promote well-connected (small and large) domains in these droplets. The presence of multiple rings is also observed in the highly concentrated lysozyme samples. A unique trend is also noticed when the protein concentration is fixed, and buffer concentration is varied (along any columns of Figure 32). The central region becomes grainy, the texture becomes darker, and some thread-like structures appear in the central region. Interestingly, a few lysozyme droplets follow different zones while moving from the droplet’s periphery to the central region. It undergoes phase separation and forms different material properties; for instance, the lysozyme concentration is highest in the peripheral ring and systematically droplets towards the central region. On the other hand, the salt concentration is almost null in the periphery but highest as we move towards the central regions. Therefore, the bio-mimetic samples containing lysozyme protein (derived from chicken) and different buffer amounts used as the drying droplets represent a multi-component and complex system in a mimetic sense, where unique pattern morphology emerges.
[0137] Figure 33 shows the time evolution of the biomimetic samples containing bovine serum albumin (derived from cows’ blood) and different buffer amounts used as the drying droplets. These samples represent a multi-component and complex system in a mimetic sense, where unique pattern morphology emerges. In the initial stages, the droplets spread uniformly, exhibiting a smooth and homogeneous appearance. As drying progresses, cracks begin to appear, primarily at the droplet periphery, indicating mechanical stress buildup due to the differential evaporation rate and protein aggregation. The propagation of these cracks varies between samples, highlighting the influence of buffer composition on the drying dynamics. In the later stages, complex crack networks and heterogeneous internal textures emerge. The fine dendrite-like structures in the central regions of these droplets reveal a granular and porous appearance, reflecting protein aggregation and buffer crystallization interactions, which are common in multi-component drying systems.
[0138] Figure 33 shows the drying droplet patterns containing fixed initial concentrations of bovine serum albumin (BSA) protein derived from cows and different amounts of PBS (in dilution factor, 1x, 2x, and 5x along the y-axis).
[0139] Figures 32-33 demonstrate the biomimetic samples showing the unique final pattern morphology. The diversity in crack propagation, central region textures, and peripheral structures underscores the intricate interplay between the compositional changes in the biological systems influenced by drying. Such patterns can serve as reference models for interpreting patterns from blood and other biofluids in diagnostic applications.
[0140] Figure 34 compares the drying patterns of a healthy person with a matched set of different bio-fluids, such as blood, serum, and urine, dried under a constant temperature and relative humidity.
[0141] Figure 34 illustrates the time evolution and final drying patterns of droplets from different biofluid (blood, serum, and urine) samples. These droplets undergo evaporation under controlled conditions, revealing varying morphologies of different bio-fluids taken from the same person. The blood droplet displays the drying progression of a droplet exhibiting an irregular and highly branched crack pattern. Initially, the droplet appears dark and uniformly distributed. As drying advances, the central region begins to lighten, and local aggregation emerges, signaling internal stresses. Cracks start to propagate radially from the center and periphery, progressively forming an intricate network of interconnected branches. The final dried pattern exhibits a striking spiderweb-like structure, with thick, radial cracks segmenting the droplet into irregular compartments.
[0142] The serum showcases a droplet with a distinctly different drying behavior, characterized by the formation of well-defined radial cracks originating from the periphery. The initial spreading stage reveals a smooth and symmetrical droplet with uniform thinning as evaporation progresses. Over time, the mechanical stress accumulates along the edges, leading to the emergence of clean, evenly spaced radial cracks extending toward the center. The crack uniformity and absence of chaotic internal textures suggest a more stable and homogeneous molecular composition. In contrast, the urine represents a droplet that dries with no visible crack formation. As drying progresses, the droplet’s central region becomes inhomogeneous. All droplets show a coffee-ring effect; however, the crack patterns of these bio-fluids are unique. This distinct approach highlights the integration of AI with drying patterns as a simple yet powerful bio-physical methodology capable of extending beyond blood to a diverse range of biofluids, including urine, tears, saliva, serum, plasma, and others.
[0143] Figure 35 compares the drying patterns of a healthy serum extracted from a person and a mouse, dried under a constant temperature and relative humidity.
[0144] Figure 35 compares the drying patterns of serum droplets obtained from a healthy human and a healthy mouse, both dried under identical environmental conditions, with temperature and relative humidity carefully controlled. The images capture the progression from the initial liquid droplet to the final dried pattern, revealing striking similarities in the initial and middle stages (i.e., the change of texture, the fluid front movement, etc.), overall crack morphology, and structural features between human and mouse serum samples. In both cases, well-defined radial cracks emerge from the periphery and extend towards the center. The inner region of the dried droplets also exhibits a comparable texture in both human and mouse samples, and the presence of the coffee-ring effect further emphasizes the resemblance in their drying behavior.
[0145] Although minor differences are observed in the timing of crack formation and subtle variations in the thickness and spacing of the cracks, these discrepancies are within an acceptable range and do not alter the fundamental pattern characteristics. This similarity holds significant implications for future validation and data collection efforts. Human serum samples are often challenging to obtain, especially from patients with specific diseases, and are subject to ethical regulations and variability due to underlying health conditions, lifestyle, or medications. In contrast, mouse models offer a more controlled environment, free from such confounding factors, allowing for the collection of large datasets under standardized conditions. The results suggest that mouse serum samples can serve as a viable substitute for human samples during the initial phases of developing image-based diagnostic tools, such as a point-of-care (POC) device powered by AI. Mouse models enable researchers to establish a baseline corpus of serum drying images, covering a wide range of initial conditions without the complexities associated with human variability. This approach would facilitate the creation of robust image training datasets, improving the accuracy and generalizability of machine-learning algorithms used in diagnostic applications. Once the model is trained and validated with this extensive corpus, human samples can be incorporated later for fine-tuning and clinical validation. Thus, the ability to rely on mouse serum for pattern reproducibility accelerates the development of image-based diagnostic platforms, rationalizing the path toward deploying AI-assisted POC devices for early disease screening.
[0146] Figure 36 displays the time evolution images that capture the drying process of serum droplets from four health conditions: healthy, diabetic Stage 1, diabetic Stage 2, and lung cancer.
[0147] Figure 36 presents a comparative time evolution sequence showing the drying behavior of serum droplets from individuals representing three different health states: a healthy individual, patients at different stages of Type II Diabetes Mellitus (T2DM), and a lung cancer patient. The initial and middle stages of the drying process look very similar, and the last stage is where the fluid front stops movement and / or the crack starts forming. A comparative analysis with healthy serum droplets reveals a higher number of peripheral cracks in the samples and less inhomogeneous structures in the central regions. In contrast, the droplets in Stage 1 of T2DM exhibit fewer fractures. However, the overall crack density is notably lower in Stage 1 T2DM droplets compared to their healthy counterparts. Finally, no crack formation occurs in Stage 2 T2DM droplets. In contrast, lung cancer samples are characterized by many well-defined radial and orthoradial cracks, with additional heterogeneous textures, hair-like cracks, aggregation, etc., developing towards the central region of the droplet.
[0148] Figure 37 shows the images capturing the dried droplets of serum collected from three different patients with four different health conditions: healthy, diabetic Stage 1, diabetic Stage 2, and lung cancer.
[0149] Figure 37 displays the dried droplet patterns of serum collected from patients representing four distinct health conditions: healthy, diabetic Stage 1, diabetic Stage 2, and lung cancer. The dataset highlights the reproducibility of characteristic pattern formation associated with each health condition, as the images represent samples from three different patients within each category. Notably, the distinct visual signatures are consistently observed---well-defined radial cracks in healthy samples, sparse and irregular cracks in Stage 1 diabetes, smooth, featureless surfaces in Stage 2 diabetes, and dense, complex crack networks in lung cancer. This consistency across multiple droplets emphasizes the reliability of the drying droplet approach as a non-invasive diagnostic tool. The patterns visually capture the biochemical and biophysical variations specific to each disease state, and when integrated with AI-based analysis, the classification accuracy exceeds 95%, further validating the method’s diagnostic potential.
[0150] The embodiments of this invention are also characterized by the following statements: (1). A rapid, accurate, and simple bio-physical methodology integrated with microscopy and a data-driven approach capable of diagnostic technology to classify biofluids.
[0151] (2). A simple bio-physical methodology of (1), wherein the circular droplets of the volume of approximately 1 μL and approximately 2 mm diameter are deposited on the solid surface, and the liquid (water) evaporates with time within the droplets, referred to as drying droplets.
[0152] (3). A bio-physical methodology of (2) takes only 5-10 mins (rapid) in which drying occurs under a fixed environmental condition (temperature (T) of 18-28 °C and relative humidity (RH) of 25-55%) and the solid surface in (2) has an average contact angle with the droplet of 30-50 degrees.
[0153] (4). A microscopy of (1) comprising: - a magnification lens ranging from 2x (minimum) to 5x (maximum) for capturing the entire 2 mm diameter droplet; - a camera with a larger sensor area is attached to the microscopy to facilitate capturing images at a given frame per second; - a computer with camera software is installed, wherein the camera is configured to save captured images on a computer; - to ensure high-resolution images, the images should not be compressed or filtered during the initial clicking of the images by the camera; however, the minimum requirement is clicking one frame in two seconds repeatedly; - microscopy under different configurations can be used, such as bright-field and crossed-polarizing configurations, depending on the biofluids of (1).
[0154] (5). The bio-physical methodology of (3) includes information on the dynamics of the pattern formation, which are stored as the images described in (4).
[0155] (6). The patterns in (5) uniquely emerge during the drying process depending on the inherent nature of the biofluids, encompassing composition (added salts, glucose, or medicines), concentrations, and patient-specific attributes (medication history, stages of a disease or disorder).
[0156] (6-2). This bio-physical methodology of (1) can be implemented over the wide ranges of temperature (T) = 10-55°C and relative humidity (RH) of 10-65%, where the pattern dynamics mentioned in (6) depend on the specific T and RH at which these droplets will be dried.
[0157] (7). The emerging patterns in (5) or (6) comprise a spectrum of features during drying, encompassing cracks of various types (radial, orthoradial, chaotic, and so forth) and their propagation dynamics (faster or slower), domains with differing crack sizes (larger or smaller), inhomogeneities, or uniform textures without any cracks.
[0158] (8). The drying evolution of the patterns in (5) shows that: - the sequential order of these emerging patterns captured in time-dependent images is crucial; - the initial stage is within a few minutes (after droplet deposition) and is characterized by a fluid-like behavior, and as water evaporates, the droplet progresses to the final stage (near the end of drying), exhibiting a solid-like behavior; - the significance of the entire drying process (not just the dried morphology), usable for implementing a data-driven approach in (1); - the complete drying process (not just the dried morphology) serves as a unique fingerprint for a pattern-recognized tool.
[0159] (9). The data-driven approach of (1), wherein the predictive model is based on different supervised machine learning algorithms (MLAs).
[0160] (10). The MLAs of (9), wherein the predictive model is a classification model.
[0161] (11). The classification model of MLAs of (10) includes two different approaches to the drying evolution of the pattern formation, which comprise: - Approach 1: a neural network-based method, specifically a convolutional neural network (CNN); - Approach 2: traditional MLAs include Decision Trees (DT), Random Forests (RF), and Support Vector Machines (SVM), Naive Bayes (NB), K-Nearest Neighbors (KNN), and Gradient Boosting (XGB).
[0162] (12). The input of CNN in (11) is based on the acquired images, which act as a feature vector, whereas the input of the traditional MLAs in (11) is based on the data, which act as a feature vector, extracted from the acquired images during the drying process.
[0163] (12-2). The input of CNN in (11) is based on the acquired images, but it can be converted into one (1D) or two dimensions (2D) depending on the nature of the feature vector.
[0164] (12-3). The neural network of transfer learning models in Approach 1 of (11) can be the base model loaded with its pre-trained weights from the ImageNet dataset, with all its layers frozen to retain the learned features, and only the final classification layers should be trained to adapt the model to the specific dataset.
[0165] (13). The data for the traditional MLAs in (11) is based on quantifying the images during drying through texture analysis, which comprises either First Order Statistics (FOS), Gray Level Co-occurrence Matrix (GLCM), Gray Level Size Zone Matrix (GLSZM), Gray Level Run Length Matrix (GLRLM), Gray Level Dependence Matrix (GLDM), or the combinations of any of these, depending on the biofluids of (1).
[0166] (13-2). The evaluation of traditional MLAs of (10) includes confusion matrix, accuracy, F1-score, Precision, and Recall. In addition to these, the training and validation / testing curves are necessary for CNN of (12).
[0167] (14). The classification model of MLAs in (11) requires an input size of 300 (minimum) images or extracted data for each droplet, ensuring optimal performance in analyzing the drying evolution of pattern formation during the bio-physical methodology.
[0168] (15). The biofluids of (1) include bio-mimetic and real samples, where pattern dynamics during drying are utilized in MLAs to predict multi-class identification, demonstrating the robustness of this bio-physical methodology coupled with microscopy and a data-driven approach.
[0169] (15-2). The evaluation metrics in (13-2) indicate that the classification of samples mentioned in (15) achieves a correct prediction rate of 90-99%.
[0170] (16). The bio-mimetic samples of (15) used as the drying droplets represent a multi-component and complex system in a mimetic sense, where unique pattern morphology emerges in the drying droplets: - containing globular proteins (lysozyme protein derived from chicken), optically active liquid crystals (thermotropic 5CB), and different amounts of phosphate saline buffer (PBS); - containing globular proteins (lysozyme protein derived from chicken) and different amounts of PBS. - containing globular proteins (bovine serum albumin protein derived from the blood of cows) and different amounts of PBS.
[0171] (17). The real biofluids of (15) used as the drying droplets are derived from humans and represent a naturally occurring system where unique pattern morphology emerges in the drying droplets: - containing the whole blood consisting of cellular components (red blood cells (RBCs), white blood cells (WBCs), and platelets), proteins, ions, and coagulant agents; - containing the whole blood fluid, excluding the described cellular components, referred to as plasma; - containing plasma fluid, excluding the coagulant agents, referred to as serum.
[0172] (18). The future diagnostic technology in (1), wherein the method is used for classification by obtaining images of the biofluid droplets that evolve a unique pattern morphology over a period of drying time and using MLAs to classify: - healthy and unhealthy samples in a bio-mimetic sense by comparing whole blood versus diluted blood prepared by adding different volumes of de-ionized water and PBS. - different abnormality stages are observed by comparing the whole blood with added volumes of de-ionized water and PBS.
[0173] (19). The diagnostic technology in (1), wherein the healthy and unhealthy real biofluids are classified based on the drying patterns.
[0174] (20). The unhealthy samples in (19) involve those obtained from patients with type 2 diabetes mellitus (T2DM), specifically serum (as detailed in (17)), showcasing distinct patterns during the drying process.
[0175] (21). The patterns in (15) are mostly radial cracks for the healthy serum sample, whereas the number of cracks decreases for the patients suffering from T2DM with HbA1C values (glucose levels measured pathologically in an average of three months) ranging from 7.7% to 9.5%.
[0176] (22). The patterns in (15) are mostly the radial and orthoradial cracks with smaller crack domains for the healthy serum sample, whereas larger crack domains appear for patients suffering from T2DM with HbA1C values ranging from 7.7% to 9.5%.
[0177] (23). The patterns in (15) are mostly radial and orthoradial cracks for the healthy serum sample, whereas no cracks appear for patients suffering from T2DM when HbA1C values increase above 14.5%.
[0178] (24). The patterns in (21), (22), or (23) are only valid when the droplets are circular and dried under similar conditions to those shown in (3).
[0179] (25). The real biofluids described in (17), specifically whole blood and plasma, contain anti-coagulants, and the patterns in (21), (22), or (23) are uniquely valid for the serum under these conditions.
[0180] (26). The real biofluids of plasma and serum of (17) are frozen and stored at -80 °C, then brought to the temperature specified in (3) before the droplet deposition, then initiated the drying process, and the unique drying patterns in (21), (22), or (23) emerge.
[0181] (27). The simple bio-physical methodology of (1) can include a range of biofluids from urine, tears, saliva, blood, serum, and plasma. extracted from humans or animals.
[0182] (28). The simple bio-physical methodology described in (1) of (15) can also be used to predict the healthy vs. unhealthy, the type of disease (diabetes and cancer), the progression of the disease (for example, stages of diabetes), and the patterns will be reproducible whether the fluid is collected from the human or animal (for e.g., healthy human serum vs. healthy mouse serum).
Claims
1. A method for biofluids classification executed by a classification device comprising a processor, the method causing the processor to: a) acquire at least one microscopy image of biofluid deposited on a predetermined solid surface, b) execute a predetermined classification processing operation based on the acquired microscopy image; and c) output a classification result corresponding to the executed classification processing.
2. The method for biofluids classification according to claim 1, the processor is further configured to: a) repeatedly acquire microscopy images of the biofluid during an evaporation process; b) execute the predetermined classification processing using the multiple acquired images; and c) output a classification result based on the classification processing.
3. The method for biofluids classification according to claim 2, wherein the microscopy images acquired by the processor are microscopy images captured while drying the biofluids under controlled temperature and humidity conditions, wherein the controlled temperature and humidity conditions are regulated to maintain substantially constant temperature and humidity during the drying process.
4. The method for biofluids classification according to claim 3, said conditions accommodate fluctuations within a range of ±5-10°C and ±5-10% relative humidity without affecting the processing method.
5. The method for biofluids classification according to claim 1, the microscopy images are captured with: (a) a magnification lens ranging from 2x to 5x for capturing the entire approximately 2 mm diameter droplet; (b) a camera is attached to the microscopy to facilitate capturing images at a given frame per second; and (c) the processer executes camera software installed, wherein the camera is configured to save captured images on a memory; and (d) the images are raw images; and captured every two seconds repeatedly;6. The method for biofluids classification according to claim 1, wherein the predetermined classification processing is done with a machine learning model.
7. The method for biofluids classification according to claim 6, the machine learning model includes one of: (a) a neural network-based method, specifically a convolutional neural network (CNN); or (b) a classifier.
8. The method for biofluids classification according to claim 7, the classifier includes one of: Decision Trees (DT), Random Forests (RF), Support Vector Machines (SVM), Naive Bayes (NB), or K-Nearest Neighbors (KNN).
9. The method for biofluids classification according to claim 7, the classifier uses texture analysis which comprises either First Order Statistics (FOS), Gray Level Co-occurrence Matrix (GLCM), Gray Level Size Zone Matrix (GLSZM), Gray Level Run Length Matrix (GLRLM), Gray Level Dependence Matrix (GLDM), or the combinations of any of these.
10. The method for biofluids classification according to claim 1, the biofluid includes one of urine, tears, saliva, blood, serum, or plasma.
11. A device for biofluid classification including a processor, and connected to a microscope which outputs captured image data; wherein the processor is configured to acquire a microscopy image of biofluids deposited on a predetermined solid surface from the microscope, execute predetermined classification process using the captured image data, and output the classification results of said classification process.