Method for measuring subvisible particles in a test sample
An automated particle classification method using deep learning and image processing enhances accuracy and efficiency in identifying sub-visible particles in biological pharmaceuticals by employing a CNN and decision tree, achieving 88% accuracy and enabling model improvement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ARES TRADING SA
- Filing Date
- 2024-05-17
- Publication Date
- 2026-05-29
AI Technical Summary
Existing methods for identifying and classifying sub-visible particles in biological pharmaceuticals are laborious, prone to errors, and time-consuming, necessitating improved accuracy and efficiency in particle classification.
An automated particle classification process using a hybrid technology combining deep learning and image processing, employing a convolutional neural network (CNN) for feature extraction and a machine learning-based classifier, with a decision tree for classification, and a user interface for correction and retraining.
Achieves 88% accuracy in particle classification with the ability to correct misclassifications and improve model performance through a retraining pipeline.
Smart Images

Figure 2026517394000009 
Figure 2026517394000010 
Figure 2026517394000011
Abstract
Description
Technical Field
[0001] The present invention is directed to a method for measuring sub-visible particles in a test sample, such as a sample from a manufacturing process for producing a biological agent.
Background Art
[0002] When producing a biological pharmaceutical, the process includes a step-by-step process including isolation and purification of the biological pharmaceutical (active ingredient). In most such processes, the active ingredient is provided in solution. The active ingredient is further prepared to remain as a stable formulation (drug product) containing the active ingredient. The drug product is also often a prescribed solution. These solutions containing the active ingredient or drug product need to comply with specific quality standards, such as not containing a specific amount of particles, including sub-visible particles. For example, these particles may be proteinaceous substances or agglomerates of fibers.
[0003] Since health authorities require accuracy and consistency in particle classification, particle classification is a highly important topic. Particle classification is a long and laborious process when done manually. Furthermore, when done manually, errors are more likely to occur when identifying such particles. Therefore, it is necessary to improve the accuracy in identifying the presence of particulate matter, including sub-visible particulate matter, in a biological sample, and it is necessary to reduce the time for analyzing such biological samples in the manufacturing process for producing a biological pharmaceutical.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This invention provides a solution to improve the accuracy and time wasted in the analysis of biological samples regarding the presence of particulate matter. An automated particle classification process based on a hybrid technology using deep learning and image processing is provided by this invention. The method of this invention includes the extraction of morphological features from images using a convolutional neural network (CNN), e.g., Inception-V3, based on transfer learning. These additional features have a significant impact on model performance. Morphological features identified by a domain expert in conjunction with features extracted from the deep learning model are input into a machine learning-based classifier (decision tree). Particles can be classified into their respective classes with 88% accuracy. Through a user interface (UI), the user is given clauses to modify or correct any misclassifications made by the model. Once a significant amount of data has been collected, a retraining pipeline can be initiated to improve model performance. [Means for solving the problem]
[0005] In one embodiment, the method is a computer-implemented method for measuring subvisible particles in a sample, and the following: a. For example, one or more data files are obtained, such as raw data format and / or images of the sample from the analytical instrument, where the data files represent image information from the analytical sample. b. Extract morphological feature information from images from data files and / or analysis samples, c. Classify each image of the analysis sample using a trained convolutional neural network within the main categories. d. Cluster images within the main category into subclusters based on image similarity feature information from sample images. e. Apply statistical analysis comparing within subclusters of major categories to morphological features and / or image similarity to obtain statistical values for a given sample, and then, f. Evaluate statistical values to determine the distribution of subvisible particles in the sample. The present invention provides the method including the above.
[0006] In another embodiment, the method is a computer implementation method for batch release based on subvisible particles in a sample from a batch, the following: a. For example, one or more data files are obtained, such as raw data format and / or images of the sample from the analytical instrument, where the data files represent image information from the analytical sample images. b. Extract morphological feature information from images from data files and / or analysis samples, c. Classify each image of the analysis sample using a trained convolutional neural network within the main categories. d. Cluster images within the main category into subclusters based on morphological feature information from sample images. e. Apply statistical analysis comparing within subclusters of major categories to morphological features and / or image similarity to obtain statistical values for a given sample. f. Compare the statistical values for a given sample with those for a reference sample, and then... g. If the sample falls within a specified confidence level when compared to a reference sample, release the batch from which the sample was obtained. The present invention provides the method including the above.
[0007] In yet another embodiment, a non-transient computer-readable medium is provided that includes machine-readable instructions arranged to cause one or more processors to perform the method of the invention when executed by one or more processors. [Brief explanation of the drawing]
[0008] [Figure 1] Morphological conditions for filtering images into a different class
[0009] [Figure 2] Decision tree architecture
[0010] [Figure 3] Retraining Pipeline
[0011] [Figure 4] Application Hosting in AWS
[0012] [Figure 5] (CNN-based) Model Performance Regarding Transfer Learning
[0013] [Figure 6] Model Performance Regarding Ensemble Learning
[0014] [Figure 7] (CNN-based) Model Performance Regarding Transfer Learning Using Deep Learning Features
[0015] [Figure 8] Model Performance Regarding Random Forest
[0016] [Figure 9] Model Performance Regarding AdaBoost
[0017] [Figure 10] Model Performance Regarding K-Nearest Neighbor
[0018] [Figure 11] Model Performance Regarding Decision Tree Classifier
[0019] [Figure 12] Comparative Study between Two Experiments
[0020] [Figure 13-1] Graphical Display of Descriptive Statistics Regarding Air and Oil
[0021] [Figure 13-2] Graphical Display of Descriptive Statistics Regarding Air and Oil
[0022] [Figure 14-1] Comparison of the distribution of morphological features related to air and oil.
[0023] [Figure 14-2] Comparison of the distribution of morphological features related to air and oil.
[0024] [Figure 15] Distribution of circularity to aspect ratio for all subvisible particles
[0025] [Figure 16] Visual representation of clusters from both experiments
[0026] [Figure 17-1] Number of samples present in each cluster
[0027] [Figure 17-2] Number of samples present in each cluster [Modes for carrying out the invention]
[0028] Detailed description of the present invention The present invention provides solutions and improvements to existing methods for measuring the presence of particulate matter in a sample analyzed using imaging techniques. Specifically, it is for samples obtained from manufacturing processes for the production of biopharmaceuticals, for example, containing active pharmaceutical ingredients or drug products. The method provided herein is a computer-implemented method.
[0029] In one embodiment, the method is a computer-implemented method for measuring subvisible particles in a sample, and the following: a. For example, one or more data files are obtained, such as raw data format and / or images of the sample from the analytical instrument, where the data files represent image information from the analytical sample. b. Extract morphological feature information from images from data files and / or analysis samples, c. Classify each image of the analysis sample using a trained convolutional neural network within the main categories. d. Cluster images within the main category into subclusters based on image similarity feature information from sample images. e. Apply statistical analysis comparing within subclusters of major categories to morphological features and / or image similarity to obtain statistical values for a given sample, and then, f. Evaluate statistical values to determine the distribution of subvisible particles in the sample. The present invention provides the method including the above.
[0030] In another embodiment, the method is a computer implementation method for batch release based on subvisible particles in a sample from a batch, the following: a. For example, one or more data files are obtained, such as raw data format and / or images of the sample from the analytical instrument, where the data files represent image information from the analytical sample images. b. Extract morphological feature information from images from data files and / or analysis samples, c. Classify each image of the analysis sample using a trained convolutional neural network within the main categories. d. Cluster images within the main category into subclusters based on morphological feature information from sample images. e. Apply statistical analysis comparing within subclusters of major categories to morphological features and / or image similarity to obtain statistical values for a given sample. f. Compare the statistical values for a given sample with those for a reference sample, and then... g. If the sample falls within a specified confidence level when compared to a reference sample, release the batch from which the sample was obtained. The present invention provides the method including the above.
[0031] The method of the present invention can be used with samples selected from samples that are active pharmaceutical ingredient samples or drug product samples.
[0032] Exemplary morphological features that can be extracted by the method of the present invention include ECD, area, perimeter, circularity, maximum feret diameter, aspect ratio, intensity, x-position, y-position, time (%), time (minutes), and any combination thereof.
[0033] In the method of the present invention, images are classified into major categories, which may be "air and oil," "dark areas and proteinaceous substances," "fibers," and "proteinaceous substances." The method further provides subclusters of images within each major category, where the number of subclusters is 2 to 10, for example, 3 to 5 subclusters in a particular embodiment, or 3 subclusters in a particular embodiment.
[0034] A convolutional neural network is used in the method of the present invention. Any such convolutional neural network, such as Inception V3, can be used.
[0035] In yet another embodiment, a non-transient computer-readable medium is provided that includes machine-readable instructions arranged to cause one or more processors to perform the method of the invention when executed by one or more processors. [Examples]
[0036] We used micro-flow imaging (MFI) datasets belonging to various biopharmaceutical subvisible particles, such as fibers, air and oil, proteinaceous materials, and dark areas and proteinaceous materials, in our research.
[0037] Table 1: Dataset [Table 1]
[0038] Data (preparation, standardization, alignment, annotation, preprocessing) Raw image data and morphological features of bioparticles were obtained from a micro-flow imaging system. The raw images obtained from the MFI device contained no class information (unlabeled). However, the reports produced by the device consisted of specific morphological features of the particles. A Python-based classification script was used to filter the images based on morphological conditions (Figure 1). Subsequently, the morphologically filtered images were validated by a domain expert to confirm their accuracy. The manually validated and corrected data was used as ground truth data to train a deep learning-based model.
[0039] Classification results obtained from analyses performed on initial morphological features were not always reliable. Therefore, a deep learning-based approach for feature extraction was used to obtain additional, unique features for various bioparticles, which are crucial for improving the performance of AI-based particle classification. The more relevant the features, the better they are for training the classification model. A total of 2065 features were obtained, including 17 features extracted from MFI devices and 2048 features extracted from the deep learning model. This morphological feature information for each image was saved in a .csv file. These features were used to train the classification model and thus improve its performance.
[0040] The obtained image sizes were not consistent. Therefore, the MFI images were standardized and resized to 299 to ensure consistency across datasets.
[0041] Deep learning models To develop an AI-based algorithm, I studied deep learning approaches. For example, I developed models using the Keras and TensorFlow frameworks, along with other libraries such as OpenCV, NumPy, scikit-learn, pandas, and matplotlib. I used a convolutional neural network (CNN)-based architecture to extract features from the second-to-last layer of the architecture. Convolutional neural networks are feedforward neural networks commonly used to analyze visual images through data processing using lattice topology. CNNs have multiple hidden layers that help extract information from images. The four key layers in a CNN are the folding layer, ReLU layer, pooling layer, and fully connected layer.
[0042] In our study, we investigated the Inception_V3 architecture for extracting features from images. A total of 2065 features were extracted, in addition to morphological features. Inception_V3 is an image recognition model that achieved an accuracy of over 78.1% on the ImageNet dataset. These features improve model training and accuracy.
[0043] Baseline Architecture Inception V3 Several architectures were explored by iteratively testing and fine-tuning the hyperparameters shown in Table 1. The performance of various models was evaluated on a fixed subset of data consisting of all four types of subparticles. Here, a transfer learning-based CNN approach was used for feature extraction. Inception-v3 was frequently applied in image recognition. The model consists of symmetric and asymmetric building blocks, including folding, mean pooling, max pooling, concatenation, dropout, and fully connected layers. The number of parameters in Inception-v3 is less than half that of AlexNet (60,000,000) and less than a quarter of that of VGGNet (140,000,000); furthermore, the total number of floating-point operations in the Inception-v3 network is approximately 5,000,000,000, which is much larger than that of Inception-v1 (approximately 1,500,000,000). These features make Inception-v3 more practical, namely it can be easily implemented on general-purpose servers to provide rapid response services. The image size input to Inception-v3 was 299x299. A total of 2048 features were extracted from the second-to-last layer of the network. While the first layer of any neural network is fundamentally involved in identifying low-level features such as edges, color, and blobs, the last layer is usually highly specific to the task for which it was trained.
[0044] Table 2: Parameters related to the baseline architecture [Table 2]
[0045] Decision tree Decision trees are nonparametric supervised learning algorithms used for both classification and regression tasks. They have a hierarchical, tree-like structure consisting of a root node, branches, internal nodes, and leaf nodes. Learning a decision tree utilizes a divide-and-conquer strategy by employing a greedy algorithm to identify the optimal split points within the tree. The splitting process is then iterated over in a top-down, recursive manner until all or most of the records are classified under a specific class marker. Pruning is typically employed to reduce complexity and prevent overfitting; this is the process of removing branches that split less important features. The fit of the model can then be evaluated by the process of cross-validation. While there are multiple ways to select the best features at each node, two methods, information gain and Gini impurity, act as popular splitting evaluation criteria for decision tree models. They help evaluate the quality of each test condition and how well it classifies the samples into classes. Entropy is a concept that measures the impurity of sample values. Its value lies between 0 and 1. Information gain represents the difference in entropy before and after a split in a given feature. The feature with the highest information gain produces the best split because it performed best in classifying the training data by its target classification. Gini impurity is the probability of an inaccurately classified random data point in the dataset, if it is labeled based on the class distribution of the dataset. Here, 2065 features, combining features obtained from a deep learning model with early morphology, are input into the machine learning classifier.
[0046] Retraining pipeline Machine learning and deep learning models are ubiquitous in modern organizations. From banking to healthcare, education, manufacturing, construction, and beyond, every industry has appropriate machine learning and deep learning applications. One of the biggest challenges in all of these ML and DL projects across various industries is model improvement. Continuous training is a form of machine learning computation that automatically and continuously retrains a machine learning capability model so that it adapts to data changes before it is redeployed. As soon as a machine learning model is deployed in manufacturing, its performance will degrade because the model is sensitive to real-world changes, and user behavior continues to change over time. All machine learning models deteriorate, but the rate of deterioration changes over time. This is largely caused by data drift. Data drift (covariate shift) is the change in the statistical distribution of production data from the baseline data used to train or build the model. Therefore, it is crucial to observe and retrain machine learning models in manufacturing.
[0047] In this study, we developed a retraining pipeline, as shown in Figure 1, that provides the user with criteria for selecting the data to be used to retrain a model. The data consists of images and .csv files containing information about the class of each image, verified by a domain expert. Once the data for retraining is selected, the images are placed in their respective folders. Features are extracted using a deep learning-based approach. The features are then input into a machine learning model for retraining purposes, which is then evaluated against a test set. Upon completion of the retraining process, the model's performance in terms of accuracy is returned, which serves as an evaluation criterion to help the user decide whether to save or discard the model.
[0048] User interface and cloud-based setup We integrated a modular software application into an extensible and open architectural system consisting of independent user interface (UI) modules built with React. Server-side logic utilized Java techniques stacked with sprint boot, and algorithmic modules were integrated into Python. These modules were combined with REST APIs to enable seamless exchange within internal services, and the application was hosted in an AWS cloud environment. The UI modules provided user access to the application via modern browsers (Chrome, Microsoft Explorer, Firefox, etc.). UI layer communication was handled using SSL. 72The system operated securely through technology. Access was authenticated via an internet-based URL and was geofencing-enabled for geofencing regions / countries. User authentication was enabled by SAML 2.0 (Security Assertion Markup Language 2.0) and SSO (Single-Sign-On) based access management. User authorization was performed using IdP (Identity Provider) and OAuth (Open Authorization) technologies. Route administrators provided access to the necessary user base by pre-designing their unique registration information (ID and email). The user interface, built on React JS, provided tracking functionality consisting of browser-based file uploads, UI-based validation, and display of results and reports. Data management and organization for the entire application were integrated into MySQL® (Amazon RDS) as the database. Raw data files in *.zip folder format, consisting of *.png and *.csv files, were uploaded and managed using Java / J2EE and Spring Boot as middleware. We configured a script to check the data integrity of uploaded files before importing them into the application. We used Amazon Elastic Compute Cloud (Amazon EC2) to build web application hosting, backup, and recovery. We built storage for source data related to application logs, model training, and predictions using AWS S3 (Simple Storage Service). We deployed the developed baseline architecture within an algorithm module using Python as the programming language. This includes artificial intelligence, machine learning, and neural network-based logic and services. We used NGINX for the web server and load balancing.
[0049] Conceptual design and project execution Understanding Exploration Research with Datasets and AI Exp 1: Transfer learning (CNN-based) Pre-trained models such as Inception_V3, ResNet50, and ResNet101 models, pre-trained weights, models, and checkpoints were collected from the TensorFlow hub (collected form). The image dataset was used as is to train the pre-trained models, where all layers are fixed (these hidden layers are not trained), but the last layer can be trained. Resnet_v1_50 was the model used to build the image classification model. Training information is shown in Table 3. Classification reports are evaluated on a scale of 0 to 1. Model performance is shown in Figure 5.
[0050] Table 3: Training information [Table 3]
[0051] Exp 2: Ensemble Learning Algorithms For example, we developed machine learning algorithms using ensemble learning algorithms with different morphological features, such as decision trees, bagging-random forests, boosting-adaboost, and gradient boosting, for data classification into four classes, such as air and oil, proteineceous, dark areas and proteineceous, and fibers. Ensemble learning is a common meta-approach for machine learning that seeks better predictive performance by combining predictions from multiple models. These models are known as weak learners. Intuitively, when several weak learners are combined, they can become a strong learner. Each weak learner is fitted to the training set and provides the resulting prediction. The final prediction result is calculated by combining the results from all the weak learners. Ensemble learning techniques have been proven to achieve better performance on machine learning problems. The performance of the models is shown in Figure 6.
[0052] Exp 3: Transfer learning (CNN-based) using morphological and deep learning features Transfer learning (CNN-based) using morphological features extracted from MFI devices and additional deep learning features extracted using image processing techniques. The fundamental premise of transfer learning is simple: obtain a model trained on a large dataset, and then transfer its knowledge to a smaller dataset. The addition of other morphological features in deep learning model training showed improvement in model performance. Training information is shown in Table 4. Classification reports are rated on a scale of 0 to 1. Model performance is shown in Figure 7.
[0053] Table 4: Training information [Table 4]
[0054] We explored the classification of particles into different classes using multiple machine learning algorithms such as decision trees, random forests, Adaboost, and K-nearest neighbors.
[0055] Exp 4: Random Forest A machine learning algorithm (random forest) using morphological features identified by domain experts and deep learning features extracted using image processing techniques. A random forest is a classification algorithm consisting of many decision trees. It attempts to create an uncorrelated forest of trees where, using bagging and randomness of features, the committee's prediction is more accurate than that of any single tree. Training information is shown in Table 4. Classification reports are evaluated on a scale of 0 to 1. Model performance is shown in Figure 8.
[0056] Exp 5: Adaboost A machine learning algorithm (ADABoost) that uses morphological features identified by domain experts and deep learning features extracted using image processing techniques. The ADABOOST algorithm, an abbreviation for adaptive boosting, is a boosting technique used as an ensemble method in machine learning. It is called adaptive boosting because weights are reassigned to each instance, with higher weights assigned to instances that are misclassified. Training information is shown in Table 4. Classification reports are evaluated on a scale of 0 to 1. Model performance is shown in Figure 9.
[0057] Exp 6:K nearest neighbor method K-Nearest Neighbors (KNN) is a deep learning feature clustering algorithm that uses morphological features identified from MFI devices and deep learning features extracted using image processing techniques. KNN is a nonparametric, supervised learning classifier that uses proximity to classify or predict groupings of individual data points. Training information is shown in Table 4. Classification reports are evaluated on a scale of 0 to 1. Model performance is shown in Figure 10.
[0058] Final architecture In addition to the morphological features obtained from the MFI device, a deep learning-based approach (Inception V3) was implemented to extract additional features. A total of 2065 features, combining morphological and deep learning features, were input into a machine learning classifier (decision tree) to classify subvisible particles into their respective classes. The training information is shown in Table 4. The classification report was evaluated on a scale of 0 to 1. The model performance is shown in Figure 11.
[0059] comparative study The presence of drug aggregates and subvisible particles in therapeutic protein products is becoming an increasingly important area for both the pharmaceutical industry and regulatory authorities. Agglutinations in the micron range are associated with adverse reactions and / or reduced efficacy of therapeutic drugs. Subvisible particles can vary in size, structure, and many other characteristics in various experiments. Various experiments include mechanical agitation (leading to exposure to hydrophobic air / water boundaries), chemical alteration, and / or temperature limits. Protein aggregates can also be produced by protein nucleation around nano / micro-scale contaminants in the product, such as silica particles detached from containers, fibers detached from filters, or metal particles detached from manufacturing equipment. Therefore, comparative studies between samples from two experiments help researchers identify new or unique samples and gain broad knowledge. Here, the comparative study consists of two parts, as shown in Figure 12.
[0060] statistical analysis Statistical analysis was performed on the obtained initial morphological features. The statistical analysis processed morphological features in .csv format obtained from the MFI device for both experiments as input. The .csv file consists of morphological features and user-corrected inference results. It includes classification results showing the volume-based subparticle count, quantity as a percentage, and concentration for each experiment, as shown in Table 5. A data table (explicitly shown in Table 4) is created. Descriptive statistics (minimum, maximum, mean, and standard deviation) for each subparticle's morphological features (ECD, area, perimeter, circularity, maximum Ferret diameter, and aspect ratio) are provided. Descriptive statistics are crucial because simply presenting raw data, especially when there is a large amount, would make it difficult to visualize what the data indicates. Therefore, descriptive statistics allow for a more meaningful presentation of the data, enabling a simpler interpretation. It also provides information on the percentage by which one study deviates from the other with respect to these four evaluation criteria, and presents the top three deviations for each subparticle, as shown in Table 7.
[0061] Table 5. Classification results of the experiment [Table 5]
[0062] Table 6. Descriptive statistical overview [Table 6]
[0063] Table 7: Deviation in descriptive statistics [Table 7]
[0064] This study creates an overview of the particle size distribution for each subparticle, providing information on particle counts for each size of features in microns (macrons). The size range is categorized into four categories: <5 mm, 5-10 mm, >=10 mm, and >=25 mm. The comparative study also provides descriptive statistical graphs, distributions of morphological features, comparisons of distributions for different combinations of morphological features, and particle size distributions for each subparticle and for all of them combined. Several comparisons are shown in Figures 13, 14, and 15.
[0065] Image analysis Image analysis involves clustering features extracted from the image for each particle. Clustering of images based on features and intensity helps identify any new or different molecules. It also helps identify outliers or any misclassifications present in each subvisible particle. Furthermore, it visualizes a graph showing the images belonging to each cluster and the count of samples present in each cluster. The image visualization is shown in Figure 16, and the graph is shown in Figure 17.
[0066] Model performance and management We built an AI model from training data and validated it using test data. Model validation defines its performance in terms of quality. Therefore, defining criteria for CBA (Model Management and Performance Monitoring) is highly relevant to the daily use of AI models.
[0067] Table 8: Evaluation Criteria for Model Performance [Table 8]
Claims
1. A computer-implemented method for measuring subvisible particles in a sample, the following: a. For example, one or more data files are obtained, such as raw data format and / or images of the sample from the analytical instrument, where the data files represent image information from the analytical sample. b. Extract morphological feature information from images from data files and / or analysis samples, c. Classify each image of the analysis sample using a trained convolutional neural network within the main categories. d. Cluster images within the main category into subclusters based on image similarity feature information from sample images. e. Apply statistical analysis comparing within subclusters of major categories to morphological features and / or image similarity to obtain statistical values for a given sample, and, f. Evaluate statistical values to determine the distribution of subvisible particles in the sample. The method, including the method described above.
2. A computer implementation method for batch release based on subvisible particles in a sample from a batch, the following: a. For example, one or more data files are obtained, such as raw data format and / or images of the sample from the analytical instrument, where the data files represent image information from the analytical sample images. b. Extract morphological feature information from images from data files and / or analysis samples, c. Classify each image of the analysis sample using a trained convolutional neural network within the main categories. d. Cluster images within the main category into subclusters based on morphological feature information from sample images. e. Apply statistical analysis comparing within subclusters of major categories to morphological features and / or image similarity to obtain statistical values for a given sample. f. Compare the statistical values for a given sample with those for a reference sample, and then... g. If the sample falls within a specified confidence level when compared to a reference sample, release the batch from which the sample was obtained. The method, including the method described above.
3. The method according to claim 1 or 2, wherein the sample is selected from samples that are active pharmaceutical ingredient samples or drug product samples.
4. The method according to any one of claims 1 to 3, wherein the morphological feature information is selected from ECD, area, perimeter, circularity, maximum feret diameter, aspect ratio, intensity, x-position, y-position, time (%), time (minutes), and any combination thereof.
5. The method according to any one of claims 1 to 4, wherein the main category is selected from "air and oil," "dark areas and proteinaceous substances," "fibers," and "proteinaceous substances."
6. The method according to any one of claims 1 to 5, wherein the number of subclusters is 2 to 10.
7. The method according to claim 6, wherein the number of subclusters is 3 to 5.
8. The method according to claim 6 or 7, wherein the number of subclusters is three.
9. The method according to any one of claims 1 to 8, wherein the convolutional neural network is Inception V3.
10. A non-transient computer-readable medium containing machine-readable instructions arranged to cause one or more processors to perform the method of the invention when executed by one or more processors.