A system and method for stratifying cancer based on radiomics.
Patent Information
- Application Number
- JP2026508949
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-08-14
- Publication Date
- 2026-09-01
Smart Images

Figure 2026529637000001_ABST
Abstract
Description
[Technical Field]
[0001] This application is directed towards using radiomic features to classify cancer states in subjects. [Background technology]
[0002] Radiomics refers to a quantitative approach to medical imaging, in which a large number of features called radiomic features are extracted from electronic medical image data using data characterization algorithms. Models are used to predict outputs based on the input radiomic features. The underlying rationale for using radiomics is the hypothesis that electronic medical image data contains information beyond visual perception that can better reflect tissue characteristics and improve the accuracy of diagnosis or prognosis. To date, radiomics has been applied to identify and quantify tumor types, assess the risk of various types of cancer, and predict the survival time of cancer patients.
[0003] Radiomics has shown promise in predicting therapy response and overall prognosis, but several questions remain. For example, radiomic models are typically trained using a large number of features (e.g., thousands). The predictive power of a single radiomic feature is low and increases with the use of a group of features, but the number of features required to learn or form a "critical number" is unknown. Furthermore, some features (or groups of features) may be more important than others depending on the characterization task. For example, one group of features may be more important for predicting a certain type of cancer, while another group of features may be more important for predicting patient survival. [Overview of the project]
[0004] As is evident from the above description, there is still a need in the art for improved methods and systems for characterizing cancer conditions using radiomics on an appropriate scale. The methods and systems described herein satisfy these and other needs by providing methods and systems for creating competitive radiomics models for stratifying cancer patients based on the characteristics of their cancer condition by utilizing individually weak radiomics features across a wide range of radiomics feature classes.
[0005] According to some embodiments disclosed herein, a radiomics model includes an ensemble model comprising multiple component models. Each component model takes the values of radiomic features within a certain individual radiomic feature class as input and outputs individual predictive components for a given cancer state, thereby obtaining multiple component predictions for a cancer state from the multiple component models. The ensemble model combines the multiple component predictions to obtain a characterization of the cancer state as the output of the ensemble model. The purpose of the ensemble model in this case is not necessarily to create a better combined model. Generalization to new data is often a challenge in radiomics models. This is especially true when training a model with a larger set of input features. The ensemble models disclosed herein have the technical advantage of keeping the training feature set small, which helps reduce overtraining. Furthermore, because individual component models tend to overlap with and correlate with other component models, combining the component predictions from these component models provides a single given model for characterizing cancer states. This allows individual models to become more robust and generalizable while retaining much of their predictive power. The output of ensemble models can also be integrated with clinical information, thereby providing useful and complementary insights for personalized medicine.
[0006] Each of the systems, methods, and apparatuses disclosed herein has several innovative aspects, and no single aspect is solely responsible for the desirable attributes disclosed herein.
[0007] According to one aspect of this disclosure, a method is provided for characterizing the cancerous state of tissue in a subject. The method includes inputting information into an ensemble model comprising multiple component models to obtain corresponding component predictions regarding the cancerous state as outputs from each individual component model within the multiple component models, thereby obtaining multiple component predictions regarding the cancerous state. The information includes, for each individual radiomic feature class within a plurality of radiomic feature classes, corresponding values for each individual radiomic feature within the corresponding radiomic features of an individual radiomic feature class obtained from a medical image dataset. The medical image dataset includes a plurality of medical images of tissue in a subject, initially obtained using a first medical image modality. In some embodiments, the plurality of medical images collectively provide a three-dimensional image of the tissue. The ensemble model includes a plurality of parameters. Input includes (i) inputting corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of a first individual radiomic feature class within a multiple radiomic feature class into the first individual component model within a multiple component model, and (ii) inputting corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of a second individual radiomic feature class within a multiple radiomic feature class into the second individual component model within a multiple component model. The corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of the first individual radiomic feature class are not input into the second component model, and the corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of the second individual radiomic feature class are not input into the first individual component model. The method includes combining the multiple component predictions to obtain a cancer state characterization as the output of an ensemble model.
[0008] In some embodiments, the characterization of the cancer state includes individual cancer types selected from a plurality of cancer types, individual cancer stages selected from a plurality of cancer stages, individual tissue origins selected from a plurality of tissue origins, individual cancer grades selected from a plurality of cancer grades, or individual prognoses selected from a plurality of prognoses.
[0009] In some embodiments, the multiple radiomic feature classes include a first subset of radiomic feature classes extracted from unfiltered versions of multiple medical images in a medical image dataset, and a second subset of radiomic feature classes extracted from filtered versions of multiple medical images in a medical image dataset filtered by a first filtering method.
[0010] In some embodiments, the first filtering method includes an image filter selected from the group consisting of wavelet transform filters, Gaussian Laplacian (LoG) filters, square transform filters, square root transform filters, logarithmic transform filters, exponential transform filters, gradient transform filters, two-dimensional local binary pattern filters, and three-dimensional local binary pattern filters.
[0011] Another aspect of this disclosure provides a computer system for characterizing the cancerous state of tissue in a subject. The computer system comprises one or more processors and one or more memory addressable by the processors. The memory stores one or more programs configured to be executed by one or more processors. One or more programs, individually or collectively, include instructions for performing any of the methods described herein.
[0012] Another aspect of this disclosure provides a non-temporary computer-readable storage medium that, when executed by a computer system, stores instructions causing the computer system to perform any of the methods described herein.
[0013] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will be apparent to those skilled in the art, particularly in view of the drawings, the specification, and the claims. Furthermore, it should be noted that the language used herein has been selected primarily for readability and instructional purposes, and may not have been selected to delineate or limit the subject matter of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In the drawings, embodiments of the system and method of the present disclosure are shown by way of example. It should be explicitly understood that the specification and drawings are merely for illustrative purposes and to assist understanding, and are not intended to define the limitations of the system and method of the present disclosure.
[0015] [Figure 1A] Figure 1A shows a computer system according to some embodiments of the present disclosure. [Figure 1B] Figure 1B shows a computer system according to some embodiments of the present disclosure. [Figure 2A] Figure 2A collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject according to some embodiments of the present disclosure. [Figure 2B] Figure 2B collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject according to some embodiments of the present disclosure. [Figure 2C] Figure 2C collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject according to some embodiments of the present disclosure. [Figure 2D] Figure 2D collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject according to some embodiments of the present disclosure. [Figure 2E]Figure 2E collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2F] Figure 2F collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2G] Figure 2G collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2H] Figure 2H collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2I] Figure 2I collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2J] Figure 2J collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2K] Figure 2K collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2L] Figure 2L collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2M] Figure 2M collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2N] Figure 2N collectively provides flowcharts of processes and features for classifying a cancer state of tissue in a subject, in accordance with some embodiments of the present disclosure. [Figure 2O]Figure 2O provides a collective flowchart of processes and characteristics for classifying the cancerous state of tissue in a subject, according to several embodiments of the present disclosure. [Figure 3] Figure 3 illustrates a process for classifying the cancerous state of tissue in a subject, according to several embodiments of the present disclosure. [Figure 4] Figure 4 illustrates the process for training a model to classify the cancerous state of tissue in a subject, according to several embodiments of the present disclosure. [Figure 5A] Figure 5A collectively shows the performance of ensemble and component models according to several embodiments of the present disclosure. [Figure 5B] Figure 5B collectively shows the performance of ensemble and component models according to several embodiments of the present disclosure. [Figure 5C] Figure 5C collectively shows the performance of ensemble and component models according to several embodiments of the present disclosure.
[0016] Similar reference numbers refer to the corresponding parts throughout several drawings within the drawing. [Modes for carrying out the invention]
[0017] Here, embodiments are shown in detail in the accompanying drawings. The following detailed description includes numerous specific details to allow for a full understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure can be practiced without these specific details. In other examples, well-known methods, procedures, components, circuits, and networks are not described in detail so as not to unnecessarily obscure the aspects of the embodiments.
[0018] A system and method for characterizing the cancerous state of tissue in a subject using radiomic features are disclosed. One of the main challenges in building risk models using radiomics is the need to deal with the large number of radiomic features that can be extracted from medical images. Many of the features are correlated with each other, and it is unclear which features or groups of features are more important than others in characterizing the cancerous state. These factors can consequently lead to training models that do not generalize well.
[0019] Advantageously, this disclosure provides methods and systems for characterizing cancer conditions using radiomics by utilizing a broad range of individually weak radiomic features to create a competing risk model that generalizes better than existing radiomic models. In some implementations, the methods and systems described herein utilize hundreds or thousands of weak radiomic features, which are divided into feature subsets. The individual feature sets are then evaluated using separate risk models, and their outputs are combined into an ensemble risk model.
[0020] In some embodiments, features are grouped based on the source and / or filtering status of the radiomics images. For example, in some embodiments, features generated from unfiltered radiomics images are grouped separately from features generated from filtered radiomics images, regardless of the feature generation method. That is, in some embodiments, a first instance of the same feature (e.g., entropy) is grouped into a first feature subgroup if it is determined to be from an unfiltered image, while a second instance of the same feature (e.g., again entropy) is grouped into a second feature subgroup if it is determined to be from a filtered version of the same image.
[0021] In some embodiments, features are additionally or alternatively grouped based on the method used to generate them. This second difference separates a large number of highly related features generated by a particular method from a separate group of features generated using different methods. For example, in some embodiments, gray-level co-occurrence matrix (GLCM) features are sorted into a separate group from gray-level run-length matrix (GLRLM) features. Similarly, in some embodiments, local binary pattern (LBP) features are separated from local terminally pattern (LTP) features, which in turn are separated from GLCM and GLRLM features.
[0022] In some embodiments, ensemble models are used to achieve better generalization. Radiomics models often do not generalize well when novel data is presented, especially when training the model with larger input feature sets. For example, as described in the embodiment, when a forty-component model was trained with a large feature set to stratify survival in non-small cell lung cancer patients in a one-time training, the model exhibited unsatisfactory statistical significance.
[0023] Advantageously, the methods and systems for training and using the ensemble models described herein improve generalization of radiomics models. In some embodiments, the methods and systems described herein achieve improved generalization, at least in part, by keeping the feature set used to train individual models relatively small and preventing overfitting. Furthermore, in some embodiments, the methods and systems described herein, by aggregating a large number of overlapping but correlated individual models, maintain the predictive power of a large number of weak radiomics features while being more robust and generalizable. For example, as also described in the Examples, when each component model is trained together as an ensemble model using the K-fold training scheme illustrated in FIG. 4, the ensemble model has better performance and statistical significance than individual component models.
[0024] FIG. 3 depicts an example process 300 that addresses these challenges and other needs, that uses radiomics features to characterize a cancer condition. Process 300 utilizes an ensemble model 302 that includes a plurality of separate component models 304 (e.g., component model 1 304-1, component model 2 304-2, and component model N 304-N (e.g., risk models)). Each individual separate component model obtains, as input, corresponding values of radiomics features from one or more individual radiomics feature classes. In the example of FIG. 3, component model 1 304-1 receives, as input, radiomics features {F within radiomics feature class C1 11 ,F 12 , ···,F 1n} corresponding values {V 11 ,V 12 , ···,V 1n} is obtained. Component model 2 304-2 receives, as input, radiomics features {F within radiomics feature class C2 21 ,F 22 , ···,F 2n} corresponding values {V 21 ,V22 ,···,V 2p} and radiomic features {F 31 ,F 32 ,···,F 3m The value corresponding to {V} 31 ,V 32 ,···,V 3m Obtain}. The component model N 304-N takes radiomic features class C as input. p Internal radiomic features {F p1 ,F p2 ,···,F pq The value corresponding to {V} p1 ,V p2 ,···,V pq} is obtained. In some embodiments, each of the component models 304 outputs an individual component prediction 306. In some embodiments, the ensemble model 302 includes a combined module 308 which combines the component predictions 306 (for example, by applying an aggregate operation 310) to output a cancer state characterization 312.
[0025] In some embodiments, radiomic features are classified into classes (e.g., groups) according to the source image (e.g., medical image, source dataset, original dataset, etc.) and its filtering status. For example, in some embodiments, radiomic features include gray-level co-occurrence matrix (GLCM) features (e.g., having a GLCM class). GLCM features generated from the original (e.g., unfiltered) medical image belong to a different category than GLCM features generated from the modified (e.g., filtered) medical image.
[0026] In some embodiments, radiomic features are classified into classes according to their individual feature generation methods. For example, GLCM features are generated by determining how frequently pairs of pixels with a specific value and a specified spatial relationship occur in an image, while gray-level run-length matrix (GLRLM) features are generated by determining the length of consecutive pixels with the same gray-level value. In this example, GLCM and GLRLM features belong to different classes. As another example, the local binary pattern (LBP) feature class is separated from the local terminally pattern (LTP) feature class, and both of these are separated from GLCM and GLRLM features.
[0027] In some embodiments, each individual radiomic feature class within a plurality of radiomic feature classes is configured to provide a quantitative evaluation of a particular medical image within a plurality of medical images in a medical dataset by transforming a medical image into a corresponding dataset, such as one or more image biomarkers. For example, in some embodiments, each individual radiomic feature class within a plurality of radiomic feature classes is configured to provide a corresponding dataset by preprocessing one or more volumes or regions of interest of a medical image to provide a unique quantitative evaluation, segmenting one or more volumes or regions of interest of a medical image, obtaining and / or reconstructing medical images, feature extraction, feature selection, statistical analysis, model development (e.g., machine learning model predictive modeling), or a combination thereof. As a non-limiting example, in some embodiments, a statistical individual radiomic feature class includes an unmodified intensity radiomic feature class, a discrete intensity radiomic feature class, a gray-level intensity radiomic feature class, or a combination thereof. In some embodiments, each individual radiomic feature class within a plurality of radiomic feature classes is a statistical radiomic feature class (e.g., a histogram-based radiomic feature class, a texture-based radiomic feature class, etc.), a model-based radiomic feature class, a transformation-based radiomic feature class, or a shape-based radiomic feature class. In some embodiments, each individual radiomic feature class within a plurality of radiomic feature classes is either a two-dimensional (2D) region of interest-based radiomic feature class or a three-dimensional (3D) volume of interest-based radiomic feature class. However, the disclosure is not limited thereto.Further details and information regarding radiomic feature classes can be found in Mayerhoefer et al., 2020, “Introduction to Radiomics,” Journal of Nuclear Medicine, 61(4), pp. 488-495, and Traverso et al., 2018, “Repeatability and Reproducibility of Radiomic Features: A Systematic Review,” International Journal of Radiation Oncology*Biology*Physics, 102(4), pp. 1143-1158, each of which is incorporated herein by reference in its entirety for all purposes.
[0028] In some embodiments, each individual component model within a multiple component model is: 2 to 8 radiomic feature classes I, multiple radiomic feature classes, 2 to 7 radiomic feature classes, 2 to 6 radiomic feature classes, 2 to 5 radiomic feature classes, 2 to 4 radiomic feature classes, 2 to 3 radiomic feature classes, 3 to 8 radiomic feature classes, 3 to 7 radiomic feature classes, 3 to 6 radiomic feature classes, 3 to 5 The radiomic feature classes are configured to utilize 3-4 radiomic feature classes, 4-8 radiomic feature classes, 4-7 radiomic feature classes, 4-6 radiomic feature classes, 4-5 radiomic feature classes, 5-8 radiomic feature classes, 5-7 radiomic feature classes, 5-6 radiomic feature classes, 6-8 radiomic feature classes, 6-7 radiomic feature classes, or 7-8 radiomic feature classes. In some embodiments, each individual component model within a multi-component model is configured to utilize at least 2 radiomic feature classes, at least 3 radiomic feature classes, at least 4 radiomic feature classes, at least 5 radiomic feature classes, at least 6 radiomic feature classes, at least 7 radiomic feature classes, or at least 8 radiomic feature classes. In some embodiments, each individual component model within a multi-component model is configured to utilize up to two, up to three, up to four, up to five, up to six, up to seven, or up to eight radiomic feature classes. For example, in some embodiments, the multi-component model includes at least two component models, and each individual component model within at least two component models is configured to utilize a unique subset of radiomic feature classes within multiple radiomic features.In some embodiments, each individual component model within at least two component models is configured to utilize a disparate subset of radiomic feature classes within a plurality of radiomic features, so that each individual component model within at least two component models has no overlaps, having non-overlapping radiomic feature classes from the plurality of radiomic feature classes. As a non-limiting example, in some embodiments, the first component model within at least two component models includes three radiomic feature classes from the plurality of radiomic feature classes (e.g., a first radiomic feature class, a second radiomic feature class, and a third radiomic feature class from the plurality of radiomic feature classes), and the second component model within at least two component models includes four radiomic feature classes from the plurality of radiomic feature classes (e.g., a fourth radiomic feature class, a fifth radiomic feature class, a sixth radiomic feature class, and a seventh radiomic feature class from the plurality of radiomic feature classes). The third component model in at least two component models includes an omics feature class, and the third component model includes one radiomic feature class from multiple radiomic feature classes (for example, an eighth radiomic feature class from multiple radiomic feature classes), where the first, second, third, fourth, fifth, sixth, seventh, and eighth radiomic feature classes from multiple radiomic feature classes are all distinct radiomic feature classes.
[0029] Having outlined improved systems and methods for structuring medical data to enable machine learning, further details of the systems, apparatus, and / or processes described herein will be explained here with reference to Figures 1, 2, 4, 5, and 6.
[0030] Figure 1A shows a computer system for characterizing the cancerous state of tissue in a subject, according to several embodiments of the present disclosure.
[0031] In a typical embodiment, computer system 100 comprises one or more computers. For illustrative purposes in Figure 1A, computer system 100 is represented as a single computer encompassing all the functions of computer system 100 of this disclosure. However, this disclosure is not limited thereto. The functions of computer system 100 can extend across any number of networked computers and / or reside on each of several networked computers and / or virtual machines. Those skilled in the art will understand that a wide range of different computer topologies are possible for computer system 100, and that all such topologies are within the scope of this disclosure.
[0032] Returning to Figure 1A with the foregoing in mind, the computer system 100 comprises one or more processing units (CPUs) 59, a network or other communication interface 84, a user interface 78 (including, for example, an optional display 82 and an optional keyboard 80 or other form of input device), memory 92 (e.g., random access memory, persistent memory, or a combination thereof), one or more magnetic disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication buses 12 for interconnecting the aforementioned components, and a power supply 79 for supplying power to the aforementioned components. Data in memory 92 can be seamlessly shared with non-volatile memory 90, or a portion of memory 92 that is non-volatile or persistent, using known computing techniques such as caching, to the extent that the components of memory 92 are not persistent. Memory 92 and / or memory 90 may include large-capacity storage located remotely from the central processing unit 59. In other words, some data stored in memory 92 and / or memory 90 may actually be outside the computer system 100 but may be hosted on a computer that can be electronically accessed by the computer system 100 via a network 102 (e.g., the Internet, an intranet, or other form of network or electronic cable) using a network interface 84. In some embodiments, the computer system 100 utilizes a model in which the data is executed from memory associated with one or more graphics processing units in order to improve the speed and performance of the system. In some alternative embodiments, the computer system 100 utilizes a model in which the data is executed from memory 92 rather than from memory associated with the graphics processing units.
[0033] The memory 92 of computer system 100 stores the following: • Operating system 30, which includes procedures for handling various basic system services. • A communication module 32 that connects to and communicates with other network devices (e.g., local networks such as an internet connection, networked storage devices, network routing devices, server systems, other computer systems 100, and / or routers providing other connected devices) connected to one or more communication networks via a network interface 84 (e.g., wired or wireless), An optional extraction module 34 for extracting radiomics data 50, such as radiomics features 64, from a medical image dataset 50, an optional filter module 36 for filtering a medical image dataset 50 through the application of one or more filtering methods and / or one or more filters 54, An optional extraction module 34 includes an optional filter 54 (e.g., a medical image filtering algorithm) applied to a medical image dataset 50, wherein in some embodiments, the filter 54 includes one or more of the following: wavelet transform filters, Gaussian Laplacian (LoG) filters, square transform filters, square root transform filters, logarithmic transform filters, exponential transform filters, gradient transform filters, 2D local binary pattern filters, and 3D local binary pattern filters. • Optional segmentation module 38 for segmenting medical images from medical image dataset 50 into segments and / or volumes. A prediction module 40 for characterizing the cancer state in a subject, wherein in some embodiments, the prediction module 40 is o Prediction model 42, in some embodiments, the prediction model 42 comprises an ensemble model (e.g., ensemble model 302, Figure 3) including a plurality of component models (e.g., component model 304, Figure 3), such as component model 1 44-1 and component model N 44-N, in some embodiments, the plurality of component models are at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 150, or more component models, in some embodiments, each of the component models 44 is configured to output individual component predictions 58 (e.g., component prediction 306, Figure 3), and the prediction model 42 comprises an ensemble model (e.g., ensemble model 302, Figure 3), in which case the plurality of component models is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 150, or more component models, and in some embodiments, each of the component models 44 is configured to output individual component predictions 58 (e.g., component prediction 306, Figure 3), Model parameters 46 (e.g., weights, biases, cluster centroids, etc.), where in some embodiments, the model parameters 46 include at least 1,000, 5,000, 10,000, or 20,000 parameters, and A prediction module 40 includes a coupling module 308 for combining multiple component predictions to obtain a cancer state characterization as the output of an ensemble model, A data store 48 for storing inputs and / or outputs from an ensemble model described herein (e.g., a predictive model 42), wherein in some embodiments, the data store 48 is oIncludes medical image datasets such as medical image dataset 1 50-1 and medical image dataset P 50-P, and in some embodiments, each medical image dataset includes corresponding radiomics data 52 (for example, medical image dataset 1 50-1 includes radiomics data 52-1, and medical image dataset P 50-P includes radiomics data 52-P), and Figure 1B is a block diagram showing radiomics data for each medical image dataset according to some embodiments, and in some embodiments, each medical image dataset includes multiple radiomics feature classes 62, such as radiomics feature class 1 62-1 and radiomics feature class Q 62-Q, and each radiomics feature class 62 includes individual multiple radiomics features 64, for example, Figure 1B shows that radiomics feature class 1 62-1 includes radiomics features 1-1 64-1-1 and radiomics features 1-M It is indicated that 64-1-M is included, and radiomic feature class Q 62-Q includes radiomic feature Q-1 64-Q-1 and radiomic feature QX 64-QX, and in some embodiments, each radiomic feature class includes at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features, and in some embodiments, each radiomic feature class includes 1000 or less, 750 or less, 500 or less, 250 or less, or 100 or less corresponding radiomic features, and each radiomic feature 64 has a corresponding value 66, medical image dataset, A data store 48 includes a model output 56, which in some embodiments includes one or more (e.g., combined) component predictions 58 from individual component models 44. • Reporting module 60 for generating reports for clinicians or patients based on, for example, an assessment of cancer status from an ensemble model (e.g., predictive model 42 and / or component model 44), and An optional training module 68 comprising labels 70 and one or more training datasets 72 for training a predictive model 42 and / or a component model 44.
[0034] In some embodiments, one or more of the data elements or modules of the computer system 100 identified above are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the functions described above. The data, modules, or programs (e.g., sets of instructions) identified above do not need to be implemented as separate software programs, procedures, or modules; therefore, various subsets of these modules can be combined or rearranged in various implementations. In some implementations, memories 92 and / or 90 optionally store a subset of the modules and data structures identified above. Furthermore, in some embodiments, memories 92 and / or 90 store additional modules and data structures not described above. Details of the modules and data structures identified above are described below with reference to Figures 2-6.
[0035] Figures 2A, 2B, 2C, 2D, 2E, 2F, 2G, 2H, 2I, 2J, 2K, 2L, 2M, 2N, and 2O collectively provide flowcharts of exemplary methods 200 for characterizing cancerous conditions in a subject, according to several embodiments. In some embodiments, methods 200 are implemented in a computer system 100 including one or more processors (e.g., CPU 59) and memory (e.g., memory 90 or memory 92). In some embodiments, the computer system 100 performs each step as shown in Figure 2.
[0036] Referring to block 202 in Figure 2A, in some embodiments, the method includes extracting corresponding values for each individual radiomic feature in each individual radiomic feature within each radiomic feature within the first subset of radiomic feature classes (e.g., via an optional extraction module 34) for each individual radiomic feature within a first subset of radiomic feature classes (e.g., a subset of radiomic feature classes extracted from a set of medical images, either processed using the same filter or unfiltered) from regions of interest (ROIs) or volumes of interest (VOIs) in multiple medical images of the tissue in the subject initially acquired using a first medical image modality (e.g., source dataset, original dataset) (e.g., medical image dataset 50), before inputting the information into an ensemble model. For example, in some embodiments, the multiple medical images include one or more two-dimensional representations of individual tissues (e.g., digital images of the first tissue) and / or one or more three-dimensional representations of individual tissues (e.g., volumetric bodies or volumetric representations of the first tissue). In some embodiments, the multiple images collectively provide three-dimensional images of the tissue. Several software packages for extracting features from medical image sets are known in the art, including the PyRadiomics package, the Cancer Imaging Phenomics Toolkit (CaPTk), and the Standardized Environment for Radiomics Analysis (SERA) package.For further information on the PyRadiomics, CaPTk, and SERA packages, please refer, for example, to van Griethuysen, JJ, et al., Computational Radiomics System to Decode the Radiographic Phenotype, Cancer Research, 77(21):e104-e107 (2017), Pati S., et al., The Cancer Imaging Phenomics Toolkit (CaPTk): Technical Overview, Springer-BrainLes 2019-LNCS, 11993:380-394 (2020), and Ashrafinia, S., Quantitative Nuclear Medicine Imaging using Advanced Image Reconstruction and Radiomics, Ph.D. Dissertation, Johns Hopkins University (2019), each of which is disclosed herein in its entirety by reference.
[0037] In some embodiments, feature extraction includes the steps of (i) segmenting images to represent regions of interest (ROIs) in two-dimensional space or volumes of interest (VOIs) in three-dimensional space, for example, determining tumor volume or cancer mass; (ii) processing images to homogenize them across the entire dataset; and (iii) extracting features from the segmented and processed images and filtered versions thereof. For a review of the feature extraction process, see, for example, van Timmeren, J., Cester, D., Tanadini-Lang, S. et al., Radiomics in medical imaging—“how-to” guide and critical reflection, Insights Imaging, 11:91 (2020), which is incorporated herein by reference in its entirety for all purposes.
[0038] Referring to block 204, in some embodiments, the method includes identifying ROIs or VOIs in multiple medical images (e.g., via an optional segmentation module 38). Referring to block 206, in some embodiments, the method includes segmenting unfiltered versions of multiple medical images into multiple segments or volumes (e.g., via an optional segmentation module 38). Several software packages for image segmentation and ROI / VOI identification are available, including 3D Slicer, MITK, ITK-SNAP, MeVisLab, LifEx, and ImageJ. For a further review of these packages, see, for example, van Timmeren, J., et al., ibid.
[0039] Referring to block 208, in some embodiments, the method includes assigning individual tissue classifications within a plurality of tissue classifications to each individual segment within a plurality of segments, or to each individual volume within a plurality of volumes, based on one or more features of the individual segment or individual volume. In some embodiments, pixel values (e.g., individual pixel values, binned pixel values, locally averaged pixel values, or normalized pixel values) are input to a model trained to distinguish between different tissue types. For example, Ferl, GZ, et al., Automated segmentation of lungs and lung tumors in mouse micro-CT scans, iScience, 25(12):105712(2022), which is disclosed in whole herein by reference, describes a two-stage automated segmentation method for healthy lungs, cancerous lungs and fibrous lungs, using a U-net of a 3D CNN trained to segment lung tissue to identify lung tissue in a set of micro-CT images, and using a support vector machine (SVM) to distinguish between healthy tissue, cancerous tissue and fibrous tissue within the identified lung tissue.
[0040] Therefore, in some embodiments, the method includes inputting the corresponding pixel values or binned pixel values of multiple pixels or binned pixels from a medical image dataset into a model, which applies multiple parameters to the corresponding pixel values or binned pixel values through multiple calculations to generate ROI or VOI identification as an output from the model.
[0041] In some embodiments, the plurality of pixels or binned pixels are at least 100 pixels or binned pixels, at least 1,000 pixels or binned pixels, at least 10,000 pixels or binned pixels, at least 100,000 pixels or binned pixels, at least 1 million pixels or binned pixels, at least 10 million pixels or binned pixels, or at least 1 billion pixels or binned pixels. In some embodiments, the plurality of pixels or binned pixels are 100 billion or fewer pixels or binned pixels, 10 billion or fewer pixels or binned pixels, 1 billion or fewer pixels or binned pixels, 100 million or fewer pixels or binned pixels, or fewer than 10 million pixels or binned pixels. In some embodiments, the multiple pixels or binned pixels are 10,000 to 100 billion, 100,000 to 100 billion, 1 million to 100 billion, 10 million to 100 billion, 10,000 to 10 billion, 1 million to 10 billion, 10 million to 10 billion, 10,000 to 1 billion, 10 million to 1 billion, 1 million to 1 billion, 10,000 to 100 million, 100,000 to 100 million, 1 million to 100 million, or 10 million to 100 million pixels or binned pixels.
[0042] In some embodiments, the multiple parameters are at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more parameters. In some embodiments, the multiple parameters are 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer parameters. In some embodiments, the multiple parameters are 10,000 to 100 billion, 100,000 to 100 billion, 1 million to 100 billion, 10 million to 100 billion, 10,000 to 10 billion, 100,000 to 10 billion, 1 million to 10 billion, 10 million to 10 billion, 10,000 to 1 billion, 100,000 to 1 billion, 1 million to 1 billion, 10,000 to 100 million, 100,000 to 100 million, 1 million to 100 million, or 10 million to 100 million parameters.
[0043] In some embodiments, the multiple calculations are at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more calculations. In some embodiments, the multiple calculations are 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer calculations. In some embodiments, the multiple calculations are performed 10,000 to 100 billion times, 100,000 to 100 billion times, 1 million to 100 billion times, 10 million to 100 billion times, 10,000 to 10 billion times, 100,000 to 10 billion times, 1 million to 10 billion times, 10 million to 10 billion times, 10,000 to 1 billion times, 10 million to 1 billion times, 10,000 to 100 million times, 100,000 to 100 million times, 1 million to 100 million times, or 10 million to 100 million times.
[0044] In some embodiments, the method involves inputting corresponding pixel values or binned pixel values for multiple pixels or binned pixels from an identified ROI or VOI into a model, where the model applies multiple parameters to the corresponding pixel values or binned pixel values through multiple calculations to generate a tissue classification for each pixel or binned pixel within the identified ROI or VOI as an output from the model. In some embodiments, one or more classifications include cancerous and non-cancerous tissues. In some embodiments, one or more classifications include cancer subtypes, cancer grades, and / or cancer stages. In some embodiments, one or more classifications include non-cancerous phenotypes, such as fibrous tissue.
[0045] In some embodiments, the number of pixels or binned pixels is at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more pixels or binned pixels. In some embodiments, the number of pixels or binned pixels is 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer pixels or binned pixels. In some embodiments, the multiple pixels or binned pixels are 10,000 to 100 billion, 100,000 to 100 billion, 1 million to 100 billion, 10 million to 100 billion, 10,000 to 10 billion, 1 million to 10 billion, 10 million to 10 billion, 10,000 to 1 billion, 10 million to 1 billion, 1 million to 1 billion, 10,000 to 100 million, 100,000 to 100 million, 1 million to 100 million, or 10 million to 100 million pixels or binned pixels.
[0046] In some embodiments, the multiple parameters are at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more parameters. In some embodiments, the multiple parameters are 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer parameters. In some embodiments, the multiple parameters are 10,000 to 100 billion, 100,000 to 100 billion, 1 million to 100 billion, 10 million to 100 billion, 10,000 to 10 billion, 100,000 to 10 billion, 1 million to 10 billion, 10 million to 10 billion, 10,000 to 1 billion, 100,000 to 1 billion, 1 million to 1 billion, 10,000 to 100 million, 100,000 to 100 million, 1 million to 100 million, or 10 million to 100 million parameters.
[0047] In some embodiments, the multiple calculations are at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more calculations. In some embodiments, the multiple calculations are 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer calculations. In some embodiments, the multiple calculations are performed 10,000 to 100 billion times, 100,000 to 100 billion times, 1 million to 100 billion times, 10 million to 100 billion times, 10,000 to 10 billion times, 100,000 to 10 billion times, 1 million to 10 billion times, 10 million to 10 billion times, 10,000 to 1 billion times, 10 million to 1 billion times, 10,000 to 100 million times, 100,000 to 100 million times, 1 million to 100 million times, or 10 million to 100 million times.
[0048] Referring to block 210, in some embodiments, the method includes grouping individual segments within multiple segments or individual volumes within multiple volumes that are assigned to a target tissue classification within multiple tissue classifications, thereby identifying an ROI or VOI. Further examples of methods and systems for segmenting medical images are disclosed in U.S. Patent No. 10,991,097, entitled "Artificial intelligence segmentation of tissue images," which is incorporated herein by reference in its entirety for all purposes.
[0049] Referring to block 211, in some embodiments, the method includes extracting corresponding values for each individual radiomic feature in
[0050] Referring to block 212, in some embodiments, the method includes inputting information into an ensemble model (e.g., prediction model 42, ensemble model 302) which includes a plurality of component models (e.g., component model 44, component 304), to obtain corresponding component predictions for cancer status as outputs from each individual component model within the plurality of component models, thereby obtaining a plurality of component predictions for cancer status (e.g., component prediction 58, component prediction 306). The information includes, for each individual radiomic feature within multiple radiomic feature classes (e.g., radiomic feature classes 62-1 to 62-Q), corresponding values (e.g., values 66-1-1, 66-1-M, 66-Q-1, 66-QM) within each corresponding radiomic feature (e.g., values 66-1-1, 66-1-M, 66-Q-1, 66-QM) obtained from a medical image dataset, where the medical image dataset includes multiple medical images of tissue in the subject initially acquired using a primary medical image modality. The ensemble model includes multiple parameters. Inputting includes (i) inputting corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of a first individual radiomic feature class within a multiple radiomic feature class into the first individual component model within a multiple component model, and (ii) inputting corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of a second individual radiomic feature class within a multiple radiomic feature class into the second individual component model within a multiple component model.The corresponding values for individual radiomic features within the corresponding multiple radiomic features of the first individual radiomic feature class are not input into the second component model, and the corresponding values for individual radiomic features within the corresponding multiple radiomic features of the second individual radiomic feature class are not input into the first individual component model.
[0051] Generally, the parameters used for feature extraction will depend on the image modality / acquisition parameters used to collect medical images, and / or the content of the images. There is a balance between filtering out noise in the data and averaging the features. For a further discussion of this balance, see, for example, van Timmeren, J., et al., ibid.
[0052] For example, in some embodiments where a medical image dataset is acquired by magnetic resonance imaging (MRI), e.g., brain MRI, the images in the dataset are normalized to a scale of, for example, 0 to 100 or 0 to 1. In some embodiments, the images (e.g., normalized images) are resampled using a fixed bin width, e.g., 5 mm, and an isotropic voxel spacing, e.g., 1 × 1 × 1 mm. In other embodiments, a fixed number of bins are used for discretization. In some embodiments, lesion ROIs are grouped based on connectivity and distance. Features are extracted from ROIs in the original image type and the transformed image type. In some embodiments, some or all of the features are also determined using the original images without normalization and with the original spacing. In some embodiments, at least shape features are also determined using the original images without normalization and with the original spacing.
[0053] As another example, in some embodiments where the medical image dataset is acquired by position emission tomography (PET) scans, e.g., whole-body PET, the image dataset is resampled using isotropic voxel spacing, e.g., 3x3x3 mm voxel spacing, using the width of a fixed bin. In some embodiments, the images are not normalized. ROIs of lesions are grouped based on connectivity. Features are extracted from the ROIs using the original image type and the transformed image type. In some embodiments, some or all of the features are also determined using the original images without normalization, with the original spacing. In some embodiments, at least shape features are also determined using the original images without normalization, with the original spacing.
[0054] As another example, in some embodiments where a medical image dataset is acquired by computed tomography (CT), e.g., lung CT, the images are resampled using a fixed bin width, e.g., 25, and an isotropic voxel spacing, e.g., 1 × 1 × 1 mm voxel spacing. Here, the term “fixed bin width” relates to the process of transforming raw CT data (attenuation values measured by the CT scanner) into a final image that represents tissue density. This process is called “binning,” and “bin width” refers to a range of attenuation values that are grouped together to create a characteristic grayscale shade within the final image. By using “fixed bin width,” the CT data is consistently mapped to a fixed number of grayscale levels. This standardizes image quality and facilitates the comparison and analysis of CT images obtained at different times or with different scanners. In some embodiments, the images are not normalized. Features are extracted from ROIs in the original image type and the transformed image type.
[0055] Referring to block 214, in some embodiments, the multiple parameters in the ensemble model are at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more parameters. In some embodiments, the multiple parameters are 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer parameters. In some embodiments, the multiple parameters are 10,000 to 100 billion, 100,000 to 100 billion, 1 million to 100 billion, 10 million to 100 billion, 10,000 to 10 billion, 100,000 to 10 billion, 1 million to 10 billion, 10 million to 10 billion, 10,000 to 1 billion, 100,000 to 1 billion, 1 million to 1 billion, 10,000 to 100 million, 100,000 to 100 million, 1 million to 100 million, or 10 million to 100 million parameters.
[0056] An ensemble model applies multiple parameters to information through multiple calculations to generate multiple component predictions about cancer status as an output from the model. In some embodiments, the multiple calculations are at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, at least 1 billion, or more. In some embodiments, the multiple calculations are 100 billion or less, 10 billion or less, 1 billion or less, 100 million or less, 10 million or less, or fewer. In some embodiments, the multiple calculations are performed 10,000 to 100 billion times, 100,000 to 100 billion times, 1 million to 100 billion times, 10 million to 100 billion times, 10,000 to 10 billion times, 100,000 to 10 billion times, 1 million to 10 billion times, 10 million to 10 billion times, 10,000 to 1 billion times, 10 million to 1 billion times, 10,000 to 100 million times, 100,000 to 100 million times, 1 million to 100 million times, or 10 million to 100 million times.
[0057] Referring to block 216 in Figure 2C, in some embodiments, cancerous conditions include carcinoma, lymphoma, blastoma, glioblastoma, sarcoma, leukemia, breast cancer, squamous cell carcinoma, lung cancer, small cell lung cancer, non-small cell lung cancer (NSCLC), adenocarcinoma of the lung, squamous cell carcinoma of the lung, head and neck cancer, peritoneal cancer, hepatocellular carcinoma, gastric or stomach cancer, pancreatic cancer, ovarian cancer, cervical cancer, liver cancer, bladder cancer, liver cancer, colon cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal cancer, liver cancer, prostate cancer, vulvar cancer, This group of cancers includes thyroid cancer, hepatocellular carcinoma, B-cell lymphoma, low-grade / follicular non-Hodgkin lymphoma (NHL), small lymphocytic (SL) NHL, intermediate-grade / follicular NHL, intermediate-grade diffuse NHL, high-grade immunoblastic NHL, high-grade lymphoblastic NHL, high-grade small undivided cell NHL, giant tumor NHL, mantle cell lymphoma, AIDS-associated lymphoma, Waldenström macroglobulinemia, chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), hairy cell leukemia, and chronic myeloblastic leukemia.
[0058] Referring to block 218, in some embodiments, the multiple component models are at least 10 component models.
[0059] Referring to block 220, in some embodiments, the multiple component models are at least 20 component models.
[0060] Referring to block 222, in some embodiments, the multiple component models are at least 40 component models.
[0061] Referring to block 224, in some embodiments, the multiple component models are 250 or fewer component models.
[0062] Referring to block 226, in some embodiments, the multiple component models are 100 or fewer component models.
[0063] Referring to block 228 in Figure 2D, in some embodiments, the number of component models is 50 or less.
[0064] Referring to block 230, in some embodiments, the multiple component models are 5 to 100 component models.
[0065] Referring to block 232, in some embodiments, the multiple component models are 10 to 75 component models.
[0066] Referring to block 234, in some embodiments, the multiple component models are 20 to 50 component models.
[0067] As a non-limiting example, in some embodiments, the multiple component models include 2 to 250 component models, 2 to 150 component models, 2 to 100 component models, 2 to 150 component models, 2 to 100 component models, 2 to 75 component models, 2 to 50 component models, 2 to 30 component models, 2 to 20 component models, 2 to 15 component models, 2 to 10 component models, 2 to 5 component models, 5 to 250 component models, 5 to 150 component models, 5 to 100 component models, 5 to 150 component models, 5 to 75 component models, 5 to 50 component models, 5 to 30 component models, 5 to 20 component models, 5 to 15 component models, and 5 to 10 component models. Includes component models, 15-150 component models, 15-100 component models, 15-150 component models, 15-100 component models, 15-75 component models, 15-50 component models, 15-30 component models, 15-20 component models, 35-150 component models, 35-100 component models, 35-150 component models, 35-100 component models, 35-75 component models, 35-50 component models, 65-150 component models, 65-100 component models, 65-150 component models, 65-100 component models, 65-75 component models, 120-150 component models, 120-250 component models, or 120-150 component models.In some embodiments, the multiple component models are at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, or more component models. In some embodiments, the multiple component models are up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, up to 10, up to 15, up to 20, up to 25, up to 30, up to 35, up to 40, up to 45, up to 50, up to 75, up to 100, up to 150, up to 200, up to 250, or more component models.
[0068] Referring to block 236, in some embodiments, the individual component models within a multi-component model are neural networks, support vector machines, naive Bayes models, nearest neighbor models, boosted tree models, random forest models, or clustering models.
[0069] As used herein, the term "model" refers to a machine learning model, algorithm, or task.
[0070] In some embodiments, the model is an unsupervised learning algorithm. One example of an unsupervised learning algorithm is cluster analysis.
[0071] In some embodiments, the model is supervised machine learning. Non-exclusive examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, naive Bayes algorithms, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosted tree algorithms, multinomial logistic regression algorithms, linear models, linear regression, gradient boosting, mixture models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, or any combination thereof. In some embodiments, the model is a multinomial classifier algorithm. In some embodiments, the model is a two-stage stochastic gradient descent (SGD) model. In some embodiments, the model is a deep neural network (e.g., a wide and deep sample-based classifier).
[0072] In some embodiments, a model is used to normalize values or datasets, such as by transforming values or sets of values into a common reference frame for comparison purposes. For example, in some embodiments, when one or more pixel values corresponding to one or more pixels in individual images are normalized against a predetermined statistic (e.g., the mean and / or standard deviation of one or more pixel values across one or more images), the pixel values of individual pixels are compared to the individual statistic to determine how much the pixel values differ from the statistic.
[0073] In some embodiments, an untrained model (e.g., an "untrained classifier" and / or an "untrained neural network") includes a machine learning model or algorithm, e.g., a classifier or a neural network, that has not been trained on a target dataset. In some embodiments, training a model (e.g., training a neural network) refers to the process of training an untrained or partially trained model (e.g., an untrained or partially trained neural network). For example, consider the case of multiple training samples, each comprising a group of corresponding medical images (e.g., a medical dataset). The multiple medical images, along with corresponding measured metrics for one or more features for each individual medical image (hereinafter referred to as the training dataset), are applied as a collective input to an untrained or partially trained model to train the untrained or partially trained model with metrics that identify features relevant to morphological classes, thereby obtaining a trained model. Furthermore, it will be understood that the term "untrained model" does not preclude the possibility that transfer learning techniques may be used for such training of untrained or partially trained models. For example, Fernandes et al., 2017, “Transfer Learning with Partial Observability Applied to Cervical Cancer Screening,” Pattern Recognition and Image Analysis: 8 thIberian Conference Proceedings, 243-250, provides a non-limiting example of such transfer learning, which is incorporated herein by reference in its entirety for all purposes. In the examples in which transfer learning is used, the untrained model described above is provided with additional data beyond the data in the primary training dataset. That is, in a non-limiting embodiment of the embodiment of transfer learning, the untrained model receives (i) multiple images and measured metrics for each individual image ("primary training dataset"), and (ii) additional data. In some embodiments, this additional data is in the form of parameters (e.g., coefficients, weights, and / or hyperparameters) learned from another auxiliary training dataset. Furthermore, while a description of a single auxiliary training dataset is disclosed, it will be understood that there is no limit to the number of auxiliary training datasets that may be used to complement the primary training dataset when training the untrained model in this disclosure. For example, in some embodiments, two or more auxiliary training datasets, three or more auxiliary training datasets, four or more auxiliary training datasets, or five or more auxiliary training datasets may be used to complement the primary training dataset through transfer learning, each of which auxiliary datasets is different from the primary training dataset. In such embodiments, any form of transfer learning may be used. For example, consider the case where, in addition to the primary training dataset, there are a first auxiliary training dataset and a second auxiliary training dataset. Parameters learned from the first auxiliary training dataset (by applying the first model to the first auxiliary training dataset) can be applied to the second auxiliary training dataset using transfer learning techniques (e.g., a second model identical or different to the first model), which may then yield a trained intermediate model whose parameters are then applied to the primary training dataset, and this, together with the primary training dataset itself, is applied to the untrained model.Alternatively, a first set of parameters learned from a first auxiliary training dataset (by applying a first model to the first auxiliary training dataset) and a second set of parameters learned from a second auxiliary training dataset (by applying a second model, identical or different from the first model, to the second auxiliary training dataset) may each be individually applied to distinct instances of the primary training dataset (e.g., by multiplication of distinct and independent matrices), and together with the primary training dataset itself (or a partially reduced form of the primary training dataset, such as principal components or regression coefficients learned from the primary training set), both such applications of these parameters may then be applied to the untrained model to train the untrained model. In some cases, additionally or alternatively, knowledge of objects related to morphological classes derived from auxiliary training datasets may be used, in conjunction with objects and / or class-labeled images in the primary training dataset, to train the untrained model.
[0074] Support Vector Machines. In some embodiments, the model is a Support Vector Machine (SVM). Suitable SVM algorithms for use as models include, for example, Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York. As described in York and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in whole for all purposes. When used for classification, SVM separates a given set of binary labeled data by a hyperplane that is as far away from the labeled data as possible. In cases where linear separation is not possible, SVM can work in combination with a “kernel” technique that automatically provides a nonlinear mapping to the feature space. The hyperplane detected by SVM in the feature space may correspond to a nonlinear decision boundary in the input space.In some embodiments, multiple parameters (e.g., weights) associated with the SVM define a hyperplane. In some embodiments, the hyperplane is defined by at least 10, at least 20, at least 50, or at least 100 parameters, and the SVM model needs to be computed by a computer because it cannot be solved by human intelligence.
[0075] Naive Bayes Algorithms. In some embodiments, the model is a naive Bayes algorithm. A naive Bayes classifier suitable for use as a model is disclosed, for example, in Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes,” Advances in Neural Information Processing Systems, 14, which is incorporated herein by reference in its entirety for all purposes. A naive Bayes classifier is a classifier of one of the “probabilistic classifier” family based on applying Bayes’ theorem with strong (naive) independence assumptions between features. In some embodiments, they are combined with kernel density estimation. For example, see Hastie et al., 2001, The elements of statistical learning: data mining, inference, and prediction, eds. Tibshirani and Friedman, Springer, New York, which is incorporated herein by reference in its entirety for all purposes.
[0076] Nearest Neighbor Algorithm. In some embodiments, the model is a nearest neighbor algorithm. The nearest neighbor model can be a memory-based model and may not contain the model being fitted. With respect to the nearest neighbor points, consider the query point x0 (first image) and the k training points x that are closest to x0. (r),r,...,k (here, the training images) are identified, and then point x0 is classified using these k nearest neighbors. In some embodiments, the distance to these neighbors is a function of the values in the discriminant set. In some embodiments, the Euclidean distance in the feature space is
number
[0077] The k-nearest neighbor model is a nonparametric machine learning method in which the input consists of the k closest training examples in the feature space. Its output is class membership. An object is classified by multiple votes from its neighbors and assigned to the most common class among its k nearest neighbors (k is typically a small positive integer). When k=1, the object is simply assigned to the class of its single nearest neighbor. See Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, which is incorporated herein by reference in its entirety for all purposes. In some embodiments, the number of distance calculations required to solve the k-nearest neighbor model is such that a computer is used to solve the model for a given input because it cannot be performed by human effort.
[0078] Random forest algorithms, decision tree algorithms, and boosted tree algorithms. In some embodiments, the model is a decision tree. Suitable decision trees for use as models are generally described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395–396, which is incorporated herein by reference in its entirety for all purposes. In the tree-based method, the feature space is divided into a set of rectangles, and then a model (such as a constant) is fitted into each rectangle. In some embodiments, the decision tree is a random forest regression. One specific algorithm that can be used is the classification and regression tree (CART). Other specific decision tree algorithms, but not limited to these, include ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396–408 and pp. 411–412, which are incorporated herein by reference in their entirety for all purposes. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which are incorporated herein by reference in their entirety for all purposes. Random forests are described in Breiman, 1999, “Random Forests--Random Features,” Technical Report 567, Statistics Department, UC Berkeley, September 1999, which are incorporated herein by reference in their entirety for all purposes.In some embodiments, the decision tree model includes at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions) that need to be computed by a computer because they cannot be solved by human intelligence.
[0079] Linear Discriminant Analysis Algorithms. Linear discriminant analysis (LDA), normal discriminant analysis (NDA), or discriminant function analysis is a generalization of Fisher's linear discriminant formula and may be a method used in statistics, pattern recognition, and machine learning to find a linear combination of features that characterize or separate two or more classes of objects or events. The resulting combination can be used as a model (e.g., a linear classifier) in some embodiments of this disclosure.
[0080] Mixed models and hidden Markov models. In some embodiments, the model is a mixed model, such as that described in McLachlan et al., Bioinformatics 18(3):413-422, 2002. In some embodiments, particularly those including a temporal component, the model is a hidden Markov model, such as that described in Schliep et al., 2003, Bioinformatics 19(1):i255-i263.
[0081] Clustering. In some embodiments, the model is an unsupervised clustering model. In some embodiments, the model is a supervised clustering model. Suitable clustering algorithms for use as models are described, for example, on pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter, "Duda 1973"), which is incorporated herein by reference in its entirety for all purposes. The task of clustering can be described as one of discovering natural classifications within a dataset. Two tasks can be addressed in order to identify natural classifications. First, a means can be determined for measuring the similarity (or dissimilarity) between two samples. This metric (e.g., a measure of similarity) can be used to ensure that samples in one cluster are more similar to each other than to samples in other clusters. Second, a mechanism can be determined for dividing the data into clusters using the measure of similarity. One means of starting a clustering investigation is to define a distance function and compute a matrix of distances between all pairs of samples in the training dataset. If distance is a good measure of similarity, then the distance between reference entities within the same cluster may be significantly shorter than the distance between reference entities in different clusters. However, clustering may not always use a distance metric. For example, a nonmetric similarity function s(x,x') can be used to compare two vectors x and x'. s(x,x') can be a symmetric function, and its value is large when x and x' are "similar" in some respects. Once a method for measuring the "similarity" or "dissimilarity" between points in a dataset has been chosen, clustering can use a criterion function to measure the clustering quality of any segment of the data. Segments of the dataset that cause the criterion function to reach extremes can be used to cluster the data.Specific exemplary clustering techniques that may be used in this disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor algorithms, longest distance algorithms, group average algorithms, centroid algorithms, or sum-of-squares algorithms), k-means clustering, fuzzy k-means clustering algorithms, and Jarvis-Patrick clustering. In some embodiments, the clustering includes unsupervised clustering (e.g., the number of clusters is not predetermined and / or cluster assignments are not predetermined).
[0082] Model ensembles and boosting. In some embodiments, an ensemble of models (two or more) is used. In some embodiments, boosting techniques such as AdaBoost are used in conjunction with many other types of learning algorithms to improve the performance of the models. In this approach, the outputs of any of the models disclosed herein, or their equivalents, are combined into a weighted sum that represents the final output of the boosted models. In some embodiments, multiple outputs from the models are combined using any measure of central tendency known in the art, including, but not limited to, the mean, median, mode, weighted mean, weighted median, weighted mode, etc. In some embodiments, multiple outputs are combined using a voting method. In some embodiments, the individual models within the ensemble of models are weighted or unweighted.
[0083] The term “classification” can refer to any number or other letter related to a particular characteristic of a sample. For example, a “+” sign (or the word “plus”) may indicate that a sample is classified as having a desired outcome or characteristic, while a “-” sign (or the word “minus”) may indicate that a sample is classified as having an undesirable outcome or characteristic. In another embodiment, the term “classification” refers to individual outcomes or characteristics (e.g., high risk, medium risk, low risk). In some embodiments, classifications are binary (e.g., plus or minus) or have more classification levels (e.g., a scale of 1 to 10 or 0 to 1). In some embodiments, the terms “cutoff” and “threshold” refer to predetermined numbers used in calculations. In one embodiment, the cutoff value refers to a value above which outcomes are excluded. In some embodiments, the threshold may be a value above or below which a particular classification is applied. Any of these terms can be used in any of these contexts.
[0084] Those skilled in the art will readily understand other models applicable to the systems and methods of this disclosure. In some embodiments, the systems, methods, and devices of this disclosure utilize two or more models to provide evaluations with higher accuracy (e.g., given one or more inputs, to arrive at a certain evaluation). For example, in some embodiments, each individual model arrives at a corresponding evaluation when individual datasets are provided. Thus, each individual model may independently arrive at a certain result, and the results of each individual model are then collectively verified by comparison or fusion of the models. This results in cumulative results being produced by the models, but this disclosure is not limited thereto.
[0085] In some embodiments, individual models are tasked with performing corresponding activities. In some embodiments, but not limited to, tasks performed by individual models include, but are not limited to, extracting corresponding values for each individual radiomic feature within an individual radiomic feature class (e.g., block 202 in Figure 2A, block 211 in Figure 2A), providing multiple component predictions (e.g., block 212 in Figure 2B), combining multiple component predictions (e.g., block 2148 in Figure 2M), assigning therapies to subjects (e.g., block 2180 in Figure 2O), or a combination thereof. In some embodiments, each individual model of the Disclosure utilizes 10 or more parameters, 100 or more parameters, 1,000 or more parameters, 10,000 or more parameters, or 100,000 or more parameters. In some embodiments, each individual model of the Disclosure is not performed using brainpower.
[0086] Referring to Block 238, in some embodiments, each individual component model within a multi-component model includes an individual neural network (e.g., a convolutional neural network and / or residual neural network). Examples of neural network algorithms, also known as artificial neural networks (ANNs), include convolutional and / or residual neural network algorithms (deep learning algorithms). A neural network can be a machine learning algorithm that can be trained to map an input dataset to an output dataset, and a neural network includes a group of interconnected nodes organized into layers of multiple nodes. For example, a neural network architecture may include at least an input layer, one or more hidden layers, and an output layer. A neural network may include any total number of layers and any number of hidden layers, the hidden layers acting as trainable feature extractors that enable mapping a set of input data to output values or a set of output values. As used herein, a deep learning algorithm (DNN) can be a neural network including multiple hidden layers, e.g., two or more hidden layers. Each layer of a neural network may include a number of nodes (or "neurons"). A node can receive input either directly from the input data or from the output of a node in the previous layer to perform a specific operation, such as addition. In some embodiments, the connection from the input to the node is associated with parameters (e.g., weights and / or weight coefficients). In some embodiments, the node receives input x iThe products of all pairs and their associated parameters can be summed. In some embodiments, the weighted sum is offset by a bias b. In some embodiments, the output of a node or neuron can be restricted using a threshold or an activation function f, which may be a linear or nonlinear function. The activation function may be, for example, a normalized linear unit (ReLU) activation function, a Leaky ReLU activation function, or other functions such as a saturated hyperbolic tangent function, an identity function, a binary step function, a logistic function, an arctangent function, a soft sine function, a parametric normalized linear unit function, an exponential linear unit function, a soft plus function, a Bent identity function, a soft exponential function, a sine wave function, a sine function, a Gaussian function, or a sigmoid function, or any combination thereof.
[0087] The weight coefficients, bias values, and thresholds, or other computational parameters of a neural network, can be "taught" or "learned" during the training phase using one or more sets of training data. For example, parameters can be trained using input data from a training dataset and gradient descent or backpropagation so that the output values computed by the ANN match the examples contained in the training dataset. Parameters can be obtained from the learning process of a backpropagation neural network.
[0088] Any of the various neural networks may be suitable for use in carrying out the methods disclosed herein. Examples, but not limited to, include feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, or any combination thereof. In some embodiments, machine learning utilizes a pre-trained and / or transfer-trained ANN or deep learning architecture. Convolutional and / or residual neural networks may be used to analyze images of interest in accordance with this disclosure.
[0089] For example, a deep neural network model includes an input layer, several individually parameterized (e.g., weighted) convolutional layers, and an output scorer. Each parameter (e.g., weight) of the convolutional layer, as well as the input layer, contributes to several parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network model. Thus, a computer is needed because the deep neural network model cannot be solved by the human brain. In other words, considering the input to the model, the model output needs to be determined using a computer rather than the human brain in such embodiments. For example, see Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 2, Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc., Zeiler, 2012 “ADADELTA: an adaptive learning rate method,” CoRR, vol. abs / 1212.5701, and Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is incorporated herein by reference in its entirety for all purposes.
[0090] Neural network algorithms, including convolutional neural network algorithms suitable for use as models, are disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408, Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40, and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference in its entirety for all purposes. Additional exemplary neural networks suitable for use as models are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is incorporated herein by reference in its entirety for all purposes. Additional exemplary neural networks suitable for use as models are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: Sequence and Genome Analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety for all purposes.
[0091] Referring to block 240, in some embodiments, each individual component model within a multi-component model includes an individual logistic regression model. In some embodiments, the component model performs a regression task, which is any type of regression. For example, in some embodiments, the regression task is a logistic regression task. In some embodiments, the regression task is Lasso regression (L1), Ridge regression (L2), or logistic regression with elastic network regularization. In some embodiments, extracted features with corresponding regression coefficients that do not satisfy a threshold are pruned (removed) from consideration. In other words, they are not included in the final trained regression model (e.g., the final trained component model). In some embodiments, a generalization of the logistic regression model that handles multi-category responses is used as the component model. For example, a component model may be a multi-category model that provides finer-grained details about cancer status, such as probability bins (e.g., bin 1: subject has a 0-20% probability of having cancer, bin 2: subject has a 20-40% probability of having cancer) or cancer stage (e.g., bin 1: subject does not have cancer, bin 2: subject has stage I cancer, bin 3: subject has stage II cancer). The logistic regression task is disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated herein by reference in its entirety for all purposes. In some embodiments, the component models of this disclosure utilize the regression task disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York.In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights) and needs to be computed by a computer because it cannot be solved by human intellect.
[0092] In some embodiments, the logistic regression model is a type I regression model. For example, in some embodiments, the type I regression model is configured (e.g., trained) to have one or more independent parameters and one dependent parameter as input parameters. Thus, in some such embodiments, the type I regression model is configured to evaluate the variation of one dependent parameter in response to changes in individual parameters among the one or more independent parameters. In some embodiments, the logistic regression model is a type II regression model. For example, in some embodiments, the type II regression model is configured to have two or more dependent parameters. Thus, in some such embodiments, the type II regression model is configured to evaluate the variation of both the first and second dependent parameters among the two or more dependent parameters in response to changes in a third parameter, such as an unknown parameter. Further details and information regarding the logistic regression model can be found in “Package 'lmode2'” by Legendre et al., 2018, which is available for print at cran.r-project.org / web / packages / lmodel2 / lmodel2.pdf (accessed July 12, 2023), and is incorporated herein by reference in its entirety for all purposes.
[0093] Referring to block 242 in Figure 2E, in some embodiments, the multiple radiomic feature classes include a first subset of radiomic feature classes extracted from unfiltered versions of multiple medical images in a medical image dataset, and a second subset of radiomic feature classes extracted from filtered versions of multiple medical images in a medical image dataset filtered by a first filtering method (e.g., via an optional filter module 36).
[0094] Referring to block 244, in some embodiments, the first filtering technique includes an image filter selected from the group consisting of wavelet transform filters, Gaussian Laplacian (LoG) filters, square transform filters, square root transform filters, logarithmic transform filters, exponential transform filters, gradient transform filters, two-dimensional local binary pattern filters, and three-dimensional local binary pattern filters.
[0095] For example, in some embodiments, each individual image filter from the group consisting of wavelet transform filters, LoG filters, square transform filters, square root transform filters, logarithmic transform filters, exponential transform filters, gradient transform filters, 2D local binary pattern filters, and 3D local binary pattern filters is configured to produce a corresponding unique filtered version of an individual medial image according to a corresponding function associated with the individual image filter. For example, in some embodiments, each image filter is configured to produce filtered versions of all or some of a plurality of medical images according to a corresponding spatial domain function applied to all or some of the individual medical images of a plurality of medical images. In some embodiments, each image filter is configured to produce filtered versions of all or some of a plurality of medical images according to a corresponding transform domain function applied to all or some of the individual medical images of a plurality of medical images. As a non-limiting example, in some embodiments, a logarithmic transform filter is configured to produce a first corresponding unique filtered medical image that expands (e.g., enhances) the values of dark pixels according to a corresponding function applied to a first medical image. As another non-limiting example, in some embodiments, the first LoG filter is configured to generate a second corresponding unique filtered medical image in which one or more points, one or more regions of interest, or one or more volumes of interest are detected based on one or more derivative representations according to a corresponding function applied to the first medical image. As yet another non-limiting example, in some embodiments, the second LoG filter is configured to generate a third corresponding unique filtered medical image in which one or more points, one or more regions of interest, or one or more volumes of interest are detected based on one or more local intensity extrema according to a corresponding function applied to the first medical image.Further details and information regarding image filters can be found in Singh, P., 2019, “Feature enhanced Speckle Reduction in Ultrasound Images: Algorithms for Scan Modelling, Speckle Filtering, Texture Analysis and Feature Improvement,” Doctoral dissertation, print; Bhoi, N, 2009, “Development of Some Novel Spatial-Domain and Transform-Domain Digital Image Filters,” Doctoral dissertation, print; and Cheng et al., 2003, “Computer-aided detection and classification of microcalcifications in mammograms: a survey,” Pattern Recognition, 36(12), pp. 2967-2991, each of which is incorporated herein by reference in whole for all purposes.
[0096] Therefore, in some embodiments, each individual image filter is configured to produce a unique filtered version of an individual medical image within a group of medical images in a medical dataset. Furthermore, in some such embodiments, the filtered versions of individual medical images uniquely generated by each individual image filter are provided as inputs to each individual component model within a group of component models of an ensemble model, thereby enabling the ensemble model to process richer information compared to processing only raw, unfiltered medical images.
[0097] Referring to block 246, in some embodiments, the first subset of radiomic feature classes includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, the first subset of radiomic feature classes includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, the first subset of radiomic feature classes includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0098] In some embodiments, the second subset of radiomic feature classes includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, the second subset of radiomic feature classes includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, the second subset of radiomic feature classes includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0099] Referring to block 248, in some embodiments, a first subset of radiomics feature classes includes radiomics feature classes selected from the group consisting of shape features, linear features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighbor gray-tone difference matrix (NGTDM) features.
[0100] Referring to block 250, in some embodiments, the first subset of the radiomics feature class includes one or more shape features, one or more linear features, one or more gray-level co-occurrence matrix (GLCM) features, one or more gray-level run-length matrix (GLRLM) features, one or more gray-level size zone matrix (GLSZM) and one or more gray-level dependency matrix (GLDM) features.
[0101] In some embodiments, each individual shape feature is configured to evaluate the n-dimensional (e.g., two-dimensional or three-dimensional) size and / or shape of the volume or region of interest in a medical image.
[0102] In some embodiments, the first subset of shape features in the radiomics feature class includes some or all of the shape features listed in Table 1. For example, in some embodiments, the first subset of shape features in the radiomics feature class includes, or at least indicates, all of the shape features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, or 44 listed in Table 1.
[0103] In some embodiments, the shape features of a second subset of the radiomics feature class include some or all of the shape features listed in Table 1. For example, in some embodiments, the shape features of a second subset of the radiomics feature class include, or at least indicate, all of the shape features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, or 44 listed in Table 1. [Table 1-1] [Table 1-2]
[0104] In some embodiments, each individual primary feature is configured to evaluate the intensity distribution within the region of interest or volume of interest of the medial image according to a corresponding function.
[0105] In some embodiments, the primary features of the first subset of the radiomics feature class include some or all of the primary features listed in Table 2. For example, in some embodiments, the primary features of the first subset of the radiomics feature class include, or at least indicate, all of the primary features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 of the primary features listed in Table 2.
[0106] In some embodiments, the primary features of a second subset of the radiomics feature class include some or all of the primary features listed in Table 2. For example, in some embodiments, the primary features of a second subset of the radiomics feature class include, or at least represent, all of the primary features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 of the primary features listed in Table 2.
[0107] [Table 2]
[0108] In some embodiments, each individual GLCM feature is configured to evaluate a quadratic joint probability function of the region of interest or volume of interest of the medial image according to the corresponding function.
[0109] In some embodiments, the GLCM features of the first subset of the radiomics feature class include some or all of the GLCM features listed in Table 3. For example, in some embodiments, the GLCM features of the first subset of the radiomics feature class include, or at least represent, all of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 of the GLCM features listed in Table 3.
[0110] In some embodiments, the GLCM features of a second subset of the radiomics feature class include some or all of the GLCM features listed in Table 3. For example, in some embodiments, the GLCM features of a second subset of the radiomics feature class include, or at least represent, all of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 of the GLCM features listed in Table 3. [Table 3]
[0111] In some embodiments, each individual GLSZM feature is configured to evaluate one or more gray-level zones or gray-level regions of a medial image according to a corresponding function.
[0112] In some embodiments, the GLSZM features of the first subset of the radiomics feature class include some or all of the GLSZM features listed in Table 4. For example, in some embodiments, the GLSZM features of the first subset of the radiomics feature class include, or at least represent, all of GLSZM features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 of the GLSZM features listed in Table 4.
[0113] In some embodiments, the GLSZM features of a second subset of the radiomics feature class include some or all of the GLSZM features listed in Table 4. For example, in some embodiments, the GLSZM features of a second subset of the radiomics feature class include, or at least represent, all of GLSZM features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 of the GLSZM features listed in Table 4. [Table 4]
[0114] In some embodiments, each individual GLRLM feature is configured to evaluate the gray run level, which is the length of multiple consecutive pixels having a first gray level value with respect to the medial image, according to a corresponding function.
[0115] In some embodiments, the GLRLM features of the first subset of the radiomics feature class include some or all of the GLRLM features listed in Table 5. For example, in some embodiments, the GLRLM features of the first subset of the radiomics feature class include, or at least represent, all of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 of the GLRLM features listed in Table 5.
[0116] In some embodiments, the GLRLM features of a second subset of the radiomics feature class include some or all of the GLRLM features listed in Table 5. For example, in some embodiments, the GLRLM features of a second subset of the radiomics feature class include, or at least represent, all of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 of the GLRLM features listed in Table 5. [Table 5]
[0117] In some embodiments, each individual GLDM feature is configured to evaluate the gray level dependency of the medial image according to a corresponding function.
[0118] In some embodiments, the GLDM features of the first subset of the radiomics feature class include some or all of the GLDM features listed in Table 6. For example, in some embodiments, the GLDM features of the first subset of the radiomics feature class include, or at least represent, all of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 of the GLDM features listed in Table 6.
[0119] In some embodiments, the GLDM features of a second subset of the radiomics feature class include some or all of the GLDM features listed in Table 6. For example, in some embodiments, the GLDM features of a second subset of the radiomics feature class include, or at least represent, all of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 of the GLDM features listed in Table 6. [Table 6]
[0120] In some embodiments, each individual NGTDM feature is configured to evaluate the difference between a given gray value and the mean gray values adjacent to it within a first distance of the medial image, according to a corresponding function.
[0121] In some embodiments, the NGTDM features of the first subset of the radiomics feature class include some or all of the NGTDM features listed in Table 7. For example, in some embodiments, the NGTDM features of the first subset of the radiomics feature class include, or at least represent, all of NGTDM features 1, 2, 3, 4, or 5 listed in Table 7.
[0122] In some embodiments, the NGTDM features of a second subset of the radiomics feature class include some or all of the NGTDM features listed in Table 7. For example, in some embodiments, the NGTDM features of a second subset of the radiomics feature class include, or at least represent, all of NGTDM features 1, 2, 3, 4, or 5 listed in Table 7. [Table 7]
[0123] In some embodiments, the first subset of the radiomics feature class includes some or all of the features listed in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, or a combination thereof. For example, in some embodiments, the first subset of the radiomics feature class includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 of the features listed in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, or a combination thereof. ,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75 ,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118, Includes all of, or at least indicates, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150.
[0124] In some embodiments, a second subset of the radiomics feature class includes some or all of the features listed in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, or a combination thereof. For example, in some embodiments, a second subset of the radiomics feature class includes features 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 of the features listed in Table 1, Table 2, Table 3, Table 4, Table 5, Table 6, Table 7, or a combination thereof. ,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75 ,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118, Includes all of, or at least indicates, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, or 150.
[0125] Additional details and information on multiple radiomic feature classes are available in van Griethuysen et al., 2017, “Computational Radiomics System to Decode the Radiographic Phenotype,” Cancer Research, 77(21), pg.e104-e107, Pyradiomics Community, “Radiomic Features,” available at pyradiomics.readthedocs.io / en / latest / features.html# (accessed July 12, 2023), Davatziko et al., 2018, “Cancer Imaging Phenomics Toolkit: Quantitative Imaging Analytics for Precision Diagnostics and Predictive Modeling of Clinical Outcome,” J.Med.Imaging, 5(1), pg.011018, Pati et al., 2019, “The Cancer Imaging Phenomics Toolkit (CaPTk): Technical Overview,” BrainLes, Springer LNCS, (11993), pg. 380-394, available at cbica.github.io / CaPTk / tr_Apps.html#appsFeatures (accessed July 12, 2023): University of Pennsylvania, Center for Biomedical Image Computing & Analytics, “Cancer Imaging Phenomics Toolkit”, Ashrafinia, S., 2019, “Quantitative Nuclear Medicine Imaging using Advanced Image Reconstruction and Radiomics”, Ph.D. Dissertation, Johns Hopkins University, print, Zwanenburg et al.See also the 2017, “Image Biomarker Standardisation Initative,” arXiv preprint arXiv:1612.07003, and the Standardized Environment for Radiomics Analysis (SERA), 2019, “SERA Feature Names and Benchmarks,” available at github.com / ashrafinia / SERA / tree / master / Feature%20Names%20and%20Benchmarks (accessed July 12, 2023), each of which is incorporated herein by reference in its entirety for all purposes.
[0126] Referring to block 252, in some embodiments, each individual radiomic feature class within a first subset of radiomic feature classes includes at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0127] Referring to block 254, in some embodiments, each individual radiomic feature class within a first subset of radiomic feature classes contains 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0128] In some embodiments, each individual radiomic feature class within a first subset of radiomic feature classes comprises 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, 10 to 50 corresponding radiomic features, 75 to 750 corresponding radiomic features, Includes 75-500 corresponding radiomic features, 75-250 corresponding radiomic features, 75-100 corresponding radiomic features, 175-750 corresponding radiomic features, 175-500 corresponding radiomic features, 175-250 corresponding radiomic features, 375-750 corresponding radiomic features, 375-500 corresponding radiomic features, or 600-750 corresponding radiomic features.
[0129] Referring to block 256 in Figure 2F, in some embodiments, the multiple radiomic feature classes include a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter.
[0130] Referring to block 258, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0131] Referring to block 260, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter includes radiomic feature classes selected from the group consisting of first-order features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
[0132] Referring to block 262, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
[0133] Referring to block 264, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter contains at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0134] Referring to block 266, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter contains 1000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0135] In some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a first wavelet transform filter contains 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, and 10 to 50 corresponding radiomic features. The quantity includes 75-750 corresponding radiomic features, 75-500 corresponding radiomic features, 75-250 corresponding radiomic features, 75-100 corresponding radiomic features, 175-750 corresponding radiomic features, 175-500 corresponding radiomic features, 175-250 corresponding radiomic features, 375-750 corresponding radiomic features, 375-500 corresponding radiomic features, or 600-750 corresponding radiomic features.
[0136] Referring to block 268 in Figure 2G, in some embodiments, the multiple radiomic feature classes include a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a second wavelet transform filter.
[0137] Referring to block 270, in some embodiments, a subset of the radiomic feature class extracted from multiple versions of medical images filtered using a second wavelet transform filter includes 3 to 5 radiomic features, 3 to 4 radiomic features, or 4 to 5 radiomic features. In some embodiments, a subset of the radiomic feature class extracted from multiple versions of medical images filtered using a second wavelet transform filter includes at least 3, at least 4, or at least 5 radiomic features. In some embodiments, a subset of the radiomic feature class extracted from multiple versions of medical images filtered using a second wavelet transform filter includes up to 3, up to 4, or up to 5 radiomic features.
[0138] Referring to block 272, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a second wavelet transform filter includes radiomic feature classes selected from the group consisting of first-order features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
[0139] Referring to block 274, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a second wavelet transform filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
[0140] Referring to block 276, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a second wavelet transform filter contains at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0141] Referring to block 278, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a second wavelet transform filter contains 1000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0142] In some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a second wavelet transform filter contains 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, and 10 to 50 corresponding radiomic features. The quantity includes 75-750 corresponding radiomic features, 75-500 corresponding radiomic features, 75-250 corresponding radiomic features, 75-100 corresponding radiomic features, 175-750 corresponding radiomic features, 175-500 corresponding radiomic features, 175-250 corresponding radiomic features, 375-750 corresponding radiomic features, 375-500 corresponding radiomic features, or 600-750 corresponding radiomic features.
[0143] Referring to block 280 in Figure 2H, in some embodiments, the multiple radiomic feature classes include a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter.
[0144] Referring to block 282, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0145] Referring to block 284, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter includes radiomic feature classes selected from the group consisting of primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
[0146] Referring to block 286, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
[0147] Referring to block 288, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter contains at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0148] Referring to block 290, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter contains 1000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0149] In some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a LoG filter contains 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, 10 to 50 corresponding radiomic features, and 75 to Includes 750 corresponding radiomic features, 75 to 500 corresponding radiomic features, 75 to 250 corresponding radiomic features, 75 to 100 corresponding radiomic features, 175 to 750 corresponding radiomic features, 175 to 500 corresponding radiomic features, 175 to 250 corresponding radiomic features, 375 to 750 corresponding radiomic features, 375 to 500 corresponding radiomic features, or 600 to 750 corresponding radiomic features.
[0150] Referring to block 292 in Figure 2I, in some embodiments, the multiple radiomic feature classes include a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter.
[0151] Referring to block 294, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0152] Referring to block 296, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter includes radiomic feature classes selected from the group consisting of first-order features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
[0153] Referring to block 298, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
[0154] Referring to block 2100, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transform filter contains at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0155] Referring to block 2102, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transform filter contains 1000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0156] In some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a logarithmic transformation filter contains 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, 10 to 50 corresponding radiomic features, and 75 Includes up to 750 corresponding radiomic features, 75 to 500 corresponding radiomic features, 75 to 250 corresponding radiomic features, 75 to 100 corresponding radiomic features, 175 to 750 corresponding radiomic features, 175 to 500 corresponding radiomic features, 175 to 250 corresponding radiomic features, 375 to 750 corresponding radiomic features, 375 to 500 corresponding radiomic features, or 600 to 750 corresponding radiomic features.
[0157] Referring to block 2104 in Figure 2J, in some embodiments, the multiple radiomic feature classes include a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential transformation filter.
[0158] Referring to block 2106, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential filter includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential filter includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential filter includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0159] Referring to block 2108, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential transformation filter includes radiomic feature classes selected from the group consisting of first-order features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighbor gray-tone difference matrix (NGTDM) features.
[0160] Referring to block 2110, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential transformation filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
[0161] Referring to block 2112, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential transform filter contains at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0162] Referring to block 2114, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential transform filter contains 1000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0163] In some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using an exponential transform filter contains 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, 10 to 50 corresponding radiomic features, and 75 Includes up to 750 corresponding radiomic features, 75 to 500 corresponding radiomic features, 75 to 250 corresponding radiomic features, 75 to 100 corresponding radiomic features, 175 to 750 corresponding radiomic features, 175 to 500 corresponding radiomic features, 175 to 250 corresponding radiomic features, 375 to 750 corresponding radiomic features, 375 to 500 corresponding radiomic features, or 600 to 750 corresponding radiomic features.
[0164] Continuing to refer to Figure 2K, in block 2116, in some embodiments, the multiple radiomic feature classes include a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter.
[0165] Referring to block 2118, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter includes 3 to 5 radiomic feature classes, 3 to 4 radiomic feature classes, or 4 to 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter includes at least 3, at least 4, or at least 5 radiomic feature classes. In some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter includes up to 3, up to 4, or up to 5 radiomic feature classes.
[0166] Referring to block 2120, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter includes radiomic feature classes selected from the group consisting of first-order features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighbor gray-tone difference matrix (NGTDM) features.
[0167] Referring to block 2122, in some embodiments, a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
[0168] Referring to block 2124, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter contains at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
[0169] Referring to block 2126, in some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter contains 1000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
[0170] In some embodiments, each individual radiomic feature class within a subset of radiomic feature classes extracted from multiple versions of medical images filtered using a squared transform filter contains 10 to 1,000 corresponding radiomic features, 10 to 750 corresponding radiomic features, 10 to 500 corresponding radiomic features, 10 to 250 corresponding radiomic features, 10 to 100 corresponding radiomic features, 10 to 50 corresponding radiomic features, and 75 Includes up to 750 corresponding radiomic features, 75 to 500 corresponding radiomic features, 75 to 250 corresponding radiomic features, 75 to 100 corresponding radiomic features, 175 to 750 corresponding radiomic features, 175 to 500 corresponding radiomic features, 175 to 250 corresponding radiomic features, 375 to 750 corresponding radiomic features, 375 to 500 corresponding radiomic features, or 600 to 750 corresponding radiomic features.
[0171] Continuing to refer to block 2128 in Figure 2L, in some embodiments, multiple radiomic feature classes include a local binary pattern (LBP) feature class.
[0172] Referring to block 2130, in some embodiments, multiple radiomic feature classes include local terminally pattern (LTP) feature classes.
[0173] Referring to block 2132, in some embodiments, multiple radiomic feature classes include upper LTP feature classes.
[0174] Referring to block 2134, in some embodiments, multiple radiomic feature classes include lower LTP feature classes.
[0175] Referring to block 2136, in some embodiments, multiple radiomic feature classes include latent feature classes.
[0176] Referring to block 2138, in some embodiments, individual latent features within a latent feature class are extracted from a segmentation model. For example, in some embodiments, individual latent features within a latent feature class originate from encodings from a layer of the segmentation model (e.g., a bottleneck layer).
[0177] Referring to block 2140, in some embodiments, the medical image dataset includes computed tomography (CT) datasets, magnetic resonance imaging (MRI) datasets, ultrasound datasets, position emission tomography (PET) datasets, or X-ray datasets.
[0178] Referring to block 2142, in some embodiments, the medical image dataset includes a brain MRI dataset.
[0179] Referring to block 2144, in some embodiments, the medical image dataset includes a whole-body PET dataset.
[0180] Referring to block 2146, in some embodiments, the medical image dataset includes lung CT.
[0181] Referring to block 2148 in Figure 2M, in some embodiments, the method includes combining multiple component predictions to obtain a cancer state characterization as the output of an ensemble model.
[0182] Referring to Block 2150, in some embodiments, the characterization of cancer status includes individual cancer types selected from multiple cancer types, individual cancer stages selected from multiple cancer stages, individual tissue origins selected from multiple tissue origins, individual cancer grades selected from multiple cancer grades, or individual prognoses selected from multiple prognoses. In some embodiments, cancer characteristics that can be distinguished using radiomics include genomic mutations, RNA expression, protein expression (e.g., immunohistochemical status), cancer molecular / histopathological subtypes, immunosignatures, treatment response, treatment monitoring (e.g., utilizing longitudinal data and delta radiomics), tumor volume, malignancy score, and adverse event risk (e.g., if the adverse event is death).
[0183] Referring to block 2152, in some embodiments, the characterization of the cancerous state includes individual tissues of multiple origins selected from among tissues of multiple origins.
[0184] Referring to Block 2154, in some embodiments, the characterization of the cancer state includes individual cancer malignancies selected from a plurality of cancer malignancies.
[0185] Referring to block 2156, in some embodiments, the characterization of the cancer state includes individual cancer stages selected from a plurality of cancer stages.
[0186] Referring to Block 2158, in some embodiments, the characterization of the cancer state includes a prognosis of the cancer state selected from a plurality of prognoses.
[0187] Referring to block 2160, in some embodiments, the multiple prognoses include a contiguous range of prognoses.
[0188] Referring to block 2162, in some embodiments, the prognosis is the expected survival time.
[0189] Referring to Block 2164, in some embodiments, the predicted survival time is cancer survival, disease-free survival, or progression-free survival.
[0190] Referring to Block 2166, in some embodiments, the predicted survival time is the cancer survival time for non-small cell lung cancer (NSCLC).
[0191] Referring to block 2168 in Figure 2N, in some embodiments, the cancer state characterization is an individual characterization selected from a plurality of separate characterizations of the cancer state. Each individual component prediction within the plurality of component predictions is an individual component characterization of the cancer state selected from a plurality of separate characterizations of the cancer state. Combining involves identifying the most representative individual component characterization within the plurality of component predictions and thereby obtaining the cancer state characterization.
[0192] Referring to block 2170, in some embodiments, the cancer state characterization is an individual characterization identified from a contiguous range of cancer state characterizations. Each individual component prediction within a group of component predictions is an individual component characterization of cancer state identified from a contiguous range of characterizations. Combining them involves determining a measure of the central trend of the group of component predictions and thereby obtaining the cancer state characterization.
[0193] Referring to block 2172, in some embodiments, the measure of central trend for multiple component predictions is a weighted measure of central trend.
[0194] Referring to block 2174, in some embodiments, the coupling involves inputting multiple component predictions of the cancer state into an aggregate model and obtaining a characterization of the cancer state as the output of the aggregate model.
[0195] Referring to block 2176, in some embodiments, the aggregation model is a neural network, a support vector machine, a naive Bayes model, a nearest neighbor model, a boosted tree model, a random forest model, or a clustering model.
[0196] Referring to block 2178, in some embodiments, the aggregate model is a voting model with adjusted thresholds.
[0197] Continuing to refer to Figure 2O and Block 2180, in some embodiments, the method includes assigning a therapy to a subject for a cancerous condition based on a characterization of the cancerous condition. Referring to Block 2182, in some embodiments, the cancerous condition is non-small cell lung cancer (NSCLC). The therapy assigned to the subject is selected from the group consisting of surgery, radiotherapy, and immunotherapy.
[0198] Referring to Block 2184, in some embodiments, the method includes administering therapy to a subject for a cancerous condition based on a characterization of the cancerous condition. Referring to Block 2186, in some embodiments, the cancerous condition is non-small cell lung cancer (NSCLC). The therapy administered to the subject is selected from the group consisting of surgery, radiotherapy, and immunotherapy.
[0199] In some embodiments, the ensemble model is intended to provide clinicians and patients with more information (e.g., without the assignment or administration of therapy). Its output is whether the patient is at high / low risk of dying while receiving standard treatment. Ultimately, the prescribed treatment is a clinician's decision and may be based on many (e.g., other) factors specific to the patient. Some hypothetical examples of how the ensemble model disclosed herein may lead to different operating strategies include (i) high-risk patients not wanting to receive treatments that would significantly reduce their remaining quality of life, such as aggressive chemotherapy regimens, or (ii) patients belonging to the high-risk group suggesting that standard treatment would be ineffective. Therefore, the output of the ensemble model may suggest seeking clinical trials or completing additional sequencing to search for novel immuno-oncology (IO) therapies.
[0200] Example: Training of an ensemble model, and comparison of the performance of the trained ensemble model with that of its trained component models.
[0201] An ensemble model was trained to characterize the cancerous state of tissues in the subjects, following the training process shown in Figure 4. In this embodiment, the full training dataset 402 consisted of images from 245 non-small cell lung cancer patients, the majority of whom were treated with standard cancer therapy. The dependent variable / label used for training was whether the patient survived two years after radiation therapy. The model output was a prediction of whether the patient was expected to survive for more than two years after radiation therapy, regardless of treatment. This was interpreted as an assessment of whether the patient had a high or low risk of experiencing an “event” (death). Patients who were not followed for at least two years and did not experience an event before the two-year time threshold were censored from the analysis. Of the training cohort, 76 patients survived the required 730 days, and 169 patients did not survive.
[0202] Each image volume (multiple images) was loaded along with its corresponding region of interest label file. These images were then transformed into several types before feature extraction, for a total of approximately 2000 radiomic features calculated for each region of interest (volume of interest, VOI). Features were extracted only from regions of interest consisting of at least five voxels. The features were divided into four categories: shape (16), intensity (18), texture (68), and filter (1892).
[0203] The 16 shape features represent the contour shape, size, major diameter, and overall sphericity (e.g., MeshVolume, SurfaceArea, SurfaceVolumeRatio, Compactness1, and Compactness2).
[0204] The 18 intensity features were derived from the voxel density statistics of the ROI and included features such as mean HU, first and second-order density histograms, and features representing the uniformity of the density histogram.
[0205] Sixty-eight texture features were extracted using mathematical matrices to capture subtle variations and patterns in the density of three-dimensional ROIs. The 68 texture features included 22 GLCM features, 16 GLRLM features, 16 GLSZM features, and 14 gray-level dependent matrix (GLDM) features.
[0206] The 22 GLCM features were autocorrelation, joint average, cluster prominence, cluster shade, cluster tendency, contrast, correlation, difference average, difference entropy, difference variance, joint energy, joint entropy, Imc1, Imc2, Idm, Idmn, Id, Idn, inverse variance, maximum probability, sum entropy, and sum squares.
[0207] The 16 gray-level run-length matrix (GLRLM) features are: GrayLevelNonUniformity, GrayLevelNonUniformityNormalized, GrayLevelVariance, HighGrayLevelRunEmphasis, LongRunEmphasis, LongRunHighGrayLevelEmphasis, LongRunLowGrayLevelEmphasis, and LowGrayLevel These were run emphasis (LowGrayLevelRunEmphasis), run entropy, run length non-uniformity, run length non-uniformity normalized, run percentage, run variance, short run emphasis (ShortRunEmphasis), short run high gray level emphasis (ShortRunHighGrayLevelEmphasis), and short run low gray level emphasis (ShortRunLowGrayLevelEmphasis).
[0208] The 16 gray-level size zone matrix (GLSZM) features are: GrayLevelNonUniformity, GrayLevelNonUniformityNormalized, GrayLevelVariance, HighGrayLevelZoneEmphasis, LargeAreaEmphasis, LargeAreaHighGrayLevelEmphasis, LargeAreaLowGrayLevelEmphasis, and LowGrayLevel Zone Emphasis. These were LowGrayLevelZoneEmphasis, SizeZoneNonUniformity, SizeZoneNonUniformityNormalized, SmallAreaEmphasis, SmallAreaHighGrayLevelEmphasis, SmallAreaLowGrayLevelEmphasis, ZoneEntropy, ZonePercentage, and ZoneVariance.
[0209] The 14 GLDM features are: Dependence Entropy, Dependence Non-Uniformity, Dependence Non-Uniformity Normalized, Dependence Variance, Gray Level Non-Uniformity, Gray Level Variance, High Gray Level Emphasis, Large Dependence Emphasis, and Large The emulations included LargeDependenceHighGrayLevelEmphasis, LargeDependenceLowGrayLevelEmphasis, LowGrayLevelEmphasis, SmallDependenceEmphasis, SmallDependenceHighGrayLevelEmphasis, and SmallDependenceLowGrayLevelEmphasis.
[0210] To arrive at the forty component models used in this embodiment, the radiomic features available for each training target were divided into groups based on a combination of two differences: (i) the source image (and its filtering status), and (ii) the feature generation method. The initial difference was fairly simple. GLCM features generated from the original unfiltered image were separated from GLCM features generated from either modified / filtered images, such as images to which a log-sigma filter was applied. Features were extracted from unfiltered images (orig) as well as from images to which various filters were applied, including a Gaussian Laplacian (log_sigma) filter, two types of wavelet filters (wavelet and wavelet2), a logarithm filter, an exponential filter, and a square filter. Further information on these filters can be found in version 6a761c4e of the pyradiomics documentation, which can be found at the following URL: pyradiomics.readthedocs.io / en / latest / index.html.
[0211] The second difference separated a large number of highly related features generated by a specific method from a separate group of features generated by a unique method. For example, GLCM features were sorted into a different group from gray-level run-length matrix (GLRLM) features. Similarly, local binary pattern (LBP) features were separated from upper local terminally pattern (LTP) features and lower LTP features, and all of them were separated from GLCM and GLRLM features. This resulted in 40 distinct feature groups, each of which was separately assigned to a logistic regression component model, resulting in the following 40 logistic regression component models. [Table 8]
[0212] The training dataset 404 was subjected to a 5-fold cross-validation step (408), where the training dataset 404 was divided into 5 parts, and cross-validation was repeated through these divisions. In each iteration, one of the 5 divisions (representing 20 percent of the randomly selected cohort) was used as the validation set, while the remaining 4 divisions (representing the remaining 80 percent of the cohort) were used as the training set.
[0213] During this five-fold training, information about the specific training target (measurements and calculations for each of approximately 2000 radiomic features) was input into the ensemble model, and the labels calculated by the ensemble model for the specific training target were compared with the actual labels of the training target. In this example, the ensemble model consisted of forty component models. The names of the forty component models are given on the X-axis in Figure 5A, and are the same as the component model names in the table above.
[0214] When information about a specific training subject (approximately 2000 extracted radiomic features) is input to an ensemble model, the output from each individual component model within the forty component models is obtained in the form of corresponding component predictions for the cancer state of the specific training subject, thereby obtaining forty component predictions for the cancer state of the training subject. In this embodiment, the information about each specific training subject includes, for each individual radiomic feature class within the set of forty radiomic feature classes, corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of the individual radiomic feature classes obtained from the medical image dataset for each specific training subject. The medical image dataset of images includes multiple medical images of tissue in the training subject initially obtained using a first medical image modality.
[0215] In this embodiment, inputting information about individual training subjects includes (i) inputting corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of a first individual radiomic feature class within a class of forty radiomic features into the first individual component model within a class of multiple component models, and (ii) inputting corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of a second individual radiomic feature class within a class of forty radiomic features into the second individual component model within a class of multiple component models, wherein the corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of the first individual radiomic feature class are not input into the second component model, and the corresponding values for each individual radiomic feature within the corresponding multiple radiomic features of the second individual radiomic feature class are not input into the first individual component model. For example, the original gray-level co-occurrence matrix (GLCM) features input to the component model "orig_glcm" were separated from the original gray-level run-length matrix (GLRLM) features input to the component model "orig_glcm".
[0216] For each training subject, forty component predictions were combined to obtain a cancer status characterization as the output of the ensemble model. In this embodiment, a single voting model was used to combine the forty component predictions, and the threshold for determining whether a patient was high-risk or low-risk was adjusted during training.
[0217] Since the model is an ensemble model, it was also possible to verify the performance of the ensemble model and the performance of each component model using a cohort of 245 non-small cell lung cancer patients. For this purpose, Figure 5A shows that nine component models had an area under the training curve (AUC) greater than 0.7, and the corresponding trial AUCs for these nine component models ranged from 0.67 to 0.54. Here, the trial AUC statistics for each component model are from instances where the component model was applied to test subjects in the holdout data (training subjects assigned to the trial dataset in any given split run, dataset 406 in Figure 4) during 5-fracture training, while the trial-training AUC is from instances where the component model was applied to training subjects in the training data (dataset 404 in Figure 4) during 5-fracture training. The difference between the training AUC and trial AUC of a given component model represented a decrease in performance. A well-trained model should have similar AUCs for both the training data (training dataset 404) and the test data (holdout test dataset 406); therefore, a significant drop in performance was an indication of overtraining.
[0218] The last entry on the X-axis in Figure 5A represents an ensemble model, which combines forty component models to provide predictions (label predictions) of the absence or presence of non-small cell lung cancer patients in a given training subject. The training and trial AUCs of the ensemble model are shown. The ensemble's label predictions coincide with the best (e.g., smallest) decline in performance exhibited by any of the nine component models having a training AUC greater than 0.7.
[0219] Figure 5B provides a further analysis of the results of model training using a cohort of 245 non-small cell lung cancer patients. Similar to Figure 5A, each component model is listed on the X-axis. In Figure 5B, the Y-axis is a measure of the hazard ratio for each component model relative to the cohort. Both training and trial hazard ratios are given for each component model; here, as shown in Figure 5A, training is for the training dataset (404), and trial is for each trial dataset in the cross-validation. The hazard ratio is an estimate of the hazard ratio between the treatment group and the control group. The hazard ratio is calculated by dividing the probability of the event in question occurring in the next time interval (if it has not yet occurred) by the length of that interval. Figure 5B shows that nine component models have hazard ratios (HR) greater than 2.0, with their corresponding trial HRs ranging from approximately 1.7 to 0.9. In contrast, the ensemble model (a combination of all forty component models) outperforms all the training component models, maintaining an HR greater than 2.0.
[0220] Figure 5C provides a further analysis of the results of training a five-component model using a cohort of 245 non-small cell lung cancer patients. Similar to Figures 5A and 5B, each component model is listed on the X-axis. In Figure 5C, the Y-axis measures the p-value of the hazard ratio for the HR of each component model relative to the cohort. The p-value of the hazard ratio for the HR of a component model is a measure of the statistical significance of the hazard ratio of the component model. Figure 5C shows that almost all component models have significant p-values for training HRs (based on training dataset 404 in Figure 4), but only 13 component models maintain significant p-values for trial (holdout dataset 406) HRs (e.g., p-values greater than 0.1). That is, in Figure 5C, statistical significance occurs when the p-value is small. Therefore, the existence of large p-values for HRs for each component model relative to the trial data indicates that the HR values determined for these models are not statistically significant. Furthermore, the 13 models that maintain significant p-values for trial HRs do not necessarily correlate with the most significant or highest training HR values. In contrast, the ensemble model maintains its significance on both the training and test datasets, as seen in the last entry on the X-axis in Figure 5C.
[0221] The following table provides the component model name (or "Ensemble" in the case of an ensemble model), as well as the training AUC, training hazard ratio (HZ), p-value of the training hazard ratio, trial HR, and p-value of the trial HR for each model after training a cohort of 245 non-small cell lung cancer patients through K-fold analysis. [Table 9]
[0222] This embodiment demonstrates that an ensemble model performs substantially better than a component model, where each component model represents a different class of radiomic features. The embodiment also demonstrates that when testing the model, statistical variability can lead to an overestimation of the performance of some component models or overtraining of some component models, resulting in an overestimation of the training performance with respect to such component models. In contrast, an ensemble of such component models maintains statistically significant and meaningful predictive power with a better combination of performance and generalizability than any single component model. Thus, an ensemble model is a better alternative to selecting a single, possibly overtrained or statistically anomalous, single component model.
[0223] Multiple entities may be provided for any component, operation, or structure described herein as a single entity. Finally, the boundaries between the various components, operations, and data stores are somewhat arbitrary, and certain operations are illustrated in the context of a particular exemplary configuration. Other forms of functionality are conceivable and may fall within the scope of implementation. In general, structures and functions presented as separate components in an exemplary configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as single components may be implemented as separate components. These and other variations, modifications, additions, and improvements may fall within the scope of implementation.
[0224] Furthermore, while terms such as "first" and "second" may be used herein to describe various elements, it will be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, the first attribute may be called the second attribute, and similarly the second attribute may be called the first attribute, without changing the meaning of the description, as long as all occurrences of "first attribute" are consistently renamed and all occurrences of "second attribute" are consistently renamed. Both the first and second attributes are attributes, but unless otherwise specified, they are not the same attribute.
[0225] The terminology used in this document is intended solely to describe a particular implementation and is not intended to limit the scope of the claims. Where used in the description of an implementation and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural form unless the context otherwise explicitly indicates. Where used herein, the terms "and / or" will be understood to refer to and encompass any possible combination of one or more of the listed items relating to the implementation. Where used herein, the terms "comprise" and / or "comprising" will be understood to specify the presence of a described feature, integer, step, operation, element, and / or component, but not to preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0226] As used herein, the term "if" may be interpreted, depending on the context, as meaning "when," "upon," "in accordance with the determination," "according to the determination," or "in response to the detection of" the precedent of the stated condition being true. Similarly, the phrases "if it is determined that (the precedent of the stated condition is true)," "when (the precedent of the stated condition is true)," or "when (the precedent of the stated condition is true)" may be interpreted, depending on the context, as meaning "when it is determined that (the precedent of the stated condition is true)," "in accordance with the determination," "according to the determination," "when (the detection)," or "in response to the detection of" the precedent of the stated condition being true.
[0227] The preceding description included exemplary systems, methods, techniques, instruction sequences, and computing machine program products that embody exemplary implementations. For illustrative purposes, numerous specific details have been provided to give an understanding of various implementations of the subject matter of the invention. However, it will be apparent to those skilled in the art that implementations of the subject matter of the invention can be carried out without these specific details. Generally, well-known instruction instances, protocols, structures, and techniques are not described in detail.
[0228] The preceding descriptions are written with reference to specific implementations for illustrative purposes. However, the above exemplary considerations are not intended to be exhaustive or to limit the scope to detailed forms that disclose implementations. Many variations and modifications are possible considering the above teachings. The implementations have been selected and described to best illustrate the principles and their practical applications, so that those skilled in the art can make the most of the various implementations with modifications suited to their specific applications.
Claims
1. A method for characterizing the cancerous state of tissue in a subject, In a computer system including one or more processors and memory, A) Inputting information into an ensemble model that includes multiple component models, obtaining corresponding component predictions regarding the cancer state as outputs from each individual component model within the multiple component models, and thereby obtaining multiple component predictions regarding the cancer state, The information includes, for each individual radiomic feature within a plurality of radiomic feature classes, corresponding values for each individual radiomic feature within a plurality of radiomic features obtained from a medical image dataset, wherein the medical image dataset includes a plurality of medical images of the tissue in the subject initially obtained using a first medical image modality. The aforementioned ensemble model applies multiple parameters to the information through multiple calculations, The input includes (i) inputting the corresponding value for each individual radiomic feature in the corresponding multiple radiomic features of the first individual radiomic feature class within the multiple radiomic feature classes into the first individual component model within the multiple component model, and (ii) inputting the corresponding value for each individual radiomic feature in the corresponding multiple radiomic features of the second individual radiomic feature class within the multiple radiomic feature classes into the second individual component model within the multiple component model, wherein the corresponding value for each individual radiomic feature in the corresponding multiple radiomic features of the first individual radiomic feature class is not input into the second component model, and the corresponding value for each individual radiomic feature in the corresponding multiple radiomic features of the second individual radiomic feature class is not input into the first individual component model, and obtaining B) A method comprising combining the multiple component predictions to obtain a characterization of the cancer state as the output of the ensemble model.
2. The characteristic evaluation of the cancerous state is, Individual cancer types selected from multiple cancer types, Individual cancer stages selected from multiple cancer stages, Individual origin organizations selected from multiple origin organizations, Individual cancer malignancies selected from multiple cancer malignancy grades, or The method according to claim 1, comprising individual prognoses selected from a plurality of prognoses.
3. The method according to claim 1 or 2, wherein the characterization of the cancerous state includes individual tissues selected from a plurality of tissues of different origins.
4. The method according to any one of claims 1 to 3, wherein the characteristic evaluation of the cancerous state includes individual cancer malignancies selected from a plurality of cancer malignancy grades.
5. The method according to any one of claims 1 to 4, wherein the characteristic evaluation of the cancerous state includes individual cancer stages selected from a plurality of cancerous stages.
6. The method according to any one of claims 1 to 5, wherein the characteristic evaluation of the cancerous state includes a prognosis of the cancerous state selected from a plurality of prognoses.
7. The method according to claim 6, wherein the plurality of prognoses include prognoses within a continuous range.
8. The method according to claim 6 or 7, wherein the prognosis is the predicted survival period.
9. The method according to claim 8, wherein the predicted survival period is cancer survival, disease-free survival, or progression-free survival.
10. The method according to claim 8, wherein the predicted survival period is the cancer survival period for non-small cell lung cancer (NSCLC).
11. The aforementioned cancerous conditions include carcinoma, lymphoma, blastoma, glioblastoma, sarcoma, leukemia, breast cancer, squamous cell carcinoma, lung cancer, small cell lung cancer, non-small cell lung cancer (NSCLC), adenocarcinoma of the lung, squamous cell carcinoma of the lung, head and neck cancer, peritoneal cancer, hepatocellular carcinoma, gastric or stomach cancer, pancreatic cancer, ovarian cancer, cervical cancer, liver cancer, bladder cancer, liver cancer, colon cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal cancer, liver cancer, prostate cancer, vulvar cancer, thyroid cancer, hepatoma, B-cell lymphoma, and low-grade / follicular cancer. The method according to any one of claims 1 to 10, wherein the cancer is selected from the group consisting of non-Hodgkin lymphoma (NHL), small lymphocytic (SL) NHL, intermediate-grade / follicular NHL, intermediate-grade diffuse NHL, high-grade immunoblastic NHL, high-grade lymphoblastic NHL, high-grade small undivided cell NHL, giant tumor NHL, mantle cell lymphoma, AIDS-associated lymphoma, Waldenström macroglobulinemia, chronic lymphocytic leukemia (CLL), acute lymphoblastic leukemia (ALL), hairy cell leukemia, and chronic myeloblastic leukemia.
12. The method according to any one of claims 1 to 11, wherein the plurality of component models are at least 10 component models.
13. The method according to any one of claims 1 to 11, wherein the plurality of component models are at least 20 component models.
14. The method according to any one of claims 1 to 11, wherein the plurality of component models are at least 40 component models.
15. The method according to any one of claims 1 to 14, wherein the plurality of component models are 250 or fewer component models.
16. The method according to any one of claims 1 to 14, wherein the plurality of component models are 100 or fewer component models.
17. The method according to any one of claims 1 to 14, wherein the plurality of component models are 50 or fewer component models.
18. The method according to any one of claims 1 to 17, wherein the plurality of component models are 5 to 100 component models.
19. The method according to any one of claims 1 to 17, wherein the plurality of component models are 10 to 75 component models.
20. The method according to any one of claims 1 to 17, wherein the plurality of component models are 20 to 50 component models.
21. The method according to any one of claims 1 to 20, wherein each component model in the plurality of component models is a neural network, a support vector machine, a naive Bayes model, a nearest neighbor model, a boosted tree model, a random forest model, or a clustering model.
22. The method according to any one of claims 1 to 20, wherein each individual component model in the plurality of component models includes an individual neural network.
23. The method according to any one of claims 1 to 22, wherein each individual component model in the plurality of component models includes an individual logistic regression model.
24. The aforementioned multiple radiomic feature classes are, A first subset of radiomic feature classes extracted from the unfiltered versions of the multiple medical images in the aforementioned medical image dataset, The method according to any one of claims 1 to 23, further comprising: a second subset of radiomic feature classes extracted from filtered versions of the plurality of medical images in the medical image dataset filtered by a first filtering method.
25. The method according to claim 24, wherein the first filtering method includes an image filter selected from the group consisting of a wavelet transform filter, a Gaussian Laplacian (LogG) filter, a squared transform filter, a square root transform filter, a logarithmic transform filter, an exponential transform filter, a gradient transform filter, a two-dimensional local binary pattern filter, and a three-dimensional local binary pattern filter.
26. The method according to claim 24 or 25, wherein the first subset of radiomic feature classes comprises at least three, at least four, or at least five radiomic feature classes.
27. The method according to claim 26, wherein the first subset of the radiomics feature class includes a radiomics feature class selected from the group consisting of shape features, linear features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
28. The method according to claim 26, wherein the first subset of the radiomics feature class includes shape features, linear features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
29. The method according to any one of claims 26 to 28, wherein each individual radiomic feature class within the first subset of radiomic feature classes comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
30. The method according to any one of claims 26 to 29, wherein each individual radiomic feature class within the first subset of radiomic feature classes includes 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
31. The method according to any one of claims 24 to 30, wherein the plurality of radiomic feature classes include a subset of radiomic feature classes extracted from the plurality of medical image versions filtered using a first wavelet transform filter.
32. The method according to claim 31, wherein the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the first wavelet transform filter comprises at least three, at least four, or at least five radiomic feature classes.
33. The method according to claim 31 or 32, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the first wavelet transform filter includes a radiomic feature class selected from the group consisting of primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
34. The method according to claim 31 or 32, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the first wavelet transform filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
35. The method according to any one of claims 31 to 34, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the first wavelet transform filter comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
36. The method according to any one of claims 31 to 35, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the first wavelet transform filter comprises 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
37. The method according to any one of claims 31 to 36, wherein the plurality of radiomic feature classes include a subset of radiomic feature classes extracted from the plurality of medical image versions filtered using a second wavelet transform filter.
38. The method according to claim 37, wherein the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the second wavelet transform filter comprises at least three, at least four, or at least five radiomic feature classes.
39. The method according to claim 37 or 38, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the second wavelet transform filter includes a radiomic feature class selected from the group consisting of primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
40. The method according to claim 37 or 38, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the second wavelet transform filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
41. The method according to any one of claims 37 to 40, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the second wavelet transform filter comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
42. The method according to any one of claims 37 to 41, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the second wavelet transform filter comprises 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
43. The method according to any one of claims 24 to 42, wherein the plurality of radiomic feature classes include a subset of radiomic feature classes extracted from the plurality of medical image versions filtered using a Log-G filter.
44. The method according to claim 43, wherein the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the Log filter comprises at least three, at least four, or at least five radiomic feature classes.
45. The method according to claim 43 or 44, wherein a subset of the radiomic feature classes extracted from the versions of the plurality of medical images filtered using the LogG filter includes a radiomic feature class selected from the group consisting of primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
46. The method according to claim 43 or 44, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the LogG filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
47. The method according to any one of claims 43 to 46, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the Log filter comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
48. The method according to any one of claims 43 to 47, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the Log filter comprises 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
49. The method according to any one of claims 24 to 48, wherein the plurality of radiomic feature classes include a subset of radiomic feature classes extracted from the plurality of medical image versions filtered using a logarithmic transform filter.
50. The method according to claim 49, wherein the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the logarithmic transformation filter comprises at least three, at least four, or at least five radiomic feature classes.
51. The method according to claim 49 or 50, wherein a subset of the radiomic feature classes extracted from the versions of the plurality of medical images filtered using the logarithmic transform filter includes a radiomic feature class selected from the group consisting of primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
52. The method according to claim 49 or 50, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the logarithmic transform filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
53. The method according to any one of claims 49 to 52, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the logarithmic transform filter comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
54. The method according to any one of claims 49 to 53, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the logarithmic transformation filter comprises 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
55. The method according to any one of claims 24 to 54, wherein the plurality of radiomic feature classes include a subset of radiomic feature classes extracted from the plurality of medical image versions filtered using an exponential transform filter.
56. The method according to claim 55, wherein the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the exponential transform filter comprises at least three, at least four, or at least five radiomic feature classes.
57. The method according to claim 55 or 56, wherein a subset of the radiomic feature classes extracted from the versions of the plurality of medical images filtered using the exponential transformation filter includes a radiomic feature class selected from the group consisting of primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
58. The method according to claim 55 or 56, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the exponential transformation filter includes primary features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), and gray-level dependency matrix (GLDM) features.
59. The method according to any one of claims 55 to 58, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the exponential transform filter comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
60. The method according to any one of claims 55 to 59, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the exponential transformation filter comprises 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
61. The method according to any one of claims 24 to 60, wherein the plurality of radiomic feature classes include a subset of radiomic feature classes extracted from the plurality of medical image versions filtered using a squared transform filter.
62. The method according to claim 61, wherein the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the squared transform filter comprises at least three, at least four, or at least five radiomic feature classes.
63. The method according to claim 61 or 62, wherein a subset of the radiomic feature classes extracted from the versions of the plurality of medical images filtered using the squared transform filter includes a radiomic feature class selected from the group consisting of first-order features, gray-level co-occurrence matrix (GLCM) features, gray-level run-length matrix (GLRLLM) features, gray-level size zone matrix (GLSZM), gray-level dependency matrix (GLDM) features, and neighboring gray-tone difference matrix (NGTDM) features.
64. The method according to claim 61 or 62, wherein a subset of the radiomic feature class extracted from the versions of the plurality of medical images filtered using the squared transform filter includes a first-order feature, a gray-level co-occurrence matrix (GLCM) feature, a gray-level run-length matrix (GLRLLM) feature, a gray-level size zone matrix (GLSZM), and a gray-level dependency matrix (GLDM) feature.
65. The method according to any one of claims 61 to 64, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the squared transform filter comprises at least 10, at least 25, at least 50, at least 100, or at least 250 corresponding radiomic features.
66. The method according to any one of claims 61 to 65, wherein each individual radiomic feature class in the subset of radiomic feature classes extracted from the versions of the plurality of medical images filtered using the squared transform filter comprises 1,000 or fewer, 750 or fewer, 500 or fewer, 250 or fewer, or 100 or fewer corresponding radiomic features.
67. The method according to any one of claims 24 to 66, wherein the plurality of radiomic feature classes include a local binary pattern (LBP) feature class.
68. The method according to any one of claims 24 to 67, wherein the plurality of radiomic feature classes include a local terminally pattern (LTP) feature class.
69. The method according to any one of claims 24 to 68, wherein the plurality of radiomic feature classes include an upper LTP feature class.
70. The method according to any one of claims 24 to 69, wherein the plurality of radiomic feature classes include a lower LTP feature class.
71. The method according to any one of claims 24 to 69, wherein the plurality of radiomic feature classes include latent feature classes.
72. The method according to claim 71, wherein each latent feature within the aforementioned latent feature class is extracted from a segmentation model.
73. Before A) inputting the aforementioned information into the ensemble model, 1) With respect to each individual radiomic feature class within the first subset of radiomic feature classes, extract the corresponding value for each individual radiomic feature within each individual radiomic feature class within the first subset of radiomic feature classes from the region of interest (ROI) or volume of interest (VOI) in the unfiltered versions of the plurality of medical images, 2) The method according to any one of claims 24 to 72, further comprising: extracting the corresponding values for each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in the second individual radiomic feature in each individual radiomic feature in the second individual radiomic feature in the second individual radio
74. The method according to claim 73, further comprising identifying the ROI or VOI within the plurality of medical images.
75. Identifying the ROI or VOI within the plurality of medical images is The unfiltered versions of the aforementioned multiple medical images are segmented into multiple segments or multiple volumes. Assigning individual tissue classifications within a plurality of tissue classifications to each individual segment of the plurality of segments, or to each individual volume of the plurality of volumes, based on one or more feature quantities of the individual segment or individual volume; The method of claim 74, comprising grouping individual segments within the plurality of segments or individual volumes within the plurality of volumes to which a target tissue classification within the plurality of tissue classifications has been assigned, thereby identifying the ROI or the VOI.
76. The method according to any one of claims 1 to 75, wherein the medical image dataset includes computed tomography (CT) datasets, magnetic resonance imaging (MRI) datasets, ultrasound datasets, position emission tomography (PET) datasets, or X-ray datasets.
77. The method according to any one of claims 1 to 75, wherein the medical image dataset includes a brain MRI dataset.
78. The method according to any one of claims 1 to 75, wherein the medical image dataset includes a whole-body PET dataset.
79. The method according to any one of claims 1 to 75, wherein the medical image dataset includes lung CT.
80. The characteristic evaluation of the cancerous state is an individual characteristic evaluation selected from a plurality of individual characteristic evaluations of the cancerous state. Each individual component prediction within the plurality of component predictions is an individual component characteristic of the cancer state selected from the plurality of individual characteristic assessments of the cancer state, The method according to any one of claims 1 to 79, wherein the combination of B) includes identifying the most representative individual component characterization within the plurality of component predictions and thereby obtaining the characterization of the cancer state.
81. The characteristic evaluation of the cancerous state is an individual characteristic evaluation identified from the characteristic evaluation of a continuous range of cancerous states. Each individual component prediction within the plurality of component predictions is an individual component characterization of the cancer state identified from the characterization of the continuous range, The method according to any one of claims 1 to 79, wherein the combination of B) comprises determining a measure of the central trend of the multiple component predictions and thereby obtaining the characteristic evaluation of the cancer state.
82. The method according to claim 81, wherein the measure of the central trend of the plurality of component predictions is a weighted measure of the central trend.
83. The method according to any one of claims 1 to 79, wherein the combination of B) includes inputting the multiple component predictions of the cancer state into an aggregation model and obtaining the characteristic evaluation of the cancer state as the output of the aggregation model.
84. The method according to claim 83, wherein the aggregation model is a neural network, a support vector machine, a naive Bayes model, a nearest neighbor model, a boosted tree model, a random forest model, or a clustering model.
85. The method according to claim 83 or 84, wherein the aggregation model is a voting model having an adjusted threshold.
86. The method according to any one of claims 1 to 79, further comprising assigning a therapy to the subject for the cancerous state based on the characteristic evaluation of the cancerous state.
87. The method according to any one of claims 1 to 79, further comprising administering therapy to the subject for the cancerous condition based on the characteristic evaluation of the cancerous condition.
88. The aforementioned cancerous condition is non-small cell lung cancer (NSCLC). The method according to claim 86 or 87, wherein the therapy assigned to or administered to the subject is selected from the group consisting of surgery, radiotherapy, and immunotherapy.
89. A computer system, One or more processors, A computer system comprising a non-temporary computer-readable medium, which includes a computer-executable instruction that, when executed by one or more processors, causes the processors to carry out the method according to any one of claims 1 to 84.
90. A non-temporary computer-readable storage medium that stores program code instructions that cause the processor to perform the method described in any one of claims 1 to 84 when executed by the processor.