Method, device and storage medium
Multi-modality data fusion using machine learning improves diagnostic accuracy by combining imaging and biomarker data, addressing the limitations of single-modality diagnostics and reducing false positives.
Patent Information
- Application Number
- JP2025010780
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-14
- Filing Date
- 2025-01-24
- Publication Date
- 2025-08-05
AI Technical Summary
Existing diagnostic methods face challenges in improving classification accuracy, particularly in distinguishing between benign and malignant lesions using medical imaging alone, leading to high false-positive rates and unnecessary biopsies.
A method involving multi-modality data fusion using machine learning, combining imaging modality data with biomarker modality data, and employing trained machine learning models to generate intermediate and classification results, including early, intermediate, and late fusion techniques to enhance diagnostic accuracy.
Enhances diagnostic accuracy by integrating data from multiple sources, reducing unnecessary biopsies and improving the distinction between benign and malignant lesions.
Smart Images

Figure 2025114522000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The embodiments disclosed herein generally relate to fusing data from multiple modalities using machine learning for the purpose of accurate diagnosis. [Background technology]
[0002] Medical imaging can produce (e.g., reconstruct) images of the internal structures of an object, such as a part inside a patient's body. For example, a computed tomography (CT) scan uses multiple X-ray images of an object taken at different angles to reconstruct a volumetric image of the object's interior.
[0003] Liquid biopsies also analyze biological analytes circulating in a patient's blood. For example, some liquid biopsies analyze biomarkers such as proteins, DNA, or RNA. Summary of the Invention [Problem to be solved by the invention]
[0004] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to improve classification accuracy in diagnosis, etc. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problem. Problems corresponding to the effects of each configuration shown in the embodiments described below can also be positioned as other problems. [Means for solving the problem]
[0005] According to an embodiment, a method includes acquiring multi-modality data, including imaging modality data and biomarker modality data, acquiring one or more trained machine learning models, generating a first multi-modality set from the multi-modality data, the first multi-modality set including at least one of the imaging modality data and the biomarker modality data, generating an intermediate data set, and generating a first classification result based on at least the intermediate data set. Generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data. Generating the first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based on at least the intermediate data set. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates an exemplary embodiment of a medical data processing system. [Figure 2] FIG. 2 illustrates an exemplary embodiment of an operational flow for generating classification results for multi-modality data. [Figure 3] FIG. 3 illustrates an exemplary embodiment of an operational flow for generating classification results from multi-modality data. [Figure 4] FIG. 4 illustrates an exemplary embodiment of an operational flow for generating classification results from multi-modality data. [Figure 5A] FIG. 5A is a diagram illustrating a multi-modality group formed by grouping data from multiple modalities and a single-modality group not yet combined into a multi-modality group. [Figure 5B] FIG. 5B shows the multi-modality group of FIG. 5A after removing data with low intra-group / inter-modality correlation. [Figure 5C]FIG. 5C shows the intermediate data group and the single modality group after some intermediate fusion has been performed on the multi-modality group of FIG. 5B. [Figure 5D] FIG. 5D shows the intermediate data group and single modality group of FIG. 5C after the intermediate data group and single modality group with high intra-group correlation have been recombined into a high correlation group. [Figure 5E] FIG. 5E is a diagram showing a first classification result and a single-modality group generated based on the high correlation group and the intermediate data group of FIG. 5D. [Figure 5F] Figure 5F shows the initial classification results generated based on the high correlation group and intermediate data group of Figure 5D, and the first classification results generated based on the single modality group of Figure 5E with the late fusion flag added. [Figure 6] FIG. 6 illustrates an exemplary embodiment of group classification and fusion of data from multiple modalities. [Figure 7] FIG. 7 illustrates an exemplary embodiment of group classification and fusion of multi-modality data. [Figure 8] FIG. 8 is a diagram illustrating the information flow in a multimodal data processing procedure performed by a data processing device. [Figure 9] FIG. 9 illustrates an exemplary embodiment of a neural network. [Figure 10] FIG. 10 illustrates an exemplary embodiment of a convolutional neural network (CNN). [Figure 11] FIG. 11 is a diagram illustrating an example of implementing a convolutional layer for one neuron node of the convolutional layer according to one embodiment. [Figure 12] FIG. 12 is a diagram illustrating the information flow in a method for generating classification results from multi-modality data. [Figure 13] FIG. 13 illustrates an exemplary embodiment of an operational flow for generating classification results from data from multiple modalities. [Figure 14] FIG. 14 illustrates an exemplary embodiment of an operational flow for generating classification results from data from multiple modalities. [Figure 15] FIG. 15 illustrates an exemplary embodiment of operations that can replace blocks B1420-B1430 of FIG. [Figure 16] FIG. 16 illustrates an exemplary embodiment of a data processing device. [Figure 17] FIG. 17 is a diagram illustrating the functional configuration of an exemplary embodiment of a data processing device. DETAILED DESCRIPTION OF THE INVENTION
[0007] The following paragraphs describe, for illustrative purposes, several embodiments. Other embodiments may include alternatives, equivalents, and variations. Furthermore, while the illustrative embodiments may include some novel features, no particular feature is essential to the device, system, or method embodiments described herein. Also, some embodiments may include two or more features of the embodiments described below. That is, features of the various embodiments may be combined or substituted as appropriate.
[0008] Additionally, the conjunction "or" used in this specification generally means an inclusive "or," but when an exclusive meaning is explicitly stated or suggested by the context, it means an exclusive "or."
[0009] Furthermore, terms such as "first," "second," etc., used in this specification do not necessarily imply a relationship such as order, sequence, or priority. These terms are used to more clearly distinguish one member, operation, element, group, collection, set, region, section, etc. from another, without expressing a relationship such as order, sequence, or priority. That is, a first member, operation, element, group, collection, set, region, section, etc. described below may be referred to as a second member, operation, element, group, collection, set, region, section, etc. without departing from the teachings of this specification.
[0010] Also, in the following description and drawings, like reference numerals designate the same or equivalent parts throughout the several views.
[0011] Further embodiments are defined in the following paragraphs.
[0012] (1) A method for generating a classification result includes acquiring multi-modality data including image modality data and biomarker modality data, acquiring one or more trained machine learning models, generating a first multi-modality set from the multi-modality data including at least one of the image modality data and the biomarker modality data, generating an intermediate data set, and generating a first classification result based on at least the intermediate data set. Generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data. Generating the first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based on at least the intermediate data set.
[0013] (2) In the method of (1), the imaging modality data includes data defining one or more of the following: ·X-ray image Computed tomography imaging Magnetic resonance imaging X-ray images Ultrasound imaging Positron emission tomography imaging
[0014] (3) In the method of (1), generating the first multi-modality group includes classifying the data of the multiple modalities into the first multi-modality group and one or more other groups based on the data of the multiple modalities and one or more group classification criteria.
[0015] (4) In the method of (3), the group classification criteria relate to at least one of the data format and the degree of correlation between the data.
[0016] (5) In the method of (1), the data of the multiple modalities further includes data of a text modality.
[0017] (6) In the method of (1), generating the first classification result further includes inputting the single-modality data set together with the intermediate data set into the second machine learning model, and the second machine learning model is trained to output a classification result based on the single-modality data set.
[0018] (7) The method of (1) further comprises generating a second classification result based on a single-modality data set from the data of the multiple modalities, and generating the second classification result includes inputting the single-modality data set to a third machine learning model trained to output a classification result based on the single-modality data set.
[0019] (8) The method of (7) further comprises generating a third classification result based on the first classification result and the second classification result.
[0020] (9) In the method of (8), the third classification result indicates one or more of the following: · Whether the tumor is benign or malignant · Tumor invasive or non-invasive cancer Luminal, Her2-enriched, or triple-negative breast cancer subtypes of cancer
[0021] (10) An apparatus includes one or more processors and one or more memories, configured to acquire multi-modality data including data of an imaging modality and data of a biomarker modality, acquire one or more trained machine learning models, generate a first multi-modality set from the multi-modality data including at least one of the data of the imaging modality and the data of the biomarker modality, generate an intermediate data set, and generate a first classification result based on at least the intermediate data set. Generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data. Generating the first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based on at least the intermediate data set.
[0022] (11) One or more computer-readable storage media store instructions that cause one or more computing devices to perform operations including acquiring multi-modality data including imaging modality data and biomarker modality data, acquiring one or more trained machine learning models, generating a first multi-modality set from the multi-modality data including at least one of the imaging modality data and the biomarker modality data, generating an intermediate data set, and generating a first classification result based on at least the intermediate data set. Generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data. Generating the first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based at least on the intermediate data set.
[0023] (12) A method includes acquiring data of a first modality and data of a second modality, acquiring one or more trained machine learning models, generating a multi-modality data set from the data of the first modality and the data of the second modality, each of the multi-modality data sets including a portion of the data of the first modality and a portion of the data of the second modality, inputting each of the multi-modality data sets to one of a first plurality of machine learning models, and inputting the intermediate data set to a second plurality of machine learning models. Each of the first plurality of machine learning models outputs an intermediate data set generated based on the input multi-modality data set. Furthermore, each of the second plurality of machine learning models outputs a classification result generated based on the input intermediate data set.
[0024] (13) A method includes obtaining data of multiple modalities, obtaining one or more group classification criteria, obtaining one or more trained machine learning models, generating multi-modality data sets from the data of the multiple modalities according to the one or more group classification criteria, each multi-modality data set including data of at least two modalities, and generating one or more first classification results by performing intermediate fusion on the one or more multi-modality data sets using one or more of the trained machine learning models.
[0025] (14) The method of (13) further comprises generating one or more first classification results by performing early fusion on at least one single-modality data group from the multiple modality data using one or more of the trained machine learning models.
[0026] (15) The method of (14) further comprises generating a second classification result by performing late fusion on the first classification result using one or more late fusion models.
[0027] (16) In the method of (15), late fusion is performed using a majority voting model, a weighted average model, or a stacking model.
[0028] (17) An apparatus includes one or more processors and one or more memories configured to acquire data of multiple modalities, acquire one or more group classification criteria, acquire one or more trained machine learning models, generate multi-modality data sets from the data of the multiple modalities according to the one or more group classification criteria, each multi-modality data set including data of at least two modalities, and generate one or more first classification results by performing intermediate fusion on the one or more multi-modality data sets using one or more of the trained machine learning models.
[0029] (18) One or more computer-readable storage media storing instructions that cause one or more computing devices to perform operations including: acquiring data of a plurality of modalities, acquiring one or more group classification criteria, acquiring one or more trained machine learning models, generating multi-modality data sets from the data of the plurality of modalities according to the one or more group classification criteria, each multi-modality data set including data of at least two modalities, and generating one or more first classification results by performing intermediate fusion on the one or more multi-modality data sets using one or more of the trained machine learning models.
[0030] Various embodiments will now be described with reference to the accompanying drawings.
[0031] 1 shows an exemplary embodiment of a medical data processing system 1. The medical data processing system 1 includes at least one data processing device 10, one or more imaging devices 20A-20C, a biomarker analysis device 30, and a server 40.
[0032] In this embodiment, the imaging devices 20A-20C include a C-arm computed tomography (CT) scanner 20A, a mammography device 20B, and a magnetic resonance imaging (MRI) scanner 20C. These are examples of imaging devices. Other embodiments may include more or fewer imaging devices, different imaging devices (e.g., ultrasound device, photoacoustic imaging device), one imaging device (e.g., only mammography device 20B), a different combination of imaging devices (e.g., mammography device 20B and ultrasound device), etc.
[0033] When imaging devices 20A-20C perform imaging operations on a subject (e.g., a patient), they generate and output image data sets 22 that define one or more images of the subject, such as two-dimensional (2D) or three-dimensional (3D) images.
[0034] The biomarker analysis device 30 performs a liquid biopsy on a patient sample (e.g., blood). A liquid biopsy analyzes biological analytes (biomarkers) circulating in a patient liquid sample, such as blood or urine. For example, liquid biopsies can analyze biomarkers such as proteins, DNA, or RNA.
[0035] In an embodiment, the biomarker analysis device 30 is an automated assay device and includes at least one controller, an assay consumable handler, a sample loader, a sealer, and an imaging system. The assay consumable handler is operably connected to the assay consumable. The assay consumable includes multiple assay locations. The sample loader places an assay sample (e.g., containing multiple analyte molecules or particles) in the assay location of the assay consumable. The sealer applies a sealing material to the surface of the assay consumable. The imaging system collects one or more images of the assay locations of the assay consumable. The biomarker analysis device 30 may further include a bead loader separate from or corresponding to the sample loader, a rinser for cleaning the surface of the assay consumable, a reagent loader for placing reagents in the assay locations of the assay consumable, a wiper for removing excess beads from the assay substrate surface, etc. One or more controllers may also control other components of the biomarker analysis device 30 (e.g., the assay consumable handler, the sample loader, the sealer, the imaging system, the bead loader, and the wiper).
[0036] Based on images collected by the imaging system of the biomarker analysis device 30, the biomarker analysis device 30 generates and outputs biomarker data 32 (e.g., multiple biomarker data groups 32) after completing biomarker analysis of a sample from a subject (e.g., a patient). The biomarker data 32 indicates the results of the biomarker analysis (e.g., one biomarker data group 32 indicates the results of each liquid biopsy). For example, the biomarker analysis device 30 can perform liquid biopsy, which determines the presence or absence of cancer in a subject in a minimally invasive manner, by analyzing biological analytes (biomarkers such as proteins, DNA, and RNA) circulating in the blood around a lesion. When performing a liquid biopsy, the biomarker analysis device 30 measures the concentration (amount) of the biological analyte, detects mutations (in the case of nucleic acid (NA) analysis), or measures copy number variation (CNV).
[0037] Examples of protein biomarkers include CA15-3, CA19-9, and CA125. Examples of tumor-associated autoantibody (TAAbs) biomarkers include HSP-47 (SERPINH1), HSP-60, HSP70, HSP-90, CA15-3, SELL, ANGPTL3, HSP-90, HGF, HER2, progranulin, c-myc, HOXD10, SOX2, endostatin, GIPC-1, GRAF-1A, GPR157, DKK1, ZNF514, and BDNF. Other examples of biomarkers that can be used are described in WO2022 / 140576. For purposes of describing biomarkers, the summary, drawings, and detailed description of WO2022 / 140576 are incorporated herein by reference, except for definitions, subject matter disclaimers, and disclaimers, and except to the extent the incorporated content is inconsistent with the specific content of this disclosure. In such cases, the language of this disclosure will control.
[0038] The server stores electronic medical record (EMR) data 42 defining one or more electronic medical records. The electronic medical records include at least some of the following: clinical findings, such as test names, test results, diagnoses, and radiology information; medication information; treatment information; dates and times of clinical findings, medications, and treatments; patient symptoms; and patient information (e.g., weight, height, age).
[0039] For example (e.g., at the time of breast cancer diagnosis), EMR data 42 may indicate at least some of the following: whether the patient has had a breast biopsy, how many biopsies the patient has had, how many pregnancies the patient has had, the patient's age at the birth of their first child, the patient's age at the birth of their last child, whether the patient has breastfed for at least one month, whether the patient has smoked at least 100 cigarettes, whether the patient has ever drunk alcohol, the patient's average weekly alcohol intake, whether the patient has had diabetes or hyperglycemia excluding pregnancy, whether the patient has ever used hormone-based contraception, and whether the patient has ever undergone hormone therapy.
[0040] The server also stores image data 22 and biomarker data 32 .
[0041] The data processing device 10 acquires data of an imaging modality (image data set 22), data of a biomarker modality (biomarker data 32), and, in one embodiment, data of another modality, such as a text modality (e.g., EMR data 42). The data processing device 10 generates one or more classification results using at least two modalities of the image data set 22, the biomarker data 32, and the EMR data 42. The classification result may be, for example, a diagnosis or a medical prediction. That is, the classification result indicates the presence or absence of a disease. The classification result may also indicate the presence or absence of multiple diagnoses. The classification result may also indicate the likelihood of each diagnosis included in the classification result. For example, the classification result may indicate at least some of the following: the presence or absence of cancer, such as breast cancer; whether the tumor is benign or malignant; the type of cancer, such as invasive or non-invasive; and the cancer molecular subtype (e.g., luminal type, Her2-enriched type, or triple-negative breast cancer subtype).
[0042] In this manner, the data processing device 10 uses (e.g., fuses) data from at least two modalities (image data 22, biomarker data 32, EMR data 42) to generate one or more classification results, such as a diagnosis. By using data from at least two modalities to generate a classification result, the data processing device 10 can generate a more accurate classification result.
[0043] For example, breast cancer diagnosis is commonly performed using medical imaging, particularly mammograms. Mammograms can detect morphological lesions within the breast that may be associated with breast cancer, such as lumps, calcifications, and architectural distortions. However, it is very difficult to distinguish benign from malignant lesions using mammograms alone.
[0044] As a result, pathological diagnosis of breast cancer cells is usually performed by performing a needle biopsy to determine whether the findings are benign or malignant. That is, if a tumor-like lesion is seen on a mammogram, a needle biopsy is performed even if the lesion appears benign. If the result is benign, the biopsy is unnecessary. Furthermore, as biopsy results show, diagnoses based solely on mammograms tend to have a very high false-positive rate (approximately 80% of biopsies are positive).
[0045] The data processing device 10 integrates and interpolates diagnostic information obtained from two or more modalities to improve classification accuracy (e.g., diagnostic accuracy, predictive accuracy) compared to classification based on mammograms alone. This can reduce unnecessary biopsies. Furthermore, although available data differs depending on the facility, the data processing device 10 can be adapted to the data available at each facility.
[0046] The data processing device 10 uses early data fusion (early fusion), intermediate data fusion (intermediate fusion), and late data fusion (late fusion) to generate classification results, and also determines the timing at which to use early fusion, intermediate fusion, or late fusion.
[0047] Early fusion involves introducing data from one or more modalities into a single information space or machine learning (ML) model that generates classification results based on the modalities of the input data. That is, the ML model used for early fusion (the early fusion model) accepts input data from one or more modalities and outputs classification results based on this input data.
[0048] Intermediate fusion involves a set of ML models, where one or more of the ML models in the set extract features from data from two or more modalities, and one or more of the other models in the set integrate (or combine) the extracted features to produce a classification result. That is, in intermediate fusion, at least one ML model accepts data from one or more modalities as input and outputs features based on the input data. At least one ML model also accepts features as input and outputs one or more classification results based on the input features.
[0049] Late fusion uses one or more late fusion models to aggregate multiple classification results into one. That is, the late fusion model accepts classification results as input and outputs one or more classification results based on the input classification results. To distinguish between the classification results output from the late fusion model and those output from the early fusion model or the intermediate fusion ML model, the classification result output from the late fusion model will be referred to as the "second classification result," and the classification results output from the early fusion model and the intermediate fusion ML model will be referred to as the "first classification result."
[0050] The output set also defines all possible classification results that can be output by the ML model (either early fusion or late fusion) and the late fusion model. For example, one output set may include only the classification results of healthy, benign, and malignant. That is, in such an embodiment, the classification result output by the ML model or late fusion model is one of healthy, benign, and malignant (or only one in some embodiments). As another example, if another output set includes different cancers such as breast cancer, lung cancer, stomach cancer, liver cancer, pancreatic cancer, and colon cancer, the classification result may include, for example, an occurrence probability of breast cancer = 75%, lung cancer = 20%, etc.
[0051] For example, if the classification goal is to diagnose the presence or absence of breast cancer, the output set may consist of healthy, benign, and malignant. Alternatively, the output set may consist of breast cancer, lung cancer, other cancer, and no cancer. Also, as mentioned above, the classification result may include each element of the output set along with its probability of occurrence.
[0052] FIG. 2 illustrates an exemplary embodiment of an operational flow for generating classification results from data from multiple modalities. Note that while this operational flow and other operational flows herein are presented in a specific order, depending on the embodiment, at least some of the operations may be performed in an order different from the presented order. Examples of different orders include parallel, concurrent, overlapping, reordered, simultaneous, incremental, and alternating. Accordingly, operational flows according to other embodiments may have blocks omitted, added, reordered, combined, or divided.
[0053] Furthermore, although the operational flow of the embodiments described herein is performed by one data processing device 10, in some embodiments the operational flow is performed by two or more data processing devices 10 or one or more specially configured computing devices.
[0054] The flow begins at block B200 and proceeds to block B205, where the data processing device 10 receives data, one or more machine learning models (ML models), one or more late fusion models, and classification settings. The data is data from multiple modalities. Different modalities have different data formats, and different modalities capture and represent different features. For example, the modalities may include imaging modalities, biomarker modalities, electronic medical record (EMR) modalities, etc.
[0055] Imaging modalities include, for example, X-ray modalities (e.g., two-dimensional (2D) X-ray images, computed tomography (CT) images), magnetic resonance imaging (MRI) modalities, fluoroscopy modalities, ultrasound modalities, tomosynthesis modalities, positron emission tomography (PET) modalities, endoscopic modalities, digital pathology scanner modalities, and the like.
[0056] Biomarker modality data (biomarker data) indicates the presence of biomarkers. For example, biomarker data indicates the quantity of each biomarker (e.g., counts or concentrations, such as mass concentration, molecular concentration, number concentration, or volume concentration). Examples of biomarkers include protein biomarkers, protein biomarkers that perform one or more specific biological roles, tumor-associated autoantibody (TAA) biomarkers, deoxyribonucleic acid (DNA) biomarkers, ribonucleic acid (RNA) biomarkers, metabolite and lipid biomarkers, circulating tumor cell biomarkers, exosome biomarkers, cytometry biomarkers, and glycobiological biomarkers. Thus, biomarker data includes one or more biomarker identifiers (e.g., names, codes) along with numerical data indicating the quantity of each of the one or more biomarkers. Examples of biomarker modalities include protein modality, DNA modality, RNA modality, metabolite modality, lipid modality, exosome modality, cytometry modality, glycobiological modality, and ion modality.
[0057] Data can also be obtained in the form of multiple single modality data sets (multiple single modality (SM) sets), where each SM set includes data from only one modality. For example, a single modality data set (SM set) defines an image or a portion of an image (e.g., a region of interest (ROI)), and the single modality data set (SM set) includes identifiers and / or numerical data for a set of biomarkers.
[0058] A classification configuration includes one or more group classification criteria (defined based on one or more group classification strategies), one or more classification objectives (e.g., determining the presence or absence of a disease (e.g., breast cancer) or pathology), or one or more output sets, and the selection of an ML model to use and a terminal fusion model.
[0059] Next, in block B210, the data processing device 10 generates one or more multi-modality data groups (multi-modality (MM) groups) based on the data and one or more grouping criteria. Each MM group includes data of two or more modalities. For example, the data processing device generates one MM group by merging two SM groups of different modalities.
[0060] For example, one or more group classification criteria (which may be included in one or more group classification strategies) define multi-omics groups based on functional and biological aspects. Each multi-omics group is associated with one or more "omes." Examples of "omes" include genomes, proteomes, transcriptomes, epigenomes, metabolomes, and microbiomes. That is, multi-omics groups include, for example, genome groups, proteome groups, transcriptome groups, epigenome groups, metabolome groups, and microbiome groups. Other examples of multi-omics groups include radiomics groups and phenotype groups.
[0061] Data can also be grouped according to its data format. For example, the format can be numeric, text, etc. Numeric data (data in numeric format) can be continuous (continuous format) or discrete (discontinuous format).
[0062] Data may also be classified into two or more multimodality groups, for example, F-fludeoxyglucose (FDG) PET images (FDG-PET images) can be classified into both radiomics and metabolomic MM groups.
[0063] That is, examples of MM groups include: a group containing different radiomics data, a group containing biomarker data from multiple multi-omics groups, a group containing both radiomics data and biomarker data belonging to a specific multi-omics group (or a specific subset (preferred subset) of a multi-omics group), and a group containing both textual information (e.g., EMR data) and numerical data (e.g., radiomics data, biomarker data).
[0064] Next, the flow proceeds to block B215, where the data processing device 10 generates an initial classification result for each of the one or more MM groups. When generating the initial classification result, the data processing device 10 uses at least two machine learning (ML) models. The data processing device 10 inputs each MM group to an ML model that outputs features based on the input. Next, the data processing device 10 inputs the features to an ML model that outputs at least one initial classification result based on the input. For example, in block B215, the data processing device uses an artificial neural network to extract features from each MM group and inputs the features to a shallow classification model or a deep neural network trained to output an initial classification result based on the input features. That is, in block B215, the data processing device performs intermediate fusion on the MM groups.
[0065] Furthermore, after acquiring the features, the data processing device 10 may generate a new group from the acquired features and other data (e.g., one or more single-modality (SM) groups) and input it into an ML model that outputs at least one initial classification result.
[0066] For each ML model, some models accept a subset of the data from one or more modalities. For example, one ML model may accept a subset of image data as input, while another model may accept a subset of biomarker data. For example, an ML model that accepts a subset of image data as input may also accept one or more of the following as input: a 2D or 3D overall image (e.g., an overall mammogram, a screening mammogram, or a diagnostic mammogram), or a region of interest (ROI) within a 2D or 3D image (e.g., an ROI within a screening mammogram, an ROI within a diagnostic mammogram). As another example, some ML models may accept other electronic data, such as EMR data, as input.
[0067] As mentioned above, the ML model may be an artificial neural network, such as a deep neural network, a feedforward neural network (e.g., a convolutional neural network), a recurrent neural network, a deep belief neural network, etc.
[0068] Each neural network operates on different data (e.g., accepts different data as input) and outputs features extracted from its respective input. For example, in one embodiment, one neural network accepts as input all biomarkers for an enzyme group, one neural network accepts as input biomarkers for a G protein group (a proper subset of the protein group), one neural network accepts as input DNA biomarkers, one neural network accepts as input an entire mammogram, and one neural network accepts as input a region-of-interest image from a mammogram (an image of a region of interest cropped from the entire image). In one embodiment, each neural network accepts the same inputs as accepted by one or more of the classification models listed in Table 1 below.
[0069] The features output from the neural network are those that are most relevant (depending on their presence or absence) to generating the first or second classification result.
[0070] Other examples of ML models include classification models (e.g., shallow classification models), including, for example, logistic regression models, K-nearest neighbor (KNN) classification models, naive Bayes classification models, decision tree classification models, support vector machine (SVM) classification models (e.g., with a linear kernel or a radial kernel), Gaussian process classification models, multilayer perceptron (MLP) classification models, ridge regression classification models, random forest classification models, quadratic discriminant analysis classification models, AdaBoost classification models, gradient boosting classification models, linear discriminant analysis (LDA) classification models, extra-tree classification models, and boosting classification models (e.g., extreme gradient boosting classification models and light gradient boosting classification models).
[0071] Next, in block B220, the data processing device 10 determines whether there are any unintegrated SM groups. An unintegrated SM group is an SM group that was not incorporated into an MM group in block B210 and was not integrated with a feature in block B215. If there are no unintegrated SM groups (B220=No), the flow proceeds to block B230. If there is at least one unintegrated SM group (B220=Yes), the flow proceeds to block B225.
[0072] In block B225, the data processing device 10 generates a classification result for each unintegrated SM group. The data processing device 10 uses an ML model as an early fusion model to generate the classification result in block B225. As described above, an early fusion model accepts data of one or more modalities (including one or more SM groups) as input and outputs one or more classification results based on the input data. Each ML model used with the unintegrated SM groups may accept input of only one data format.
[0073] The data processing device 10 may input each uncombined SM group to only one early fusion model, or may input uncombined SM groups one at a time. Thus, one classification result is generated for each uncombined SM group. The data processing device 10 may also input different uncombined SM groups to different early fusion models. The data processing device 10 may also input two or more uncombined SM groups to several early fusion models, such as an early fusion model that accepts data from multiple modalities. If one or more of the two or more uncombined SM groups is not in a format acceptable to the early fusion model, the data processing device 10 converts the uncombined SM group into a format acceptable to the early fusion model. In this way, the data of two or more uncombined SM groups input to the early fusion model is in the same format even if the modalities are different.
[0074] Additionally, the early fusion model may be a classification model that acts as a classifier and outputs a first classification result.
[0075] Each classification model outputs a classification result. Each classification model operates on different data (e.g., accepts different SM groups as input). For example, Table 1 shows the data that 23 classification models operate on (accept as input). The biomarker modalities of the input data in Table 1 include the following: growth factor modality, enzyme modality, cell proliferation modality, G protein modality, phosphorylation modality, inflammation modality, tumor-associated autoantibody (TAAbs) modality, DNA modality, RNA modality, metabolite / lipid modality, circulating tumor cell modality, and exosome / EV modality. The image modality of the input data in Table 1 is the mammogram modality.
[0076] [Table 1]
[0077] TIFF2025114522000003.tif94162
[0078] For example, if initial classification is performed on 23 applicable unintegrated SM groups using the 23 classifiers in Table 1 in block B225, the data processing device 10 will generate 23 first classification results in block B225.
[0079] Flow then proceeds to block B230, where the data processing apparatus 10 generates a second classification result based on some or all of the first classification result generated in block B215, or, if block B225 is executed, in block B225. In generating the second classification result, the data processing apparatus 10 inputs the first classification result to at least one late fusion model, which outputs a second classification result based on the input. The second classification result may include one element of the output set, or may include multiple elements of the output set along with a probability for each element.
[0080] A late fusion model may or may not be an ML model. Examples of late fusion models include majority voting models, weighted average models, and stacking models.
[0081] TIFF2025114522000004.tif28162
[0082]
number
[0083] TIFF2025114522000006.tif41162
[0084] TIFF2025114522000007.tif15162
[0085]
number
[0086] TIFF2025114522000009.tif28163
[0087] TIFF2025114522000010.tif21163
[0088]
number
[0089] TIFF2025114522000012.tif22163
[0090] TIFF2025114522000013.tif21163
[0091]
number
[0092] TIFF2025114522000015.tif13163
[0093] TIFF2025114522000016.tif22163
[0094]
number
[0095] TIFF2025114522000018.tif24163
[0096] Next, in block B235, the data processing device 10 stores or outputs the second classification result. In block B235, the first classification result may also be stored or output. The flow then ends in block B240.
[0097] 3 illustrates an exemplary embodiment of an operational flow for generating classification results from multi-modality data. The flow begins at block B300 and proceeds to block B305, where the data processing device 10 obtains multi-modality data, one or more group classification criteria (and other classification settings), one or more machine learning models (ML models), and one or more end-stage fusion models. The multi-modality data consists of multiple single-modality (SM) groups.
[0098] Next, in block B310, the data processing device 10 generates one or more multi-modality data sets (MM sets) based on the data and one or more grouping criteria (eg, at least one grouping strategy).
[0099] Flow continues to block B315, where the data processing device 10 generates an intermediate data set for each of one or more MM sets. Block B315 involves inputting each MM set to an ML model trained to extract and output features in response to the input. The output features constitute intermediate data in the intermediate data set. All of the MM sets may be input to the same ML model, or subsets of the MM sets may be input to different ML models. Block B315 is also part of intermediate fusion.
[0100] In block B320, the data processing device 10 generates one or more high-correlation groups based on the intermediate data groups and the SM groups that were not integrated into the MM groups in block B310. The high-correlation groups are generated by integrating each intermediate data group with other intermediate data groups that have a high correlation with that intermediate data group (a correlation exceeding a threshold or satisfying one or more other criteria) and by combining each intermediate data group with one or more SM groups that were not integrated into the MM groups in block B310 and have a high correlation with that intermediate data group. Furthermore, if there are no intermediate data groups that are highly correlated with other intermediate data groups or if there are no SM groups that were not integrated into the MM groups in block B310, the high-correlation group includes only the intermediate data group. That is, the high-correlation group may include (i) only one intermediate data group, (ii) multiple intermediate data groups, or (iii) one or more intermediate data groups and one or more SM groups that were not integrated into the MM groups in block B310.
[0101] Furthermore, the embodiment of Figure 3 (and some of the other embodiments described herein) uses other similarity or dissimilarity measures (e.g., Euclidean distance) in addition to or instead of correlation (intra-group / inter-modality correlation, inter-group correlation). Thus, correlation is an example of a similarity measure, and other embodiments utilize other similarity measures.
[0102] For example, in one embodiment, when comparing data of the same size (e.g., when comparing two types of data of the same data size), only correlation or Euclidean distance is used. When comparing different data sizes (e.g., data sets of different data sizes, e.g., different numbers of records), such an embodiment uses cosine similarity and / or maximum mean discrepancy (MMD). For example, such an embodiment uses cosine similarity and / or MMD to compare 50 data sets with 5 features to 100 data sets with 5 features. Additionally, regardless of whether the data is image data or numerical data (e.g., from blood tests), data similarity can be evaluated by extracting feature vectors from the target data and aligning them to the same scale using a data standardization process (e.g., -1.0). <x<1.0へのスケーリング)。
[0103] TIFF2025114522000019.tif15163
[0104]
number
[0105] A result of "1" indicates perfect similarity, a result of "0" indicates neutrality, and a result of "-1" indicates no similarity.
[0106] TIFF2025114522000021.tif13159
[0107]
number
[0108] The flow continues to block B325, where the data processing device 10 generates a first classification result for each of the one or more highly correlated groups. Block B325 includes inputting the highly correlated groups to an ML model trained to extract and output features in response to input. All of the highly correlated groups may be input to the same ML model, or subsets of the highly correlated groups may be input to different ML models. Block B325 is also part of intermediate fusion.
[0109] The flow then proceeds to block B330, where the data processing device 10 determines whether there are any unintegrated SM groups. An unintegrated SM group is an SM group that was not integrated into an MM group in block B310 and was not integrated into a high correlation group in block B320. If there are no unintegrated SM groups (B330=No), the flow proceeds to block B340. If there is at least one unintegrated SM group (B330=Yes), the flow proceeds to block B335.
[0110] In block B335, the data processing device 10 generates one or more first classification results based on the unintegrated SM groups. The data processing device 10 uses an early fusion model when generating the first classification results. Generating the first classification results includes inputting one or more of the unintegrated single-modality groups into the early fusion model. The data processing device 10 inputs each unintegrated single-modality group into the early fusion model, but does not input another unintegrated single-modality group together with the input unintegrated single-modality group (one-to-one mapping). However, the data processing device 10 may input multiple unintegrated single-modality groups together into the same early fusion model (one-to-many mapping). In this way, in block B335, the data processing device 10 obtains one or more first classification results.
[0111] Flow proceeds to block B340, where the data processing device 10 generates a second classification result based on at least some (e.g., all) of the first classification results generated in block B325 and, if block B335 is executed, generated in block B335. To generate the second classification result, the data processing device 10 inputs the first classification results into at least one late fusion model, which outputs a second classification result based on the input.
[0112] Next, in block B345, the data processing device 10 stores or outputs the second classification result. The flow ends in block B350.
[0113] 4 shows an exemplary embodiment of an operational flow for generating classification results from data of multiple modalities. The flow starts at block B400 and proceeds to block B405, where the data processing device 10 acquires data, one or more group classification criteria (and other classification settings), one or more machine learning models (ML models), and one or more final fusion models. Next, at block B410, the data processing device 10 determines whether the acquired data includes data of multiple modalities.
[0114] If the acquired data does not include data of multiple modalities (B410=No), the flow proceeds to block B412, where the data processing device 10 generates a classification result based on the acquired data (including data of only one modality). Further, early fusion may be performed in block B412. For example, if the acquired data are all biomarker modalities, but some of the data are protein biomarkers A, B, and C, and another part of the data are protein biomarkers D, E, and F, then in block B412, this data is fused by early fusion. The flow then ends in block B475. If the acquired data includes data from multiple modalities (B410=Yes), the flow proceeds to block B415.
[0115] In block B415, the data processing device 10 determines whether any of the modalities contain incomplete data (e.g., whether there are any SM groups containing incomplete data). For example, in one SM group, one or more biomarkers are not measured, the amount of one or more biomarkers is invalid (e.g., negative), the data of one or more biomarkers indicates an error, or some image data is omitted. If one or more of the modalities contain incomplete data (B415=Yes), the flow proceeds to block B417, where the data processing device 10 adds a late fusion flag to the incomplete data (e.g., to the SM group containing incomplete data). The flow proceeds to block B420. If there are no modalities containing incomplete data (B415=No), the flow proceeds directly to block B420.
[0116] In block B420, the data processing device 10 generates an MM group by classifying the acquired data (e.g., an SM group) according to a group classification policy including one or more group classification criteria. The group classification may exclude data flagged as late fusion, i.e., incomplete data.
[0117] For example, Figure 5A shows an example of an MM group formed by classifying data from multiple modalities and an example of an SM group that was not integrated into the MM group. MM group G1 includes data from modalities M1, M2, and M3, which in this example are SM groups SM1, SM2, and SM3, respectively. MM group G2 includes data from modalities M4, M5, and M6, which in this example are SM groups SM4, SM5, and SM6, respectively. Data from modality M6 is included in SM group SM7, which is not integrated with data from other modalities. SM group SM1 in MM group G1 may be the same as or different from SM group SM4 in MM group G2.
[0118] Next, in block B425, the data processing device 10 calculates the intra-group / inter-modality correlation between data of different modalities within the same MM group. For example, in the embodiment of FIG. 5A, the data processing device 10 calculates the intra-group / inter-modality correlation between data of modalities M1, M2, and M3 within MM group G1, and calculates the intra-group / inter-modality correlation between data of modalities M1, M4, and M5 within MM group G2. For example, the correlation between SM groups SM1, SM2, and SM3 is calculated, and the correlation between SM groups SM4, SM5, and SM6 is calculated.
[0119] Next, the flow proceeds to block B430, where the data processing device 10 determines whether or not there is data with low intra-group / inter-modality correlation in the MM group (i.e., whether or not there is data in one modality data in the MM group (e.g., data of the SM group) that shows inter-modality correlations with data of other modalities in the MM group (e.g., data of the SM group of other modalities) that are all below a threshold). If there is data with low intra-group / inter-modality correlation in the MM group data (B430=Yes), the flow proceeds to block B432. In block B432, the data processing device 10 deletes data of a modality (e.g., the SM group) that has low correlation with the corresponding MM group, and sets a late fusion flag for the deleted data (e.g., the deleted SM group).
[0120] For example, Figure 5B shows the MM group of Figure 5A after removing data with low intra-group / inter-modality correlations. Because there is no data with low intra-group / inter-modality correlations in MM group G1, no data is removed from MM group G1. However, SM group SM5 (data from modality M4) shows low intra-group / inter-modality correlations with SM groups SM4 and SM6 (data from modalities M1 and M5), but SM groups SM4 and SM6 (data from modalities M1 and M5) do not show low intra-group / inter-modality correlations with each other. Therefore, SM group SM5 (data from modality M4) is removed from MM group G2. Additionally, a late fusion flag is set for SM group SM5 (data from modality M4).
[0121] Flow continues to block B435.
[0122] If there is no data with low intra-group / inter-modality correlation in the MM group (B430=No), the flow proceeds to block B435.
[0123] In block B435, the data processing device 10 determines whether there are any more groups in the MM group that include data of multiple modalities. If there are no MM groups that include data of multiple modalities (B435=No), the flow proceeds to block B465.
[0124] If one or more of the MM groups includes data of multiple modalities (B435=Yes), the flow proceeds to block B437. In block B437, the data processing device 10 performs a portion of intermediate fusion using one or more of the ML models for each MM group including data of multiple modalities. This portion of intermediate fusion outputs an intermediate data group (intermediate data (ID) group).
[0125] For example, FIG. 5C shows the intermediate data group and SM group after a portion of the intermediate fusion has been performed on the MM group of FIG. 5B. The data of MM group G1, which includes data from modalities M1, M2, and M3, is input to one or more ML models. The ML models generate data group I1 based on the data of MM group G1. For example, one or more ML models extract features from the data from modalities M1, M2, and M3, and the extracted features constitute intermediate data group I1. Furthermore, the data of MM group G2 is input to one or more ML models, which output intermediate data group I2 based on the data of MM group G2. SM groups SM5 and SM7 remain unchanged.
[0126] Flow continues to block B440.
[0127] In block B440, the data processing device 10 calculates the intra-group correlation between the intermediate data group, the result of block B420, and the SM group that was not merged into the MM group. The data processing device 10 calculates the correlation for some or all pairs of combinations. Each pair includes an intermediate data group and another intermediate data group or an SM group. That is, each pair includes (i) two intermediate data groups or (ii) one intermediate data group and one SM group. For example, if there are three intermediate data groups and two SM groups, the data processing device 10 calculates a maximum of nine correlations.
[0128] Next, in block B445, the data processing device 10 determines whether all intra-group correlations are high (exceeding a threshold). If there is at least one intra-group correlation that is not high (B445=No), the flow proceeds to block B447, where the data processing device 10 sets a terminal fusion flag for the SM group that does not show a high correlation with the intermediate data group. The flow then proceeds to block B450. If all intra-group correlations are high (B445=Yes), the flow proceeds to block B450.
[0129] In block B450, the data processing device 10 reintegrates the intermediate data group and the SM group with high intra-group correlation into one or more high correlation groups. In one embodiment, all intermediate data groups showing at least one high intra-group correlation are integrated into one high correlation group. Also, in one embodiment, the high correlation group includes only the intermediate data group and SM groups with high intra-group correlation with each other. The high correlation group does not include SM groups with low intra-group correlation with the intermediate data group. A late fusion flag is set for each SM group that does not show high intra-group correlation.
[0130] For example, Figure 5D shows the intermediate data group and SM group of Figure 5C after the intermediate data group and the SM group with high intra-group correlation have been reintegrated into one or more high-correlation groups. In this example, intermediate data group I2 is highly correlated with SM group SM7, but there are no intermediate data groups or SM groups highly correlated with intermediate data group I1. Therefore, intermediate data group I2 and SM group SM7 are integrated into high-correlation group G4 (updated group G4), but intermediate data group I1 is not integrated with other intermediate data groups or SM groups.
[0131] In block B455, the data processing device 10 performs a portion of intermediate fusion on one or more high correlation groups by inputting the one or more high correlation groups into the ML model. The ML model outputs each first classification result based on the input high correlation groups. The data processing device 10 also performs a portion of intermediate fusion on the remaining intermediate data groups (intermediate data groups not integrated into the high correlation groups) by inputting the remaining intermediate data groups into the ML model. The ML model outputs each first classification result based on the input intermediate data groups. This ends the intermediate fusion started in block B437 in block B455.
[0132] For example, FIG. 5E shows a first classification result generated based on the high-correlation group and intermediate data group of FIG. 5D, along with an SM group. Data in the high-correlation group G4, which includes the intermediate data group I2 and the SM group SM7, is input to an ML model, and the ML model generates and outputs a first classification result Cl1 based on the data in the high-correlation group G4. Before inputting the SM group to the ML model in block B455, the data processing device 10 may perform preprocessing on the SM group to convert the data into a format acceptable to the ML model. Alternatively, the data processing device 10 may use an ML model that accepts the data as is. The intermediate data group I1 is input to an ML model that generates a first classification result Cl2 based on the data in the intermediate data group I1.
[0133] The flow proceeds to block B460, where the data processing device 10 determines whether there is a group of SMs for which the late fusion flag has been set. If there is at least one group of SMs for which the late fusion flag has been set (B460=Yes), the flow proceeds to block B465.
[0134] In block B465, the data processing device 10 performs early fusion on the group of SMs to which the late fusion flag has been added. That is, the data processing device 10 uses one or more early fusion models to generate one or more first classification results based on the group of SMs to which the late fusion flag has been added. For example, FIG. 5F shows first classification results Cl1 and Cl2 generated based on the high correlation group and intermediate data group of FIG. 5D, and a first classification result Cl3 generated based on the group of SMs to which the late fusion flag has been added of FIG. 5E. Because the group of SMs SM5 has been added with a late fusion flag (which identifies an unintegrated group of SMs), the data processing device 10 uses an early fusion model to generate the first classification result Cl3 based on the group of SMs SM5.
[0135] If none of the SMs have the terminal fusion flag set (B460=No), the flow proceeds to block B470.
[0136] In block B470, the data processing device 10 performs late fusion using one or more late fusion models, the classification results generated in block B455, and the classification results generated in block B465. For example, the data processing device generates a second classification result based on the first classification results Cl1, Cl2, and Cl3 of FIG. 5F.
[0137] Finally, in block B475, the data processing device 10 saves or outputs the second classification result (and in some embodiments the first classification result), and the flow ends.
[0138] 6 illustrates an exemplary embodiment of group classification and fusion for data from multiple modalities. First, data from the following modalities are acquired: protein biomarkers in modality M1, electronic medical records (EMR) in modality M2, mammograms in modality M3, and tumor-associated autoantibody biomarkers in modality M4. The data from modality M1 is included in SM group SM1, and the data from modality M2 is included in SM group SM2. The data from modality M3 is included in SM group SM3, and the data from modality M4 is included in SM group SM4.
[0139] The group classification policy includes group classification criteria that integrate M1 (protein biomarkers) and M4 (tumor-associated autoantibody biomarkers) that belong to the same proteomics category. Therefore, the data processing device 10 integrates SM group SM1 and SM group SM4 to create MM group G1 (B210 in Figure 2, B310 in Figure 3, B420 in Figure 4). In addition, the intra-group / inter-modality correlation degree of the data of M1 and M4 exceeds a threshold (B425-B432 in Figure 4).
[0140] The data processing device 10 performs intermediate fusion on the MM group G1. The intermediate fusion consists of two parts. The data processing device 10 performs the first part of intermediate fusion on the MM group G1 to create the intermediate data group I1 (B215 in FIG. 2, B315 in FIG. 3, B437 in FIG. 5). The data processing device 10 also reintegrates the intermediate data group I1 and the SM group SM2 into the high correlation group G2 (B215 in FIG. 2, B320 in FIG. 3, B450 in FIG. 4). As part of the reintegration, the data processing device 10 also determines that the SM group SM2 including the intermediate data group I1 and data of modality M2 has a high intra-group correlation (B320 in FIG. 3, B440 in FIG. 4). Furthermore, the data processing device 10 sets a final fusion flag on the SM group SM3 including data of modality M3 (B247 in FIG. 2). The final fusion flag indicates an unintegrated SM group.
[0141] Furthermore, the data processing device 10 executes the second part of intermediate fusion on the data of the high correlation group G2 to generate a first classification result Cl1 (B215 in FIG. 2, B325 in FIG. 3, B455 in FIG. 4).
[0142] The data processing device 10 generates a first classification result for each unintegrated SM group, i.e., generates classification result Cl2 for SM group SM3 (early fusion) (B220-B225 in FIG. 2, B330-B335 in FIG. 3, B460-B465 in FIG. 4). For example, if there is an SM group for which the late fusion flag is set, the data processing device 10 determines that the SM group is unintegrated.
[0143] Finally, the data processing device 10 generates a second classification result Cl3 based on the first classification results Cl1 and Cl2 (terminal fusion) (B230 in FIG. 2, B340 in FIG. 3, B470 in FIG. 4).
[0144] 7 illustrates an exemplary embodiment of grouping and fusing data from multiple modalities. First, data from the following modalities are acquired: 2D mammogram (modality M1), PET (modality M2), protein biomarker (modality M3), and EMR (modality M4). The data from modality M1 is included in SM group SM1, the data from modality M2 is included in SM group SM2, the data from modality M3 is included in SM group SM3, and the data from modality M4 is included in SM group SM4.
[0145] The group classification strategy integrates M1 (2D mammogram) and M2 (PET), both of which are radiomics groups. Therefore, the data processing device 10 integrates SM group SM1 and SM group SM2 to create MM group G1 (B210 in FIG. 2, B310 in FIG. 3, B420 in FIG. 4). In addition, the intra-group / inter-modality correlation between SM group SM1 and SM group SM2 exceeds a threshold (B425-B432 in FIG. 4).
[0146] The data processing device 10 performs a two-part intermediate fusion based on the MM group G1. The data processing device 10 performs the first part of the intermediate fusion on the MM group G1 to create an intermediate data group I1 (B215 in FIG. 2, B315 in FIG. 3, B437 in FIG. 5).
[0147] The data processing device 10 also reintegrates the intermediate data group I1 and the SM group SM4 into the high correlation group G2 (B215 in FIG. 2, B320 in FIG. 3, B450 in FIG. 4). As part of the reintegration, the data processing device 10 also determines that the intra-group correlation between the intermediate data group I1 and the SM group SM4 including data of modality M4 is high (B320 in FIG. 3, B440 in FIG. 4). Furthermore, the data processing device 10 sets a late fusion flag for the SM group SM3 including data of modality M3 (B247 in FIG. 2). The late fusion flag indicates an unintegrated SM group.
[0148] Furthermore, the data processing device 10 executes the second part of intermediate fusion on the data of the high correlation group G2 to generate a first classification result Cl1 (B215 in FIG. 2, B325 in FIG. 3, B455 in FIG. 4).
[0149] The data processing device 10 generates a first classification result for each unintegrated group of SMs, i.e., generates classification result Cl2 for SM group SM3 (early fusion) (B220-B225 in Figure 2, B330-B335 in Figure 3, B460-B465 in Figure 4).
[0150] Finally, the data processing device 10 generates a second classification result Cl3 based on the first classification results Cl1 and Cl2 (terminal fusion) (B230 in FIG. 2, B340 in FIG. 3, B470 in FIG. 4).
[0151] 8 illustrates an information flow for a method executable by a data processing device for generating classification results from data of multiple modalities. For example, data of multiple modalities, including image data 22, biomarker data 32, EMR data 42, etc., are acquired in block B800. In block B801, grouping of the data is performed based on the acquired data and grouping strategy 70. The grouping of the data in block B801 generates a MM group 80 based on the input data and grouping strategy 70. Also, after block B801, an unintegrated SM group 81 may remain.
[0152] In block B802, an intra-group / inter-modality correlation 82 is calculated based on the MM group 80. Next, in block B803, data of modalities with low intra-group / inter-modality correlation 82 (for example, an SM group) is deleted from the MM group 80, and an SM group 81 is generated that includes an updated MM group 83 (a group that still includes data of multiple modalities after the data deletion) and the SM group with low intra-group / inter-modality correlation 82 that has been deleted from the MM group.
[0153] In block B804, a portion of the intermediate fusion is performed on the updated MM groups 83 to generate one or more intermediate data groups 84 from the updated MM groups 83.
[0154] Next, in block B805, an intra-group correlation 85 is calculated based on the SM group 81 and one or more intermediate data groups 84. In block B806, based on the intra-group correlation 85, the highly correlated intermediate data group 84 and the SM group 81 are reintegrated into one or more highly correlated groups 87. A terminal fusion flag is set for the unintegrated SM group 81 (an SM group that has not been integrated with other groups).
[0155] In block B807, early fusion is performed on the unmerge SM group. Through the early fusion, one or more first classification results 88 are output based on the unmerge SM group 81. Also, in block B808, part of intermediate fusion is performed on the high correlation group 86, and part of intermediate fusion is performed on any of the intermediate data groups 84 not included in the high correlation group, thereby generating first classification results 88 (e.g., classification results 88 of each high correlation group 86, classification results 88 of each intermediate data group 84 not included in the high correlation group).
[0156] Finally, in block B809, a second classification result is generated by performing late fusion on all of the first classification results 88. Furthermore, late fusion is performed only if there are two or more first classification results. If there is only one first classification result 88, that one first classification result 88 is used as the second classification result 89.
[0157] TIFF2025114522000023.tif107159
[0158] TIFF2025114522000024.tif56159
[0159] In Figure 9 (and similarly in Figure 10), neurons (i.e., nodes) are represented as circles surrounding a threshold function. As a non-limiting example shown in Figure 9, inputs are depicted as circles surrounding a linear function, with arrows indicating directed connections between neurons. In one embodiment, the neural network is a feedforward network (e.g., represented as a directed acyclic graph) as illustrated in Figures 9 and 10.
[0160] TIFF2025114522000025.tif71162
[0161] In one embodiment, the neural network is a convolutional neural network (CNN). Figure 10 shows an exemplary embodiment of a CNN. A CNN uses a feedforward ANN, where the connection pattern between neurons is convolutional. For example, a CNN can be used to optimize image processing by using multiple layers of small neuron ensembles, called receptive fields, that process portions of the input data (e.g., projection data). The outputs of these ensembles are tiled so that they overlap. This processing pattern is repeated across multiple layers, which alternate between convolutional and pooling layers.
[0162] TIFF2025114522000026.tif31162
[0163] A CNN may include local or global pooling layers after convolutional layers that combine the outputs of neuron clusters within the convolutional layers, and in some embodiments, a CNN may also include various combinations of convolutional and fully connected layers with pointwise nonlinearities added at the end of or after each layer.
[0164] FIG. 12 shows the information flow in a method for generating classification results from multi-modality data.
[0165] In block B1201, the data processing device 10 acquires data of multiple modalities.
[0166] In block B1202, the data processing device 10 performs a first model selection. In the first model selection, the data processing device 10 selects one or more machine learning (ML) models 90 from a model repository 127 of the data processing device 10. Alternatively, the model repository 127 may be located in an external device (external to the data processing device 10), and the data processing device 10 may select and obtain the one or more ML models 90 from the external device via one or more electrical connections, such as a network connection.
[0167] In block B1203, the data processing device 10 inputs at least a portion of the image data 22, biomarker data 32, and EMR data 42 into selected ML models 90. Each ML model 90 accepts data input from only one modality (e.g., an early fusion model). From the data from the accepted modalities, some ML models 90 accept a subset of the data. For example, one of the ML models 90 accepts a subset of the image data as input, and another ML model 90 accepts a subset of the biomarker data as input. As another example, some of the ML models 90 accept the same input as one of the classification models listed in Table 1 above.
[0168] Additionally, the ML model 90 may be an early fusion model, such as a classification model, that operates as a classifier and outputs a first classification result 88. The ML model 90 may also be an artificial neural network (neural network) that outputs features 92.
[0169] Next, in block B1204, the data processing device 10 performs a second model selection. In the second model selection, the data processing device 10 selects an ML model 90 or a late fusion model 93 from the model repository 127. The data processing device 10 selects the ML model 90 or the late fusion model 93 based on one or more of the following: the acquired data, the ML model 90 selected in the first model selection in block B1202, the output (first classification result 88 or feature 92) of the ML model 90 selected in the first model selection in block B1202, and the purpose of the classification (e.g., the symptom being diagnosed). For example, if the output of the ML model 90 selected in the first model selection in block B1202 is the feature 92, the data processing device 10 selects an ML model that accepts the feature 92 as input in block B1204. On the other hand, if the output of the ML model 90 selected in the first model selection in block B1202 is the first classification result 88, the data processing device 10 selects the final fusion model 93 in block B1204.
[0170] Subsequently, in block B1205, the data processing device 10 performs classification using the features 92 generated in block B1203 as input to the ML model 90 selected in block B1204, or performs final classification using the first classification result 88 generated in block B1203 as input to the final fusion model 93 selected in block B1204. If the final fusion model 93 is used in block B1205, a second classification result 89 is output by the classification in block B1205, and if the ML model 90 is selected in block B1204, the first classification result 88 is output.
[0171] 13 illustrates an exemplary embodiment of an operational flow for generating classification results from data of multiple modalities. The flow begins at block B1300 and proceeds to block B1305, where a data processing device obtains data of multiple modalities (e.g., image data and biomarker data) and a classification configuration indicating at least one classification objective (e.g., a disease or medical condition being diagnosed). The flow then proceeds to block B1310, where the data processing device selects one or more first ML models based on one or more of the classification objective and the obtained data. The first ML models may be, for example, early fusion models or intermediate fusion models that output features.
[0172] Next, in block B1315, the data processing device inputs data accepted by each first ML model from the acquired data into the first ML model, and acquires an output from each first ML model.
[0173] In block B1320, the data processing device selects a second model based on one or more of the at least one classification objective, the acquired data, the first ML model, and the output of the first ML model. The second model may be a late fusion model or an ML model that accepts features as input, such as an ML model that is part of an intermediate fusion model.
[0174] The flow then proceeds to block B1325, where the data processing device inputs the output of the first ML model into a second model to obtain a classification result output from the second model. If the second model is a terminal fusion model, the classification result is the second classification result. If the second model is an ML model as part of an intermediate fusion model, the classification result is the first classification result. In block B1330, the data processing device saves or outputs the second classification result, and the flow ends at block B1335.
[0175] 14 illustrates an exemplary embodiment of an operational flow for generating classification results from data of multiple modalities. The flow begins at block B1400 and proceeds to block B1405, where the data processing device acquires data of multiple modalities (e.g., image data and biomarker data) and a classification configuration indicating at least one classification objective (e.g., a symptom being diagnosed). The flow proceeds to block B1410, where the data processing device selects one or more first ML models based on one or more of the at least one classification objective (e.g., a symptom being diagnosed) and the acquired data. The one or more first ML models may be early fusion models or intermediate fusion models, for example.
[0176] The flow proceeds to block B1415, where the data processing device determines whether the one or more first ML models output a first classification result or feature. If the one or more first ML models output a first classification result (B1415 = first classification result), they are classification models (early fusion models), and the flow proceeds to block B1420. If the one or more first ML models output features, they are neural networks trained for feature extraction (and intermediate fusion models), and the flow proceeds to block B1435.
[0177] In block B1420, the data processing device inputs data accepted by each classification model from the acquired data into the classification models and obtains a first classification result output from the classification models. Then, in block B1425, the classification models select a terminal fusion model as a second model. The second model is selected based on one or more of at least one classification purpose, the acquired data, the selected classification model, and the first classification result.
[0178] Next, in block B1430, the data processing device inputs the first classification result into a second model and obtains a second classification result output from the second model, after which flow proceeds to block B1460.
[0179] If the flow proceeds from block B1415 to block B1435 (B1415 = features), in block B1435 the data processing device inputs data that each of the one or more neural networks accepts from the acquired data into the one or more neural networks and obtains features output from the one or more neural networks.
[0180] In block B1440, the data processing device determines whether all of the features expected to be output from the one or more neural networks in block B1435 have been obtained. For example, the data processing device determines whether all of the features required as inputs to the second model (i.e., the ML model) have been obtained. If all of such features have not been obtained (B1440=No), flow proceeds to block B1445, where the data processing device selects one or more classification models as the first ML model. The one or more classification models are selected based on one or more of at least one classification objective and the obtained data. Flow then proceeds to block B1420. If all of the expected features have been obtained (B1440=Yes), flow proceeds to block B1450.
[0181] In block B1450, the data processing device selects a shallow machine learning model or a deep neural network as the second model based on one or more of the at least one classification objective, the acquired data, the neural network used in block B1435, and the features acquired in block B1435. In block B1455, the data processing device inputs the features into the second model and obtains a first classification result output from the second model. Flow then proceeds to block B1460.
[0182] In block B1460, the data processing device saves or outputs the classification result, which is the second classification result if block B1430 is executed, or the first classification result if block B1455 is executed. The flow ends in block B1465.
[0183] In this way, when executing blocks B1435 to B1455, the data processing device performs intermediate fusion, and the classification result of the intermediate fusion becomes the classification result. However, if it is determined in block B1440 that all expected features have not been obtained (indicating that features required by the second ML model in the intermediate fusion have not been obtained), the data processing device switches to processing of blocks B1420 to B1430 to perform early fusion and then late fusion.
[0184] Figure 15 shows an exemplary embodiment of operations using, for example, a hard voting majority model as the final fusion model, which replaces blocks B1420 to B1430 in Figure 14. The flow moves from block B1415 or block B1445 to block B1520. In block B1520, the data processing device inputs data that each classification model can accept from the acquired data into the classification models, and obtains a first classification result output from the classification model.
[0185] Next, in block B1522, the data processing device determines whether the number of first classification results is an odd number.
[0186] If the number of first classification results is odd (B1522=Yes), the flow proceeds to block B1524. Also, in one embodiment, if the number is an odd number greater than or equal to 3, the flow proceeds to block B1524.
[0187] In block B1524, the data processing device selects a majority model, a weighted average model, or a second-level classification model as a second model (final fusion model) based on one or more of the at least one classification objective, the acquired data, the selected classification model, and the first classification result, and flow proceeds to block B1530.
[0188] If the number of first classification results is even (B1522=No) (or in one embodiment, an odd number less than 3), flow proceeds to block B1526. In block B1526, the data processing device selects a weighted average model or a second-level classification model as a second model (terminal fusion model) based on one or more of at least one classification objective, the acquired data, the selected classification model, and the first classification results. Flow proceeds to block B1530.
[0189] In block B1530, the data processing device inputs the first classification result into a second model and obtains a second classification result output from the second model. Flow proceeds to block B1460.
[0190] Thus, in FIG. 15, if the number of first classification results is not an odd number, the data processing device does not select the majority model as the second model.
[0191] Also, in block B1522, if the number of classification results is less than a threshold (for example, 2, 3, 4, 5), the data processing device according to the embodiment generates an error notification and ends the operational flow.
[0192] 16 illustrates an exemplary embodiment of a data processing device 10. The data processing device 10 includes processing circuitry 14, one or more input interface circuits 13, memory 11, and storage device 12. The hardware components of the data processing device 10 communicate via one or more buses 19 or other electronic connections. Examples of buses 19 include a Universal Serial Bus (USB), an IEEE 1394 bus, a Peripheral Component Interconnect (PCI) bus, an Accelerated Graphics Port (AGP) bus, a Serial AT Attachment (SATA) bus, and a Small Computer System Interface (SCSI).
[0193] One or more input interface circuits 13 include communication components (e.g., GPU, network interface controller, electronic interface) for communicating with a display, gantry, network, or other input or output device (not shown), such as a keyboard, mouse, printing device, touch screen, light pen, optical storage device, scanner, microphone, drive, joystick, control pad, etc.
[0194] The storage device 12 includes one or more computer-readable storage media. In this specification, a computer-readable storage medium refers to a computer-readable medium including products such as magnetic disks (e.g., flexible disks (FDs), hard disk drives (HDDs)), optical disks (e.g., CDs, DVDs, Blu-ray®), magneto-optical disks, magnetic tapes, and semiconductor memories (e.g., non-volatile memory cards, flash memories, solid-state drives, SRAMs, DRAMs, EPROMs, EEPROMs). The storage device 12 may include both ROMs and RAMs and can store computer-readable data and computer-executable instructions. The storage device 12 is an example of a storage unit.
[0195] Memory 11 includes one or more computer-readable storage media (eg, RAM) and provides working storage for processing circuitry 14 .
[0196] The data processing device 10 further includes a data collection module 121, a data group classification module 122, an ML model selection module 123, an ML model execution module 124, a calculation module 125, and a communication module 126. In the embodiment shown in FIG. 16, each module is implemented as software (e.g., Assembly, C, C++, C#, Java, BASIC, Perl, Visual Basic, Python, etc.). However, in one embodiment, each module is implemented as hardware (e.g., customized circuitry) or a combination of software and hardware. When a module is implemented at least partially in software, the software can be stored in the storage device 12. Also, in one embodiment, the data processing device 10 includes more or fewer modules, and the modules are combined into fewer modules or divided into more modules. Furthermore, the storage device 12 includes a model repository 127 that stores machine learning models and final fusion models, a group classification criteria repository 128 that stores group classification criteria (e.g., group classification strategies), and a data / classification result / feature repository 129 that stores acquired or generated data, classification results, and features extracted from data using ML models.
[0197] The data acquisition module 121 includes instructions for causing applicable components of the data processing device 10 (e.g., the processing circuitry 14, the input interface circuitry 13, and the memory 11) to acquire data of multiple modalities. For example, the data acquisition module 121 according to one embodiment includes instructions for causing applicable components of the data processing device 10 to perform at least some of the operations described in block B205 of Figure 2, block B305 of Figure 3, block B405 of Figure 4, block B1201 of Figure 12, block B1305 of Figure 13, and block B1405 of Figure 14. The applicable components operating in accordance with the data acquisition module 121 also implement an example of a data acquisition unit.
[0198] The data group classification module 122 includes instructions for applicable components of the data processing device 10 (e.g., the processing circuitry 14, the input interface circuitry 13, and the memory 11) to classify data into multi-modality (MM) groups according to one or more grouping criteria included in the grouping strategy, remove data with low intra-group / inter-modality correlations from the MM groups, and generate groups including multiple intermediate data groups or groups including at least one intermediate data group and at least one single-modality (SM) group (e.g., high-correlation groups). For example, the data group classification module 122 according to one embodiment includes instructions for causing applicable components of the data processing device 10 to perform at least some of the operations described in block B210 of FIG. 2, blocks B310 and B320 of FIG. 3, blocks B420, B432, and B450 of FIG. 4, and blocks B801, B803, and B806 of FIG. 8. Together, the applicable components operating in accordance with the data group classification module 122 implement an example of a data group classifier.
[0199] The model selection model 123 includes instructions for causing applicable components of the data processing apparatus 10 (e.g., the processing circuitry 14, the input interface circuitry 13, the memory 11) to select one or more machine learning models or one or more late fusion models based on one or more criteria, and to mark data as late fusion based on one or more criteria, such as the acquired data (e.g., the modality of the acquired data, the content of the acquired data), the classification objective, previously selected ML models, the output of other ML models, etc. Model selection model 123 according to one embodiment includes instructions that cause applicable components of data processing apparatus 10 to perform at least some of the operations described in blocks B220-B225 of Figure 2, blocks B330-B335 of Figure 3, blocks B410-B417, B432, B447, B455-B470 of Figure 4, blocks B1202 and B1204 of Figure 12, blocks B1310 and B1320 of Figure 13, blocks B1410, B1425, B1445, B1450 of Figure 14, and blocks B1522, B1524, B1526 of Figure 15. Together, the applicable components operating in accordance with model selection model 123 implement an example of a model selector.
[0200] The model execution module 124 includes instructions for causing applicable components of the data processing apparatus 10 (e.g., the processing circuitry 14, the input interface circuitry 13, the memory 11) to input data into one or more ML models or one or more late-stage fusion models, run the ML models or late-stage fusion models, obtain outputs of the ML models or late-stage fusion models, and store the outputs in the data / classification results / feature repository 129. For example, model execution module 124 according to one embodiment includes instructions to cause applicable components of data processing device 10 to perform at least some of the operations described in blocks B215, B225-B235 of Figure 2, blocks B315, B325, B335, B340, and B345 of Figure 3, blocks B412, B437, B455, B465, and B470 of Figure 4, blocks B804, B807, B808, and B809 of Figure 8, blocks B1203 and B1205 of Figure 12, B1315 and B1325 of Figure 13, blocks B1420, B1430, B1435, and B1455 of Figure 14, and blocks B1520 and B1530 of Figure 15. Together, the applicable components operating in accordance with model execution module 124 implement an example of a model executor.
[0201] The calculation module 125 includes instructions for causing applicable components of the data processing device 10 (e.g., the processing circuitry 14, the input interface circuitry 13, and the memory 11) to calculate intra-group / inter-modality correlations between data groups and intra-group correlations between different data groups (the intermediate data group and the single-modality group). For example, the calculation module 125 according to one embodiment includes instructions for causing applicable components of the data processing device 10 to perform at least some of the operations described in block B320 of FIG. 3, blocks B425 and B440 of FIG. 4, and blocks B802 and B805 of FIG. 8. The applicable components operating in accordance with the calculation module 125 implement an example of a calculation unit.
[0202] The communication module 126 includes instructions for causing applicable components (e.g., processing circuitry 14, input interface circuitry 13, memory 11) of the data processing device 10 to communicate with other devices, such as other computing devices (e.g., consoles, servers, databases, terminals), input devices, and output devices, to obtain group classification criteria, ML models, late fusion models, and classification settings, and output second classification results. The applicable components operating in accordance with the communication module 126 also implement an example of a communication unit. When performing display control on the display device (e.g., displaying the classification results), the applicable components operating in accordance with the communication module 126 also implement an example of a display control unit. When receiving classification settings (e.g., group classification criteria, classification purpose) from the input device, the applicable components operating in accordance with the communication module 126 also implement an example of a configuration unit.
[0203] 17 shows the functional configuration of an exemplary embodiment of the data processing device 10. The data processing device 10 includes a data collection unit 1721, a data group classification unit 1722, a model selection unit 1723, a model execution unit 1724, a calculation unit 1725, a configuration unit 1730, a display control unit 1731, and a model repository 127.
[0204] The data collector 1721 acquires data of multiple modalities. For example, the data collector 1721 according to one embodiment performs at least some of the operations described in block B205 of Figure 2, block B305 of Figure 3, block B405 of Figure 4, block B1201 of Figure 12, block B1305 of Figure 13, and block B1405 of Figure 14.
[0205] The data grouping unit 1722 groups data into one or more MM groups according to one or more grouping criteria included in the grouping policy, removes data with low intra-group / inter-modality correlation from the groups, and generates a group including multiple intermediate data groups or a group including at least one intermediate data group and at least one SM group (e.g., a high-correlation group). For example, the data grouping unit 1722 according to one embodiment performs at least some of the operations described in block B210 of FIG. 2, blocks B310 and B320 of FIG. 3, blocks B420, B432, and B450 of FIG. 4, and blocks B801, B803, and B806 of FIG. 8.
[0206] The model selector 1723 selects one or more machine learning models or one or more late-fusion models based on one or more criteria and marks data with a late-fusion flag according to the one or more criteria. For example, the model selector 1723 according to one embodiment performs at least some of the operations described in blocks B220-B225 of Figure 2, blocks B330-B335 of Figure 3, blocks B410-B417, B432, B447, and B455-B470 of Figure 4, blocks B1202 and B1204 of Figure 12, blocks B1310 and B1320 of Figure 13, blocks B1410, B1425, B1445, and B1450 of Figure 14, and blocks B1522, B1524, and B1526 of Figure 15.
[0207] The model executor 1724 inputs data into one or more ML models or one or more late-stage fusion models, runs the ML models or late-stage fusion models, obtains outputs of the ML models or late-stage fusion models, and stores the outputs in the data / classification results / feature repository 129. For example, the model execution module 124 of one embodiment performs at least some of the operations described in blocks B215, B225-B235 of FIG. 2, blocks B315, B325, B335, B340, and B345 of FIG. 3, blocks B412, B437, B455, B465, and B470 of FIG. 4, blocks B804, B807, B808, and B809 of FIG. 8, blocks B1203 and B1205 of FIG. 12, B1315 and B1325 of FIG. 13, blocks B1420, B1430, B1435, and B1455 of FIG. 14, and blocks B1520 and B1530 of FIG. 15.
[0208] The calculation unit 1725 calculates the intra-group / inter-modality correlation between the data groups and the intra-group correlation between different data groups (the intermediate data group and the single-modality group). For example, the calculation module 125 according to one embodiment performs at least some of the operations described in block B320 of FIG. 3, blocks B425 and B440 of FIG. 4, and blocks B802 and B805 of FIG. 8.
[0209] The configuration unit 1730 receives classification settings (eg, group classification criteria, classification purpose) from an input device.
[0210] The display control unit 1731 causes the display device to display the classification result (for example, the second classification result) by, for example, displaying a graphic user interface on the display device.
[0211] Although specific details are set forth herein to provide a thorough understanding of the disclosed embodiments, well-known methods, procedures, components, and circuits may not have been described in detail in order to avoid unnecessarily lengthening the disclosure.
[0212] Furthermore, when an element (e.g., an element, part, or component) is described as being "on," "to," "connected," or "coupled" to another element, the element may be directly "on," "to," "connected," or "coupled" to another element, although there may be intervening elements between the element and the other element. On the other hand, when an element is described as being "directly on," "directly" to, "directly connected," or "directly coupled" to another element, there are no intervening elements between the element and the other element.
[0213] Additionally, terms such as "comprise," "have," "include," "includes," and the like are considered open-ended terms unless otherwise indicated. When used herein, these terms specify the presence of stated features, integers, steps, operations, elements, materials, components, etc., but do not exclude the presence or addition of unspecified features, integers, steps, operations, elements, materials, components, etc.
[0214] Although several embodiments have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. The novel methods, devices, and systems may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. These embodiments and modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the invention and its equivalents as defined in the claims.
Claims
1. acquiring data from multiple modalities, including data from an imaging modality and data from a biomarker modality; Obtaining one or more trained machine learning models; generating a first multi-modality set from the data of the plurality of modalities, the first multi-modality set including at least one of data of the imaging modality and data of the biomarker modality; generating an intermediate data set, wherein generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data; generating a first classification result based on at least the intermediate data set, wherein generating the first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based on at least the intermediate data set. method.
2. The imaging modality data includes data defining one or more of the following: ・X-ray image ・Computer tomography images Magnetic resonance imaging ・X-ray fluoroscopic images Ultrasound images ・Positron emission tomography imaging The method of claim 1.
3. generating the first multi-modality group includes classifying the data of the multiple modalities into the first multi-modality group and one or more other groups based on the data of the multiple modalities and one or more grouping criteria; The method of claim 1.
4. the group classification criteria relate to at least one of a data type and a degree of correlation between data; The method of claim 3.
5. The data of the plurality of modalities further includes data of a text modality. The method of claim 1.
6. generating the first classification result further includes inputting the single-modality data set together with the intermediate data set into the second machine learning model; the second machine learning model is further trained to output a classification result based on the single-modality data set. The method of claim 1.
7. generating a second classification result based on a single-modality data set from the data of the multiple modalities; generating the second classification result includes inputting the single-modality data set to a third machine learning model trained to output a classification result based on the single-modality data set; The method of claim 1.
8. generating a third classification result based on the first classification result and the second classification result. The method of claim 7.
9. The third classification result indicates one or more of the following: - Whether the tumor is benign or malignant - Invasive or non-invasive cancer of the tumor - luminal, Her2-enriched, or triple-negative breast cancer subtype of the cancer; The method of claim 8.
10. one or more processors; and one or more memories, The one or more processors and the one or more memories Acquire data from multiple modalities, including data from an imaging modality and data from a biomarker modality; Obtain one or more trained machine learning models; generating a first multi-modality set from the data of the plurality of modalities, the first multi-modality set including at least one of data of the imaging modality and data of the biomarker modality; generating an intermediate data set, wherein generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data; configured to generate a first classification result based on at least the intermediate data set; generating a first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based on at least the intermediate data set; Device.
11. One or more computer-readable storage media storing instructions that cause one or more computing devices to perform operations, the operations including: acquiring data from multiple modalities, including data from an imaging modality and data from a biomarker modality; Obtaining one or more trained machine learning models; generating a first multi-modality set from the data of the plurality of modalities, the first multi-modality set including at least one of data of the imaging modality and data of the biomarker modality; generating an intermediate data set, wherein generating the intermediate data set includes inputting the first multi-modality set to a first machine learning model trained to extract features from the input multi-modality data and output the features as intermediate data; generating a first classification result based on at least the intermediate data set, wherein generating the first classification result includes inputting at least the intermediate data set to a second machine learning model trained to output a classification result based on at least the intermediate data set. storage medium.
12. acquiring first modality data and second modality data; Obtaining one or more trained machine learning models; generating a multi-modality data set from the first modality data and the second modality data, each multi-modality data set including a portion of the first modality data and a portion of the second modality data; inputting each of the multi-modality data sets into one of a first plurality of machine learning models, each of the first plurality of machine learning models outputting a respective intermediate data set generated based on the input multi-modality data set; inputting the intermediate data set into a second plurality of machine learning models; each of the second plurality of machine learning models outputs a classification result generated based on the input intermediate data group; method.
13. acquiring data from multiple modalities; obtaining one or more group classification criteria; Obtaining one or more trained machine learning models; generating multi-modality data sets from the data of the plurality of modalities according to the one or more grouping criteria, each multi-modality data set including data of at least two modalities; generating one or more first classification results by performing intermediate fusion on one or more sets of multi-modality data using one or more of the trained machine learning models. method.
14. generating one or more first classification results by performing early fusion on at least one single-modality data group from the multiple modality data using one or more of the trained machine learning models. The method of claim 13.
15. generating a second classification result by performing late fusion on the first classification result using one or more late fusion models.
15. The method of claim 14.
16. The late fusion is performed using a majority voting model, a weighted average model, or a stacking model.
16. The method of claim 15.
17. one or more processors; one or more memories, The one or more processors and the one or more memories Acquire data from multiple modalities Obtaining one or more grouping criteria; Obtain one or more trained machine learning models; generating multi-modality data sets from the data of the plurality of modalities according to the one or more grouping criteria, each multi-modality data set including data of at least two modalities; and generating one or more first classification results by performing intermediate fusion on one or more sets of multi-modality data using one or more of the trained machine learning models. Device.
18. One or more computer-readable storage media storing instructions that cause one or more computing devices to perform operations, the operations including: acquiring data from multiple modalities; obtaining one or more group classification criteria; Obtaining one or more trained machine learning models; generating multi-modality data sets from the data of the plurality of modalities according to the one or more grouping criteria, each multi-modality data set including data of at least two modalities; generating one or more first classification results by performing intermediate fusion on one or more sets of multi-modality data using one or more of the trained machine learning models; storage medium.