A blood tumor prediction method and system based on attention multi-instance learning of mass spectrometry flow data

CN122822290APending Publication Date: 2026-09-25ZHEJIANG PULUOTING HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610877837.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]针对现有技术中需要昂贵细胞级标注、仅支持单任务分类、可解释性不足的问题,本发明提供一种基于质谱流式数据的注意力多示例学习的血液肿瘤预测方法及系统,通过注意力多示例学习框架、门控注意力机制和细胞类型先验融合,实现高效、准确、可解释的多种血液肿瘤自动诊断

Benefits of technology

[0023]第一,本发明采用注意力多示例学习框架,仅需样本级标签,无需昂贵的单细胞标注,大幅降低了数据准备成本,使模型易于在临床环境中推广。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822290A_ABST
    Figure CN122822290A_ABST
Patent Text Reader

Abstract

The application discloses a blood tumor prediction method and system based on attention multi-instance learning of mass spectrometry flow data. First, mass spectrometry flow cytometry data of a bone marrow sample is acquired and preprocessed. Then, all cells in the sample form an instance bag, a pre-trained attention multi-instance classifier is used to process the bag, a high-dimensional feature of each cell is obtained through a feature extractor, cell features are weighted and aggregated through an attention aggregation network, a bag overall feature vector is obtained, the bag overall feature vector is input into a multi-task classification head, and blood tumor type probabilities are output. Finally, a training set is constructed to train the attention multi-instance classifier, and the trained attention multi-instance classifier is used to predict a to-be-tested sample, and disease prediction probabilities are output. The application does not require single-cell labeling, can simultaneously predict multiple blood tumors such as leukemia, lymphoma and plasma cell myeloma, and has high accuracy and good interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biomedical data analysis and artificial intelligence-assisted diagnosis, and in particular to a method and system for predicting hematologic malignancies based on attentional multi-instance learning of mass spectrometry flow cytometry data. Background Technology

[0002] Hematologic malignancies are malignant tumors originating from the hematopoietic system, and their early and accurate diagnosis is crucial for treatment selection and prognostic assessment. Currently, commonly used clinical diagnostic methods include bone marrow smear microscopy, flow cytometry immunophenotyping, and cytogenetic and molecular biological testing. Among these, flow cytometry can rapidly measure the expression of multi-parameter markers in individual cells, providing high-dimensional single-cell quantitative data, and has become one of the core technologies for the diagnosis of hematologic malignancies and the monitoring of minimal residual disease.

[0003] However, existing flow cytometry-based hematologic malignancy analysis has the following technical problems:

[0004] (1) Traditional analysis methods rely on manual gates and fixed threshold interpretation, which require operators to have rich clinical experience and professional training. The consistency between different laboratories is poor, making it difficult to achieve automated and standardized diagnosis.

[0005] (2) Most existing machine learning models adopt supervised learning frameworks, which require precise category labeling (e.g., normal / abnormal, specific subtype, etc.) for each cell. Single-cell level labeling is extremely difficult: on the one hand, experts need to perform microscopic discrimination on each cell, and building a large dataset containing tens of thousands or even hundreds of thousands of labeled images is costly and time-consuming; on the other hand, the preparation of bone marrow smears is affected by the differences in doctors' techniques and imaging equipment, which will introduce systematic errors.

[0006] (3) Most existing models are binary classification (leukemia / non-leukemia) or designed for a single disease type. They cannot handle classification tasks of multiple hematologic tumor subtypes (such as lymphoma, plasma cell myeloma, etc.) at the same time, which limits their application value in comprehensive clinical diagnosis.

[0007] (4) Existing methods do not make full use of the population structure information of cells in mass spectrometry flow cytometry data, ignore the differences in the contribution of different cell subpopulations to diagnosis, and have poor interpretability.

[0008] Therefore, there is an urgent need for an automated diagnostic method for hematological malignancies that does not require cell-level annotation, can utilize sample-level weak labels, can predict multiple diseases simultaneously, and has good interpretability. Summary of the Invention

[0009] To address the problems of existing technologies, such as the need for expensive cell-level annotation, support for only single-task classification, and insufficient interpretability, this invention provides a method and system for predicting hematologic malignancies based on attentional multi-instance learning of mass spectrometry streaming data. By using an attentional multi-instance learning framework, a gated attention mechanism, and cell type prior fusion, it achieves efficient, accurate, and interpretable automatic diagnosis of various hematologic malignancies.

[0010] In a first aspect, the present invention provides a method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data, comprising the following steps:

[0011] S1: Mass cytometry was used to analyze bone marrow samples and obtain mass cytometry data of the bone marrow samples.

[0012] S2: Preprocess the mass cytometry data to obtain a cell feature matrix;

[0013] S3: Construct an example package using sample cells from the cell feature matrix, input it into a pre-trained attention-based multi-instance classifier, and output the predicted probability of each category through each module of the attention-based multi-instance classifier;

[0014] S4: Construct a training set to train the attention-based multi-instance classifier;

[0015] S5: Use the trained attention multi-instance classifier to predict the test sample and output the predicted probability of each category of hematological malignancy.

[0016] Secondly, the present invention provides a hematologic malignancy prediction system based on attention-based multi-instance learning of mass spectrometry streaming data, comprising:

[0017] The data acquisition module is used to detect bone marrow samples using mass flow cytometry and acquire mass flow cytometry data of the bone marrow samples.

[0018] The preprocessing module is used to preprocess the mass spectrometry flow cytometry cell data to obtain a cell feature matrix;

[0019] The model storage module is used to store the structure file and parameters of the pre-trained attention-based multi-instance classifier;

[0020] The inference module is used to construct an example package from the sample cells in the cell feature matrix, input it into a pre-trained attention-based multi-example classifier, perform inference through each module of the attention-based multi-example classifier, and output the predicted probability of each category.

[0021] The results display module is used to show the prediction results on the monitor and generate diagnostic reports.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] First, the present invention adopts an attention-based multi-instance learning framework, which only requires sample-level labels and does not require expensive single-cell annotation, greatly reducing the cost of data preparation and making the model easy to promote in clinical settings.

[0024] Second, this invention introduces a gated attention mechanism, which can automatically learn the importance of different cells in disease diagnosis and output multiple disease types, thereby improving diagnostic efficiency and comprehensive capabilities.

[0025] Third, by visualizing attention weights, this invention allows clinicians to understand which cell subpopulations the model focuses on, thus enhancing the model's interpretability.

[0026] Fourth, this invention employs an optional cell type prior fusion strategy, utilizing existing labeling threshold knowledge in mass spectrometry flow cytometry data to guide the model to focus on cell populations closely related to specific diseases, thereby further improving diagnostic accuracy.

[0027] Fifth, the present invention uses Poly-1 loss to handle class imbalance, which improves the ability to identify rare classes (such as plasma cell myeloma). Attached Figure Description

[0028] Figure 1 This is a structural block diagram of a hematologic tumor prediction system provided in an embodiment of the present invention;

[0029] Figure 2 A schematic diagram of the network structure of the attention-based multi-instance classifier provided in an embodiment of the present invention;

[0030] Figure 3 A detailed computational block diagram of the attention aggregation network provided in this embodiment of the invention;

[0031] Figure 4 This is an example diagram of the confusion matrix of the test dataset provided in an embodiment of the present invention. Detailed Implementation

[0032] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention. It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the present invention.

[0033] Furthermore, regarding the numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, are also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0034] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention.

[0035] Example 1:

[0036] This invention discloses a method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data. The hematologic malignancies include leukemia, lymphoma involvement, and plasma cell myeloma. The method includes the following steps:

[0037] Step S1: Obtain mass cytometry data of the bone marrow sample to be tested.

[0038] Specifically, in this embodiment, a mass flow cytometer is used to detect bone marrow samples, obtaining mass flow cytometry data in FCS format. The intensity values ​​of multiple fluorescent labeling channels are recorded for each cell. The labeling channels are determined according to a preset labeling panel, which contains multiple labels, such as CD45, CD19, CD138, CD34, CD117, CD14, and CD16.

[0039] Step S2: Preprocess the mass cytometry data.

[0040] Furthermore, the sample dataset in this embodiment comprises 500 cases, including 80 healthy controls, 260 cases of leukemia, 100 cases of lymphoma, and 60 cases of plasma cell myeloma. These are randomly divided into training, validation, and test sets in a 7:1:2 ratio. The training set includes: 56 healthy controls, 182 cases of leukemia, 70 cases of lymphoma, and 42 cases of plasma cell myeloma; the validation set includes: 8 healthy controls, 26 cases of leukemia, 10 cases of lymphoma, and 6 cases of plasma cell myeloma; the test set includes: 16 healthy controls, 52 cases of leukemia, 20 cases of lymphoma, and 12 cases of plasma cell myeloma.

[0041] Furthermore, the preprocessing step specifically includes:

[0042] Step S21: Filter out the target marker channel according to the preset marker panel and extract the data of the preset marker channel;

[0043] Step S22: Perform intensity value calculation for each labeled channel. The transformation is performed to stabilize the variance and approximate a normal distribution, resulting in the transformed cell feature matrix.

[0044] Step S23 (optional): As shown in Table 1 below, the cell type of each cell is automatically labeled according to the label expression threshold, and the priority of cell type determination is as follows: plasma cells, lymphocytes, immature granulocytes, monocytes and CD34+ cells.

[0045] For example, if a cell's CD138 expression is greater than 1000, it is labeled as a plasma cell; if CD45 > 100 and CD19 > 10, it is labeled as a lymphocyte; if Lactoferrin, CD45, and Lysozyme are all less than 100, it is labeled as an immature granulocyte; if CD14 or CD16 is greater than 10, it is labeled as a monocyte; if CD34 is greater than 10, it is labeled as a CD34+ cell. Cell types are mapped to integer indices (0: other, 1: plasma cell, 2: lymphocyte, 3: immature granulocyte, 4: monocyte, 5: CD34+ cell).

[0046] Specifically, in this embodiment, after preprocessing, each sample is represented as a matrix of cell number × label number and a cell type vector of cell number × 1.

[0047] Table 1. Cell type labeling rules (example)

[0048]

[0049] Step S3: Input the preprocessed cell feature matrix into the pre-trained attention multi-instance classifier.

[0050] Specifically, in this embodiment, all cells obtained after preprocessing each sample are integrated into an "example package". Only the disease label of the sample is needed as the package-level label, and no cell-level label is required.

[0051] In this embodiment, as Figure 2 As shown, the attention-based multi-instance classifier includes a feature extractor, an attention aggregation network, and a multi-task classification head, with the following structure:

[0052] Step S31: Feature Extractor: Composed of multiple fully connected layers, it maps the original label vector of each cell (dimension M, representing the number of labels) into a high-dimensional feature vector (dimension D). Each fully connected layer is followed by a batch normalization layer and... Layers are added to alleviate overfitting.

[0053] Step S32: Attention Convergence Network: such as Figure 3 As shown, an attention mechanism based on gating is used to automatically learn the importance weight of each cell for diagnosis. The specific calculation method is as follows:

[0054] For each cell Based on its high-dimensional feature vector First, calculate the raw attention score:

[0055] )

[0056] Where ⊙ represents element-wise multiplication. This is a trainable weight matrix.

[0057] Then through Normalization yields the attention weights:

[0058]

[0059] Finally, calculate the overall package feature vector:

[0060]

[0061] In this embodiment, the gating-based attention mechanism can automatically focus on cells closely related to disease diagnosis while suppressing interference from irrelevant cells.

[0062] Step S33 Multi-task classification head: Input the overall packet feature vector Z into one or more fully connected layers, and finally output the number of classes. of (C represents the number of preset disease categories).

[0063] In this embodiment, It can cover non-leukemia, leukemia, lymphoma involvement, plasma cell myeloma, etc., through Will Convert to a probability distribution.

[0064] Furthermore, if cell type labels are obtained in step S2, they are used as bias fusion before attention calculation: the cell type embedding vector (which maps the type index to a low-dimensional vector through the embedding layer) is projected and then fused with... The attention scores are added together, or multiplied by a mask, to give key cell subpopulations (such as plasma cells and CD34+ cells) higher attention weights.

[0065] Step S4: Train the attention-based multi-instance classifier

[0066] Furthermore, training the attention-based multi-instance classifier includes the following steps:

[0067] Step S41: Construct a training set by collecting bone marrow samples (FCS files) from healthy donors and patients with confirmed hematologic malignancies. Each sample should be labeled with a disease type, and the training set should cover all preset categories and contain a sufficient number of samples.

[0068] Step S42: Perform the preprocessing operation of step S2 on each training sample to obtain the cell feature matrix and optional cell type labels;

[0069] Step S43: Input each sample as a bag into the attention multi-instance classifier and calculate the model's prediction probability for the sample;

[0070] Step S44: Based on the predicted probability, calculate the loss function for model training. In this embodiment, we use... Classification loss function:

[0071]

[0072] in, For sparse cross-entropy loss, This represents the model's predicted probability of the true class. This loss performs better than standard cross-entropy when dealing with class imbalance.

[0073] In this embodiment, if the attention aggregation network contains multiple attention branches (e.g., multiple independent attention computation units designed to increase model richness), a differential regularization term can be further introduced to encourage different branches to focus on different cell subgroups, thereby enhancing the diversity of feature representations.

[0074] Step S45: Update network parameters based on the loss function through backpropagation and the optimizer. During training, a validation set can be used for monitoring, and early stopping and learning rate decay strategies can be employed.

[0075] Step S5: Use the trained classifier to predict the test sample. Obtain the cell feature matrix of the test sample according to steps S1 and S2, input it into the classifier, obtain the probability of each category, take the category corresponding to the highest probability as the prediction result, and output the probability of each disease for clinical reference.

[0076] Example 2: Model Building and Training

[0077] Furthermore, the specific parameters of the attention multi-instance classifier in this embodiment are as follows (the following parameters are merely examples and do not constitute a limitation on the scope of the present invention):

[0078] Feature extractor: The fully connected layer structure is: 128 (ReLU) → BN → Dropout(0.2) → 64 (ReLU) → BN → Dropout(0.2) → 64 (ReLU) → BN. Output dimension D=64.

[0079] Attention convergence network: It adopts a gated attention mechanism and has a set of trainable parameters W, V, U.

[0080] Classification Header: Fully connected layer D → 128 (ReLU) → Dropout (0.3) → C (output, C is the number of classes).

[0081] Optimizer: Adam, initial learning rate 0.001.

[0082] Batch size: 32 (can be adjusted according to hardware).

[0083] Number of training epochs: determined by early stopping (stop if the validation loss does not decrease after 10 epochs), and save the model with the minimum validation loss.

[0084] Loss function: loss, .

[0085] Class weights: Inverse frequency weights are calculated based on the number of samples in each class in the training set to alleviate class imbalance.

[0086] During training, for each batch, the model's forward propagation yields... and attention weight .calculate - After applying the classification loss, backpropagation is performed. The training process is completed on the GPU.

[0087] Example 3: Model Evaluation

[0088] The model performance was evaluated on an independent test set, and the test results are as follows: Figure 4 The confusion matrix is ​​shown (the results are based on actual experimental data). The results show that, considering that supervised models require tens of thousands of single-cell annotations, this invention has a significant advantage in annotation cost.

[0089] Example 4: Visualization and Explainability

[0090] A significant advantage of this invention is that it can output the attention weight distribution, showing which cells the model focuses on during diagnosis. These cells can be further visualized, and clinicians can verify the rationality of the model's decisions through the visualization results.

[0091] Example 5:

[0092] like Figure 1 As shown, a hematologic tumor prediction system based on attention-based multi-instance learning of mass spectrometry streaming data is also provided, including:

[0093] The data acquisition module is used to read FCS files and can acquire flow mass spectrometry cell data through the local file system or network interface.

[0094] The preprocessing module is used to perform channel filtering, Transformation and optional automatic cell type labeling;

[0095] The model storage module is used to store the structure file (such as .h5 or .pb format) and parameters of the pre-trained attention multi-instance classifier;

[0096] The inference module is used to load the pre-trained model, input the pre-processed cell feature matrix into the model for forward propagation, and output the probability of each category.

[0097] The results display module is used to show the prediction results (tumor type and probability) on the display screen, visualize the specific marker expression of high-weight cells, and generate a diagnostic report.

[0098] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data, characterized in that, Includes the following steps: S1: Mass cytometry was used to analyze bone marrow samples and obtain mass cytometry data of the bone marrow samples. S2: Preprocess the mass cytometry data to obtain a cell feature matrix; S3: Construct an example package using sample cells from the cell feature matrix, input it into a pre-trained attention-based multi-instance classifier, and output the predicted probability of each category through each module of the attention-based multi-instance classifier; S4: Construct a training set to train the attention-based multi-instance classifier; S5: Use the trained attention multi-instance classifier to predict the test sample and output the predicted probability of each category of hematological malignancy.

2. The method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 1, characterized in that, The mass spectrometry flow cytometry records the intensity values ​​of fluorescent labeling channels, which are determined based on a preset labeling panel, including but not limited to CD45, CD19, CD138, CD34, CD117, CD14, and CD16.

3. A method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data, as described in claim 1 or 2, is characterized in that... The preprocessing includes the following steps: Select the target marking channel based on the preset marking panel; Intensity values ​​of each marker channel were analyzed. Transformation; Each cell type is automatically labeled based on the marker expression threshold.

4. The method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 3, characterized in that, The labeling rules for each cell type are as follows: A CD138 count greater than 1000 indicates a plasma cell. If CD45 > 100 and CD19 > 10, it is labeled as a lymphocyte; Lactoferrin, CD45, and Lysozyme were all less than 100, indicating immature granulocytes; A CD14 or CD16 value greater than 10 indicates a monocyte; Cells with a CD34 count greater than 10 are labeled as CD34+ cells.

5. The method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 1, characterized in that, The attention-based multi-instance classifier includes a feature extractor, an attention aggregation network, and a multi-task classification head, wherein: The feature extractor consists of multiple fully connected layers, each followed by a batch normalization and dropout layer, which maps the original label vector of each cell into a high-dimensional feature vector. The attention aggregation network employs a gating-based attention mechanism to calculate the original attention score of each cell's high-dimensional feature vector, obtains the attention weight through normalization, and finally calculates the overall packet feature vector. The multi-task classification head inputs the overall package feature vector into the fully connected layer and outputs the probability of each category.

6. A method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 1 or 5, characterized in that, When preprocessing the mass cytometry data, if cell type labels are obtained, they are used as bias fusion before attention calculation. The projected cell type embedding vector is added to the high-dimensional feature vector, or the attention score is multiplied by a mask to increase the attention weight of key cell subpopulations.

7. The method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 6, characterized in that, The training steps for the attention-based multi-instance classifier are as follows: Collect bone marrow samples from healthy donors and patients diagnosed with hematologic malignancies, and label each sample with a disease type label; Perform the preprocessing described in S2 on each sample to obtain the cell feature matrix; Each sample is fed into the attention-based multi-instance classifier as a bag, and the model's predicted probability for the sample is calculated. The loss function for model training is calculated based on the predicted probabilities. Based on the loss function, the network parameters are updated through backpropagation and an optimizer.

8. The method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 7, characterized in that, The loss function adopts The loss function addresses class imbalance and improves the ability to identify rare classes.

9. A method for predicting hematologic malignancies based on attention-based multi-instance learning of mass spectrometry streaming data according to claim 5, characterized in that, When the attention aggregation network contains multiple attention branches, a differential regularization term is introduced to guide each branch to focus on different cell subgroups, thereby enhancing the diversity of feature representations.

10. A hematologic tumor prediction system based on attention-based multi-instance learning of mass spectrometry streaming data, characterized in that, include: The data acquisition module is used to detect bone marrow samples using mass flow cytometry and acquire mass flow cytometry data of the bone marrow samples. The preprocessing module is used to preprocess the mass spectrometry flow cytometry cell data to obtain a cell feature matrix; The model storage module is used to store the structure file and parameters of the pre-trained attention-based multi-instance classifier; The inference module is used to construct an example package from the sample cells in the cell feature matrix, input it into a pre-trained attention-based multi-example classifier, perform inference through each module of the attention-based multi-example classifier, and output the predicted probability of each category. The results display module is used to show the prediction results on the monitor and generate diagnostic reports.