A marker composition, kit, diagnostic model construction method and application for identifying different types of breast fibroepithelial tumors

The diagnostic model constructed using marker-based compositions and machine learning algorithms has solved the challenge of preoperative differentiation of breast fibroepithelial tumors, enabling accurate differentiation and grading of phyllodes tumors and fibroadenomas, improving diagnostic accuracy and consistency, guiding surgical decisions, and reducing the risk of recurrence.

CN120082653BActive Publication Date: 2026-04-17THE FIRST AFFILIATED HOSPITAL OF MEDICAL COLLEGE OF XIAN JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE FIRST AFFILIATED HOSPITAL OF MEDICAL COLLEGE OF XIAN JIAOTONG UNIV
Filing Date
2025-03-04
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Current technology makes it difficult to accurately identify the type of breast fibroepithelial tumors before surgery, especially benign, borderline, and malignant phyllodes tumors, leading to inaccurate surgical planning and increasing the risk of recurrence for patients.

Method used

Using a combination of markers, including protein markers such as MAL2, KRT8, and DSP, and DNA methylation biomarkers such as cg00515756, a diagnostic model is constructed through high-throughput omics detection and machine learning algorithms to achieve the differentiation of phyllodes tumors and fibroadenomas of the breast, as well as the differentiation of phyllodes tumors of different grades.

Benefits of technology

It improves the accuracy and consistency of preoperative diagnosis of breast fibroepithelial tumors, outperforming previous diagnostic models. It can effectively distinguish different types of breast fibroepithelial tumors, guide surgical decisions, and reduce the risk of recurrence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120082653B_ABST
    Figure CN120082653B_ABST
Patent Text Reader

Abstract

The application provides a marker composition, a kit, a diagnostic model construction method and application for identifying different types of breast fibroepithelial tumors, and specifically belongs to the technical field of pathological molecular diagnostic products. The marker composition comprises any one or two or more of the following markers: MAL2, KRT8, DSP, cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, cg00819233 and cg04404381. The diagnostic model formed by the marker composition provided by the application can well distinguish PT and FA and different grades of PT, and is superior to the previously reported diagnostic model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of pathological molecular diagnostic products, specifically relating to a marker composition, kit, diagnostic model construction method, and application for different types of breast fibroepithelial tumors. Background Technology

[0002] Phyllodes tumors of the breast (PT) belong to the fibroepithelial lesions (FELs), composed of fibrous connective tissue and epithelial tissue of the breast. PT is not a single disease, but a collective term for a series of fibroepithelial tumors with different clinical courses and pathological histological characteristics. Pathologically, they are classified into three types: benign phyllodes tumors (benign PT), borderline phyllodes tumors (borderline PT), and malignant phyllodes tumors (malignant PT). All types of PT have a certain risk of recurrence; malignant PT has a particularly high risk of metastasis, approximately 16-25%, and once distant metastasis occurs, the prognosis is extremely poor, with an average overall survival of only 10.7-11.5 months.

[0003] The primary treatment for PT is surgery—extensive resection of the tumor with a negative margin of at least 1 cm. There is no evidence that radiotherapy and chemotherapy reduce the risk of recurrence or death. However, PT presents similarly to fibroadenoma (FA), a fibroepithelial tumor that does not require extensive resection, and is often difficult to differentiate using preoperative core needle biopsy (CNB), making it difficult for surgeons to design an accurate surgical scope preoperatively. Patients diagnosed with PT postoperatively require a second surgery to achieve the guideline-defined margins and reduce the risk of recurrence. Fibrous epithelial tumors include FA, benign PT, borderline PT, and malignant PT (benign phyllodes tumors). The diagnosis of fibroepithelial tumors mainly relies on morphological features obtained from HE staining, which has a degree of subjectivity. Differential diagnosis between fibroadenoma and benign phyllodes tumors, as well as histological grading of phyllodes tumors, are challenging aspects of preoperative and intraoperative diagnosis of PT, and pathologists with different levels of experience often arrive at different diagnoses. Furthermore, the differential diagnosis between benign, borderline, and malignant PT often results in inconsistencies between institutions or pathologists due to the lack of effective molecular markers. An analysis of 285 patients with febrile elutitis (FELs) who underwent preoperative biopsy and postoperative pathological diagnosis revealed a diagnostic consistency of only 37.6%. Among these patients, 31.2% were diagnosed with febrile elutitis (FA) preoperatively and subsequently diagnosed with pterygium (PT) postoperatively, requiring a second surgery to achieve adequate margins. Therefore, new diagnostic strategies and methods are urgently needed for accurate preoperative diagnosis of PT to guide surgical decisions, optimize treatment processes, and facilitate clinical work and patient management.

[0004] In recent years, an increasing number of studies have focused on biomarkers for the differential diagnosis of prostatitis (PT). Cani et al. and Tan et al., through targeted sequencing and exon sequencing of breast fecal elutions (FELs), found that FELs are a disease characterized by recurrent somatic mutations in exon 2 of MED12, but this characteristic cannot differentiate between the four types of FELs. Piscuoglio et al. and Otsuji et al., using MSK-IMPACT sequencing and ddPCR, respectively, sequenced and validated TERTp mutations, finding that TERT alterations may promote PT progression, but this mutation still cannot completely distinguish between PT and breast cancer (FA). Zhang et al., through proteomic sequencing, found no significant differences in proteomic features between low-grade PT and FA, but their limitation was the absence of malignant PT samples. Hench et al. and Meyer et al., using DNA methylation microarray analysis of PT, found that PT has different methylation patterns from breast cancer and normal breast tissue, identifying methylation features for differentiating malignant PT and metaplastic breast cancer, but their research on the differential diagnosis between PT and FA has not been in-depth. In summary, no single omics study has resolved the clinical dilemma of preoperative differential diagnosis between PT and FA, or between different types of PT. Summary of the Invention

[0005] The purpose of this invention is to provide a marker composition, kit, diagnostic model construction method, and application for different types of breast fibroepithelial neoplasia. The diagnostic model constructed from the marker composition provided by this invention can effectively distinguish between PT and FA, as well as different grades of PT, and this diagnostic model is superior to previously reported diagnostic models.

[0006] This invention provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors, borderline phyllodes tumors, and malignant phyllodes tumors. The marker composition comprises any one or more of the following markers: MAL2, KRT8, DSP, cg00515756, cg24852561, and cg006868. 80, cg22449901, cg09025210, cg00819233 and cg04404381; among them, MAL2, KRT8 and DSP are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, cg00819233 and cg04404381 ​​are DNA methylation biomarkers in the human genome.

[0007] This invention also provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors, borderline phyllodes tumors, and malignant phyllodes tumors of the breast, and the marker composition includes any one or more of the following markers: MAL2, KRT8, cg00515756, cg24852561, cg006868. 80, cg22449901, cg09025210, cg00819233 and cg04404381; among them, MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, cg00819233 and cg04404381 ​​are DNA methylation biomarkers in the human genome.

[0008] This invention also provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors, borderline phyllodes tumors, and malignant phyllodes tumors of the breast, and the marker composition includes any one or more of the following markers: MAL2, KRT8, cg00515756, cg248525. 61, cg00686880, cg22449901, cg09025210 and cg04404381; among them, MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, cg09025210 and cg04404381 ​​are DNA methylation biomarkers in the human genome.

[0009] The present invention also provides a marker composition for identifying phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors of the breast, borderline phyllodes tumors of the breast, and malignant phyllodes tumors of the breast, and the marker composition includes any one or more of the following markers: MAL2, KRT8, cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210; wherein MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210 are DNA methylation biomarkers in the human genome.

[0010] The present invention also provides the use of reagents for detecting the marker compositions described in the above technical solutions in the preparation of kits for identifying phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast.

[0011] This invention also provides a kit for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, the kit comprising reagents for detecting the marker composition described in the above-described technical solution; the reagents for detecting the marker composition described in the above-described technical solution include reagents for detecting protein markers and reagents for detecting methylated biomarkers;

[0012] The reagents used to detect protein markers include those used in any of the following methods: PRM targeted protein sequencing, selected reaction monitoring targeted sequencing, multiple reaction monitoring targeted sequencing, tandem mass spectrometry tag quantitative protein sequencing, isotope tag quantitative protein sequencing, Western blotting, enzyme-linked immunosorbent assay (ELISA), and immunohistochemical staining.

[0013] The reagents used to detect DNA methylation biomarkers include those used in any of the following methods: MethylTarget methylation sequencing, methylated DNA immunoprecipitation sequencing, degenerate representative bisulfite sequencing, whole-genome bisulfite sequencing, deep target region DNA methylation sequencing, and pyrosequencing.

[0014] This invention also provides a method for constructing a diagnostic model for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, comprising the following steps:

[0015] A diagnostic model for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast is constructed using the marker composition described in the above technical solution.

[0016] The diagnostic model includes one or more of the following types: Naive Bayes model, Support Vector Machine model, Multivariate Logistic Regression model, Decision Tree model, Random Forest model, K-Nearest Neighbors algorithm model, Neural Network model, Adaptive Boosting Algorithm model, Gradient Boosting Machine model, Extreme Gradient Boosting Algorithm model, Minimum Absolute Shrinkage and Selection Operator model, Elastic Network Regression model, and Ridge Regression model.

[0017] This invention also provides a diagnostic model device for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast. The device includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, they are used to perform the following operations:

[0018] The detection data of the marker composition of the sample to be tested is obtained, and the detection data is input into the diagnostic model constructed by the marker composition to distinguish between phyllodes tumors and fibroadenomas of the breast and / or to distinguish between phyllodes tumors of different grades of the breast, so as to obtain the diagnostic results of the sample to be tested.

[0019] The marker composition is the marker composition described in the above technical solution.

[0020] This invention also provides a diagnostic system for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, comprising:

[0021] Acquisition unit, used to acquire detection data of the marker composition of the sample to be tested;

[0022] The processing unit inputs the detection data of the marker composition into a diagnostic model constructed from the marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, to obtain the diagnostic results of the sample to be tested.

[0023] The marker composition is the marker composition described in the above technical solution.

[0024] The present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program implementing the following method when executed by a processor:

[0025] The detection data of the marker composition of the sample to be tested is obtained, and the detection data of the marker composition is input into a diagnostic model constructed by the marker composition to differentiate between phyllodes tumors and fibroadenomas of the breast and / or to differentiate between phyllodes tumors of different grades of the breast, so as to obtain the diagnostic results of the sample to be tested.

[0026] The marker composition is the marker composition described in the above technical solution.

[0027] This invention provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors. This invention determines the optimal diagnostic markers for PT (proliferative fibroadenoma) through joint analysis of high-throughput omics data from multiple levels and sources. Specifically, this invention identifies molecular markers for differentiating four types of fibroepithelial tumors through multi-omics detection, including protein and methylation site markers. The diagnostic model constructed from these markers can effectively distinguish between PT and FA (proliferative fibroadenoma), as well as different grades of PT, and this diagnostic model is superior to previously reported diagnostic models. Experimental results show that this invention first screened 30 proteins and 20 methylation sites from high-throughput proteomics and methylometabolics data of 40 FELs patients as candidate diagnostic biomarkers for different subtypes of FELs. Then, the varImpPlot function in the random forest algorithm was used to rank these 50 candidate biomarkers by importance, and the top 10 biomarkers were validated by targeted sequencing. Through a new independent cohort, 74 FELs samples underwent simultaneous PRM-targeted proteomics and MethylTarget-targeted methylometabolics sequencing. The results show that the biomarkers described in this invention, used alone, can achieve FA vs. benign PT vs. borderline PT vs. malignant PT. When the biomarkers of this invention are used in combination, they can also achieve FA vs. benign PT vs. borderline PT. When differentiating between PT and malignant PT, the results are even better. Specifically, combinations of 5, 6, 7, 8, 9, and 10 biomarkers show superior diagnostic efficacy; the combination of 7 biomarkers yields the best results. In two independent XJTU cohorts, the four-category diagnostic panel for FELs constructed using a multivariate logistic regression (MLR) machine learning algorithm exhibits the best diagnostic efficacy, with accuracy, specificity, and sensitivity all reaching 100%. Furthermore, compared to previously reported binary diagnostic panels, this invention, using a binary logistic regression method, found that a diagnostic panel composed of 7 key biomarkers also demonstrates high specificity (100%) and sensitivity (100%) in distinguishing between PT and FA, and between benign PT and borderline / malignant PT. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1The volcano plot of differentially expressed proteins between multiple groups provided by the present invention; wherein, each point in the volcano plot represents a gene, red represents upregulated proteins (P<0.05, log2FC≥1), blue represents downregulated proteins (P<0.05, log2FC≤1), and gray represents proteins that do not change significantly between the two groups;

[0030] Figure 2 This invention provides a hierarchical clustering heatmap of differentially expressed proteins; each row represents a protein, and the expression level of each protein is converted using a Z-score. Blue, which gradually deepens from 0 to -2, represents a lower protein expression level, while red, which gradually deepens from 0 to 3, represents a higher protein expression level; each column represents a sample, and the top of the column contains the corresponding clinical information for each sample.

[0031] Figure 3 Volcano plot of differentially methylated sites among multiple groups provided by this invention; wherein each red dot represents a CpG site, only DMPs with P < 0.05 and |Δβ| ≥ 0.2 after calibration are shown; Δβ ≥ 0.2 are upregulated DMPs; Δβ ≤ -0.2 are downregulated DMPs; Δβ, delta;

[0032] Figure 4 This invention provides a clustering heatmap of differentially methylated sites; each row represents a CpG site, and the methylation level of each CpG site is converted using a Z-score. Blue, which gradually deepens from 0 to -3, represents a lower methylation level, and red, which gradually deepens from 0 to 3, represents a higher methylation level; each column represents a sample, and the top of the column contains the corresponding clinical information for each sample.

[0033] Figure 5 A flowchart for screening diagnostic candidate biomarkers provided by the present invention;

[0034] Figure 6 A flowchart illustrating the screening process for diagnostic biomarkers and patient cohorts in the validation cohort provided by this invention;

[0035] Figure 7 The graph shows the importance ranking results of 50 candidate markers using the random forest algorithm in XJTU cohort2 provided by this invention.

[0036] Figure 8The graph shows the cumulative AUC values ​​of three machine learning models in XJTU cohort 2 provided by this invention. The horizontal axis represents the number of features included, and the vertical axis represents the cumulative AUC value. The data table shows the cumulative AUC values ​​for each machine learning algorithm as features are included sequentially. The red box indicates that the AUC of all three machine learning models reaches its highest point when the first seven markers are included. NB stands for Naive Bayes; SVM for Support Vector Machine; and MLR for Multivariate Logistic Regression.

[0037] Figure 9 The graph shows the cumulative AUC values ​​of three machine learning models in XJTU cohort 1 provided by this invention. The horizontal axis represents the number of features included, and the vertical axis represents the cumulative AUC value. The data table shows the cumulative AUC values ​​for each machine learning algorithm as features are included sequentially. The red box indicates that the AUC of all three machine learning models reaches its highest point when the first seven markers are included. NB stands for Naive Bayes; SVM for Support Vector Machine; and MLR for Multivariate Logistic Regression.

[0038] Figure 10 The diagram shows the diagnostic efficacy (AUC, specificity, and sensitivity) of the seven biomarkers provided by this invention for differentiating different outcomes; where the values ​​range from 0 to 1, starting from the center point 0 and extending outwards to the outermost circle; different colors represent different diagnostic indicators.

[0039] Figure 11 The graph shows the differential expression results of the two diagnostic proteins provided by this invention in different PT grades; where (a) is the expression box plot of the two diagnostic proteins in XJTU cohort 1; (b) is the expression box plot of the two diagnostic proteins in XJTU cohort 2; ns P>0.05, *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001;

[0040] Figure 12 The results of the differences in methylation levels of the five diagnostic methylation sites provided by this invention in different grade PTs are shown in the figure; (a) box plot of methylation levels of the five diagnostic methylation sites in XJTU cohort 1; (b) box plot of methylation levels of the five diagnostic methylation sites in XJTU cohort 2; ns P>0.05, *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001;

[0041] Figure 13This is a schematic diagram of a decision tree for the future clinical application of the 7-molecule diagnostic panel provided by this invention; the AUC values ​​are from XJTU Cohort 2, the red "hyper" indicates that the site is highly methylated in FA compared to PT; the green "low" indicates that the protein is poorly expressed in malignant PT compared to other FEL subtypes;

[0042] Figure 14 The diagram shows the comparison results between the multi-omics diagnostic model provided by this invention and traditional models; wherein, (a) is a comparison diagram between the diagnostic model of this invention and the previous diagnostic model for differentiating FA and PT; and (b) is a comparison diagram between the diagnostic model of this invention and the previous diagnostic model for differentiating benign PT and borderline / malignant PT. Detailed Implementation

[0043] This invention provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors, borderline phyllodes tumors, and malignant phyllodes tumors. The marker composition comprises any one or more of the following markers: MAL2 (myelin and lymphocyte protein 2) and KRT8 (keratin 8). 8) DSP (Desmoplakin), cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, cg00819233 and cg04404381; among them, MAL2, KRT8 and DSP are protein markers, while cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, cg00819233 and cg04404381 ​​are DNA methylation biomarkers in the human genome. The results showed that the methylation sites of cg22449901 (XJTU cohort 1: sensitivity 100% and specificity 100%; XJTU cohort 2: sensitivity 93.3% and specificity 95%) and cg24852561 (XJTU cohort 1: sensitivity 100% and specificity 100%; XJTU cohort 2: sensitivity 93.3% and specificity 90%) could be used alone as diagnostic markers for FA vs. benign PT; the methylation site of cg09025210 (XJTU cohort 1: sensitivity 80% and specificity 100%; XJTU cohort 2: sensitivity 90% and specificity 95%) could be used alone as diagnostic markers for benign PT vs. borderline PT; the KRT8 protein (XJTU cohort 1: sensitivity 100% and specificity 80%; XJTU...) could also be used alone as a diagnostic marker for FA vs. borderline PT. Both cohort 2 (sensitivity 94.7% and specificity 95%) and MAL2 protein (XJTU cohort 1: sensitivity 100% and specificity 100%; XJTU cohort 2: sensitivity 100% and specificity 95%) can be used alone as a differential diagnostic marker for malignant PT vs. borderline PT; MAL2 protein (XJTU cohort 1: sensitivity 96.3% and specificity 100%; XJTU cohort 2: sensitivity 100% and specificity 94.5%) can be used alone as a differential diagnostic marker for malignant PT vs. Other FELs (FA, BPT, and BorPT).When 10 biomarkers were used in combination to diagnose FA vs. benign PT vs. borderline PT vs. malignant PT, the diagnostic AUCs of the three machine learning algorithms were NB = 0.861, SVM = 0.917, MLR = 1 (XJTU cohort 1) and NB = 0.945, SVM = 0.947, MLR = 0.956 (XJTU cohort 2), respectively. Differential diagnosis of PT and FA, as well as histological grading of PT, are challenges in the preoperative diagnosis of PT. This invention focuses on the clinical problem of difficult differential diagnosis of PT. Based on previous multi-omics sequencing data, a molecular diagnostic model is constructed to improve diagnostic efficacy and guide surgical decisions. Based on previous multi-omics detection results, this invention found that integrin and methylomes can better distinguish different subtypes of FELs, thus initially screening candidate diagnostic biomarkers for FELs.

[0044] This invention also provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors, borderline phyllodes tumors, and malignant phyllodes tumors of the breast, and the marker composition includes any one or more of the following markers: MAL2, KRT8, cg00515756, cg24852561, cg006868. 80, cg22449901, cg09025210, cg00819233 and cg04404381; among them, MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, cg00819233 and cg04404381 ​​are DNA methylation biomarkers in the human genome. When nine biomarkers were used in combination to diagnose FA vs. benign PT vs. borderline PT vs. malignant PT, the diagnostic AUCs of the three machine learning algorithms were NB = 0.908, SVM = 0.917, MLR = 1 (XJTU cohort 1) and NB = 0.967, SVM = 0.969, MLR = 0.956 (XJTU cohort 2), respectively.

[0045] This invention also provides a marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors, borderline phyllodes tumors, and malignant phyllodes tumors of the breast, and the marker composition includes any one or more of the following markers: MAL2, KRT8, cg00515756, cg248525. 61, cg00686880, cg22449901, cg09025210, and cg04404381; among them, MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, cg09025210, and cg04404381 ​​are DNA methylation biomarkers in the human genome. When the eight biomarkers were used in combination for FA vs. benign PT vs. borderline PT vs. malignant PT, the diagnostic AUCs of the three machine learning algorithms were NB=0.861, SVM=0.908, MLR=1 (XJTU cohort 1) and NB=0.967, SVM=0.975, MLR=0.915 (XJTU cohort 2), respectively.

[0046] The present invention also provides a marker composition for identifying phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, wherein the different grades of phyllodes tumors of the breast include benign phyllodes tumors of the breast, borderline phyllodes tumors of the breast, and malignant phyllodes tumors of the breast, and the marker composition includes any one or more of the following markers: MAL2, KRT8, cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210; wherein MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210 are DNA methylation biomarkers in the human genome. This invention further utilizes PRM and MethlTraget targeted sequencing technologies and three machine learning algorithms to validate a diagnostic model consisting of seven biomarkers in an expanded external cohort (XJTUcohort 2), including two proteins (MAL2 and KRT8) and five methylation sites (cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210). The results showed that cg24852561 and cg22449901 made significant contributions when used alone to differentiate between FA and PT or FA and benign PT, with AUC, specificity, and sensitivity all greater than 0.9 in both the discovery and validation cohorts; cg00686880 made significant contributions when used alone to differentiate between malignant PT and Other FELs, with AUC, specificity, and sensitivity all greater than 0.7 in both the discovery and validation cohorts; cg09025210 made significant contributions when used alone to differentiate between benign and malignant PT, with AUC, specificity, and sensitivity all greater than 0.8 in both the discovery and validation cohorts; and cg00515756, KRT8, and MAL2 made significant contributions when used alone to differentiate between malignant PT and borderline PT, and between malignant PT and Other FELs, with AUC, specificity, and sensitivity all greater than 0.778 in both the discovery and validation cohorts. Furthermore, this invention found that when seven biomarkers are used in combination for FA vs. benign PT vs. borderline PT vs. malignant PT, the diagnostic AUCs of the three machine learning algorithms are NB=0.944, SVM=0.95, MLR=1 (XJTU cohort 1) and NB=0.971, SVM=0.975, MLR=1 (XJTU cohort 2), respectively, which are superior to the differential diagnostic efficacy when 10, 9 or 8 biomarkers are used in combination.

[0047] The present invention also provides the use of reagents for detecting the marker compositions described in the above technical solutions in the preparation of kits for identifying phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast.

[0048] This invention also provides a kit for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, the kit comprising reagents for detecting the marker composition described in the above-described technical solution; the reagents for detecting the marker composition described in the above-described technical solution include reagents for detecting protein markers and reagents for detecting methylated biomarkers;

[0049] The reagents for detecting protein markers include those used in any of the following methods: PRM targeted protein sequencing, selected reaction monitoring targeted sequencing, multiple reaction monitoring targeted sequencing, tandem mass spectrometry-tagged quantitative protein sequencing (TMT), isotope-tagged quantitative protein sequencing (iTRAQ), Western blotting (WB), enzyme-linked immunosorbent assay (ELSIA), and immunohistochemical staining (IHC); the reagents for detecting DNA methylation biomarkers include those used in any of the following methods: MethylTarget targeted methylation sequencing, methylated DNA immunoprecipitation sequencing, degenerate representative bisulfite sequencing, whole-genome bisulfite sequencing, deep target region DNA methylation sequencing, and pyrosequencing. In a specific embodiment, when the reagent for detecting DNA methylation biomarkers is a reagent used in the MethylTarget targeted methylation method, the reagent used in the MethylTarget targeted methylation method includes targeted methylation sequencing primers; the targeted methylation sequencing primers include any one or more of the following primer pairs: a forward primer with the nucleotide sequence cg00515756 as shown in SEQ ID NO.1 and a reverse primer with the nucleotide sequence cg00515756 as shown in SEQ ID NO.2; a forward primer with the nucleotide sequence cg24852561 as shown in SEQ ID NO.3 and a reverse primer with the nucleotide sequence cg24852561 as shown in SEQ ID NO.4; a forward primer with the nucleotide sequence cg00686880 as shown in SEQ ID NO.5 and a reverse primer with the nucleotide sequence cg00686880 as shown in SEQ ID NO.6; a forward primer with the nucleotide sequence cg22449901 as shown in SEQ ID NO.7 and a reverse primer with the nucleotide sequence cg00686880 as shown in SEQ ID NO.4; and a reverse primer with the nucleotide sequence cg00686880 as shown in SEQ ID NO.5. The reverse primer cg22449901 shown in NO.8, the forward primer cg09025210 with the nucleotide sequence shown in SEQ ID NO.9, the reverse primer cg09025210 with the nucleotide sequence shown in SEQ ID NO.10, the forward primer cg00819233 with the nucleotide sequence shown in SEQ ID NO.11, the reverse primer cg00819233 with the nucleotide sequence shown in SEQ ID NO.12, the forward primer cg04404381 ​​with the nucleotide sequence shown in SEQ ID NO.13, and the reverse primer cg04404381 ​​with the nucleotide sequence shown in SEQ ID NO.14.

[0050] This invention also provides a method for constructing a diagnostic model for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, comprising the following steps:

[0051] A diagnostic model for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast is constructed using the marker composition described in the above technical solution.

[0052] The diagnostic model includes one or more of the following types: Naive Bayes model, Support Vector Machine model, Multivariate Logistic Regression model, Decision Tree model, Random Forest model, K-Nearest Neighbors (KNN) model, Neural Networks model, Adaptive Boosting (AdaBoost) model, Gradient Boosting Machine (GBM) model, Extreme Gradient Boosting (XGBoost) model, Least Absolute Shrinkage and Selection Operator (lasso) model, Elastic Network (EN) model, and Ridge Regression model. This invention, by detecting proteins and methylation sites in fibroepithelial tumors and employing a machine learning model, can accurately diagnose tumor types. The specificity, sensitivity, accuracy, and consistency of this molecular diagnostic model are all excellent. The seven biomarkers of this invention can differentiate between FA vs. benign PT, FA vs. PT, benign PT vs. borderline PT, malignant PT vs. borderline PT, and malignant PT vs. Other FELs tumors.Specifically: cg00515756, cg00686880, cg22449901, cg24852561, and KRT8 can be used together to differentiate between FA and PT; cg00515756, cg00686880, cg09025210, cg22449901, cg24852561, and KRT8 can be used together to differentiate between FA and benign PT; cg00515756, cg00686880, cg09025210, cg22449901, cg24852561, and MAL2 can be used together to differentiate between benign PT and borderline PT; cg00515756, cg00686880, cg09025210, KRT8, and MAL2 can be used together to differentiate between malignant PT. vs. borderline PT; cg00515756, cg00686880, cg09025210, cg22449901, cg24852561, KRT8 and MAL2 can be used together to differentiate malignant PT vs. other FELs (the bolded markers can also differentiate between two types of tumors on their own, i.e., cg22449901 and cg24852561 can differentiate between FA and PT on their own; cg22449901 and cg24852561 can differentiate between FA and benign PT on their own; cg09025210 can differentiate between benign PT and borderline PT on their own; cg00515756, KRT8 and MAL2 can differentiate between malignant PT and borderline PT on their own; KRT8 and MAL2 can differentiate between malignant PT and other FELs on their own; however, the efficacy is best when multiple markers are used together for differentiation). Furthermore, compared to previous diagnostic models or panels for differentiating between FA and PT or benign PT and borderline / malignant PT, the combination of seven biomarkers in this invention has higher diagnostic efficacy, with AUC, specificity, and sensitivity of 1, 100%, and 100%, respectively.

[0053] This invention also provides a diagnostic model device for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast. The device includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, they are used to perform the following operations:

[0054] The detection data of the marker composition of the sample to be tested is obtained, and the detection data is input into the diagnostic model constructed by the marker composition to distinguish between phyllodes tumors and fibroadenomas of the breast and / or to distinguish between phyllodes tumors of different grades of the breast, so as to obtain the diagnostic results of the sample to be tested.

[0055] The marker composition is the marker composition described in the above technical solution.

[0056] This invention also provides a diagnostic system for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, comprising:

[0057] Acquisition unit, used to acquire detection data of the marker composition of the sample to be tested;

[0058] The processing unit inputs the detection data of the marker composition into a diagnostic model constructed from the marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, to obtain the diagnostic results of the sample to be tested.

[0059] The marker composition is the marker composition described in the above technical solution.

[0060] The present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program implementing the following method when executed by a processor:

[0061] The detection data of the marker composition of the sample to be tested is obtained, and the detection data of the marker composition is input into a diagnostic model constructed by the marker composition to differentiate between phyllodes tumors and fibroadenomas of the breast and / or to differentiate between phyllodes tumors of different grades of the breast, so as to obtain the diagnostic results of the sample to be tested.

[0062] The marker composition is the marker composition described in the above technical solution.

[0063] To further illustrate the present invention, the following detailed description, in conjunction with the accompanying drawings and embodiments, describes a marker composition, kit, diagnostic model construction method, and application for different types of breast fibroepithelial tumors provided by the present invention, but these descriptions should not be construed as limiting the scope of protection of the present invention.

[0064] Example 1

[0065] Screening of diagnostic molecules for phyllodes tumors of the breast based on high-throughput DIA proteomics and 850k microarray methylomics sequencing (XJTU Cohort 1)

[0066] 1. Patient clinical data and sample collection

[0067] 1) Sample Collection: Paraffin-embedded tumor specimens were collected from 40 patients with breast lumps (FELs) who underwent breast lumpectomy, extended breast lumpectomy, or total mastectomy at our hospital from January 2015 to December 2020 (designated XJTU cohort 1). Ten paraffin-embedded tumor blocks each were identified as FA, benign PT, borderline PT, and malignant PT. Inclusion Criteria: Cases with consistent diagnoses after double-blind review by two senior pathologists were included in the study. Exclusion Criteria: Inconsistent diagnoses from two pathologists, or samples that did not meet sequencing quality control requirements. After quality assessment, all 40 samples met proteomics requirements, and 36 samples met DNA methylation sequencing requirements (8 FA, 9 benign PT, 10 borderline PT, and 9 malignant PT).

[0068] 2) Collection of clinical data: Collect the patient's hospital number, age, height, weight, side of tumor occurrence, tumor size in the ultrasound report at the time of initial diagnosis, pathological diagnosis at the time of initial diagnosis, surgical details after initial diagnosis (operation time, surgical method, surgical margins), and postoperative specimen pathology report (pathological diagnosis, tumor size, tumor boundaries, stromal cell density, cell atypia, mitotic figures, stromal hyperplasia, whether there are heterogeneous components, whether there is intratumoral hemorrhage, and whether there is intratumoral necrosis).

[0069] 3) Patient follow-up: Through telephone follow-ups and case review, the main focus is on recording whether there is recurrence or metastasis. If there is recurrence / metastasis, the time and location of the recurrence / metastasis, whether the recurrence / metastasis was biopsied or surgically treated, and the pathological diagnosis in the biopsy or surgical pathology report should be recorded to provide clinical pathological and prognostic information for subsequent analysis.

[0070] 2GEO Dataset Download and Processing

[0071] This invention downloaded the methylation dataset for breast FELs—GSE231574 (including 2 cases of FA, 2 cases of benign PT, 9 cases of borderline PT, 18 cases of malignant PT, and 2 cases of breast metaplastic carcinoma) from the GEO database website (https: / / www.ncbi.nlm.nih.gov / geo / ). Since this invention does not involve the study of breast metaplastic carcinoma, these two samples were excluded, and the remaining 31 FEL patients were used as an external validation cohort for screening candidate diagnostic methylation sites. Preliminary data cleaning (e.g., removing samples or probes with excessive missing values), annotation, and normalization were performed using the ChAMP R package in R software. Subsequently, using the probe IDs of differentially methylated sites with AUC ≥ 0.9 selected from XJTU cohort 1 as unique identifiers, the methylation data of the samples corresponding to the aforementioned probe IDs were extracted from the GSE231574 dataset, and the diagnostic AUC value for each methylation site was calculated. This data, combined with 850k methylation data from 40 FEL patients in this invention, was used for further screening and validation of candidate diagnostic methylation sites.

[0072] 3DIA Quantitative Proteomics Detection

[0073] The DIA quantitative proteomics detection method of this invention was provided by Beijing Novogene Bioinformatics Technology Co., Ltd.

[0074] Specific experimental steps:

[0075] 1) Extraction of tissue protein

[0076] Tissue samples were removed from a -80°C freezer and ground into powder while maintaining the low temperature. The powder was then immediately transferred to centrifuge tubes pre-cooled with liquid nitrogen. PASP protein lysis buffer, prepared in a specific ratio containing 100 mM ammonium bicarbonate and 8 M urea, with the pH adjusted to 8, was added to the centrifuge tubes. The mixture was thoroughly vortexed and sonicated in an ice-water bath for 5 minutes to ensure complete lysis. Subsequently, the samples were centrifuged at 12,000 rpm for 15 minutes at 4°C.

[0077] 2) Protein reduction and alkylation

[0078] Obtain the supernatant after centrifugation, add 10 mM DTT, and allow it to react at 56°C for 1 hour to ensure protein reduction. Next, add sufficient IAM and continue the reaction at room temperature in a light-protected environment for 1 hour to complete the protein alkylation process.

[0079] 3) Protein precipitation and purification

[0080] After the reaction, add four times the volume of acetone pre-chilled at -20°C to the solution, then place the mixture at -20°C to allow precipitation for at least 2 hours. After precipitation, centrifuge at 12,000 rpm for 15 minutes at 4°C and collect the precipitate. Resuspend the precipitate in 1 mL of pre-chilled acetone at -20°C and wash it. Then, centrifuge again under the same conditions, collect the precipitate, and allow it to air dry. Finally, add an appropriate amount of protein lysis buffer (containing 8M urea and 100mM TEAB, with pH adjusted to 8.5) to dissolve the air-dried protein precipitate.

[0081] 4) Protein concentration determination

[0082] Prepare BSA standard protein solution (0–0.5 μg / μL) according to the Bradford Protein Quantitative Reagent Kit instructions. Add 20 μL of the standard solution and the sample to be tested to a 96-well plate, repeating 3 times. Add G250 staining solution, let stand for 5 min, and measure the absorbance at 595 nm. Calculate the protein concentration of the sample to be tested based on the standard curve.

[0083] 5) Gel electrophoresis analysis

[0084] In this invention, a 20 μg protein sample was selected and analyzed using 10% SDS-PAGE gel electrophoresis. After electrophoresis, Coomassie Brilliant Blue was used for staining until the bands were clearly visible.

[0085] 6) Proteolytic treatment

[0086] Dissolve an appropriate amount of protein sample in DB solution, add trypsin and TEAB buffer, and digest at 37°C for 4 hours; then add trypsin and CaCl2 again for overnight digestion; next, adjust the pH to <3, centrifuge at 12,000 rpm for 5 minutes; collect the supernatant after centrifugation and pass it through a C... 18 The desalting column was washed three times and eluted; the filtrate after passing through the desalting column was collected and freeze-dried.

[0087] 7) Construction of DDA spectral library and DIA mass spectrometry analysis

[0088] In the spectral library construction stage, the dissolved lyophilized samples were fractionated using an HPLC system, combined, lyophilized, and then redissolved in 0.1% formic acid. DDA mode was used for LC-MS detection, and raw data was generated by combining HPLC and mass spectrometry for constructing the DDA spectral library. Subsequently, DDA mode LC-MS analysis was performed. A specific mobile phase was prepared, the sample was processed, and iRT reagent was added. Detection was then performed using an advanced nano-scale UHPLC system and a Q Exactive™ HF-X mass spectrometer. During the detection process, various parameters of the mass spectrometer, including scan range, resolution, and C-trap capacity, were set in detail, and secondary mass spectrometry analysis was performed, ultimately generating raw mass spectrometry detection data.

[0089] 4. Database search and information analysis process

[0090] On one hand, the raw spectra obtained from DDA scanning are analyzed using Proteome Discoverer to convert spectral information into protein information. Simultaneously, DDA data quality control is performed, and the quality-controlled data is imported into Spectronaut to construct a real spectral library, DDAlibrary. On the other hand, Spectronaut is used to align the raw spectra obtained from DDA scanning into the DDAlibrary for protein identification.

[0091] Specific experimental steps:

[0092] 1) Database search and parameter settings

[0093] First, based on the homo_sapiens_uniprot_2020_7_2.fasta protein database (containing 192,320 sequences), a comprehensive search was performed on the mass spectrometry data acquired in DDA scanning mode using Proteome Discoverer 2.2 (PD2.2, provided by Thermo). During the search, strict mass tolerances were set for precursor ions and fragment ions; the mass tolerance for precursor ions was 10 ppm, while the mass tolerance for fragment ions was 0.02 Da.

[0094] 2) Results Filtering and Quality Control

[0095] To enhance the accuracy and reliability of the analysis results, this invention employed PD2.2 software for in-depth filtering and screening of the search results. During this process, only peptides with a reliability exceeding 99%, and proteins containing at least one unique peptide segment, were retained as reliable proteins. Furthermore, peptides and proteins with an FDR > 1% were removed. Through this series of rigorous screening and validation procedures, 34,379 peptides and 6,870 proteins were successfully identified in DDA mode.

[0096] 3) Spectral library generation and DIA data analysis

[0097] Search results from PD2.2 software were imported into Spectronaut software (version 14.0, provided by Biognosys) to generate a spectral library. Peptides and daughter ions were screened according to standards to generate a target list. Then, Spectronaut software was used to process the DIA data, extracting chromatographic peaks for matching and calculation, completing the qualitative and quantitative analysis of peptides. Simultaneously, iRT standards were used to correct retention times, improving accuracy. During the Spectronaut analysis, the PrecursorQ value cutoff parameter was set to 0.01 to ensure the reliability of the analytical results. Ultimately, in 40 samples of breast FELs, 29,305 peptides and 5,541 proteins were successfully identified using DIA mode.

[0098] 5. Protein differential expression analysis and cluster analysis

[0099] To determine the significance of protein differences between groups, the linear model method built into the limma software package was used to perform t-tests on the relative quantitative data of each protein in the two groups, and the corresponding P values ​​were calculated one by one. The criteria for screening differentially expressed proteins are as follows: if the Fold Change (FC) value of a protein is ≥2.0 and P<0.05, it is considered an upregulated protein; if the FC value of a protein is ≤0.5 and P<0.05, it is considered a downregulated protein. The upregulated and downregulated proteins between groups screened according to the above criteria were visualized by plotting volcano plots between multiple groups using ggplot2. The results show that as the histological grade difference between the two groups of tumors increases, the number of DEPs screened also increases, and vice versa. Figure 1A volcano plot was used to visualize differentially expressed proteins among multiple groups. In the volcano plot, each point represents a gene; red indicates upregulated proteins (P < 0.05, log2FC ≥ 1), blue indicates downregulated proteins (P < 0.05, log2FC ≤ 1), and gray indicates proteins with no significant change between groups. Furthermore, hierarchical clustering was implemented and visualized using the R package `hclust` and `pheatmap`, with color changes visually indicating protein expression levels. The results showed that FA and benign PT had similar protein expression levels, while borderline PT and malignant PT had similar protein expression levels. Figure 2 A hierarchical clustering heatmap of differentially expressed proteins; where each row represents a protein, and the expression level of each protein is converted using Z-score, with blue from 0 to -2 representing lower protein expression levels and red from 0 to 3 representing higher protein expression levels; each column represents a sample, and the top of the image shows the corresponding clinical information for each sample.

[0100] 6850k DNA methylomics sequencing

[0101] The 850k DNA methylomics sequencing (Illumina Infinium Methylation EPIC BeadChip) of this invention was provided by Beijing Novogene Bioinformatics Technology Co., Ltd.

[0102] The specific experimental steps are as follows:

[0103] 1) Extracting DNA

[0104] The QIAGEN QIAamp DNAFFPE Tissue Kit is a professional reagent kit for the precise extraction of DNA from FFPE (paraffin-embedded tissue) samples. This process encompasses dewaxing, protease digestion, DNA extraction, and subsequent purification steps to ensure the acquisition of high-quality DNA samples.

[0105] 2) DNA quantification and quality control

[0106] The concentration and volume of the extracted DNA were determined using a UV spectrophotometer. Furthermore, agarose gel electrophoresis was used to analyze the degree of DNA degradation and the presence of impurities such as RNA and proteins, ensuring the purity and integrity of the sample.

[0107] 3) DNA bisulfite conversion

[0108] According to EZ DNA Methylation-Gold TMThe kit guide converts unmethylated cytosine in DNA into uracil, while methylated cytosine remains unchanged, laying the foundation for subsequent methylation sequencing.

[0109] 4) Post-transformation DNA processing and instrumental detection

[0110] (1) Starting amount requirement: Ensure at least 200 ng of transformed DNA as the starting sample;

[0111] (2) DNA amplification: Denaturation and neutralization pretreatment are performed to prepare for the amplification reaction;

[0112] (3) Isothermal amplification: Amplification is performed overnight under suitable conditions to achieve uniform amplification of the whole genome;

[0113] (4) Enzyme-controlled fragmentation: Precisely process amplification products to avoid excessive fragmentation;

[0114] (5) DNA precipitation and resuspension: DNA was precipitated with isopropanol and resuspended in hybridization buffer;

[0115] (6) Chip hybridization: The processed DNA sample is added to the chip and annealed with a specific probe;

[0116] (7) Chip washing: Removes unhybridized and non-specifically hybridized DNA to ensure a clean environment;

[0117] (8) Extension and staining: Perform a single-base extension reaction on the chip and add detection markers;

[0118] (9) Scanning imaging: Using laser to excite fluorescent groups, images are recorded by a high-resolution scanning system.

[0119] 5) Data Analysis

[0120] (1) Data Quality Control and Standardization: After loading the original idat file generated by EPIC, the first step is to strictly control the raw data and standardize it using BMIQ to ensure the accuracy and comparability of the data. The probe site quality control includes the following: First, remove probes with P≥0.01, as well as probes with <3 magnetic beads and a proportion >5% in the sample; second, remove all non-CpG site probes in the dataset, as well as multi-hit probes with multiple matches; in addition, it is necessary to filter out SNPs within 5 base pairs around CpG sites, as well as all probes located on the X and Y chromosomes, to eliminate sex-specific effects.

[0121] (2) Calculation and evaluation of methylation level: Using the original signal data, this invention can calculate the methylation level of each site. The beta value (β) is usually used as the evaluation standard, ranging from 0 to 1. Judgment criteria: β>0.8 is high methylation, β<0.2 is low methylation, and 0.2≤β≤0.8 is partial methylation.

[0122] 7. Differential methylation site analysis

[0123] Differentially methylated CpG positions (DMPs) are a key component of methylation studies and are crucial for identifying subsequent biomarkers. This invention utilizes the `champ.DMP` function from the `champ` package for DMP analysis. This function internally calls the `limma` package to perform linear regression analysis and t-tests, while also conducting multiple hypothesis testing to obtain corrected p-values. The screening criteria for differentially methylated positions are: calibrated p-value < 0.05 and |Δβ| ≥ 0.2. Intergroup up- and down-regulated methylation positions selected based on these criteria are visualized using volcano plots generated using the `ggplot2` package in R software. The results show that the greater the histological grade difference between the two tumor groups, the more DMPs are screened, and vice versa. Figure 3 Volcano plots of differentially methylated sites across multiple groups were generated; each red dot represents a CpG site, and only DMPs with a calibration P < 0.05 and |Δβ| ≥ 0.2 are shown. Δβ ≥ 0.2 indicates upregulated DMPs; Δβ ≤ -0.2 indicates downregulated DMPs. (Δβ, delta). Furthermore, hierarchical clustering and visualization were performed using the hclust function and pheatmap in R software, and the methylation level of methylation sites can be intuitively displayed through color changes. The results showed that benign PTs and borderline PTs had similar methylation levels, while FAs and malignant PTs had their own distinct methylation patterns (…). Figure 4 A clustering heatmap of differentially methylated sites; where each row represents a CpG site, and the methylation level of each CpG site is converted using a Z-score, with blue gradually deepening from 0 to -3 indicating a lower methylation level, and red gradually deepening from 0 to 3 indicating a higher methylation level; each column represents a sample, and the top of the image shows the corresponding clinical information for each sample.

[0124] 8. Screening for diagnostic markers for phyllodes tumors

[0125] Proteomics based on the aforementioned 40 FEL samples ( Figure 2 ) and DNA methylomics ( Figure 4 Based on the difference analysis results, this invention found that integrating the two can more effectively distinguish between FA, benign PT, borderline PT, and malignant PT. Therefore, according to... Figure 5 The screening process (flowchart of diagnostic candidate biomarker screening) shows that, from the previously identified 1072 differentially expressed proteins (DEPs) and 94507 differentially methylated sites (DMPs), the DEPs with AUC ≥ 0.9 were ranked, and the top 30 were selected as candidate diagnostic proteins. However, during the screening of diagnostic methylation sites, the DMPs with AUC ≥ 0.9 of this invention were also validated using 850k methylation data from GSE231574. The DMPs that simultaneously met the AUC ≥ 0.9 were ranked, and the top 20 were selected as candidate diagnostic methylation sites. Finally, the varImpPlot function in the random forest algorithm was used to rank these 50 candidate biomarkers by importance, and the top 10 biomarkers were validated by targeted sequencing.

[0126] Example 2

[0127] Validation of candidate diagnostic molecules for phyllodes tumors of the breast based on PRM and MethlyTarget targeting technologies

[0128] 1. Collection of patient sequencing samples

[0129] Paraffin-embedded tumor specimens were collected from 88 patients with breast lumps who underwent excision, extended excision of breast lumps, or total mastectomy at the First Affiliated Hospital of Xi'an Jiaotong University from January 2021 to September 2023 (designated XJTU cohort 2). Among them, there were 23 cases of breast fibroadenomas (FA), 21 cases of benign superficial lumps (PT), 21 cases of borderline PT, and 23 cases of malignant PT (due to the large number of FA and benign PT patients, only information from patients hospitalized for surgery in the first week of each year was selected). The inclusion and exclusion criteria were the same as in step 1(1) of Example 1. The specific screening process is as follows: Figure 6 (Flowchart of patient screening in the validation cohort) As shown, due to the more stringent sample quality control requirements of MethlyTarget targeted methylation sequencing, after quality control of 88 samples, the remaining 82 samples met the MethlyTarget sequencing requirements (18 FA, 21 benign PT, 21 borderline PT, and 22 malignant PT). Additionally, 80 FELs samples were selected for PRM targeted protein quantitative sequencing (20 FA, 20 benign PT, 20 borderline PT, and 20 malignant PT). Therefore, due to sequencing quality control and sample size limitations, only 74 FELs samples (including 15 FA, 20 benign PT, 20 borderline PT, and 19 malignant PT) underwent both PRM targeted proteomics and MethylTarget targeted methylomics sequencing to further validate 10 diagnostic candidate biomarkers (including 3 proteins and 7 methylation sites).

[0130] 2PRM Targeted Quantitative Proteomics Detection

[0131] The PRM targeted quantitative proteomics of this invention was provided by Beijing Novogene Bioinformatics Technology Co., Ltd.

[0132] Specific experimental steps:

[0133] 1) Extraction of tissue protein

[0134] The steps are the same as step 3(1) in Example 1.

[0135] 2) Protein reduction and alkylation

[0136] The steps are the same as step 3 (2) in Example 1.

[0137] 3) Protein precipitation and purification

[0138] The steps are the same as step 3 in Example 1.

[0139] 4) Protein concentration determination

[0140] The steps are the same as step 3 (4) in Example 1.

[0141] 5) Gel electrophoresis analysis

[0142] The steps are the same as step 3 (5) in Example 1.

[0143] 6) Proteolytic treatment

[0144] The steps are the same as step 3 (6) in Example 1.

[0145] 7) Liquid chromatography-mass spectrometry (preliminary experiment)

[0146] (1) Preparation of mobile phase: Prepare solution A (made by mixing water and 0.1% formic acid) and solution B (made by mixing 80% acetonitrile and 0.1% formic acid);

[0147] (2) Sample preparation and injection: Mix equal amounts of peptide fragments, take 1 μg for label-free liquid chromatography-mass spectrometry detection, and use a pre-column and analytical column of specific specifications;

[0148] (3) Mass spectrometry parameter settings: Set the ion spray voltage and temperature, collect data in data-dependent mode, and set parameters such as mass range and resolution;

[0149] (4) Secondary mass spectrometry analysis: Select the top 20 high-intensity precursor ions for fragmentation, set the secondary mass spectrometry resolution, and optimize parameters to generate high-quality data;

[0150] (5) Data processing and peptide selection: Use PD2.2 software to search the library and select 1 to 3 unique peptides for each protein. Perform three PRM collections.

[0151] (6) Adjustment of chromatographic and mass spectrometric conditions: Replace the pre-column, repeat the liquid chromatography process, adjust the mass spectrometry to full scan plus PRM scan, and optimize the PRM parameters;

[0152] (7) Data analysis: Use Skyline software to analyze mass spectrometry data and evaluate peptide suitability.

[0153] 8) Liquid crystal mass spectrometry (formal experiment)

[0154] (1) Mobile phase remixing: Repeat the preparation of mobile phase A and mobile phase B to ensure that the component ratio is accurate;

[0155] (2) Sample processing and internal standard addition: Take equal amounts of peptide fragments after enzymatic digestion of each sample, mix them together, and uniformly incorporate a specific amount of relabeled peptide fragments as internal standards;

[0156] (3) Optimization of chromatographic conditions: The same EASY-nLCTM 1200 system and pre-column and analytical column configuration as the pre-experiment were used.

[0157] (4) Fine-tuning of mass spectrometry parameters: The mass spectrometer and ion source settings are consistent with the preliminary experiment. The key adjustment is to use full scan + PRM scan for mass spectrometry, and the peptide information in the inclusion list is updated based on the preliminary experimental results.

[0158] (5) Data acquisition: Perform full scan and PRM scan to ensure that parameters such as resolution, C-trap capacity and injection time are consistent with the preset, and keep the peptide fragmentation conditions unchanged in order to generate accurate raw mass spectrometry data.

[0159] (6) Data correction and analysis: Finally, the Skyline software was used to perform in-depth analysis of the data after the test, and the peak area was precisely corrected with the help of the internal standard peptide to ensure the accuracy of the quantitative results.

[0160] 3PRM Data Analysis

[0161] 1) Relative quantification of the target peptide in the sample

[0162] The target peptide chromatographic peaks were extracted from the raw data (.raw) of the PRM formal experiment using Skyline. Three daughter ions with high peptide abundance were selected for quantitative analysis. The peak area results of each target peptide segment after Skyline analysis were exported, including the target protein name, target peptide sequence, parent ion charge, selected daughter ions, daughter ion charge, and the original peak area of ​​each daughter ion used for quantification.

[0163] The peak areas of daughter ions of peptides in the target protein were integrated to obtain the original peak area of ​​the peptide in each sample. The relative expression level of each peptide in the sample was corrected using the peak area of ​​the heavily isotopically labeled internal standard peptide. The correction value was the ratio between the original peak area of ​​the peptide and the peak area of ​​the heavily labeled peptide incorporated into that sample. Finally, the relative quantification result of the target peptide in the sample was obtained.

[0164] 2) Relative expression level of the target protein in the sample

[0165] Based on the relative expression levels of the corresponding peptides of each target protein, the expression level of the target protein in the sample is further calculated.

[0166] 4MethylTarget targeted methylation sequencing

[0167] The MethylTarget targeted methylation sequencing of this invention was provided by Shanghai Tianhao Biotechnology Co., Ltd.

[0168] Specific experimental steps:

[0169] 1) DNA sample quality control

[0170] DNA sample concentration was detected using a NanoDrop 2000, with OD260 / 280 ranging from 1.7 to 1.9 to ensure DNA integrity and freedom from contamination. Additionally, a DNA concentration ≥20 ng / μL and a total volume ≥400 ng were required to meet the needs of downstream experiments. Finally, the main band was clearly visible during agarose gel electrophoresis.

[0171] 2) Primer design and PCR condition optimization

[0172] Methylation site sequencing primers were designed using Methylation Fast Target V4.1 software. The primer design included an Illumina adapter and specific amplification sequences. To ensure the success of downstream experiments, this invention provided one sample each of paraffin-embedded tissue samples of different PT subtypes as standards, which were then subjected to bisulfite treatment. Subsequently, the products were used to validate the primers, determining the accuracy and uniformity of the amplified product size.

[0173] 3) Optimization process for multiplex PCR primer models

[0174] The optimized primers (2) were used as a standard multiplex PCR panel. Amplification was performed using a unique multiplex PCR technique with bisulfite-treated standards as templates. Through repeated adjustments, a balance between primer amplification efficiency and specificity was achieved, and the composition and concentration of primers in the model were determined. Primer 3 was used for primer design on the bisulfite-treated sequences. http: / / primer3.ut.ee / .

[0175] 4) Bisulfite treatment process

[0176] In accordance with EZ DNA Methylation-Gold TM The Kit guides you through processing DNA samples to accurately convert unmethylated cytosine into uracil.

[0177] 5) Multiplex PCR reaction of the target fragment of the sample

[0178] (1) Amplification steps: The target region was amplified using an optimized primer model, and the amplification effectiveness was verified by agarose gel electrophoresis.

[0179] (2) Product mixing and dilution: For different models of the same sample, the PCR products were quantitatively mixed and diluted as Index PCR templates for adding specific tag sequences. The reaction system is shown in Table 1 below and the reaction conditions are shown in Table 2 below.

[0180] Table 1 Multiplex PCR reaction system

[0181] Components Volume (μL) 10× buffer solution 2 dNTP (2.5mM) 24. <![CDATA[MgCl2(25mM)]]> 12. DNA samples treated with bisulfite 1 HotTaq (5U / μL) 03. Multiplex PCR panel primers (1 μM) 2 Deionized water 111. Total volume 20

[0182] Table 2 Multiplex PCR reaction conditions

[0183]

[0184]

[0185] 6) Add specific tag sequences to samples

[0186] Prepare the reaction mixture as shown in Table 3 in a 96-well plate, add primers containing the Index sequence and tag sequences matching the Illumina platform, and amplify according to the reaction conditions shown in Table 4. Finally, mix the Index PCR products in equal proportions according to the amplification efficiency.

[0187] Table 3 Reaction System

[0188]

[0189] Table 4 Reaction System

[0190] step transsexual annealing extend Keep Cycle number Step 1 95℃, 2min 1× Step 2 95℃,20s 60℃,30s 72℃,30s 11× Step 3 72℃, 3min 1× Step 4 4℃ Keep

[0191] 7) Sample mixing, gel cutting and recycling

[0192] Prepare a 2% agarose gel, take the product from step 6) for electrophoresis (conditions: 120V, 35min), cut the target band according to DNA Marker B, and finally use TIANGEN Gel Extraction Kit for gel recovery.

[0193] 8) Library quantification and sequencing preparation

[0194] First, the library fragment length was validated using an Agilent 2100 Bioanalyzer to ensure library quality. Then, paired-end sequencing (2×150bp) was performed using an Illumina Hiseq or Nova seq platform to obtain raw FastQ data.

[0195] 5MetlyTarget Data Analysis

[0196] First, the filtered R1 and R2 reads were assembled using the software FLASH (FLASH: Fast length adjustment of short reads to improve genome assemblies). Then, FastX was used... http: / / hannonlab.cshl.edu / fastx_toolkit / index.html The spliced ​​FASTQ file was processed to obtain the FA format sequence. Then, BLAST+ (Camacho C, (2009) "BLAST+: architecture and applications") was used to align all reads in the FA with the target region reference sequence. Reads that covered 90% of their target sequence, or whose bases completely covered 90% of their target sequence, were selected as valid reads and statistically analyzed. The efficiency of C-to-T conversion after bisulfite treatment in the valid sequencing data of each sample was calculated. Finally, the methylation level of CpG sites in the amplicon was calculated.

[0197] Integrated analysis of 6PRM protein and MetlyTarget methylation data

[0198] Data from 74 patients in the XJTU cohort 2 study, including three proteins and seven methylation sites, were integrated. The mean Gini index was then calculated using the `randomForest` function in the `randomForest` R package. Generally, a higher Gini index indicates greater importance for the feature. The importance ranking of each feature was visualized using the `varImpPlot` function. The results are shown below. Figure 7(The graph shows the ranking of the importance of 10 candidate biomarkers using the random forest algorithm in XJTU cohort 2; the mean decrease in Gini index for each feature is calculated by averaging the decrease in Gini impurity of that feature in the entire model. The larger this value, the higher the importance of the feature.) As shown, this invention uses the random forest algorithm in XJTU cohort 2 to rank the diagnostic importance of 10 candidate biomarkers.

[0199] 7. Three machine learning algorithms to construct a phyllodes tumor diagnostic model

[0200] The top 10 diagnostic biomarkers were sequentially incorporated into three machine learning models. To ensure the reliability of the final selected diagnostic biomarkers, this invention employed three machine learning models: Naive Bayes (NB), Support Vector Machine (SVM), and Multinomial Logistic Regression (MLR). This invention used the Naive Bayes and SVM functions from the e1071R package to construct the NB and SVM models, respectively, and the multinom function from the nnet R package to construct the MLR model; the multiclass.roc function was used to calculate the model's AUC. The results show ( Figure 8 The graph shows the cumulative AUC values ​​of three machine learning models in XJTU Cohort 2. The horizontal axis represents the number of features included in each of the three models when the top 10 markers are sequentially incorporated (e.g., 5 represents the top 5 markers). The vertical axis represents the cumulative AUC value. The data table shows the cumulative AUC values ​​for each machine learning algorithm when features are sequentially incorporated. The red box indicates that the AUC of all three machine learning models reaches its highest point when the top 7 markers are incorporated. (NB, Naive Bayes; SVM, Support Vector Machine; MLR, Multivariate Logistic Regression) For the first seven features—including two proteins (A0A024R9E4_MAL2 and P05787_KRT8) and five methylation sites (cg00515756_DAGLA, cg24852561_CLYBL, cg00686880_RIPK4, cg22449901_PEX19, and cg09025210_ZBTB38)—all three machine learning models achieved optimal AUC values, with NB, SVM, and MLR AUCs of 0.971, 0.975, and 1, respectively. This result was also validated in XJTU cohort 1. Figure 9The graph shows the cumulative AUC values ​​of three machine learning models in XJTU cohort 1. The horizontal axis represents the number of features that were included in the top 10 markers in the three machine learning models, and the vertical axis represents the cumulative AUC value. The data table shows the cumulative AUC values ​​corresponding to each machine learning algorithm when the features were included in sequence. The red box indicates that the AUC of all three machine learning models reached its highest point when the first 7 markers were included. (NB, Naive Bayes; SVM, Support Vector Machine; MLR, Multivariate Logistic Regression). The AUCs of NB, SVM, and MLR are 0.944, 0.950, and 1, respectively. However, other combinations also showed excellent diagnostic efficacy. For example, when 10 biomarkers were used for joint diagnosis, the AUCs of the three machine learning algorithms were 0.945, 0.947, 0.956 (XJTU cohort 2) and 0.861, 0.917, 1 (XJTU cohort 1), respectively; when 9 biomarkers were used for joint diagnosis, the AUCs of the three machine learning algorithms were 0.967, 0.969, 0.956 (XJTU cohort 2) and 0.908, 0.917, 1 (XJTU cohort 1), respectively; and when 8 biomarkers were used for joint diagnosis, the AUCs of the three machine learning algorithms were 0.967, 0.975, 0.915 (XJTU cohort 2) and 0.861, 0.908, 1 (XJTU cohort 1), respectively. 1) When 6 biomarkers are used for joint diagnosis, the AUCs of the three machine learning algorithms are 0.959, 0.927, 0.852 (XJTU cohort2) and 0.894, 0.878, 0.991 (XJTU cohort1), respectively; when 5 biomarkers are used for joint diagnosis, the AUCs of the three machine learning algorithms are 0.949, 0.908, 0.852 (XJTU cohort 2) and 0.885, 0.862, 0.938 (XJTU cohort 1), respectively.Furthermore, individual biomarkers also demonstrated excellent efficacy in differential diagnosis. For example, cg24852561 and cg22449901 made outstanding contributions when used alone to differentiate between FA and PT or FA and benign PT, with AUC, specificity, and sensitivity all greater than 0.9 in both the detection and validation cohorts; cg00686880 made outstanding contributions when used alone to differentiate between malignant PT and other FELs, with AUC, specificity, and sensitivity all greater than 0.7 in both the detection and validation cohorts; cg09025210 made outstanding contributions when used alone to differentiate between benign PT and borderline PT, with AUC, specificity, and sensitivity all greater than 0.8 in both the detection and validation cohorts; and cg00515756, KRT8, and MAL2 made outstanding contributions when used alone to differentiate between malignant PT and borderline PT, and between malignant PT and other FELs, with AUC, specificity, and sensitivity all greater than 0.8 in both the detection and validation cohorts.

[0201] The peptide sequence information of the three candidate diagnostic proteins and the primer information for the seven methylation sites are shown in Tables 5 and 6.

[0202] Table 5. Mass-to-charge ratio (m / z) and peptide list of PRM-targeted sequencing for the three diagnostic proteins.

[0203] protein name Target peptide sequence (SEQ ID NO.) Mass-to-charge ratio (m / z) P05787(KRT8) ISSSSFSR(1) 435.719424 A0A024R9E4(MAL2) VTLPAGPDILR(2) 576.34278 P15924 (DSP) AELIVQPELK(3) 570.337164

[0204] Table 6. MethlyTarget targeted methylation sequencing primer information for 7 diagnostic methylation sites.

[0205] CpGs PrimerF (SEQ ID NO.) PrimerR (SEQ ID NO.) cg00515756 TGGTTYGTTGTTGTGGTTAGGAG(4) TAACACCACCTAAAACTCAACCAC(5) cg00686880 GTGGGTTTTTGGGTTTTGTG(6) CCCACTACCAAACCCRAACC(7) cg09025210 GGAAGATTTAGTYGGTYGAGAGTTTTAGTT(8) CCAAAATCACAAATACCCACAAAA(9) cg22449901 GYGAGAGTAAAGAGTGGTTAATGAATG(10) TTCCTTTTCTTCCTAACACTAAACAA(11) cg24852561 AATTGAGAAATGTTTTTGTAGTGTGAGAAG(12) CAACACAATAAACCAAATATCTAATACTCAA(13) cg00819233 GGGAGAGTTTAGTTTTAGGTTTTTATG(14) CCTACCCCACACCTCATTTTT(15) cg04404381 GTTTAAAGGGATTTTTGAGATTATAAATAGGA(16) TCCTACTTACTAAACRCCCCRAAAC(17)

[0206] The confusion matrix function was further used to output the confusion matrix, accuracy, consistency, specificity, sensitivity, positive predictive value, and negative predictive value of the three machine learning diagnostic models. As shown in Table 7, the confusion matrix and diagnostic efficacy metrics (specificity, sensitivity, accuracy, and consistency) of the three machine learning models constructed from the seven biomarkers in both cohorts are excellent. In summary, the seven-molecule diagnostic model integrating proteomics and methylmics can achieve differential diagnosis between FA and PT, as well as between different grades of PT, across multiple machine learning models.

[0207] Table 7 Evaluation metrics for three machine learning models based on seven markers in two cohorts.

[0208]

[0209] Note: NB stands for Naive Bayes; SVM stands for Support Vector Machine; MLR stands for Multivariate Logistic Regression; AUC stands for Area Under the Curve; Accuracy stands for Accuracy; Kappa stands for Consistency; PPV stands for Positive Predictive Value; NPV stands for Negative Predictive Value. Generally, a diagnostic efficacy of the model is considered good if the evaluation index is greater than 0.9 or 90%.

[0210] 8. Diagnostic contribution of diagnostic biomarkers to different classification outcomes

[0211] This invention further clarifies the diagnostic contribution of seven biomarkers to different classification outcomes and their expression trends in different graded PTs. For example... Figure 10 (Radar graph showing the diagnostic efficacy (AUC, specificity, and sensitivity) of seven biomarkers for different outcome classifications; where values ​​range from 0 to 1, extending outward from the center point 0; different colors represent different diagnostic indicators.) As shown, among these seven diagnostic biomarkers, the three methylation sites cg24852561, cg00686880, and cg22449901 are mainly used to differentiate between febrile pulmonary artery disease (FA) and total thrombotic PT (PT); cg09025210 is mainly used to distinguish between benign PT and borderline PT; MAL2, KRT8, cg00686880, and cg00515756 can effectively identify malignant PT from febrile pulmonary artery disease (FELs). These results indicate that each of the seven diagnostic biomarkers has its own advantages in individually differentiating different outcomes. While a single biomarker can achieve differentiation, it often fails to simultaneously possess both high specificity and sensitivity, leading to a high rate of missed or misdiagnosed diseases. Therefore, to avoid this drawback, clinical practice often employs combined diagnostic methods using multiple biomarkers to achieve optimal sensitivity, specificity, and accuracy. In this invention, biomarkers that simultaneously meet the criteria of specificity or sensitivity ≥0.7 in both cohorts are preferably used as diagnostic panels for two tumor subtypes, with specific combinations shown in Table 8.

[0212] Table 87 Diagnostic combinations of diagnostic markers in differentiating different outcomes

[0213]

[0214] Note: Taking the diagnostic combination of FA vs PT as an example, for the differential diagnosis of FA and PT, any one or any combination of the markers in the last five rows can be selected. The same applies to the differential diagnoses in the last four rows; "-" indicates that the value is less than 0.7. "Other" refers to the other three subtypes of FELs, including FA, benign PT, and borderline PT.

[0215] 9. Correlation between diagnostic markers and PT histological grading

[0216] t-tests were used to compare statistical differences between the two groups, and box plots of protein expression or methylation levels between groups were visualized using the ggboxplot function in the ggplot2 R package. The results showed that the protein expression levels of MAL2 and KRT8 decreased with increasing PT histological grade. Figure 11 The graph shows the differential expression of two diagnostic proteins in different PT grades; (a) expression box plot of the two diagnostic proteins in XJTU cohort 1; (b) expression box plot of the two diagnostic proteins in XJTU cohort 2; nsP>0.05, *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001); the DNA methylation levels at the three sites cg00515756, cg00686880, and cg09025210 increased with increasing PT histological grade. Figure 12 The results of the differences in methylation levels of 5 diagnostic methylation sites in different grade PTs are shown in the figure; (a) box plot of methylation levels of 5 diagnostic methylation sites in XJTU cohort 1; (b) box plot of methylation levels of 5 diagnostic methylation sites in XJTU cohort 2; ns P>0.05, *P<0.05, **P<0.01, ***P<0.001, ****P<0.0001); however, cg24852561 and cg22449901 showed the opposite trend ( Figure 12 The above results indicate that the levels of diagnostic marker proteins or methylation are closely related to the histological grading of PT. The diagnostic contribution of diagnostic markers to different classification outcomes (…) Figure 10 ) and its correlation with PT histological grading ( Figure 11 and Figure 12 This invention presents a schematic diagram of a decision tree for the future clinical application of a 7-molecule diagnostic panel. Figure 13 The AUC values ​​indicated here are from XJTU Cohort2. The red "hyper" indicates that this site is highly methylated in FA compared to PT; the green "low" indicates that this protein is poorly expressed in malignant PT compared to other FEL isoforms.

[0217] 10. Comparison of the multi-omics diagnostic model of this invention with traditional models

[0218] Previous single-atom traditional models focused on distinguishing between febrile apnea (FA) and persistent thromboembolic (PT), and between benign PT and borderline / malignant PT. The diagnostic model of this invention can effectively distinguish between FA, benign PT, borderline PT, and malignant PT. To compare with previous single-atom binary classification models, this invention used a binary logistic regression model to evaluate the diagnostic performance of its 7-molecule diagnostic model in distinguishing between FA and PT, and between benign PT and borderline / malignant PT. The results show that the model can effectively distinguish between FA and PT, and between benign PT and borderline / malignant PT, in both XJTU cohort 1 and XJTU cohort 2 (binary logistic regression: AUC = 1, sensitivity, specificity, positive predictive value, and negative predictive value are all 100%) (Table 9).

[0219] like Figure 14 (Comparison of multi-omics diagnostic models with traditional models; where (a) is a comparison of the diagnostic model of the present invention with previous diagnostic models for differentiating FA and PT; (b) is a comparison of the diagnostic model of the present invention with previous diagnostic models for differentiating benign PT and borderline / malignant PT. This figure is a lollipop chart, with the horizontal axis representing the AUC value of the model and the vertical axis representing the diagnostic models from different studies. When the same study uses different methods to construct models, each model constructed by each method is treated as an independent model and compared with the research model of the present invention; in addition, when multiple indicators are used to construct diagnostic models separately in the same study, each model constructed by each indicator is treated as an independent model and compared with the research model of the present invention.) As shown, the multi-omics diagnostic model of the present invention is superior to the reported models, including the somatic mutation model (TERTpromoter / MED12 / FLNA / SETD2), the radiomics model based on magnetic resonance imaging (MRI), the full-view digital slice (WSI) model of pathological HE, and the ultrasound (US) texture feature model (Table 9). In summary, the 7-feature multi-omics diagnostic model of the present invention performs excellently in distinguishing between FA, benign PT, borderline PT and malignant PT, and is superior to many previous diagnostic models.

[0220] Table 9. Diagnostic models of previously reported breast fibroepithelial neoplasia.

[0221]

[0222]

[0223] Note: NA indicates not reported in the literature; MRI, magnetic resonance imaging; PR, progesterone receptor; WSI, full-field digital slice; SVM, support vector machine; XGB, extreme gradient boosting; RF, random forest; CNN, convolutional neural network; RNN, recurrent neural network; Ridge-RFE, ridge regression combined with recursive feature elimination; AUC, area under the curve; PPV, positive predictive value; NPV, negative predictive value; FA, fibroadenoma; BPT, benign phyllodes tumor; BorPT, borderline phyllodes tumor; MPT, malignant phyllodes tumor; PT, phyllodes tumor.

[0224] Although the above embodiments have provided a detailed description of the present invention, they are only some embodiments of the present invention, and not all embodiments. People can obtain other embodiments based on these embodiments without creative effort, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A marker composition for differentiating between fibroadenoma and phyllodes tumor of the breast and / or for differentiating between different grades of phyllodes tumor of the breast, including benign, borderline and malignant phyllodes tumor of the breast, characterized in that, The marker composition comprises the following markers: MAL2, KRT8, cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210; wherein MAL2 and KRT8 are protein markers, and cg00515756, cg24852561, cg00686880, cg22449901, and cg09025210 are DNA methylation biomarkers in the human genome. ​ 2. The use of the reagent for detecting the marker composition of claim 1 in the preparation of a kit for identifying phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast.

3. A method for constructing a diagnostic model for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, characterized in that, Includes the following steps: A diagnostic model for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast is constructed using the marker composition of claim 1. The diagnostic models include: Naive Bayes model, Support Vector Machine model, or Multivariate Logistic Regression model.

4. A diagnostic model device for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, characterized in that, The device includes: a memory and a processor; the memory is used to store program instructions; the processor is used to invoke the program instructions, and when the program instructions are executed, it is used to perform the following operations: The detection data of the marker composition of the sample to be tested is obtained, and the detection data is input into the diagnostic model constructed by the marker composition to distinguish between phyllodes tumors and fibroadenomas of the breast and / or to distinguish between phyllodes tumors of different grades of the breast, so as to obtain the diagnostic results of the sample to be tested. The diagnostic model is the diagnostic model constructed by the construction method described in claim 3; The marker composition is the marker composition according to claim 1.

5. A diagnostic system for differentiating between phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, comprising: Acquisition unit, used to acquire detection data of the marker composition of the sample to be tested; The processing unit inputs the detection data of the marker composition into a diagnostic model constructed from the marker composition for differentiating phyllodes tumors and fibroadenomas of the breast and / or different grades of phyllodes tumors of the breast, to obtain the diagnostic results of the sample to be tested. The diagnostic model is the diagnostic model constructed by the construction method described in claim 3; The marker composition is the marker composition according to claim 1.

6. A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the following methods: The detection data of the marker composition of the sample to be tested is obtained, and the detection data of the marker composition is input into a diagnostic model constructed by the marker composition to differentiate between phyllodes tumors and fibroadenomas of the breast and / or to differentiate between phyllodes tumors of different grades of the breast, so as to obtain the diagnostic results of the sample to be tested. The diagnostic model is the diagnostic model constructed by the construction method described in claim 3; The marker composition is the marker composition according to claim 1.

Citation Information

Patent Citations

  • Method for carrying out in vitro molecular diagnosis of ovarian tumor and kit

    US20230203593A1