Tumor prediction system, method and application based on tongue coating microorganisms

A deep learning-based tumor prediction system analyzing tongue coating microorganisms effectively predicts tumors with high accuracy and sensitivity, addressing the limitations of existing gastric cancer diagnosis methods by leveraging microbial prevalence data.

JP2025524022APending Publication Date: 2025-07-25ZHEJIANG CANCER HOSPITAL
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025503354
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-22
Filing Date
2023-06-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Current gastric cancer diagnosis methods are invasive, costly, and lack specificity and sensitivity, with existing AI-based systems requiring specialized devices and high tester standards, and there is a lack of research on the correlation between tongue coating microbial populations and tumors.

Method used

A tumor prediction system using deep learning to analyze tongue coating microorganisms, specifically through a microbial information acquisition module and data processing module to predict tumor positivity based on discriminative features of microbial prevalence, employing a multi-layer perceptron neural network.

Benefits of technology

The system achieves high sensitivity (0.914/0.929), specificity (0.947), and accuracy (0.929/0.937) in tumor prediction, outperforming AI models based on blood tumor markers, providing a non-invasive and cost-effective early tumor screening solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524022000001_ABST
    Figure 2025524022000001_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of tumor diagnosis, prediction, and evaluation, and particularly relates to a tumor prediction system, method, and its application based on tongue coating microorganisms. The system includes a microorganism information acquisition module configured to acquire tongue coating microorganism information of a test sample, and a data processing module configured to obtain the probability that the test sample belongs to tumor positive by the following operation. The data processing module predicts the probability that the test sample belongs to positive based on the discriminative features on the tongue coating microorganism information obtained by automatic learning. Based on the abundance of different species and genera in tongue coating microorganisms, the probability that different test samples belong to tumor positive is automatically predicted, and this is used as an economic, non-invasive, highly efficient, and accurate early tumor screening policy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tumor diagnosis, prediction, and evaluation. More specifically, it relates to a tumor prediction system, method, and its application based on tongue coating microorganisms. By analyzing the correlation between tongue coating microorganisms and oncology, economic, non-invasive, and highly accurate tumor prediction is realized.

Background Art

[0002] According to the latest data, gastric cancer (GC) is the third leading cause of cancer-related death in the world. In 2020 alone, there were 1.09 million new GC cases and 770,000 deaths. Among them, in China, there were 480,000 new cases and 370,000 deaths, accounting for about half of the world's cases. China is a country with relatively high incidence and mortality rates of gastric cancer. Early detection, early diagnosis, and early treatment are crucial for reducing the mortality rate of gastric cancer. However, the current national early gastric cancer diagnosis rate is still less than 20%. The diagnosis and screening of GC still rely on gastroscopy, but its invasiveness is strong, the cost is high, and a professional endoscopist is required, so its application is greatly limited. In addition, due to the lack of specific symptoms in the early stage of gastric cancer, the specificity and sensitivity of clinical disease markers are relatively poor, and more than 60% of patients have local or distant metastases at the time of definitive diagnosis. The 5-year survival rate of patients with locally early GC exceeds 60%, while the 5-year survival rates of patients with local and distant metastases have significantly decreased to 30% and 5% respectively. Therefore, in order to improve the early diagnosis rate and prognostic effect of this group, a new GC diagnosis or screening method is urgently needed.

[0003] Traditional Chinese medicine is a medical science and cultural heritage that has been applied and preserved by the Chinese people for thousands of years. Tongue diagnosis is one of the important bases for traditional Chinese medicine to diagnose diseases. According to the theory of traditional Chinese medicine, the changes in tongue images (the color, size and shape of the tongue, the color, thickness and water content of the tongue coating) can reflect the health status of the human body, and are particularly closely related to stomach diseases. In the diagnosis of stomach diseases by traditional Chinese medicine practitioners, they often adopt empirical or dialectical methods based on tongue image information to obtain the specific manifestation forms of stomach diseases, and research has also revealed that the oral cavity or tongue coating microbiota is closely related.

[0004] Artificial intelligence (AI) can be used for the screening, diagnosis, and treatment of various diseases. The paper by scholars Cheung CY et al. (Cheung CY, Xu D, Cheng CY, et al. A deep-learning system for the assessment of cardiovascular disease risk via the measurement of retinal-vessel calibre. Nature biomedical engineering 2021;5(6):498-508. doi:10.1038 / s41551-020-00626-4 [published Online First:2020 / 10 / 14]) discloses a deep learning system that can measure the caliber of retinal blood vessels to assess the risk of cardiovascular disease and effectively predict the risk of cardiovascular disease. The paper by scholars Takenaka K et al. (Takenaka K, Ohtsuka K, Fujii T, et al. Development and Validation of a Deep Neural Network for Accurate Evaluation of Endoscopic Images From Patients With Ulcerative Colitis. Gastroenterology 2020;158(8):2150-57. doi:10.1053 / j.gastro.2020.02.012 [published Online First:2020 / 02 / 16]) develops a deep neural network (see reference) for evaluating endoscopic images of patients with ulcerative colitis. The network can recognize patients with endoscopic remission and histological remission with an accuracy of 90.1%, and the positive rate is 92.9%.

[0005] The patent CN110251084A of Fuzhou Data Technology Research Institute Co., Ltd. solves the problems of real-time detection, shooting, storage, and uploading of the tongue body in the tongue image collection process, and provides an artificial intelligence-based tongue image detection and identification method for identifying tongue image tongue color, tongue shape, coating texture, and coating color. Its solution is mainly related to the collection and identification technology of tongue images. Among them, tongue image identification focuses on extracting characteristics such as the color, texture, coating area, or coating thickness of the tongue image. However, these operations have not established a corresponding relationship between tongue images, coating information, and certain special stomach diseases, such as gastric cancer.

[0006] The patent CN111710394A of Shenyang Zhilang Technology Co., Ltd. proposes an artificial intelligence-assisted early gastric cancer screening system, which solves the problem of a large workload for determining gastric cancer positivity by automatically analyzing gastric camera rice images instead of manually. However, such a policy based on gastric camera image analysis first requires obtaining gastric camera images collected by a large number of specialized devices for model learning, and still needs to make decisions based on the gastric camera images of each tester during the test stage. There are still defects such as high time consumption, high material cost, and high tester standards for obtaining gastric camera images, making it difficult to achieve a comprehensive screening across the country.

[0007] The patent CN112133427A of Jiangsu Tianrui Precision Medical Technology Co., Ltd. provides an artificial intelligence-based gastric cancer auxiliary diagnosis system including a diagnosis selection module, a data collection module, a preprocessing module, a diagnosis module, and a display output module. The system can give personalized diagnosis results based on the collected data of the examinee. The data on which the diagnosis of the diagnosis system depends includes the basic information, living diet, infection history, disease history, family history, clinical symptoms, and examination items of the examinee. Among them, the data such as clinical symptoms and examination items are relatively difficult to collect, and the information such as basic information, living diet, infection history, disease history, and family history alone affects the early screening diagnosis effect.

[0008] The patent CN114203256A of Shanghai Rendon Medical Laboratory Co., Ltd. provides a method for constructing a MIBC classification and prognosis prediction model based on the prevalence of microorganisms. The method mainly analyzes the microorganism data of MIBC (muscle-invasive bladder cancer) patients from the MIBC transcriptome RNA seq data in the TCGA database. After obtaining the microorganism data, NMF clustering is performed using the prevalence spectrum of microorganisms as a feature to establish the molecular classification at the MIBC microorganism level. The correlation between microorganisms and MIBC is deeply analyzed from the microorganism level of tumor tissues, and a MIBC prognosis prediction model is established. This model helps to accurately predict the 1- to 5-year survival rate of MIBC patients. Therefore, the inventive solution is mainly for establishing a molecular classification for MIBC at the microorganism level and aims to accurately predict the prognosis survival rate of patients.

[0009] However, unfortunately, there is still no research demonstrating the corresponding relationship between the changes in the tongue coating microbial population and tumors, and the value of the changes in the tongue coating microbial population in tumor diagnosis and screening.

[0010] The present invention aims to solve these and other unsolved needs in the art.

Summary of the Invention

Problems to be Solved by the Invention

[0011] In order to solve at least one of the technical problems mentioned in the above background art, the present invention aims to design a tumor screening system based on deep learning using computer-aided means. The system automatically predicts the probability that different test samples belong to tumor-positive based on the prevalence of different species and genera in the tongue coating microorganisms, and uses this as an economic, non-invasive, highly efficient, and accurate early tumor screening policy.

Means for Solving the Problems

[0012] A tumor prediction system based on tongue coating microorganisms, wherein the system is A microbial information acquisition module configured to acquire tongue coating microbial information of a test sample, A data processing module configured to acquire the probability that a test sample belongs to tumor positive by the following operations, A data processing module that predicts the probability that a test sample belongs to positive based on discriminative features on the tongue coating microbial information obtained by automatic learning, and includes.

[0013] In one specific embodiment, the tumor is at least one of gastric cancer, breast cancer, colorectal cancer, esophageal cancer, hepatobiliary pancreatic cancer, lung cancer, prostate cancer, thyroid cancer, ovarian cancer, neuroblastoma, trophoblastic tumor or head and neck squamous cell carcinoma.

[0014] In one specific embodiment, the tumor is at least one of gastric cancer, breast cancer, colorectal cancer, esophageal cancer, hepatobiliary pancreatic cancer, lung cancer.

[0015] In one specific embodiment, the tumor is gastric cancer.

[0016] In one specific embodiment, the system further includes an output module configured to output a prediction result.

[0017] In one specific embodiment, the output module is configured to output tongue coating microbial information and a prediction result.

[0018] In one specific embodiment, the output module is output in at least one mode of electronic display, voice broadcast, printing, and network transmission.

[0019] In one specific embodiment, the tongue coating microbial information includes the prevalence of the genus and species of tongue coating microorganisms.

[0020] In one specific embodiment, the discriminative features are derived from the prevalence of the genus and species of microorganisms.

[0021] In a specific embodiment, the discriminative feature is derived from the high-dimensional features of the prevalence of genera and species of microorganisms. By sufficiently comparing, analyzing, and learning the commonalities and differences regarding the tongue coating microbial information of positive and negative patients, the discriminative features between the prevalences of genera and species of the tongue coating microorganisms in the test sample are deeply discriminated, and by obtaining the differences between positive and negative patients, the probability that the test sample belongs to tumor-positive can be determined, and through the comparison, analysis, and learning of the tongue coating microbial information, the diagnosis and prediction of tumors in the test sample can be realized, aiming to provide a non-invasive, non-human tissue-derived, highly economical, and highly accurate tumor diagnosis and prediction system.

[0022] In a specific embodiment, the data processing module obtains the probability that the test sample belongs to tumor-positive by the following operations: The trained neural network extracts high-dimensional features for the prevalence of genera and species of microorganisms input therein and then predicts the probability that the test sample belongs to positive.

[0023] In a specific embodiment, the neural network is a multi-layer perceptron (MLP).

[0024] In a specific embodiment, the neural network is trained in the following steps: 1) Input the prevalence of genera and / or species of tongue coating microorganisms collected from tumor-positive patients and / or tumor-negative populations as input vectors into the input layer of the model. 2) The hidden layer of the model extracts high-dimensional features of the prevalence of species and / or genera of microorganisms. 3) The softmax classifier in the output layer outputs the probability distribution of whether the genera and / or species of tongue coating microorganisms belong to positive and negative.

[0025] In a specific embodiment, the length of the input vector is 706 / 1339, representing the genera / species of microorganisms respectively.

[0026] In a specific embodiment, each element of the input vector represents the prevalence of microorganisms on a specific genus or species.

[0027] In a specific embodiment, the input layer corresponds to the length of the input vector.

[0028] In a specific embodiment, the number of neurons in the hidden layer is set to 512. Since the number of neurons in the hidden layer is larger than the number of neurons in the input layer, during forward calculation, the occupancy rate feature of the input layer is non-linearly mapped to a high-dimensional space to form high-dimensional features. After extracting the high-dimensional features, the probabilities belonging to positive and negative can be output via the input layer.

[0029] In a specific embodiment, each layer of the neural network except the output layer is attached with an activation function and normalization.

[0030] In a specific embodiment, the output layer includes two neurons, namely tumor positive and tumor negative.

[0031] In a specific embodiment, the neurons in each layer of the neural network are interconnected with each other, and the cross-entropy objective function is minimized using the probability distribution of the output:

Number

[0032] After verification with a large number of actual patients, applying tongue coating microorganisms as a means of non-invasive diagnosis and tumor screening has shown to be significantly superior to ordinary blood tumor markers. The tumor prediction system based on tongue coating microorganisms has better genus / species sensitivity (0.914 / 0.929 VS 0.283 - 0.566, 0.362 - 0.539), specificity (0.947 / 0.947 VS 0.688 - 0.976, 0.759 - 0.938), and accuracy (0.929 / 0.937 VS 0.603 - 0.622, 0.645 - 0.662) than the artificial intelligence models based on ordinary blood tumor markers, and its AUC value is also higher (0.945 / 0.975 VS 0.682 - 0.715, 0.694 - 0.760). Considering the huge burden of tumor examinations in China and the world, the extensive use of tongue coating microorganisms combined with artificial intelligence deep learning methods is the most economical, non-invasive, and acceptable way to screen and predict early tumors, and it is also considered to have a huge socio-economic impact.

[0033] By increasing the number of hidden layers or adjusting the number of units in the hidden layers, similar discrimination accuracy can be achieved. That is, neural networks composed of fully connected layers with different hyperparameters are all applied to the tumor positive discrimination task based on the prevalence of microorganisms. Through deep learning technology, the probability of tumor positivity is automatically judged to screen people with multiple tumors.

[0034] A tumor prediction method based on tongue coating microorganisms, the method comprising: obtaining the tongue coating microorganism information of the test sample; inputting the tongue coating microorganism information of the test sample into the system to obtain the tumor positive probability of the test sample.

[0035] The tongue coating microorganism information includes the prevalence of the genus and species of tongue coating microorganisms.

[0036] An application of the aforementioned tumor prediction system and / or method based on tongue coating microorganisms, the application comprising: applying the system and / or method to perform tumor prediction on a test sample.

[0037] Based on the common knowledge in the art, the above preferred conditions can be combined with each other to obtain specific embodiments.

Advantages of the Invention

[0038] The beneficial effects of the present invention are as follows: The present invention aims to provide a tumor screening system based on deep learning. By sufficiently comparing, analyzing, and learning the commonalities and differences in the tongue coating microbial information of positive and negative patients, the differences between positive and negative patients are obtained. By deeply discriminating the discriminative features between the prevalence rates of genera and species of the tongue coating microorganisms in the test sample, the probability that the test sample belongs to tumor positive can be determined. Thereby, through learning the tongue coating microbial information, the diagnosis and prediction of tumors in the test sample can be realized. Through verification, the sensitivity of genera / species when the tongue coating microorganisms predict gastric cancer reaches 0.914 / 0.929, the specificity reaches 0.947 for both, the accuracy reaches 0.929 / 0.937, and its AUC value reaches 0.945 / 0.975. The prediction effect is superior to the artificial intelligence model based on ordinary blood tumor markers. As an economic, non-invasive, highly efficient, and accurate early tumor screening policy, it will surely bring huge social and economic effects.

[0039] The present invention adopts the above technical solutions to achieve the above objectives, make up for the deficiencies of the prior art, and has a reasonable design and easy operation.

Brief Description of the Drawings

[0040] To enable those skilled in the art to more quickly and clearly understand the above and / or other objectives, features, advantages, and examples of the present application, a part of the accompanying drawings is provided. It should be pointed out that the accompanying drawings, schematic embodiments, and their descriptions constituting the specification of the present application are for providing a further understanding of the present application and do not constitute an undue limitation to the present application.

[0041]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0042] Those skilled in the art can, referring to the content of this specification, realize appropriate substitutions and / or changes of process parameters. In particular, it is necessary to point out that all similar substitutions and / or changes are obvious to those skilled in the art, and all of these are considered to be included in the present invention. The content described in the present invention has been explained by preferred embodiments, but it is obvious to those concerned that the technology of the present invention can be realized and applied by changing or making appropriate changes and combinations of the content described in this specification without departing from the content, spirit and scope of the present invention.

[0043] It should be noted that the following detailed descriptions are all illustrative and are intended to further explain the present application. Unless otherwise limited, all technical terms and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art.

[0044] Note that the terms used herein are for the purpose of describing specific embodiments and are not intended to limit the technical solutions of the present application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly dictates otherwise. Further, when the terms "comprising" and / or "including" are used in this specification, it should also be understood that they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0045] In the present invention, the tongue coating microbial information specifically includes the genus, species of microorganisms and their abundance rates. Any method capable of obtaining the genus, species of tongue coating microorganisms in a sample and their abundance rates can be used to obtain relevant information. In the present application, the following method was adopted to obtain microbial information.

[0046] The samples (including subjects, training sets and test sets) were gargled three times with sterile water before eating breakfast and drinking water. A professional operator collected the tongue coating samples using a tongue swab. Specifically, the throat swab was used to roll the swab simultaneously, 30 times from the root of the tongue to the tip (each swab was rolled 5 times, a total of 6 swabs). Immediately after sampling was completed, the swab was placed in a freezer tube and the sample was transferred to a refrigerator at -80°C. According to the manufacturer's instructions, the E.Z.N.A. Tissue DNA Extraction Kit (D3396-01, Omega, Norcross, Georgia, USA) was applied to extract microbial DNA. The AxyPrep PCR Purification Kit (AP-PCR-500 G, Corning, NY, USA) was used to separate, extract and purify the PCR products. The Quant-iT PicoGreen ds DNA reagent (P 7581, Thermo Scientific, Waltham, MA, USA) was used to quantitatively measure the products. Finally, 2×250bp paired-end sequencing was performed using the Novaseq sequencer of LC-Bio Co., Ltd.

[0047] The present invention will be described in more detail below.

[0048] <Clinical Specimen> A national multi-center clinical study was conducted to eliminate the influence of differences in region, diet, and center on the study. It included 11 centers in 8 cities: Hangzhou, Wenzhou, and Shanghai in the east, Fuzhou in the south, Chengdu in the west, Liaoning and Heilongjiang in the north, and Taiyuan in the central region.

[0049] As shown in Figure 1, from January 2020 to October 2021, 1,111 gastric cancer (GC) patients were recruited from 8 centers, 1,519 non-gastric cancer (NGC) patients were recruited from 3 centers, including 169 healthy controls (HCs), 648 superficial gastritis (SGs), and 702 atrophic gastritis (AGs). Among the gastric cancer (GC) patients, 865 cases were randomly selected from the gastric cancer (GC) patients, and 1,287 cases were randomly selected from the non-gastric cancer (NGC) patients for training and verification of the system. Among them, there were 448 cases of early GC (TNMI+II stage), 417 cases of advanced GC (TNMIII+IV stage), 141 cases in the healthy control group (HC), 547 cases of superficial gastritis (SG), and 599 cases of atrophic gastritis (AG). Approximately 80% of the cases were used as the training dataset, and approximately 20% of the cases were used as the internal validation dataset. Additionally, 246 cases of GC and 232 cases of NGC from 3 centers were used as an independent external validation dataset, including 162 cases of early GC, 84 cases of advanced GC, 28 cases of HC, 101 cases of SG, and 103 cases of AG. These gastric cancer (GC) patients were all newly diagnosed gastric cancers, had never received treatment for the disease in the past, and had not undergone surgery, chemotherapy, radiotherapy, targeted therapy, or biological therapy for the disease. All gastric cancer (GC) patients had single tumors, that is, patients with two or more malignant tumors were excluded. HCs, SGs, and AGs were confirmed by gastroscopy.

[0050] Clinical information of all participants such as age, gender, height, weight, family history, smoking, drinking, TNM staging, and blood tumor markers was collected. The pathological staging was based on the 8th edition, No. 23 of the American Joint Committee on Cancer. Table 1 shows the general patient information such as age, gender, BMI, smoking, and drinking in the GC group and the NGC group, which was also in good agreement in the training, internal validation, and independent external validation datasets. JPEG2025524022000004.jpg95170JPEG2025524022000005.jpg147170

[0051] <Statistical analysis> All statistical analyses were performed using SPSS 23.0 software (SPSS Inc., Chicago, IL, USA). The results were expressed as mean ± SD or mean ± SEM. Parametric or non-parametric tests were used according to whether the data were orthogonally distributed. Count data were analyzed using the chi-square test. P < 0.05 was considered statistically significant.

[0052] <Ethical approval> The study of this application obtained approval from the centralized ethics committee used by 11 participating centers, such as the Research Ethics Committee of Zhejiang Cancer Hospital, the First Affiliated Hospital of Wenzhou Medical University, Liaoning Cancer Hospital, Renji Hospital Affiliated to Shanghai Jiao Tong University, Fujian Cancer Hospital, Tumor Hospital Affiliated to Harbin Medical University, Sichuan Cancer Hospital, Shanxi Cancer Hospital, Tongde Hospital of Zhejiang Province, Zhejiang Provincial Hospital of Traditional Chinese Medicine, and Yuhang People's Hospital.

[0053] <Clinical verification> Example 1 Using computer-aided means, a tumor prediction system based on deep learning was designed, and the system automatically predicted the probability that different test samples belong to tumor positive based on the prevalence of microorganisms in different genera and species.

[0054] We obtained the prevalence data of the genera and species of tongue coating microorganisms in positive and negative samples from gastric cancer patients and non-gastric cancer groups, respectively.

[0055] Based on the above prevalence data of the genera and species of tongue coating microorganisms, a neural network based on the multi-layer perceptron (MLP) was designed, with the prevalence of different genera and species of tongue coating microorganisms as the input, and the probability that the test sample belongs to positive was predicted based on the finally extracted high-dimensional features.

[0056] As shown in Figure 2 for the overall algorithm structure, the input of the entire model is a vector of length 706 or 1339 respectively representing the number of genera or species of microorganisms, and each element of the input vector represents the prevalence of microorganisms in a specific genus or species. The input layer of the multi-layer perceptron corresponded to the length of the input vector. High-dimensional features of the prevalence of microorganisms were extracted by adding hidden layers. The output layer included two neurons and a softmax classifier, which output the probabilities that the input sample was positive and negative for gastric cancer respectively. Neurons between each layer were interconnected respectively, and the cross-entropy objective function was minimized using the output probability distribution:

Number

[0057] Tongue coating microorganism information data of 328 GCs (gastric cancers) and 304 NGCs (non-gastric cancers, including 155 HCs (health) and 149 AGs (atrophic gastritis)) were collected from the Zhejiang Cancer Hospital and the First Affiliated Hospital of Zhejiang Chinese Medical University. The clinical information of gastric cancer participants and non-gastric cancer participants is shown in Table 2, and it was found that the general clinical information such as age, gender, smoking, and drinking in both groups was well matched. JPEG2025524022000007.jpg40170

[0058] Gastric cancer participants and non-gastric cancer participants divided the data into a training set and a test set at a ratio of 8:2. The sensitivity, specificity, and accuracy for the test set are shown in Tables 3 and 4. JPEG2025524022000008.jpg17170JPEG2025524022000009.jpg17170

[0059] As can be seen from Table 3 and Table 4, regarding the diagnostic prediction of gastric cancer, the tongue coating microbiota shows the same high specificity at both the genus and species levels. The possible reason for the same data is that the sample size is limited. It must be noted that the microbiota shows higher sensitivity and accuracy at the species level than at the genus level, and shows high sensitivity, specificity, and accuracy for the diagnostic prediction of gastric cancer. As a control, a tumor detection system based on tongue coating microbiota was compared with blood tumor markers for normal clinical applications, and the prediction of tumors was verified using a combination of multiple classical blood tumor markers. The selectable blood tumor markers are at least one selected from alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), cancer antigen 125 (CA125), cancer antigen 15-3 (CA15-3), cancer antigen 199 (CA199), cancer antigen 72-4 (CA72-4), cancer antigen 242 (CA242), cancer antigen 50 (CA50), non-small cell lung cancer-related antigen (CYFRA21-1), small cell lung cancer-related antigen (neuron-specific enolase, NSE), squamous cell carcinoma antigen (SCC), total prostate-specific antigen (TPSA), free prostate-specific antigen (FPSA), alpha-L-fucosidase (AFU), Epstein-Barr virus antibody (EBV-VCA), tumor-associated substance (TSGF), ferritin (Ferritin), beta2-microglobulin (β2-MG), pancreatic embryo antigen (POA), or gastrin-releasing peptide precursor (PROGRP), and particularly at least one selected from CEA, CA242, CA72-4, CA125, CA199, CA50, AFP, or Ferritin, and more particularly a combination of the above eight blood tumor markers was selected. The prediction method based on the above blood tumor marker index includes the following steps.

[0060] 1) Data preprocessing: Since there are varying degrees of missing values in the serum indicators of all cases, the training data needs to be complete. Therefore, before model training, it is necessary to first perform complementation on the data. In this application, the K-nearest neighbor missing value interpolation method is used to perform complementation on the data. Specifically, the missing serum indicator complementation value is the average value of the values of two adjacent neighbors.

[0061] 2) Model training: The present invention adopts three types of machine learning classification methods, namely, Support Vector Machine (SVM), Decision Tree (DT), and K-Nearest Neighbor Classifier (KNN). Specifically, eight types of blood tumor marker indicators (CEA, CA242, CA72-4, CA125, CA199, CA50, AFP, and Ferritin) of cases correspond to the characteristics of the samples, and the positive and negative diagnoses of the cases correspond to the tags of the samples. All the complemented samples are sent to the three types of classifiers for fitting.

[0062] 3) Model evaluation: This application evaluates the model using internal validation and external validation. Internal validation adopts data of different cases from the same hospital as the training data, and external validation adopts data of case from hospitals different from the training data. The model is predicted using three indicators including sensitivity, specificity, and accuracy.

[0063] As shown in Table 5 for the clinical information of blood tumor markers of related GC patients, it was found that the concentrations of blood tumor markers such as CEA, CA424, CA724, CA125, CA199, CA50, AFP, and Ferritin in GC patients were significantly increased compared with those in NGC patients. JPEG2025524022000010.jpg114170

[0064] A total of 246 cases of GC and 232 cases of NGC from three centers were used as an independent external validation dataset. Approximately 80% of the cases were used as the training dataset, and approximately 20% of the cases were used as the internal validation dataset. As shown in Table 6, the verification results of the sensitivity, specificity, and accuracy of the GC diagnosis using the machine learning classification method based on hematological tumor markers showed that the sensitivity of the internal validation of the three models SVM, DT, and KNN based on hematological tumor markers was 0.283 - 0.566, the specificity was 0.688 - 0.976, the accuracy was 0.603 - 0.622, the external validation sensitivity was 0.362 - 0.539, the specificity was 0.759 - 0.938, and the accuracy was 0.645 - 0.662. Referring to Figure 3, the ROC and AUC of the internal and external validation of the GC diagnosis using the machine learning classification method based on hematological tumor markers showed that the range of the AUC value of the internal validation was 0.682 - 0.715, and the range of the AUC value of the external validation was 0.694 - 0.760. In the SVM algorithm, the specificities of both the internal and external validations reached over 90%, indicating that the algorithm can provide valuable information for gastric cancer diagnosis. However, among DT and KNN, the specificity decreased to some extent, and both the sensitivity and accuracy improved to varying degrees, providing comprehensive information for gastric cancer diagnosis. JPEG2025524022000011.jpg40170

[0065] It should be clearly noted that the above comparison scheme of the present application selects 8 serum indicators including CEA, CA242, CA72 - 4, CA125, CA199, CA50, AFP, and Ferritin. Any increase, decrease, or replacement of some of the serum indicators can predict the positive and negative of tumors, especially gastric cancer. The above comparison scheme adopts three machine learning classifiers SVM, DT, and KNN, and other machine learning classifier methods such as logistic regression and random forest can also achieve the corresponding purpose.

[0066] Comparing and analyzing Tables 3-4 and Table 6, the tumor prediction system based on tongue coating microorganisms according to the present application has significantly better sensitivity and accuracy than the model based on the above blood tumor markers, and its specificity is better than that of models DT and KNN. It was found that the solution of the present application is highly forward-looking, has good economy, is non-invasive, and has higher sensitivity, specificity and accuracy.

[0067] Principal Coordinate Analysis (PcoA) is a visualization method for dimensionalizing multi-dimensional data to examine the similarity and difference of data. It searches for principal coordinates based on the distance matrix, rearranges them with a series of eigenvalues and eigenvectors, and then mainly selects the upper eigenvalues to effectively find the most "major" elements and structures in the data, describe the relationship between samples. First, sampling is randomly performed to calculate the one-sided distance between each sample, and a two-dimensional PcoA diagram is created based on the distance matrix. The closer the distance between samples, that is, the more similar the presence rate and composition of microorganisms, the closer their distances in the PcoA diagram. The principal coordinate analysis diagram of microorganisms in the GCs group and the NGCs group is as shown in Figure 4, and the heat map of the top 30 microorganisms in the GCs group and the NGCs group is as shown in Figure 5. Analyzing Figures 4 and 5, it was found that there are significant differences between the microbial communities of the GCs group and the NGCs group.

[0068] The Shannon diversity index is an index for examining the diversity (α-diversity) in the local growth environment of a community, and its value shows a positive correlation with species diversity. As shown in Figure 6, the Shannon diversity index of the present application indicates that in terms of α-diversity, the presence rate of species in GCs is significantly higher than that in NGCs, and the number of species in the tongue coating of GCs is relatively large.

[0069] In addition, 80% of the samples were used as the training set, 20% of the samples were used as the validation set, and an MLP model was constructed. Figure 7 shows the ROC (Receiver Operating Characteristic) and AUC (Area Under roc Curve) of the validation set. It can be seen that the system of the present application has an ROC curve relatively far from the line of (0,0)-(1,1). By observation, the value of AUC genus is 0.945, the value of species is 0.975, and the AUC values of species / genus are significantly higher than the internal validation AUC values (0.682 - 0.715) and external validation AUC values (0.694 - 0.760) of the SVM, DT, and KNN models of eight types of blood tumor markers. It is found that the tumor prediction system based on the tongue coating microorganisms of the present application is a prediction model with good performance, and the diagnostic value for GC is significantly better than that of the model simply applying the combination of eight types of blood tumor markers.

[0070] Comparing the species distributions of the GCs group and the NGCs group, as shown in E of Figure 8 for the display at the genus level and F of Figure 8 for the display at the species level, the tumor prediction system based on the tongue coating microorganisms of the present application can accurately perform diagnostic prediction for GCs by clearly distinguishing at the genus and species levels of the tongue coating microorganisms, and it is found that it provides a forward, economic, non-invasive, and effective screening and diagnostic prediction method for tumors.

[0071] Since the prior art in the above embodiments is the prior art known to those skilled in the art, detailed descriptions are omitted here.

[0072] The specific embodiments described in this specification are only illustrative of the spirit of the present invention. Those skilled in the art can make various changes or supplements to the described specific embodiments, or can replace them in a similar way, but they will not deviate from the spirit of the present invention or exceed the scope defined in the appended claims.

[0073] Although the present invention has been described in detail and several specific embodiments have been cited, it will be apparent to those skilled in the art that various changes or modifications are possible without departing from the spirit and scope of the present invention.

[0074] The above description is only a preferred embodiment of the present application and is not used to limit the present application. For those skilled in the art, various changes and variations are possible to the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application should all be included within the protection scope of the present application.

[0075] Matters not described in the present invention are all well-known techniques.

Claims

1. A tumor prediction system based on tongue coating microorganisms, comprising: a microorganism information acquisition module configured to acquire tongue coating microorganism information of a test sample; a data processing module configured to obtain the probability that the test sample belongs to tumor positive by the following operations, the data processing module for predicting the probability that the test sample belongs to positive based on discriminative features on the tongue coating microorganism information obtained by automatic learning;

2. The tumor is at least one of gastric cancer, breast cancer, colorectal cancer, esophageal cancer, hepatobiliary pancreatic cancer, lung cancer, prostate cancer, thyroid cancer, ovarian cancer, neuroblastoma, trophoblastic tumor or head and neck squamous cell carcinoma. The system according to claim 1, characterized in that.

3. The tongue coating microorganism information includes the prevalence of genera and species of tongue coating microorganisms. The system according to claim 1 or 2, characterized in that.

4. The discriminative features are derived from the prevalence of genera and species of microorganisms. The system according to claim 1 or 2, characterized in that.

5. The data processing module obtains the probability that the test sample belongs to tumor positive by the following operations: The trained neural network predicts the probability that the test sample belongs to positive after extracting high-dimensional features for the prevalence of genera and species of microorganisms input therein. The system according to claim 1 or 2, characterized in that.

6. The neural network is trained in the following steps: 1) Input the prevalence of genera and / or species of tongue coating microorganisms collected from tumor positive patients and / or tumor negative populations into the input layer of the model as input vectors; 2) The hidden layer of the model extracts high-dimensional features of the prevalence of genera and / or genera of microorganisms; 3) Output the probability distribution of the genera and / or species of tongue coating microorganisms belonging to positive and negative by the softmax classifier of the output layer. The system according to claim 5, characterized in that.

7. Each element of the input vector represents the prevalence of microorganisms on a specific genus or species. The system according to claim 6, characterized in that.

8. The output layer includes two neurons, namely tumor positive and tumor negative. The system according to claim 6 or 7, characterized in that.

9. Obtaining the tongue coating microorganism information of the test sample; Input the tongue coating microbial information of the test sample into the system according to any one of claims 1 to 8 to obtain the tumor positive probability of the test sample, and a tumor prediction method based on tongue coating microorganisms, characterized in that it includes this.

10. An application of the system and / or method according to any one of claims 1 to 8 and / or the method according to claim 9, characterized in that it includes performing tumor prediction on a test sample by applying the system and / or method.

Citation Information

Patent Citations

  • Method and equipment for analyzing evolutionary relationship and abundance information of sample-based microbial population

    CN114093411A

  • MIBC typing and prognosis prediction model construction method based on microbial abundance

    CN114203256A

  • Tumor prediction system, method and its application based on tongue image

    JP2025524023A

  • Methods and systems for oral microbiome analysis

    US20210142906A1

  • Machine learning device, learned model, data structure, periodontal disease inspection method, periodontal disease inspection system and periodontal disease inspection kit

    WO2018159712A1