Tumor prediction system, method and its application based on tongue image
An AI-based tumor prediction system using tongue image analysis addresses the limitations of invasive gastric cancer diagnosis by achieving high accuracy and cost-effectiveness in screening and predicting various cancers.
Patent Information
- Application Number
- JP2025503357
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-22
- Filing Date
- 2023-06-29
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-29
AI Technical Summary
Current gastric cancer diagnosis methods are invasive, costly, and lack sensitivity and specificity, especially in early stages, necessitating a non-invasive and cost-effective screening method.
A tumor prediction system using AI deep learning to analyze tongue images, deriving discriminative features between positive and negative categories to predict the likelihood of gastric cancer based on tongue image analysis.
The system achieves high accuracy in tumor prediction, with sensitivity and specificity significantly better than conventional blood tumor markers, providing a non-invasive and economical screening solution for gastric cancer and other malignancies.
Smart Images

Figure 2025524023000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tumor diagnosis, prediction, and evaluation. More specifically, it relates to a tumor prediction system, method, and its application based on tongue image, and realizes economic, non-invasive, and highly accurate tumor prediction by analyzing the correlation between tongue image and oncology.
Background Art
[0002] According to the latest data, gastric cancer (GC) is the third leading cause of cancer-related death in the world. In 2020 alone, there were 1.09 million new GC cases and 770,000 deaths. Among them, in China, there were 480,000 new cases and 370,000 deaths, accounting for about half of the world's cases. The diagnosis and screening of GC still rely on gastroscopy, but due to its strong invasiveness, high cost, and the need for professional endoscopists, its application is greatly limited. In addition, due to the lack of specific symptoms in the early stage of gastric cancer, the specificity and sensitivity of clinical disease markers are relatively poor, and more than 60% of patients have local or distant metastasis at the time of definitive diagnosis. The 5-year survival rate of patients with locally early GC exceeds 60%, while the 5-year survival rates of patients with local and distant metastases decreased significantly to 30% and 5% respectively. Therefore, there is an urgent need for a new GC diagnosis or screening method to improve the early diagnosis rate and prognosis of this group of patients.
[0003] Traditional Chinese medicine is a medical science and cultural heritage that has been applied and preserved by the Chinese people for thousands of years. Tongue diagnosis is one of the important bases for traditional Chinese medicine to diagnose diseases. According to the theory of traditional Chinese medicine, the changes in tongue image (the color, size and shape of the tongue, the color, thickness and water content of the tongue coating) can reflect the health status of the human body, and is especially closely related to stomach diseases. However, there is still no research demonstrating the corresponding relationship between tongue image changes and GC, and the value of tongue image changes in the diagnosis and screening of GC.
[0004] Artificial intelligence (AI) can be used for the screening, diagnosis, and treatment of various diseases. Scholars such as Cheung CY et al. developed a deep learning system (refer to the reference) to measure the caliber of retinal blood vessels, evaluate the risk of cardiovascular diseases, and effectively predict the risk of cardiovascular diseases. Scholars such as Takenaka K et al. developed a deep neural network (refer to the reference) for evaluating endoscopic images of patients with ulcerative colitis. The network can recognize patients with endoscopic remission and histological remission with an accuracy of 90.1%, and the positive rate is 92.9%.
[0005] Patent CN110251084A of Fuzhou Data Technology Research Institute Co., Ltd. solves the problems of real-time detection, shooting, storage, and uploading of the tongue body in the tongue image collection process, and provides an artificial intelligence-based tongue image detection and recognition method for recognizing tongue image tongue color, tongue shape, coating quality, and coating color. Its solution is mainly related to the collection and recognition technology of tongue images. Among them, tongue image recognition focuses on extracting characteristics such as the color, texture, coating area, or thickness of the coating of the tongue image. However, these operations do not establish a correspondence between tongue image information and certain special stomach diseases, such as gastric cancer.
[0006] Patent CN111710394A of Shenyang Zhilang Technology Co., Ltd. proposes an artificial intelligence-assisted early gastric cancer screening system, which solves the problem of a large workload for determining gastric cancer positivity by automatically analyzing gastric camera slice images instead of manually. However, such a policy based on gastric camera image analysis first requires obtaining gastric camera images collected by a large number of specialized devices for model learning, and still needs to make decisions based on the gastric camera images of each tester during the test stage. There are still defects such as high time consumption, high material cost, and high tester standards for obtaining gastric camera images, making it difficult to achieve a comprehensive screening across the country.
[0007] Jiangsu Tianrui Precision Medical Technology Co., Ltd. CN112133427A provides an artificial intelligence-based gastric cancer auxiliary diagnosis system including a diagnosis selection module, a data collection module, a preprocessing module, a diagnosis module and a display output module, and the system can give personalized diagnosis results based on the collected data of the examinee. The data relied on by the diagnosis of the diagnosis system includes the basic information, living diet, infection history, disease history, family history, clinical symptoms and examination items of the examinee. Among them, the data such as clinical symptoms and examination items are relatively difficult to collect, and the information such as basic information, living diet, infection history, disease history, family history alone does not affect the early screening diagnosis effect.
[0008] References: Cheung CY, Xu D, Cheng CY, et al. A deep-learning system for the assessment of cardiovascular disease risk via the measurement of retinal-vessel calibre. Nature biomedical engineering 2021;5(6):498-508. doi:10.1038 / s41551-020-00626-4 [published Online First:2020 / 10 / 14]; Takenaka K, Ohtsuka K, Fujii T, et al. Development and Validation of a Deep Neural Network for Accurate Evaluation of Endoscopic Images From Patients With Ulcerative Colitis. Gastroenterology 2020;158(8):2150-57. doi:10.1053 / j.gastro.2020.02.012[published Online First:2020 / 02 / 16].
[0009] The present invention aims to solve these and other needs to be solved in the art.
Summary of the Invention
Problems to be Solved by the Invention
[0010] In order to solve at least one technical problem mentioned in the above background art, an object of the present invention is to provide a tumor prediction system based on tongue image, and the system is intended to apply AI deep learning and make a diagnostic prediction for tumors based on tongue images. The tumor prediction system is easy to operate, low in cost, painless, non-invasive, and through a large number of test cases, it has been demonstrated that the prediction system is a forward-looking, economical, non-invasive, and effective screening system for tumors.
Means for Solving the Problems
[0011] Currently, in the application of artificial intelligence to tongue in traditional Chinese medicine image diagnosis, the focus is mainly on the standardization of tongue feature extraction to eliminate differences due to artificial interpretation. For the first time, AI deep learning is applied to construct a GC diagnosis model based on tongue image, evaluate its value in GC diagnosis, and provide a scientific basis for the tongue image diagnosis theory of traditional Chinese medicine doctors.
[0012] A tumor prediction system based on tongue image of the present invention, wherein the system A tongue image acquisition module configured to acquire a tongue image of a test sample, and a data processing module configured to acquire the probability that the test sample belongs to positive by the following operations, A data processing module that predicts the probability that a test sample belongs to positive based on discriminative features on the tongue image obtained by automatic learning.
[0013] In a specific embodiment, the tumor is at least one of gastric cancer, breast cancer, colorectal cancer, esophageal cancer, hepatobiliary pancreatic cancer, lung cancer, prostate cancer, thyroid cancer, ovarian cancer, neuroblastoma, trophoblastic tumor or head and neck squamous cell carcinoma.
[0014] In a specific embodiment, the tumor is at least one of gastric cancer, breast cancer, colorectal cancer, esophageal cancer, hepatobiliary pancreatic cancer, and lung cancer.
[0015] In a specific embodiment, the system further includes an output module configured to output a prediction result.
[0016] In a specific embodiment, the output module is configured to output a tongue image and a prediction result.
[0017] In a specific embodiment, the output module outputs in at least one mode of electronic display, voice broadcast, printing, and network transmission.
[0018] In a specific embodiment, the discriminative feature is derived between the positive category and the negative category on the tongue image. By sufficiently comparing, analyzing, and learning the commonalities and differences between and within the positive tongue image and / or negative tongue image, the purpose is to obtain the discriminative feature between the positive category and the negative category. By deeply discriminating the discriminative feature between the positive category and the negative category on the tongue image of the test sample, the probability that the test sample belongs to the positive category can be determined, thereby realizing the tumor prediction of the test sample based on the tongue image. The discriminative feature can be derived from the commonalities and differences between the positive tongue image and the negative tongue image, and can also be derived from the commonalities and differences between the positive category and the negative category on a single tongue image. That is, obtaining the discriminative feature between the positive category and the negative category from the tongue image can be used to predict whether the test sample belongs to the positive category or the negative category.
[0019] The discriminative feature is derived from the positive tongue image and the negative tongue image of the interactive deep learning model with paired inputs.
[0020] In a specific embodiment, the data processing module is specifically configured to predict the probability that the test sample belongs to the positive category by the following operations: The positive tongue image and the negative tongue image simultaneously input into the interactive deep learning model are compared sufficiently, the commonalities and differences between the positive category and the negative category on the tongue image are automatically learned, and the probability that the test sample belongs to the positive category is predicted based on the discriminative features between the positive category and the negative category. The solution in this part is to obtain the discriminative features between the positive category and the negative category by sufficiently comparing, analyzing, and learning the commonalities and differences between the positive tongue image and the negative tongue image, and based on the discriminative features, the probability that the tongue image of the test sample input into the model belongs to the positive category can be predicted. Therefore, any model that can obtain the discriminative features between the positive category and the negative category by comparing, analyzing, and learning the commonalities and differences between the positive and negative tongue images can be applied to the solution in this part and is also included in the protection scope of the solution in this part. In particular, the present application is not limited to taking the APINet model as an example for illustration.
[0021] In a specific embodiment, the positive tongue image is collected from a tumor-positive patient.
[0022] In a specific embodiment, the negative tongue image is collected from a tumor-negative patient.
[0023] In a specific embodiment, the interactive deep learning model is an APINet model.
[0024] In a specific embodiment, the data processing module is specifically configured to obtain the probability that the test sample belongs to the positive category by the following operations: 1) Extract and obtain positive features and negative features from the pre-acquired positive tongue image and negative tongue image. 2) Train the model with the positive features and negative features and output the probability that the features belong to each category. 3) Input the tongue image of the test sample into the trained model and output the probability that the test sample belongs to the positive category.
[0025] In one specific embodiment, the step 1) of extracting and obtaining positive features and negative features described above includes: The encoder extracts a feature vector of an image and outputs positive features f1 and negative features f2. f1 and f2 and the combined feature f m input simultaneously to the MLP in the feature selection domain and output two control vectors g1 and g2 correspondingly; g1 activates f1 and f2 respectively and selects the feature f1 + and f2 - g2 activates f1 and f2 respectively to form the selected feature f1 - and f2 + and two positive features f1 + and F1 - and two negative features f2 + and f2 - and obtaining the
[0026] In one specific embodiment, the MLP of the feature selection region is m The control vector g1 is output by learning the commonalities and differences of f and f. m It learns the commonalities and differences between them and outputs the control vector g2.
[0027] In one specific embodiment, the step 2) of training the model with positive and negative features described above specifically involves inputting the positive and negative features into a fully connected layer classifier, which outputs the probability that each of these features belongs to each category.
[0028] In one specific embodiment, in step 2), when outputting the probability that a feature belongs to each category, a cross-entropy loss function is minimized according to the categories of the four features:
number
[0029] f1 + is activated by the control vector g1 corresponding to the positive feature, so it contains positive feature information, f1 - is activated by the control vector g2 corresponding to the negative feature, so it contains negative feature information, f2 + and f2 - Note that the same applies to
[0030] In a specific embodiment, in step 2), when outputting the probability that the feature belongs to each category, considering that the confidence of the model for the output of feature f i + should be higher than that of feature f i - minimize the sorting loss function:
Equation
[0031] In a specific embodiment, step 3) of inputting the tongue image of the test sample into the trained model described above refers to inputting the tongue image of a single test sample.
[0032] In a specific embodiment, step 3) of outputting the probability that the test sample belongs to the category described above finally outputs the probability distribution on each category of the corresponding test sample, and the corresponding category with the maximum probability is taken as the predicted category.
[0033] In a specific embodiment, by applying and training only the circumscribed rectangle part of the tongue surface area in the tongue image, the influence on the model of the image background can be effectively eliminated.
[0034] In a specific embodiment, in the training process, in order to enrich the sample space of the training set, the samples in the training set are randomly flipped with a certain probability, then subgraphs are cut out at random positions on the image, and finally linearly interpolated into images of a fixed size, normalized, and input into the interactive deep learning model.
[0035] High-quality sample data is a prerequisite for obtaining a highly generalized depth model. Therefore, positive and negative tongue image data are obtained in advance from tumor patients and non-tumor groups respectively. In this part of the solution, only by sufficiently comparing a pair of samples (including positive and negative tongue images) can their commonalities and differences be discovered. Simulating the actual scene with the pair of images as input, after the encoder extracts the image feature vectors, it outputs positive and negative features, further combines the combined features, and finally outputs a pair of positive features and a pair of negative features. If input into the fully connected layer classifier, the probabilities of these features belonging to each category can be output, and at the same time, the cross-entropy loss function and the sort loss function are minimized to achieve the purpose of the training model. During the test, if the tongue image of the test sample is input into the system, the probability of belonging to tumor positive can be obtained. By deeply analyzing the differences between positive and negative of the tongue image, the internal relationship between tumors and tongue image information is learned based on deep learning technology, and for problems such as low accuracy of early tumor screening and high cost of diagnosis policies, the probability of tumor positive is automatically judged to screen out multiple tumor patients.
[0036] The discriminative features are derived from a single positive or negative tongue image.
[0037] In a specific embodiment, the discriminative feature is obtained by forming an input sequence after cutting the tongue image into n small blocks for feature extraction, so as to obtain deep features that are advantageous for classification.
[0038] In a specific embodiment, the data processing module is specifically configured to obtain the probability that the test sample belongs to the positive by the following operations: Cut the tongue image of the test sample into small blocks, form an input vector by linear mapping and add a position index, introduce a deep learning model that has completed training to perform feature extraction and feature fusion, output deep features that are advantageous for classification after selection, and obtain the probability of belonging to each category.
[0039] In a specific embodiment, the deep learning model completes training in the following steps: a) After cutting the tongue surface image into n small blocks, the n cut small blocks are sequentially used to form an input sequence, an input sequence with a length of n is formed, an input vector is formed by linear mapping, and position indexes 0, 1, 2,..., n - 1 are added. b) Perform feature extraction and feature fusion with an encoder based on the TransFG model, output deep features that are advantageous for classification after selection, and finally output the probability distribution of the deep features belonging to each category with a softmax classifier.
[0040] In a specific embodiment, step a) of cutting the tongue surface image into n small blocks as described above means cutting the tongue image into n square regions that do not overlap each other.
[0041] In a specific embodiment, in step b), when the encoder performs feature extraction, it includes a total of L + 1 Transformer layers, and each layer internally includes a self-attention mechanism.
[0042] In a specific embodiment, when the encoder performs feature extraction and feature fusion in step b), in order to remove redundant features, before inputting the deep features into the last layer, region selection is performed through a feature selection module including a multi-head attention mechanism. The feature selection module returns the index of the top features with the largest attention weights, and inputs the selected top features into the last Transformer layer for feature fusion.
[0043] In a specific embodiment, the top features are the previous K features, and k is one of 1, 2, 3, ……, 20.
[0044] In a specific embodiment, k = 12.
[0045] In a specific embodiment, when the deep features in step b) output the probability distribution belonging to each category, the cross-entropy loss function is minimized:
Equation
[0046] In a specific embodiment, when the deep features in step b) output the probability distribution belonging to each category, the contrastive loss function is minimized:
Equation
[0047] In the solution of this part, after cutting the tongue image into non-overlapping small regions, forming an input vector by linear mapping after sequencing them in order, inputting it into the TransFG model for feature extraction and feature fusion, generating deep features advantageous for classification, and outputting the probability of belonging to each category by a softmax classifier, the prediction of belonging to the category of the tongue image is completed. Through the automatic learning mode of the deep learning model, the tumor positive probability of the test is automatically predicted and screened. For problems such as the low accuracy of conventional early tumor screening and the high cost of the diagnosis policy, the solution of this part automatically determines the probability of tumor positivity based on the tongue surface image and deep learning technology, screens the multiple tumor group. The solution of this part is simple to operate, low in cost, and high in test accuracy.
[0048] The discriminative features are derived from each pixel of the tongue image.
[0049] In a specific embodiment, the data processing module is specifically configured to obtain the probability that the test sample belongs to positive by the following operations: Input the tongue image of the test sample into the deep learning model that has completed training, output the probability that each pixel belongs to positive, negative, and background respectively, and set the maximum probability category as the predicted category of the pixel. The number of pixels predicted as positive in the test sample / (the number of pixels predicted as positive + the number of pixels predicted as negative) is the probability that the test sample belongs to positive.
[0050] In a specific embodiment, in the process of obtaining the probability that the test sample belongs to positive, if the number of pixels predicted as positive is greater than the number of pixels predicted as negative, the test sample is finally predicted as positive, and conversely, as negative.
[0051] In a specific embodiment, the deep learning model is trained in the following steps: Perform pixel-by-pixel labeling on positive tongue image or negative tongue image. Specifically, perform labeling on positive tongue surface area pixels, negative tongue surface area pixels, and background area pixels respectively. The overall algorithm framework adopts an auto-encoding and decoding structure. The image encoder is used to encode the features of the whole image, and the feature decoder outputs as the probability map of the whole image.
[0052] Calculate the loss value of each pixel according to the true label of each pixel and the predicted probability in the probability map, and update the model parameters until the training is completed.
[0053] In a specific embodiment, the positive tongue surface area pixels are labeled as 2, the negative tongue surface area pixels are labeled as 1, and the background area pixels are labeled as 0.
[0054] In a specific embodiment, in the auto-encoding and decoding structure, the DeeplabV3+ image segmentation network structure and / or the Unet series network structure are adopted. Any auto-encoding and decoding framework that can generate a probability map can be applied to the present invention. Since there are many selectable deep network structures, the DeeplabV3+ model is preferentially selected in the present invention. Other frameworks that generate probability maps by auto-encoding and decoding can also achieve the invention purpose, such as the Unet series network structure commonly used in medical image processing.
[0055] In a specific embodiment, a category judgment module is added after the output layer of the network structure in the auto-encoding and decoding structure to determine the final test result based on the probability map of the whole image.
[0056] In a specific embodiment, the category judgment module probability map adopts the judgment policy shown in the following formula:
Number
[0057] In a specific embodiment, t = 0.5 in the decoding policy formula.
[0058] In a specific embodiment, in the deep learning model training process, the cross-entropy cost function predicted for each pixel is adopted:
Number
[0059] Based on the determined clinical diagnosis results, it is necessary for the optimization of the learning model and prediction accuracy to label the tongue image of different cases collected as tumor positive and negative respectively to obtain sufficient label data. After adopting a per-pixel labeling method for the tongue surface image and using the existing label tool to draw the tongue surface area, tags are assigned to each positive region pixel, negative region pixel, and background region pixel, and after outputting the probability map through the autoencoder-decoder structure, the loss difference between the true label of each pixel and the predicted probability in the probability map is calculated, the model parameters are updated to complete the model training, and by inputting the tongue image of the test sample into the model, the probability of tumor positivity can be automatically determined, thereby screening the multiple tumor group and overcoming the defects such as the high cost of collecting data for the basis of conventional early tumor diagnosis and the difficulty in realizing a wide-range comprehensive survey. The accuracy of the internal test reaches 86.6%, so it has high clinical application value.
[0060] A tumor prediction method based on a tongue image, the method comprising: Obtaining a tongue image of a test sample, Inputting the tongue image of the test sample into the system to obtain the tumor positive probability of the test sample, and
[0061] An application of the tumor prediction system and / or method based on the tongue image described above, wherein the application Includes performing tumor prediction on a test sample by applying the system and / or method.
[0062] Based on the common knowledge in the industry, the above preferred conditions can be combined with each other to obtain specific embodiments.
Advantages of the Invention
[0063] The beneficial effects of the present invention are as follows: The present invention provides a plurality of tumor prediction systems based on tongue images, directly using the tongue images of non-biological samples as the implementation objects, and analyzing and learning the commonalities and differences between positive and negative features in the tongue images, so as to exert an excellent diagnostic and prediction function for various tumors. After analyzing and verifying a large number of actual patient samples, the accuracy of gastric cancer test prediction can reach about 80%, the sensitivity during internal testing can reach 0.741 - 0.826, the accuracy can reach 0.785 - 0.806, the sensitivity during external testing can reach 0.841 - 0.862, and the accuracy can reach 0.709 - 0.734. Both the sensitivity and accuracy of the test are significantly better than the sensitivity and accuracy of the machine learning model applying conventional blood tumor markers. In addition, the tumor prediction system based on tongue images can show excellent diagnostic and prediction value for various malignant tumors including breast cancer, colorectal cancer, esophageal cancer, hepatobiliary and pancreatic cancer, lung cancer, etc., and is clearly superior to the combination of conventional blood tumor markers, providing a forward-looking, economical, non-invasive, and effective screening and diagnostic prediction system and method for tumors.
[0064] The present invention adopts the above technical solution to achieve the above object, makes up for the deficiencies of the prior art, and has a reasonable design and easy operation.
Brief Description of the Drawings
[0065] To enable those skilled in the art to more quickly and clearly understand the above and / or other objects, features, advantages and examples of this application, a part of the accompanying drawings is provided. It should be pointed out that the accompanying drawings, schematic embodiments and their descriptions of the specification constituting this application are for providing a further understanding of this application and do not constitute an undue limitation to this application.
[0066]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Mode for Carrying Out the Invention
[0067] Those skilled in the art can, with reference to the content of this specification, realize appropriate substitutions and / or changes of process parameters. In particular, it is necessary to point out that all similar substitutions and / or changes are obvious to those skilled in the art, and all of these are considered to be included in the present invention. The content described in the present invention has been explained by preferred embodiments, but it is obvious to those concerned that the technology of the present invention can be realized and applied by changing or making appropriate changes and combinations to the content described in this specification without departing from the content, spirit and scope of the present invention.
[0068] It should be noted that the following detailed descriptions are all exemplary and are intended to further explain the present application. Unless otherwise limited, all technical terms and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art.
[0069] It should be noted that the terms used here are for explaining specific embodiments and are not intended to limit the technical solutions of the present application. As used in this specification, the singular form is intended to include the plural form as well unless the context clearly dictates otherwise. Furthermore, when the terms "comprising" and / or "including" are used in this specification, it should also be understood that they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0070] APINet model: The APINet model, that is, the attentive pairwise interaction neural network (APINet) model.
[0071] TransFG model: The TransFG model, that is, the transformer architecture for fine-grained recognition (TransFG) model.
[0072] The present invention will be described in more detail below.
[0073] <Clinical specimen> A national multi-center clinical study was conducted, excluding the influence of differences in regions, diets, and centers. It included 11 centers in 8 cities, namely Hangzhou, Wenzhou, and Shanghai in the east, Fuzhou in the south, Chengdu in the west, Liaoning and Heilongjiang in the north, and Taiyuan in the central region.
[0074] As shown in Figure 1, from January 2020 to October 2021, 1,111 gastric cancer (GC) patients from 8 centers, 1,519 non-gastric cancer (NGC) patients from 3 centers, 169 healthy controls (HCs), 648 superficial gastritis (SGs), and 702 atrophic gastritis (AGs) were recruited. Among the gastric cancer (GC) patients, 865 cases were selected, and 1,287 cases were randomly selected from the non-gastric cancer (NGC) patients for training and verification of the system. Among them, there were 448 cases of early GC (TNM I+II stage), 417 cases of advanced GC (TNM III+IV stage), 141 cases in the healthy control group (HC), 547 cases of superficial gastritis (SG), and 599 cases of atrophic gastritis (AG). Approximately 80% of the cases were used as the training dataset, and approximately 20% of the cases were used as the internal validation dataset. Furthermore, 246 cases of GC and 232 cases of NGC from 3 centers were used as an independent external validation dataset, including 162 cases of early GC, 84 cases of advanced GC, 28 cases of HC, 101 cases of SG, and 103 cases of AG. These gastric cancer (GC) patients were all newly diagnosed gastric cancers, had never received treatment for the disease in the past, and had not undergone surgery, chemotherapy, radiotherapy, targeted therapy, or biological therapy for the disease. All gastric cancer (GC) patients had solitary tumors, that is, patients with two or more malignant tumors were excluded. HCs, SGs, and AGs were confirmed by gastroscopy.
[0075] The tongue image and clinical information of all participants, such as age, gender, height, weight, family history, smoking, drinking, TNM staging, and blood tumor markers, were collected. The pathological staging was based on the 8th edition, No. 23 of the American Joint Committee on Cancer. The tongue image collection time for all GC participants was in the morning of gastric surgery, and the tongue image collection time for NGC participants was in the morning of gastroscopy, with a fasting time exceeding 8 hours, excluding the influence of diet on tongue images. Table 1 shows the general patient information such as age, gender, BMI, smoking, and drinking in the GC group and NGC group, which were also in good agreement in the training, internal validation, and independent external validation datasets. JPEG2025524023000008.jpg195170
[0076] In addition, 104 esophageal cancer (EC) patients, 129 hepatobiliary pancreatic cancer (HBPC) patients, 116 colorectal cancer (CRC) patients, 260 lung cancer (LC) patients, and 154 breast cancer (BC) patients were recruited from the Zhejiang Cancer Hospital. Table 2 shows the clinical information of other cancer participants. Except for BC, it was found that the general information between GC and other cancers, such as age, gender, BMI, smoking, and drinking, was in good agreement. JPEG2025524023000009.jpg111170
[0077] <Statistical analysis> All statistical analyses were performed using SPSS 23.0 software (SPSS Inc., Chicago, IL, USA). The results were expressed as mean ± SD or mean ± SEM. Parametric or non-parametric tests were used depending on whether the data were normally distributed. Count data were analyzed using the chi-square test. P < 0.05 was considered statistically significant.
[0078] <Ethical approval> The research of this application was approved by the centralized ethics committees used by 11 participating centers, including the Research Ethics Committee of Zhejiang Cancer Hospital, the First Affiliated Hospital of Wenzhou Medical University, Liaoning Cancer Hospital, Renji Hospital Affiliated to Shanghai Jiao Tong University, Fujian Cancer Hospital, Tumor Hospital Affiliated to Harbin Medical University, Sichuan Cancer Hospital, Shanxi Cancer Hospital, Tongde Hospital of Zhejiang Province, Zhejiang Provincial Hospital of Traditional Chinese Medicine, and Yuhang People's Hospital.
[0079] <Clinical verification> Example 1 Verification was performed using the APINet model. Specifically, it is a tumor diagnosis system based on tongue image, and the system includes a tongue image acquisition module configured to acquire the tongue image of the test sample, and a data processing module configured to obtain the probability that the test sample belongs to the positive by the following operations, a data processing module that predicts the probability that the test sample belongs to the positive based on the discriminative features on the tongue image obtained by automatic learning.
[0080] Design an interactive deep learning model based on comparison, and by sufficiently comparing a pair of tongue surface images input simultaneously, automatically learn the commonalities and differences between the positive and negative categories on the tongue surface image, and finally predict the probability that the test sample belongs to the positive based on the discriminative features. As shown in Figure 2, the overall algorithm framework is divided into three modules: a feature fusion module, a feature selection module, and a classification module.
[0081] Feature fusion module: Input a pair of tongue image pairs belonging to the positive and negative categories simultaneously. First, the encoder extracts the feature vectors of the images and outputs the positive feature f1 and the negative feature f2.
[0082] Feature selection module: Input f1, f2, and the combined feature f m into the MLP in the feature selection area simultaneously, and corresponding to f1 and f2 respectively, output two control vectors g1 and g2 (f1 + f m → g1, f2 + f m → g2). Activate f1 and f2 with the control vector g1 respectively to form the selected features f1 + and f2 - respectively, and g2 acts on f1 and f2 in the same way to form the selected features f1 - and f2 + respectively. Finally, output two positive features f1 + and f1 - and two negative features f2 + and f2 - respectively.
[0083] Classification module: Input the selected features into a classifier (fully connected layer), and finally output the probabilities that these features belong to each category respectively.
[0084] In the training process, based on the categories to which the four features belong, minimize the cross-entropy cost function:
Equation
[0085] Note that the reliability of the highly generalized model for the output of feature f i + must be higher than that of feature f i - Therefore, the sorting cost function is minimized simultaneously:
Equation
[0086] A total of 905 related patients were tested, among which 427 internal test cases from the same center as the training set and 478 cases of data from different centers were used for external testing, and the test results are shown in Tables 3 and 4 below. JPEG2025524023000012.jpg31170JPEG2025524023000013.jpg31170
[0087] In Table 3, the actual number of negative cases is 162 + 52 = 214 cases, and the actual number of positive cases is 37 + 176 = 213 cases. As a result of the prediction, 162 cases are accurately predicted as negative, and 52 cases are wrongly predicted as positive. Therefore, the prediction accuracy in the internal test is (the number of cases accurately predicted as negative + the number of cases accurately predicted as positive) / the total number of test samples = (162 + 176) / (162 + 52 + 37 + 176) = 79%. As can be seen from Table 4 similarly, the accuracy of the external test can reach 71%. As can be seen from the results of the internal test and the external test, the tumor diagnosis system in this part has relatively good prediction accuracy for gastric cancer.
[0088] Figure 3 is a visualization schematic diagram of the model classification basis. The three test samples in the first row on the left side of the dashed line are positive tongue surface images, the second row is the area mainly used as the basis when the model recognizes based on the tongue surface image, and the right side of the dashed line is the visualization image of the tongue surface recognition basis corresponding to the negative samples. The darker the color in the image of the second row, the more the model pays attention to that area. From the display results, it can be seen that the area on which the model recognition process is based mainly concentrates on the tongue surface and is not affected by the background regardless of the black background.
[0089] Clinical symptoms are hidden, diagnosis and screening rely on gastrointestinal endoscopy, the early diagnosis rate of gastrointestinal tumors is low, the prognosis is poor, which brings a heavy burden to society and economy. In order to improve the early diagnosis rate of digestive system tumors, it is urgently necessary to develop non-invasive and effective screening and diagnosis methods. Artificial intelligence is showing a clear path for the evolving medical system with higher accuracy and computing power, which is playing an increasingly important role in cancer screening and diagnosis. In our research, an observational, prospective, multi-center clinical trial was conducted to evaluate the value of tongue image in the screening and diagnosis of GC and other tumors.
[0090] To further evaluate the value of tongue image as a means for tumor diagnosis and screening, a blood tumor marker with clinical application was compared with the tongue image. As a control, the prediction of tumors was verified using a combination of multiple classical blood tumor markers. The selectable blood tumor markers are at least one selected from alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), cancer antigen 125 (CA125), cancer antigen 15-3 (CA15-3), cancer antigen 199 (CA199), cancer antigen 72-4 (CA72-4), cancer antigen 242 (CA242), cancer antigen 50 (CA50), non-small cell lung cancer-related antigen (CYFRA21-1), small cell lung cancer-related antigen (neuron-specific enolase, NSE), squamous cell carcinoma antigen (SCC), total prostate-specific antigen (TPSA), free prostate-specific antigen (FPSA), alpha-L-fucosidase (AFU), Epstein-Barr virus antibody (EBV-VCA), tumor-associated substance (TSGF), ferritin (Ferritin), beta2-microglobulin (β2-MG), pancreatic embryonic antigen (POA) or gastrin-releasing peptide precursor (PROGRP), particularly at least one selected from CEA, CA242, CA72-4, CA125, CA199, CA50, AFP or Ferritin, and more particularly a combination of the above eight blood tumor markers was selected. The prediction method based on the above blood tumor marker index includes the following steps.
[0091] 1) Data preprocessing: Since there are varying degrees of missing values in the serum indicators of all cases, the training data needs to be complete. Therefore, before model training, it is necessary to first perform data complementation. In this application, the K-nearest neighbor missing value interpolation method is used to perform data complementation. Specifically, the complemented value of the missing serum indicator is the average value of the values of two adjacent neighbors.
[0092] 2) Model training: The present invention employs three types of machine learning classification methods, namely, Support Vector Machine (SVM), Decision Tree (DT), and K-Nearest Neighbor classifier (KNN). Specifically, eight types of blood tumor marker indicators (CEA, CA242, CA72-4, CA125, CA199, CA50, AFP, and Ferritin) of the cases correspond to the characteristics of the samples, and the positive / negative diagnosis of the cases corresponds to the labels of the samples. All the complemented samples are sent to the three types of classifiers for fitting.
[0093] 3) Model evaluation: This application evaluates the model using internal validation and external validation. For internal validation, data of different cases from the same hospital as the training data are adopted. For external validation, data of case from hospitals different from the training data are adopted. Three indicators including sensitivity, specificity, and accuracy are used to predict the model.
[0094] As shown in Table 5 for the clinical information of blood tumor markers of related GC patients, it was found that the concentrations of blood tumor markers such as CEA, CA424, CA724, CA125, CA199, CA50, AFP, and Ferritin in GC patients were significantly increased compared with NGC patients. JPEG2025524023000014.jpg131170
[0095] The training, internal validation, and external validation datasets of the model are consistent with the tongue image model (except when blood indicators are missing). As shown in Table 6, the verification results of the sensitivity, specificity, and accuracy of the GC diagnosis of blood tumor markers based on three machine learning classification methods, and for the ROC and AUC of internal validation and external validation, refer to Figure 4. It can be seen that the range of the AUC value of internal validation is 0.682 - 0.715, and the range of the AUC value of external validation is 0.694 - 0.760. In the SVM algorithm, the specificity of both internal validation and external validation reaches over 90%, indicating that the algorithm can provide valuable information for gastric cancer diagnosis. However, among DT and KNN, the specificity decreases to a certain extent, while the sensitivity and accuracy both show different degrees of improvement, providing comprehensive information for gastric cancer diagnosis. JPEG2025524023000015.jpg42170
[0096] It should be clearly stated that the above comparison scheme of the present application selects 8 serum indicators including CEA, CA242, CA72 - 4, CA125, CA199, CA50, AFP, and Ferritin. Any increase, decrease, or replacement of some serum indicators can predict the positive and negative of tumors, especially gastric cancer. The above comparison scheme adopts three machine learning classifiers, SVM, DT, and KNN, and the corresponding objectives can also be achieved by adopting other machine learning classifier methods such as logistic regression and random forest.
[0097] Compared with the above SVM, DT, and KNN, the APINet model of this embodiment has different degrees of improvement or changes in terms of sensitivity, specificity, and accuracy for GC diagnosis, as shown in Table 7. JPEG2025524023000016.jpg34170
[0098] Table 7 shows the sensitivity, specificity, and accuracy data of the APINet model based on tongue image for GC diagnosis. It is found that the APINet model has significantly higher sensitivity and accuracy for GC diagnosis than the aforementioned SVM, DT, and KNN models based on eight types of blood tumor markers in internal validation and external validation, providing a forward, economic, non-invasive, and effective screening and diagnostic prediction method for tumors.
[0099] The APINet model fully compares the input data through pair-wise interactions and recognizes the contrast cues for classification. Figure 5 shows the ROC (Receiver Operating Characteristic) and AUC (Area Under roc Curve) of the internal validation and external validation of the APINet model. As can be seen from Figure 5, the APINet model in Figure 5 has an ROC curve that is relatively far from the line of (0,0)-(1,1) both in internal validation and external validation compared to the SVM, DT, and KNN models in Figure 4. Its internal validation AUC value reaches 0.875, and its external validation AUC value reaches 0.792, which is higher than the internal validation AUC values (0.682 - 0.715) and external validation AUC values (0.694 - 0.760) of the SVM, DT, and KNN models of eight types of blood tumor markers. It is found that the APINet model is a good predictive model. The diagnostic value of the AI diagnostic model based on tongue image for GC is clearly superior to the model that simply applies the combination of eight-item blood tumor markers.
[0100] Analyze the correlation between the accuracy of the model and clinical information. For the correlation between the accuracy of the specific APINet model and the clinical information of GC patients, refer to Table 8. For the correlation between the accuracy of the APINet model and the clinical information of NGC patients, refer to Table 9. It is found that in the discrimination of NGC, the accuracy of the APINet model is related to smoking, drinking, and blood tumor marker indicators, while in the discrimination of GC, the accuracy of the APINet model is only related to gender. That is, the function of the APINet model to distinguish GC and NGC is less affected by clinical information. JPEG2025524023000017.jpg122170JPEG2025524023000018.jpg71170
[0101] To observe the specificity and effectiveness of the GC diagnosis model APINet based on tongue image, 104 cases of EC, 129 cases of HBPC, 116 cases of CRC, 260 cases of LC and 154 cases of BC patients were selected to evaluate the diagnostic value. As shown in Table 10 for the specificity results of the APINet model for GC and other tumors, the tongue image model APINet is the most useful for GC diagnosis, and it was found that it plays a certain role in the diagnosis of gastrointestinal tumors such as EC, HBPC, CRC, and LC. As shown in Figure 6 for the ROC and AUC of the APINet model for GC and other tumors, the diagnostic effect of the APINet model for GC is the best, and for EC, HBPC, CRC, LC, etc., there are different diagnostic effects, and it was found that the APINet model is positive in the diagnostic prediction for the above-mentioned various tumors. JPEG2025524023000019.jpg25170
[0102] Example 2 Verify with the TransFG model. Specifically, it is a tumor diagnosis system based on tongue images, and the system includes a tongue image acquisition module configured to acquire the tongue image of the test sample, and a data processing module configured to obtain the probability that the test sample belongs to the positive by the following operations, a data processing module that predicts the probability that the test sample belongs to the positive based on the discriminative features on the tongue image obtained by automatic learning.
[0103] Design a deep learning model based on Transformer, divide the input tongue surface image into small blocks without duplication, and sequentially form the divided small blocks into a sequence and input it into a multi-layer neural network. Finally, predict the probability that the test sample belongs to the positive based on the extracted high-discriminative features.
[0104] The overall algorithm structure is shown in Figure 7. The input of the entire model is the tongue surface image. First, the tongue surface image is cut into n small blocks, and then the n cut small blocks are used to form an input sequence in order to form an input sequence with a length of n. The small image blocks are used to form input vectors through linear mapping, and position indexes 0, 1, 2, …, n - 1 are added. The present invention performs feature extraction based on the encoder part of the Transformer model, including a total of L (L = 9) + 1 Transformer layers, and each layer contains a self-attention mechanism. In order to remove redundant features, before inputting the deep features into the last layer, region selection is first performed through a feature selection module. This module includes a multi-head attention mechanism, returns the indexes of the features of the previous k (k = 12) blocks with the largest attention weights, inputs the selected k features into the last Transformer layer for feature fusion, outputs deep features that are advantageous for classification after selection, and finally outputs the probability distribution of each category by a softmax classifier.
[0105] When the deep features output the probability distribution belonging to each category, the cross-entropy loss function is minimized respectively:
Number
Number
[0106] In order to make the features within the class more concentrated and the feature differences between classes larger, the prediction accuracy is improved thereby.
[0107] A total of 905 related patients were tested. Among them, 427 cases of internal tests from the same center as the training set and 478 cases of data from different centers were used for external tests. The test results are shown in Tables 11 and 12 below. Here, the accuracies of the internal test and the external test can reach 81% and 73% respectively. As can be seen from the results of the internal test and the external test, the tumor diagnosis system in this part has relatively good prediction accuracy for gastric cancer. JPEG2025524023000022.jpg31170JPEG2025524023000023.jpg31170
[0108] Figure 8 is a visualization diagram based on which the model classification is grounded. The three test samples in the first row on the left side of the dashed line are positive tongue surface images. The small yellow blocks in the images in the second row are the regions corresponding to the feature indexes returned by the region selection module in the original image. The right side of the dashed line is the negative sample and the region selection result. From the results shown, it was found that the regions on which the model recognition process is grounded mainly concentrate on the region with thick tongue coating in the upper half of the tongue surface, and have low correlation with the black background and the lower half of the tongue surface.
[0109] Referring to the above, in order to further evaluate the value of tongue image as a diagnostic and screening means, the tongue image was compared with a blood tumor marker with clinical application. Specifically, the TransFG model based on the tongue image was compared with the SVM, DT, and KNN models based on the blood tumor marker. As a result, the TransFG model in this embodiment has different degrees of improvement or change in terms of sensitivity, specificity, and accuracy for GC diagnosis, as shown in Table 13. JPEG2025524023000024.jpg28170
[0110] Table 13 shows the sensitivity, specificity, and accuracy data of the TransFG model based on tongue image for GC diagnosis. The TransFG model has significantly higher sensitivity and accuracy than the aforementioned SVM, DT, and KNN models based on eight types of hematological tumor markers (CEA, CA242, CA72-4, CA125, CA199, CA50, AFP, and Ferritin) for GC diagnosis in internal validation and external validation, providing a forward, economical, non-invasive, and effective screening and diagnostic prediction method for tumors.
[0111] The TransFG model is data-driven and automatically selects regions advantageous for classification. Figure 5 shows the ROC and AUC of the internal validation and external validation of the TransFG model. As can be seen from Figure 5, the internal validation AUC of the TransFG model is 0.859, the external validation AUC is 0.815, which are significantly higher than the internal validation AUC values (0.682 - 0.715) and external validation AUC values (0.694 - 0.760) of the SVM, DT, and KNN models based on eight types of hematological tumor markers. It was found that the TransFG model is a prediction model with relatively good performance. The diagnostic value of the AI diagnosis model based on tongue image for GC is clearly superior to the combination of eight-item hematological tumor markers.
[0112] Analyze the correlation between the accuracy of the model and clinical information. For the correlation between the accuracy of the specific TransFG model and the clinical information of GC patients, refer to Table 14. For the correlation between the accuracy of the TransFG model and the clinical information of NGC patients, refer to Table 15. In the discrimination of NGC, the accuracy of the TransFG model is related to general situations such as age, gender, BMI, smoking, and drinking. While in the discrimination of GC, the TransFG model is only related to gender. Therefore, the function of the TransFG model to distinguish GC from NGC is not affected by clinical information. JPEG2025524023000025.jpg137170JPEG2025524023000026.jpg70170
[0113] To observe the specificity and effectiveness of the GC diagnosis model TransFG based on tongue image, the aforementioned EC, HBPC, CRC, LC, and BC patients were selected for the purpose of evaluating the diagnostic value. As shown in Table 16 for the specificity results of the APINet model for GC and other tumors, similar to the APINet model, the TransFG model is also most useful for GC diagnosis and was found to have various effects on the diagnosis of tumors such as EC, HBPC, CRC, and LC. JPEG2025524023000027.jpg34170
[0114] As shown in Figure 9, for the ROC and AUC of the TransFG model for GC and other tumors, the diagnostic effect of the TransFG model for GC is the best, with its AUC = 0.815. The AUC for the diagnosis of tumors such as EC, HBPC, CRC, and LC all exceed 0.5, indicating a certain diagnostic effect. Therefore, it was found that the TransFG model is positive for the diagnostic prediction of various tumors including GC.
[0115] Example 3 Verification was performed using the DeeplabV3+ model. Specifically, it is a tumor diagnosis system based on tongue image, and the system includes a tongue image acquisition module configured to acquire the tongue image of the test sample, and a data processing module configured to obtain the probability that the test sample belongs to the positive category by the following operations, a data processing module that predicts the probability that the test sample belongs to the positive category based on the discriminative features on the tongue image obtained by automatic learning.
[0116] Using computer-aided means, a bottom-up deep learning decision framework was designed, and the determination of whether different test subjects belong to tumor positive or negative based on the pre-acquired tongue images is automatically performed.
[0117] To learn a better deep learning model for this task, it is first necessary to obtain sufficient labeled data. As shown in Figure 10, a per-pixel labeling method is adopted for tongue image. Using existing labeling tools, the tongue surface area is drawn, tags are assigned to each pixel. If the sample is a tumor positive sample, the tongue surface area pixel tag is labeled as 2, if the sample is a tumor negative sample, the tongue surface area pixel is labeled as 1, and all background area pixels are set to 0. It should be made clear that labeling the above-mentioned tongue surface area pixels as 2, 1, and 0 is only exemplary, and all labelings that can distinguish positive sample pixels, negative sample pixels, and background area pixels are acceptable. For example, A, B, C, Jia, Yi, Bing, I, II, III, (1), (2), (3), one, two, three, etc.
[0118] Based on the above labeling method of tongue images, the present invention designs a bottom-up deep learning model for tongue images, automatically learns the tongue surface features of the tongue image, and finally outputs the tumor positive probability of the corresponding sample based on the tongue image information. The overall algorithm framework adopts an autoencoding-decoding structure. As shown in Figure 11, the Encoder is an image encoder and is used to encode the features of the entire image. The Decoder is a feature decoder. The decoder outputs as the probability map of the entire image, and each pixel represents the probability of being classified into the specified category. The total number of categories for this task is the number of layers of the model output probability map. Specifically, among all autoencoding-decoding structures, we adopt the DeeplabV3+ image segmentation network structure. To effectively determine the positive probability of the test sample, we add a category judgment module after the output layer of the DeeplabV3+ network, that is, determine the final test result based on the probability map of the entire image.
[0119] Let the pixel set of the input image be M={i|i = 1, 2, ……, m}, the category set be C={c|c = 0, 1, 2}, where 0, 1, 2 in the category set represent background, negative, and positive respectively, and P c(i) represents the probability that a certain pixel belongs to category c. In the category discrimination module, based on the intermediate result probability map, the discrimination policy shown in the following formula is adopted: [Number] Here, the function I is an indicator function, and when the condition is satisfied, the function value is 1, and when it is not, the function value is 0. t is the pixel category determination threshold, where t ∈ [0, 1], and in the present invention, the value is 0.5. Therefore, r represents the ratio of the area predicted as positive by the model in one tongue surface image to the tongue surface area, and the probability that the sample is positive for gastric cancer is predicted using the ratio r as the overall framework.
[0120] In the model training process, the cross-entropy cost function predicted for each pixel is adopted. For the input tongue surface image, the corresponding cost function L is as follows: [Number] Here, the true label category of pixel i is c, and P c (i) indicates the probability that pixel i is predicted to be category c.
[0121] The detailed steps for predicting the positive probability of the tongue image by applying the DeeplabV3+ model are as follows:
[0122] 1) During prediction, the model outputs the probability that each pixel belongs to three categories (background, negative, positive) respectively, corresponding to 0, 1, and 2 in the category discrimination module. By comparing the magnitudes of the probabilities that each pixel belongs to each category, the category with the maximum probability is selected as the predicted category. For example, when the probabilities that one pixel belongs to the background, negative, and positive are 0.3, 0.5, and 0.2 respectively, the pixel is predicted to be the negative category.
[0123] 2) Count the number of pixels predicted as positive and negative categories respectively in the input image. If the number of pixels predicted as positive is more than the number of pixels predicted as negative, the input image is finally predicted as positive, that is, if the number of positive pixels / (the number of positive pixels + the number of negative pixels) is greater than 0.5, the input image is predicted as positive. The number of positive pixels / (the number of positive pixels + the number of negative pixels) is the probability that the input image belongs to the positive category.
[0124] To eliminate the influence of the image background on the experiment, we only apply the circumscribed rectangle part of the tongue surface area in the image for training and testing. During the training process, to enrich the sample space of the training set, we randomly flip the samples in the training set, blur the flipped tongue image at a specified ratio, and the ratio takes random values in the range of 0 to 0.5 within each training cycle. We collected a total of 678 tongue images, of which 544 were used for training and 134 were used for testing. As shown in Table 17, among the test results, 10 positive cases were misjudged as negative, 8 negative cases were misjudged as positive, and the prediction accuracy of the model was 87%. It was found that both negative samples and positive samples had high prediction accuracy. JPEG2025524023000030.jpg31170
[0125] Figure 12 shows the test results after model visualization. The four test samples were respectively collected from the tongue surface images of negative, positive, positive and negative cases. However, our final tongue surface image category determination result is the same as the direct display result. The first column is the input tongue surface image to be predicted, and the second column is the prediction result of each pixel of the input image. Here, the pixels in the green area (G area indicated by the arrow in the figure) are predicted to be negative, the pixels corresponding to the yellow area (Y area indicated by the arrow in the figure) are predicted to be positive, and the purple area (P area indicated by the arrow in the figure) is the background area of the model prediction. Therefore, both the negative and positive determination areas are in the tongue surface area, and the prediction of the image category is not affected by the background area. The third column is the area corresponding to the predicted result in the original figure. By setting the above formula, the ratio of the yellow area in the whole tongue surface is regarded as the probability that the model predicts the sample to be positive for gastric cancer. When the probability is greater than 0.5, the test sample is positive for gastric cancer. Here, the tongue surface area predicted by the model is considered to be the sum of the positive area and the negative area.
[0126] Figure 13 shows the semantic segmentation results of some samples of the DeeplabV3+ model based on tongue image. The three test samples in the first row on the left side of the dashed line are positive tongue images. The whole tongue surface area in the second row image is the area corresponding to the feature index returned in the original image. It can be seen that the pixels marked in yellow are predicted to be positive. Similarly, the three test samples in the first row on the right side of the dashed line are negative tongue images. The whole tongue surface area in the second row image is the area corresponding to the feature index returned in the original image. It can be seen that the pixels marked in green are predicted to be negative. Therefore, both the positive and negative judgment areas are in the tongue surface area, and the prediction of the image category is not affected by the background area.
[0127] Referring to the above, in order to further evaluate the value of tongue image as a means for diagnosing and screening the test sample tumor, a blood tumor marker with clinical application was compared with the tongue image. Specifically, the DeeplabV3+ model based on the tongue image was compared with the SVM, DT, and KNN models based on the blood tumor marker. As a result, the DeeplabV3+ in this example has different degrees of improvement or changes in terms of sensitivity, specificity, and accuracy for GC diagnosis, as shown in Table 18. JPEG2025524023000031.jpg34170
[0128] Table 18 shows that the DeeplabV3+ model based on the tongue image has excellent sensitivity and accuracy for GC diagnosis, and is superior to the sensitivity (0.283 - 0.566, 0.362 - 0.539) and accuracy (0.603 - 0.622, 0.645 - 0.662) of the SVM, DT, and KNN models based on eight blood tumor markers (CEA, CA242, CA72-4, CA125, CA199, CA50, AFP, and Ferritin) in both internal verification and external verification, enriching the forward, economic, non-invasive, and effective tumor screening and diagnostic prediction methods.
[0129] Figure 5 shows the ROC and AUC of the internal verification and external verification of the DeeplabV3+ model. As can be seen from Figure 5, the internal verification AUC of the DeeplabV3+ model is 0.836, and the external verification AUC is 0.801, which is significantly higher than the internal verification AUC values (0.682 - 0.715) and external verification AUC values (0.694 - 0.760) of the SVM, DT, and KNN models based on eight blood tumor markers. It was found that the DeeplabV3+ model is a prediction model with relatively good performance. The diagnostic value of the AI diagnosis model based on the tongue image for GC is significantly superior to the combination of eight blood tumor markers.
[0130] Analyze the correlation between the accuracy of the model and clinical information. For the correlation between the accuracy of the specific DeeplabV3+ model and the clinical information of GC patients, refer to Table 19. For the correlation between the accuracy of the DeeplabV3+ model and the clinical information of NGC patients, refer to Table 20. In the discrimination of NGC, the accuracy of the DeepLabV3+ model is only related to gender. On the other hand, in the discrimination of GC, it was found that the DeeplabV3+ model is related to BMI and tumor location, but not related to other clinical information. Therefore, the function of the DeeplabV3+ model to distinguish GC and NGC is less affected by clinical information. JPEG2025524023000032.jpg135170JPEG2025524023000033.jpg70170
[0131] In order to observe the specificity and effectiveness of the GC diagnosis model DeeplabV3+ based on tongue image, the above-mentioned EC, HBPC, CRC, LC and BC patients are selected for the purpose of evaluating the diagnostic value. As shown in Table 21 for the specificity results of the DeeplabV3+ model for GC and other tumors, similar to the APINet model and the TransFG model, the DeeplabV3+ model is also most useful for GC diagnosis, and it was found that its effectiveness for tumor diagnoses such as EC, HBPC, CRC, LC is different. JPEG2025524023000034.jpg24170
[0132] As shown in Figure 14, the ROC and AUC of the DeeplabV3+ model for GC and other tumors show that the effect of the DeeplabV3+ model for GC diagnosis is the best, with its AUC = 0.801. The AUC for tumor diagnoses such as EC, HBPC, CRC, LC all exceed 0.5 and show a certain diagnostic effect. Therefore, the DeeplabV3+ model is positive for the diagnostic prediction of various tumors including GC.
[0133] Upon comprehensively analyzing the APINet model, TransFG model, and DeeplabV3+ model according to the foregoing Examples 1 to 3, FIG. 15 shows the external verification probability distribution (gastric cancer) of the three tongue image models. Most cases are distributed on both sides. That is, the three judgments for positive and negative cases are relatively certain, and it can be seen that there are relatively few cases with ambiguous diagnoses between 0.41 and 0.60, indicating that the diagnostic prediction results of the model for tumors are reliable. FIG. 16 shows representative tongue image (gastric cancer, intersection of the three models) with different probabilities. It can be seen that it is difficult to distinguish the positive probability intuitively from the tongue image without the intervention of the automatic learning model. Therefore, the present application provides a tumor diagnosis method based on tongue images that can exert excellent diagnostic prediction value for various tumors including gastric cancer, and provides a scientific basis for the tongue image diagnosis theory of traditional Chinese medicine.
[0134] The prior art in the above embodiments is prior art known to those skilled in the art, so detailed description is omitted here.
[0135] The specific embodiments described in this specification are merely illustrative of the spirit of the present invention. Those skilled in the art can make various changes or supplements to the described specific embodiments, or can substitute them in a similar way, but will not deviate from the spirit of the present invention or exceed the scope defined in the appended claims.
[0136] Although the present invention has been described in detail and several specific embodiments have been cited, it is obvious to those skilled in the art that various changes or modifications are possible without departing from the spirit and scope of the present invention.
[0137] What is described above is only a preferred embodiment of the present application and is not used to limit the present application. For those skilled in the art, various changes and variations are possible for the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should all be included within the protection scope of the present application.
[0138] All matters not described in the present invention are known technologies.
Claims
1. A tongue image collection module configured to acquire a tongue image of a test sample, A data processing module configured to obtain the probability that the test sample belongs to the positive by the following operations, A data processing module that predicts the probability that the test sample belongs to the positive based on the discriminative features on the tongue image obtained by automatic learning, and a tumor prediction system based on the tongue image, characterized in that it includes the above.
2. The discriminative feature is derived from the positive tongue image and the negative tongue image of the interactive deep learning model of paired input, and the system according to claim 1 is characterized in that.
3. Specifically, the data processing module is configured to obtain the probability that the test sample belongs to the positive by the following operations: 1) Extract and obtain positive features and negative features from the previously obtained positive tongue image and negative tongue image, 2) Train the model with positive features and negative features, and output the probability that the features belong to each category, 3) Input the tongue image of the test sample into the trained model, and output the probability that the test sample belongs to the positive, and the system according to claim 2 is characterized in that.
4. Step 1) of extracting and obtaining the above-mentioned positive features and negative features is The encoder extracts the feature vector of the image and outputs the positive feature f 1 and the negative feature f 2 and f 1 and f 2 and the combined feature f m are simultaneously input into the MLP of the feature selection region, and correspondingly two control vectors g 1 and g 2 are output, g 1 activates f 1 and f 2 respectively, and after activation and selection, the features f 1 + and f 2 - are formed. g 2 activates f 1 and f 2 respectively, and after activation and selection, the features f 1 - and f 2 + are formed, obtaining two positive features f 1 + and f 1 - and two negative features f 2 + and f 2 - The system according to claim 3, characterized by including the above.
5. The discriminative feature is derived from a single positive tongue image or negative tongue image, and the system according to claim 1 is characterized in that.
6. Specifically, the data processing module is configured to obtain the probability that the test sample belongs to the positive by the following operations: Cut the tongue image of the test sample into small blocks, form an input vector by linear mapping and add a position index, introduce a trained deep learning model to perform feature extraction and feature fusion, output deep features advantageous for classification after selection, and obtain the probability of belonging to each category, and the system according to claim 5 is characterized in that.
7. The discriminative feature is derived from each pixel of the tongue image, and the system according to claim 1 is characterized in that.
8. Specifically, the data processing module is configured to obtain the probability that the test sample belongs to the positive by the following operations: Input the tongue image of the test sample into a trained deep learning model, output the probability that each pixel belongs to positive, negative, and background respectively, and use the maximum probability category as the predicted category of the pixel. The system according to claim 7, wherein the number of pixels predicted as positive in the test sample / (the number of pixels predicted as positive + the number of pixels predicted as negative) is the probability that the test sample belongs to the positive category.
9. Obtaining a tongue image of the test sample, Inputting the tongue image of the test sample into the system according to any one of claims 1 to 8 to obtain the tumor positive probability of the test sample, characterized by a tumor prediction method based on the tongue image.
10. Application of the system according to any one of claims 1 to 8 and / or the method according to claim 9, including performing tumor prediction on a test sample by applying the system and / or method.
Citation Information
Patent Citations
Systems and methods for detecting digestive disorders
JP2023543255A
System and method for detecting gastrointestinal disorders
WO2022074644A1
Cited By
Tumor prediction system, method and application based on tongue coating microorganisms
JP2025524022A