Oral cancer diagnosis model construction method, electronic equipment, program product and system
Through the dual-branch network structure of the OCMS-Net model and the improved DeepLabv3+ network, the problem of inaccurate multi-category identification in oral cancer histopathological image segmentation is solved, and more efficient and accurate tissue segmentation and diagnosis is achieved, and reliable diagnostic auxiliary tools are provided.
Patent Information
- Application Number
- CN202510338472.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art has the problem of inaccurate identification of various tissue types in the segmentation of oral cancer histopathological image of oral cancer, especially the effect of segmentation of tissue categories with similar morphology is not ideal, and the lack of large public data sets and sufficient tissue category annotations, which limits the development of the segmentation field.
The OCMS-Net model adopts a dual-branch network structure, including the primary branch and the secondary branch. The primary branch is used to identify semantic information, the secondary branch is used to extract nuclear information, and combined with the SCConv attention mechanism and the CellViT module, multi-category organization segmentation is performed through the improved DeepLabv3+ network, and trained using the ORA_Nine and ORA_TCGA datasets.
It improves the segmentation accuracy of oral cancer pathological images, can more accurately identify and segment different types of tissues, generate more accurate segmentation masks, reduces the work burden of pathologists, and improves diagnostic efficiency and accuracy.
Smart Images

Figure CN120356683A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical diagnosis, and particularly relates to a method for constructing an oral cancer diagnosis model, an electronic device, a program product and a system. Background Art
[0002] Oral cancer refers to the general term for malignant tumors that occur in the oral cavity, mainly including squamous cell carcinoma, that is, the mucosa undergoes variation [1] 。It is one of the more common malignant tumors in the head and neck [2] 。Oral cancer is the seventh most common malignant tumor in the world [3] 。Globally, the incidence rate of oral cancer remains high, and its incidence rate even shows a continuous growth trend. According to the data of the World Health Organization, about 300,000 people die from oral cancer every year in the world, and about 80% of them occur in developing countries [4] 。Currently, it is the sixth most common cancer. Therefore, the urgency of preventing and treating oral cancer is extremely urgent.
[0003] The pathogenesis of oral cancer is a complex and variable biological process. It is deeply affected by a variety of internal biological factors and external environmental factors. There are many risk factors. Ethanol, smoking, and bad eating habits can directly damage oral cells and promote abnormal cell proliferation. Infectious HPV (human papillomavirus) can directly damage oral cells and promote abnormal cell proliferation. Genetic susceptibility is also an important factor. At the same time, changes in the oral microbial community are also considered a potential factor in the development of oral cancer. In addition, there are some other factors. If these factors interact, it is easy to cause DNA damage and gene mutations in oral cells. Among them, oral squamous cell carcinoma is the most common type of oral cancer, and its main pathogenic factors include smoking and drinking. If both smoking and drinking are excessive, the incidence rate of women is much higher than that of men. In addition, chronic stimuli such as chewing tobacco or betel nuts may also lead to the appearance of squamous cell carcinoma [5] 。
[0004] In the early stage, oral cancer may manifest as oral mucosal ulcers or masses. After the ulcer enlarges, central necrosis forms a depression, the edge bulges and turns outwards, and bleeding and infection may occur. The enlargement of the mass may cause soft tissue or bone destruction, resulting in pain, difficulty in opening the mouth, etc. When the cancer affects the surrounding nerves, it may cause facial numbness, limited tongue movement, etc. [6] 。However, the success of treating oral cancer lies in whether early detection and accurate early diagnosis can be achieved. Emphasizing the importance of early detection and diagnosis is the key to grasping the golden period of treatment [7] 。
[0005] In the diagnostic process of oral cancer, it is crucial to comprehensively utilize various examination methods. Physical examination, as a preliminary screening, can visually identify abnormal changes in the oral cavity; biopsy is the gold standard for diagnosis. It precisely excises a small piece of tissue from the suspected lesion area and sends it to the laboratory for detailed microscopic analysis to confirm the presence and type of cancer cells. Although this process may be somewhat painful for patients and the results take time to obtain, its accuracy is irreplaceable and serves as the basis for formulating subsequent treatment plans. In addition, imaging tests such as CT and MRI can generally show the state of the tumor in the human body and provide important references for surgical planning; endoscopy can directly observe areas in the oral cavity and pharynx that are difficult to reach, improving the comprehensiveness of diagnosis. These examination methods each have their own advantages, complement each other, and jointly construct a complete system for oral cancer diagnosis.
[0006] In addition, pathological diagnosis is an important basis for the diagnosis of oral and oropharyngeal mucosal squamous cell carcinoma and for clinicians to formulate treatment plans. [8] The standardized pathological diagnosis of oral and oropharyngeal mucosal squamous cell carcinoma should not only provide accurate histopathological diagnosis for the clinic, but also provide pathological elements related to prognosis assessment, treatment strategy selection, etc. [9] Among various diagnostic bases, histopathological images can accurately diagnose the type of cancer in patients. Pathological diagnosis is a complex and time-consuming task that requires pathologists to observe, analyze, and judge a large number of sections.
[0007] During the diagnostic process, doctors often use some special staining methods to enhance the visualization effect of tissue sections. In this way, clinicians systematically observe oral mucosal tissue sections through microscopic morphological analysis techniques to accurately identify the pathological features at the histological level and the morphological changes of the substantial structure, and thereby screen for initial pathological evidence of malignant tumor lesions.
[0008] Quantitatively evaluating the degree of cellular atypia and organizational structure disorder based on high-resolution microscopic imaging technology can significantly improve the pathological diagnosis accuracy of common malignant tumors such as oral squamous cell carcinoma. These features include, but are not limited to, abnormal cell morphology, enlarged cell nuclei, disordered cell arrangement, etc., which are all important bases for cancer diagnosis.
[10] The histopathological images of oral cancer contain intricate organizational structure information and variable cancer cell aggregation patterns, posing a major challenge in the field of diagnostic technology. Their accurate identification is crucial for the treatment plan and prognosis assessment of patients.
[0009] In clinical practice, the accurate judgment of images relies on doctors' professional experience and vision, and the limitations brought about by this are imaginable. This requires doctors to participate in strict diagnostic training and then continuously enrich their experience in practice to ensure that the diagnostic conclusions are relatively reliable. However, it takes a large amount of resources in the hospital for each doctor to progress from entry-level to an experienced doctor with excellent diagnostic capabilities. Moreover, due to differences in the professional backgrounds, training experiences, and clinical practice accumulations of different histopathologists, there is a certain subjective judgment space in the diagnostic process. In addition, complex factors at the individual level cannot be ignored, including subjective state changes such as physical and mental fatigue, various negative emotions, and psychological fluctuations induced by life events. These factors may potentially risk the objectivity and accuracy of pathological diagnosis by affecting the cognitive function and judgment stability of doctors. Summary of the Invention
[0010] One embodiment of the present disclosure is a method for constructing an oral cancer diagnosis model. The diagnosis model is based on DeepLabv3+, adds an auxiliary branch to provide position information for the diagnosis model, and adds SCConv attention to achieve the segmentation processing of pathological images of multiple tissue categories of oral cancer.
[0011] The diagnosis model is based on OCMS-Net and includes a main branch A and an auxiliary branch B. The input of the main branch A is an oral cancer pathological image, and the output is a segmentation mask, where different colors in the segmentation mask represent different tissue types;
[0012] The input of the auxiliary branch B is an oral cancer pathological image, and the output is a binary segmentation map of cell nuclei and a nuclear pixel distance map.
[0013] One embodiment of the present disclosure is an oral cancer pathological diagnosis system, including a tissue segmentation module and a diagnosis module, where
[0014] The tissue segmentation module has:
[0015] Picture selection function - allowing users to select or upload oral cancer pathological images that need to be analyzed,
[0016] Prediction function - performing tissue segmentation on the selected oral cancer pathological image to identify and distinguish different tissue types;
[0017] The diagnosis module has:
[0018] Picture selection function - the function of selecting from the oral cancer pathological images output by the tissue segmentation module,
[0019] Tumor stroma ratio function - analyzing the ratio of tumor tissue to surrounding normal tissue,
[0020] Vascular and perineural invasion function - Detect whether the tumor has invaded the surrounding blood vessels and nerves. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary but non-limiting manner, wherein:
[0022] Figure 1 Overall framework diagram of an oral cancer pathological diagnosis model according to one embodiment of the present invention.
[0023] Figure 2 Schematic diagram of a picture of the segmentation result of oral cancer according to one embodiment of the present invention.
[0024] Figure 3 Schematic diagram of the composition of an oral cancer pathological diagnosis system according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] With the progress of computer vision and pattern recognition technologies, artificial intelligence has proven to be helpful in making clinical workflows more efficient
[11] . The emergence of quantitative analysis methods in digital pathology
[12] indicates that computer-aided diagnosis (CAD) based on digital whole-slide images (WSI) is becoming increasingly widespread in the field of medical image analysis, and computer-aided automatic image analysis of oral cancer has become feasible. It can effectively avoid cognitive biases and experience dependencies in manual interpretation. In this process, machine vision algorithms
[13] are powerful assistants for doctors to analyze medical images, capable of quickly classifying oral cancer tissue pathological images, simplifying the cumbersome operation process, and significantly improving the accuracy of diagnostic results.
[0026] In the field of medical images, convolutional neural networks (CNN)
[14] are the most commonly used learning models in deep learning. With its unique structural design and algorithmic advantages, the CNN model can automatically extract multi-scale and interesting features from complex medical images, which are crucial for subsequent tasks such as image analysis and disease diagnosis. When dealing with complex medical image recognition tasks such as tumor detection and lesion identification, the CNN model can demonstrate excellent classification performance, with its accuracy and efficiency far exceeding traditional methods. This benefits from the delicate design of its deep neural network structure combined with various efficient training strategies, such as data augmentation techniques
[15] and transfer learning methods
[16] etc. In the research on the segmentation of oral cancer tissue pathological images based on deep learning, relevant research in recent years [17-22] has shown the great advantages and potential of CNN in the segmentation of oral cancer tissue pathological images.
[0027] Currently, in the task of cancer tissue pathological image segmentation, there is still very little research on oral cancer. The main reason is that there are few large publicly available datasets, and the tissue categories labeled in the datasets are too few, which further restricts the development of the field of oral cancer tissue pathological image segmentation. ORCA
[37] dataset is a relatively large oral cancer tissue pathological image dataset created in 2020, aiming to help more researchers study oral cancer tissue pathological images with the help of deep learning technology. The original WSI samples belong to The Cancer Genome Atlas (TCGA), and all samples are divided into tumor and non-tumor regions. This dataset can only be used for binary classification segmentation tasks. Martino
[37] made the first attempt to apply well-known deep learning-based segmentation methods on publicly available TCGA images. Four different network architectures were selected from the most commonly used semantic segmentation network architectures and trained on the ORCA dataset, namely: SegNet
[38] , UNet
[39] , UNet with VGG16 encoder, UNet with ResNet50 encoder. Experiments show that deeper networks, such as UNet modified with ResNet50 as the encoder, perform better than the original UNet (with a shallower encoder). The H&E-stained WSI dataset was constructed using tissue specimens collected from human patients diagnosed with OSCC. OCDC dataset
[40] is an oral cancer tissue pathological image dataset created in 2021. This H&E-stained WSI dataset was constructed using tissue specimens collected from human patients diagnosed with OSCC. All OSCC cases were from the Department of Oral and Medical Pathology Services between 2006 and 2013. This dataset can only be used for binary classification segmentation tasks. Dos et al
[40] were the first to use the FCN method to study OSCC detection at the WSI level. This method uses a color-based tissue detector in the preprocessing step to remove background and scanning artifacts, and uses a UNet-based FCN architecture to locate and segment the tumor regions of 640×640 image patches. The OCDC dataset was used to train and validate the model. Albishri
[41] et al. pioneered a new model, OCU-Net, which combines advanced deep learning modules such as the Channel and Spatial Attention Fusion (CSAF) module, a novel and innovative feature that emphasizes important channels and spatial regions in H&E images while exploring context information. Additionally, OCU-Net integrates other innovative components such as the Squeeze-and-Excite (SE) attention module, the Atrous Spatial Pyramid Pooling (ASPP) module, residual blocks, and multi-scale fusion. These modules demonstrated excellent performance in oral cancer segmentation using the ORCA dataset and the OCDC dataset. Although the above was attempted on public datasets, it only satisfied the binary classification segmentation task.
[0028] Currently, in the research related to the segmentation of oral cancer histopathological images, most are private datasets except for a small number of publicly available datasets. Li Lianbing
[42] et al. used a convolutional neural network with a DenseNet architecture to classify images as normal or cancerous. They trained on high-resolution image patches using the idea of image patching and adopted transfer learning and data augmentation methods to reduce overfitting. After classification, they used a UNet++ segmentation network with a DenseNet network as the encoding structure to localize cancerous regions in the images judged to be cancerous, obtaining relatively ideal segmentation results. In Pennisi et al.
[43] proposed an improved UNet architecture called multi-encoder UNet for segmenting OSCC in whole-slide image samples. The method divides the input image into tiles, each tile is encoded by a separate encoder and merged using convolutional layers. The resulting merged layer is decoded to obtain the segmented image. Multi-encoder UNet can better extract effective features to obtain better segmentation results. Wu
[44] et al. developed a UNet-based CNN model to segment the epithelial region of digitized slides and conducted rigorous training, validation, and independent testing using a total of 810 tissue microarrays (TMAs) and WSIs from 787 OC-SCC patients from five different institutions. The effectiveness and applicability of the CNN model in performing epithelial segmentation using digitized H&E-stained diagnostic slides from OC-SCC patients in a multi-center setting were studied. The research showed that the developed DL model could provide consistent segmentation performance on TMAs and WSIs from five different institutions. Liu
[45] et al. found that previous studies have shown that even when performed by trained experts, the scoring task of OED is affected by reader variability and misdiagnosis rates. Therefore, they compared the OED segmentation and classification metrics of two mature CNN architectures for medical imaging, Deep-Labv3+ and UNet++, and identified a convolutional neural network (CNN) model that can identify regions of suspicious OED whole-slide pathology images. Musulin
[46] et al. proposed an automated multi-class grading system for oral tissue pathology images based on a two-stage AI (the first stage) and segmentation of epithelial and mesenchymal tissues (the second stage) to assist clinicians in the diagnosis of oral squamous cell carcinoma. In the first stage, it can potentially improve the objectivity and reproducibility of histopathological examinations and reduce the time required for pathological examinations. The AI-based system segments tumors in the epithelial and mesenchymal regions in the second stage, which can help clinicians discover new informative features. They also proposed a preprocessing method of stationary wavelet transform (SWT) to enhance high-frequency components in the case of multi-class classification and extract low-level features in the case of semantic segmentation. This method allows for more effective prediction and improves the robustness of the entire AI-based system. FRAZ
[47] et al. developed a new model, FABnet, using Deeplabv3+ as the baseline. A feature attention module is used for skip connections from the encoder to the decoder and upsampling in the decoder, enabling the network to focus on more significant features. The results show better performance in segmenting microvessels and nerves.
[0029] Through research on existing technical solutions, it can be found that most studies use the UNet network and Deeplabv3 network as the main architectures for tissue region segmentation. Research has been conducted on improving the encoder extraction ability and enabling the model to focus on important features, and various encoder structures and various attention mechanisms have been tried. Although certain progress has been made, they are all for the study of one or two tissues such as epithelium and mesenchyme, microvessels and nerves, and no experiments have been conducted under multiple tissue types, and there are still problems such as inaccurate recognition of each target tissue.
[0030] According to one or more embodiments, as Figure 3 shown, an oral cancer pathological diagnosis system mainly includes a tissue segmentation module, a diagnosis module, and an instruction manual module.
[0031] The tissue segmentation module has:
[0032] A select picture function - allowing users to select or upload images to be analyzed from the system.
[0033] A prediction function - using advanced image processing and machine learning technologies to perform tissue segmentation on the selected images, identifying and differentiating different tissue types.
[0034] Download function - Users can download the processed images or analysis results to the local for further research or record.
[0035] The diagnostic module has:
[0036] Picture selection function - A function to select from the pictures output by the tissue segmentation module.
[0037] Tumor stroma ratio function - Analyze the ratio of tumor tissue to surrounding normal tissue, which is of great significance for evaluating the malignancy and prognosis of tumors.
[0038] Vascular and perineural invasion function - Detect whether the tumor has invaded the surrounding blood vessels and nerves, which is crucial for formulating surgical plans and evaluating treatment effects.
[0039] Download module - Users can download the diagnostic results to the local for subsequent medical decision-making and patient management.
[0040] Instruction module:
[0041] Provide the operation manual and guidelines of the system to help users understand and master the usage method of the system, including how to upload images, how to use the prediction function, etc.
[0042] Among them, the tissue segmentation module realizes the precise segmentation of oral tissue pictures to obtain the predicted images of each tissue. This module loads the original picture of the oral cancer pathological image and obtains the predicted picture of the corresponding oral cancer tissue pathological image through diagnostic model prediction. The diagnostic module is used to diagnose the pictures that have been precisely segmented to obtain corresponding clinical indicators.
[0043] In the embodiments of the present disclosure, the oral cancer pathological diagnosis system detects and identifies each tissue through a trained oral cancer pathological diagnosis model, and statistically analyzes the tumor stroma ratio, the situation of vascular and perineural invasion, and the proportion score data of each tissue.
[0044] According to one or more embodiments, the method for constructing an oral cancer pathological diagnosis model is to improve on the segmentation model DeepLabv3+ which has been proven to have good segmentation ability through the experiments of the present disclosure, and then continue with the oral cancer tissue segmentation task. The ORA_Nine dataset and the ORA_TCGA dataset are used as the datasets for this experiment.
[0045] In the existing DeepLabv3+ network, the ability to extract key information is insufficient, and often these key information can lead to inaccurate segmentation results. Due to the morphological similarities among different tissue categories. For example, tumors originating from epithelium, especially well-differentiated / moderately differentiated CTRs, have great morphological similarities with normal epithelium, resulting in unsatisfactory segmentation effects. To address the above problems, the embodiments of the present disclosure add two modules on the basis of DeepLabv3+, effectively solving this problem, and propose an improved oral cancer multi-class segmentation network OCMS-Net based on DeepLabv3+ as Figure 1 shown.
[0046] Figure 1 The Chinese explanations of Chinese and English terms are as follows:
[0047] Backbone - The backbone network is the underlying network structure in a deep learning model used to extract features. PixelDecoder - The pixel decoder is used to convert the features extracted by the backbone network into pixel-level prediction results.
[0048] NPD - Nuclear Pixel Classification, used to identify whether each pixel in the image belongs to the cell nucleus.
[0049] HVD - HoVer Branch, used to predict the horizontal and vertical distances of each nuclear pixel to the centroid of its corresponding cell nucleus.
[0050] True labels - The true labels represent the actual tissue category annotations in the image.
[0051] Figure 1 Shown in [Figure] is the schematic diagram of the oral cancer pathological diagnosis model based on deep learning of the present disclosure. This model adopts a dual-branch network structure, including a main branch and an auxiliary branch, to achieve accurate segmentation and diagnosis of oral cancer pathological images. Among them, the main branch is responsible for identifying and classifying different semantic information, and its structure includes:
[0052] 1. Backbone (the backbone network): Used to extract high-level semantic features of the image. Common backbone networks include ResNet, EfficientNet, etc.
[0053] 2. Pixel Decoder (the pixel decoder): Converts the features extracted by the backbone network into pixel-level prediction results for generating the final segmentation mask.
[0054] The auxiliary branch is used to extract the cell nucleus information of the labeled categories to assist tissue segmentation, and its structure includes:
[0055] 1. Backbone: Shares the backbone network with the main branch, which is used to extract image features.
[0056] 2. NPD (Nuclear Pixel Classification Branch): Used to identify whether each pixel in the image belongs to the nucleus, and distinguish nuclear pixels from background pixels through binary classification.
[0057] 3. HVD (Horizontal and Vertical Distance of Nuclear Pixel Branch): Used to predict the horizontal and vertical distances of each nuclear pixel to the centroid of its nucleus, and this distance information is used to distinguish adjacent nuclei.
[0058] Through the dual-branch network structure, the model makes full use of the feature extraction and segmentation capabilities of deep learning models. The main branch is responsible for semantic segmentation of the image to generate tissue-level segmentation results; the auxiliary branch focuses on nuclear-level feature extraction to provide more detailed auxiliary information for the segmentation task. Through the collaborative work of the main branch and the auxiliary branch, the model can more accurately identify and segment different types of tissues, thus providing a reliable diagnostic assistance tool for pathologists.
[0059] Therefore, OCMS-Net is a dual-branch network structure. The main branch is responsible for identifying and classifying different semantic information, and the auxiliary branch is responsible for extracting nuclear information of the labeled categories to assist tissue segmentation. In the encoder part, EfficientNet-B4 with stronger feature extraction ability is adopted to improve the network's ability to obtain global information. In the decoder part, the Atrous Spatial Pyramid Pooling (ASPP) structure can collect global features at multiple scales, which helps the model maintain performance when processing images at different magnifications. The Spatial Reconstruction Unit (SRU) and the Reconstruction Unit (CRU) included in SCConv
[61] can pay more attention to multi-scale information and effective information in the feature space, improving the feature expression ability of the decoder. Compared with the existing DeepLabv3+, OCMS-Net not only has stronger capabilities in feature extraction and expression, but also has good performance in identifying morphologically similar tissues. Next, each module will be introduced one by one.
[0060] (1) ASPP Structure
[0061] The structure of ASPP (Atrous Spatial Pyramid Pooling) usually includes the following parts:
[0062] 1x1 Convolution: Use 1x1 convolution to extract features and then perform dimensionality reduction.
[0063] Three or more dilated convolutions: Apply multiple dilated convolutions in parallel, each with a different dilation rate, which can increase the receptive field without increasing the computational cost.
[0064] Global average pooling: Apply global average pooling, which helps the model capture more discriminative global features and then restores to the size of the original feature map through upsampling.
[0065] Finally, perform a concatenation operation, then fuse the features through 1x1 convolution and reduce the number of channels.
[0066] (2) SCConv structure
[0067] In the field of deep neural network feature optimization, the Spatial-Channel Cooperative Convolution module (SCConv) constructs a dual-path cooperative optimization mechanism. It consists of two units, the Spatial Reconstruction Unit (SRU) and the Channel Reconstruction Unit (CRU), placed in a sequential manner. Specifically, for the intermediate input feature X in the bottleneck residual block, the present disclosure first performs decoupled learning through the Spatial Reconstruction Unit (SRU) operation, simulating the dual visual processing in pathologist diagnosis (local detail observation and global structure analysis), and then uses the Channel Reconstruction Unit (CRU) operation for dynamic resource allocation, adaptively adjusting the computing resources according to the tissue region complexity. According to the characteristics of the SCConv module, it can be seamlessly integrated into any CNN architecture with little impact on the entire network.
[0068] Specifically, the SRU uses a separation and reconstruction operation. It mainly uses the scale factor in the Group Normalization (GN) layer to evaluate the information content of different feature maps. Then a threshold operation is performed before reconstruction to obtain two weighted features. The CRU uses a separation-fusion operation to first split the channels of X into a dual-segment configuration. Feature interaction enhancement is performed, and dual-path information fusion is achieved through a cross-flow gating mechanism. Then end-to-end optimization is achieved through relaxed integer constraints. After the above operations, the embodiment of the present disclosure splits the spatially refined feature X into an upper half and a lower half X. Then these two parts of the feature map are fused to obtain the feature map Y. In short, the CRU is used to further reduce the redundancy of the spatially defined feature map X along the channel dimension using a splitting transformation and fusion strategy. In addition, the CRU is good at extracting robust representative features through a streamlined convolution operation, while managing redundant features through a cost-effective process and feature recycling scheme.
[0069] (3) Auxiliary branch
[0070] In designing auxiliary branches, CellViT outperforms existing methods on the PanNuke dataset and provides better nucleus instance segmentation results. The present disclosure uses the nucleus morphological features obtained by CellVit on the PanNuke dataset to assist the segmentation task. The three branches of CellViT refer to the three different strategies it adopts when processing the nucleus instance segmentation task, which work together to improve the accuracy and efficiency of segmentation. Each branch is responsible for different tasks:
[0071] 1. Nuclear Pixel Classification (NPC): This branch is responsible for identifying whether each pixel in the image belongs to a cell nucleus. It distinguishes nuclear pixels from background pixels by binary classification.
[0072] 2. Nuclear pixel distance branch (HoVer Branch): This branch predicts the horizontal and vertical distances of each nuclear pixel to the centroid of the cell nucleus to which it belongs. This distance information is used to distinguish adjacent cell nuclei, even if they are closely connected in the image.
[0073] 3. Nuclear Type Classification (NTC): If a label of the nucleus type is provided, this branch will classify the nucleus. It can identify different types of nuclei and assign them to corresponding categories. For the nuclear pixel classification branch and the nuclear pixel distance branch, the nuclear pixel classification branch can predict whether each pixel belongs to the nucleus, thereby achieving binary segmentation of the nucleus; the nuclear pixel distance branch is used to separate overlapping or stacked nuclei and provide location information for each category. The dual-branch structure achieves accurate segmentation of the nucleus, ensuring that the network can fully extract position features while retaining fine-grained features at the nuclear level.
[0074] Because there are certain differences in the morphology of the cell nucleus arrangement and size among the labeled categories, the auxiliary branch can assist in the segmentation of pathological tissues based on morphological information to a certain extent. After combining the loss functions of the main branch and the auxiliary branch according to a certain weight, the network parameters can be updated through back propagation to jointly guide the network to learn better features.
[0075] In order to evaluate the impact of each module on the network, the ablation experiment of the tissue segmentation diagnosis model of the present disclosure uses the validation OCMS-Net model selected from the ORA_Nine dataset and the ORA_TCGA dataset. The following is an explanation of the ablation experiment of the model of the present disclosure.
[0076] (1) Experimental environment, experimental parameter settings, data preprocessing, experimental procedures and model evaluation indicators: the same as those described in Chapter 3.
[0077] (2) Real loss function
[0078] To train faster and achieve better network convergence, the present disclosure uses a weighted combination of different loss functions for each network branch as the total loss, as shown in formula (4-1).
[0079] L total = 0.6·L dice + 0.2·L NP + 0.2·L HV (4-1)
[0080] where L dice represents the loss of the main branch, L NP represents the loss of the NP branch, and L HV represents the loss of the HV branch. The loss of each branch consists of the following weighted loss functions, as shown in formulas (4-2), (4-3), and (4-4):
[0081] L dice = L Dice (4-2)
[0082]
[0083] where the formulas are as shown in (4-5), (4-6), and (4-7):
[0084]
[0085]
[0086] The contribution of each loss pair to the total loss (4-1) controlled by the i-th hyperparameter λ i . represents the mean squared error of the horizontal and vertical distance maps, represents the mean squared error of the gradient of the horizontal and vertical distance maps, y ic is the ground-truth, the i-th pixel belonging to class c, N px is the total number of pixels, ε is the smoothing factor, α FT , β FT and γ FTis a hyperparameter of the Focal Tversky loss. Cross-entropy loss (4-5) and Dice loss (4-6) are commonly used in semantic segmentation. To address the challenge of underrepresented instance classes, the Focal Tversky loss (4-7) is used. The Focal Tversky loss places more emphasis on accurately classifying underrepresented instances by assigning higher weights to these samples. This weighting enhances the model's ability to handle class imbalance and focuses its learning on the more challenging regions of the segmentation task.
[0087] To evaluate the impact of each component on the network, the present disclosure conducted ablation experiments on the ORA-Nine and ORA-TCGA datasets as shown in Table 4-1. The present disclosure first used the DeepLabv3+ model with the encoder being EfficientNet-b4 for segmentation. On the ORA-Nine dataset, the average Dice value was 0.726, the Precision value was 0.730, the Accuracy value was 0.865, and the MioU value was 0.718. On the ORA-TCGA dataset, the average Dice value was 0.766, the Precision value was 0.771, the Accuracy value was 0.894, and the MioU value was 0.759. Then, the present disclosure added the SCConv module to the model to improve the efficiency and effectiveness of feature extraction, making the model pay more attention to important features. On the ORA-Nine dataset, the average Dice value increased by 0.02 to reach 0.742, the Precision value was 0.737, the Accuracy value was 0.884, and the MioU value was 0.727. On the ORA-TCGA dataset, the average Dice value increased by 0.058 to reach 0.810, the Precision value was 0.816, the Accuracy value was 0.908, and the MioU value was 0.792. Next, considering the morphological differences in the nuclei among the five categories, the present disclosure added an auxiliary branch for extracting nuclear information. The HV branch and the NP branch were added to DeepLabv3+. On the ORA-Nine dataset, after adding the auxiliary branch, the average Dice value further increased to 0.801, the Precision value was 0.812, the Accuracy value was 0.901, and the MioU value was 0.781. On the ORA-TCGA dataset, after adding the auxiliary branch, the average Dice value further increased to 0.838, the Precision value was 0.843, the Accuracy value was 0.924, and the MioU value was 0.814. Finally, in Experiment 4, all components were integrated. As expected, the framework obtained higher metric values compared to the baseline. On the ORA-Nine dataset, the average Dice value was 0.810, the Precision value was 0.822, the Accuracy value was 0.912, and the MioU value was 0.805. On the ORA-TCGA dataset, the average Dice value was 0.842, the Precision value was 0.853, the Accuracy value was 0.943, and the MioU value was 0.836.
[0088] Meanwhile, as shown in Table 4-2, in the ablation experiment, the present disclosure also recorded the Dice values of each tissue. After adding the SCConv module, on the ORA-Nine dataset, there was no improvement in the ORER class, but the Dice values of the remaining tissue classes all increased. On the ORA-TCGA dataset, the Dice values of all tissue classes increased. After adding the auxiliary branch, on the ORA-Nine dataset, the Dice values of all tissue classes increased to a certain extent. On the ORA-TCGA dataset, the Dice values of all tissue classes also increased to a certain extent. After integrating all components, higher Dice values were obtained for each class.
[0089] Table 4-1 Ablation Experiment
[0090]
[0091] Table 4-2 Ablation Experiment of Each Tissue
[0092]
[0093]
[0094] The OCMS-Net architecture proposed by the present disclosure achieved satisfactory results on the ORA-TCGA dataset and the ORA-Nine dataset. As can be seen from Figure 2 the figure shown, OCMS-Net can more accurately describe each tissue class and generate better segmentation masks through semantic information at different resolutions and the specificities of each tissue class. The generated segmentation map is superior to other architectures in capturing shape information, indicating that the segmentation masks generated by OCMS-Net show more accurate information within the target area than existing models.
[0095] OCMS-Net achieved higher accuracy than other models in almost all metrics. As can be seen from the ablation experiments in Tables 4-1 and 4-2, after adding these modules, the segmentation performance of the algorithm improved to a certain extent. Since SCConv focuses on the expression of more important high-level semantic features in some local multi-scale feature integrations and cannot extract semantic features. After adding SCConv to the ASPP structure, it can reduce some redundant features brought by the ASPP structure while enhancing important features. On two datasets, this brought a certain improvement to the experimental results. This improvement may benefit from the function of the SCConv structure to highlight important features in multi-scale feature integration. The auxiliary branch only uses the pre-trained weights of CellVit on the PanNuke dataset, which greatly enhances the segmentation performance of the network. On two datasets, the average Dice value has a large improvement. This excellent performance may benefit from the function that the auxiliary branch can provide the positions of various types of cell nuclei. According to the different morphological characteristics of cell nuclei in different tissue categories, the network can effectively promote the recognition of various tissue categories of oral cancer lesions.
[0096] In summary, the present disclosure proposes an oral cancer tissue pathological diagnosis model, a model construction method and its system composition. It is used to assist doctors in diagnosing clinical indicators related to oral cancer and improve the work efficiency of oral cancer pathological examinations.
[0097] During the model training process of the present disclosure, for the dataset, for the first time, by annotating the dataset, the dataset is divided into tumor (CTR), epithelium (ORER), SR (stroma), LV (lymphatic vessels) and skeletal muscle and vasculature (MN). The model trained therefrom can more accurately re-diagnose the tumor-stroma ratio (TSR), the worst pattern of invasion-5 (WPOI-5), vascular and perineural invasion, etc.
[0098] For the oral cancer tissue pathological image segmentation task. The present disclosure first selects a basic model with better segmentation performance from three basic models with good performance. Through experiments, DeepLabv3+ is selected as the improved basic model, and further Efficienct-b4 is selected as the encoder of OCMS-Net. The improved model proposed by the present disclosure uses the SCConv attention mechanism and the auxiliary branch module to solve the problems such as unclear recognition of similar morphologies of each tissue in the original DeepLabv3+ network. In order to train faster and achieve better network convergence, the total loss is weighted by using different loss functions for each network branch. The results prove that OCMS-Net has strong segmentation performance.
[0099] Therefore, the technical effects of the present disclosure include:
[0100] 1. The diagnostic model improves the segmentation accuracy of oral cancer pathological images. The average Dice values of the OCMS-Net model on the ORA-TCGA and ORA-Nine datasets reach 0.81 and 0.842 respectively, which are 10% and 9% higher than those of DeepLabv3+. The experimental results show that OCMS-Net has good comprehensive segmentation ability.
[0101] 2. It can more accurately re-diagnose clinical indicators such as tumor stroma ratio (TSR), worst pattern of invasion-5 (WPOI-5), vascular and perineural invasion, etc., providing a key basis for clinical treatment.
[0102] 3. The intelligent pathological diagnosis system for oral cancer, through integrating the segmentation function, can quickly predict each tissue of oral cancer pathological images, realizing preliminary screening and diagnosis, effectively reducing the workload of pathologists and improving work efficiency. It effectively solves the problems of efficiency, cost and accuracy of traditional pathological diagnosis methods.
[0103] The references related to this disclosure are as follows.
[0104] [1] Ya Gao. What is oral cancer and how to detect and treat it early? [J]. Anti-Cancer Window, 2023
[0105] [2] Tingting Zhang. Discussion on the impact of health education on oral health knowledge and behaviors of oral cancer patients [J]. Smart Healthcare, 2022.
[0106] [3] Johnson DE, Burtness B, Leemans CR, Lui VWY, Bauman JE, Grandis JR. Head and neck squamous cell carcinoma. Nat Rev Dis Primers 2020.
[0107] [4] Fanhao Meng, Yu Tian, Bo Qiao, etc. Construction of an artificial intelligence prevention and diagnosis platform for oral diseases based on deep learning of large-scale clinical data [J]. Journal of Precision Medicine, 2020.
[0108] [5] Shilu Zhao, Fuqiang Cheng, Fanhua Shi. Current status of clinical epidemiology research on oral squamous cell carcinoma [J]. China Healthcare & Nutrition, 2013.
[0109] [6] Fu Zhong. What are the symptoms before oral cancer? [J]. Family Life Guide, 2020.
[0110] [7] Xing Zhu. Early detection is the key to treating oral cancer [J]. Anti-Cancer Window, 2023.
[0111] [8] Wang Liming, Wang Yue. Correlation analysis of early clinical diagnosis and pathological diagnosis of oral mucosal squamous cell carcinoma [J]. Shanghai Journal of Biomedical Engineering, 2002.
[0112] [9] Li Jiang, Zhang Chunye. Diagnostic criteria for oral and oropharyngeal cancer pathology [J]. China Journal of Oral and Maxillofacial Surgery, 2020.
[10] Kumar A, Singh S K, Saxena S, Lakshmanan K, Sangaiah A K, Chauhan H, Shrivastava S, Singh R K. Deep feature learning for histopathological image classification of canine mammary tumors and human breast cancer [J]. Information Sciences, 2020.
[0113]
[11] LIU J L, LI S H, CAI Y M, et al. Automated Radiographic Evaluation of Adenoid Hypertrophy Based on VGG-Lite [J / OL]. Journal of Dental Research, 2021.
[0114]
[12] BAXI V, EDWARDS R, MONTALTO M, et al. Digital pathology and artificial intelligence in translational medicine and clinical practice [J].
[0115]
[13] Zhang Y, Chan S, Park V Y, Chang K-T, Mehta S, Kim M J, Combs F J, Chang P, Chow D, Parajuli R, Mehta R S, Lin C-Y, Chien S-H, Chen J-H, Su M-Y. Automatic detection and segmentation of breast cancer on MRI using mask R-CNN trained on non–fat-sat images and tested on fat-sat images[J]. Academic Radiology, 2022.
[0116]
[14] KIM Y. Convolutional Neural Networks for Sentence Classification[C / OL] / / Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing(EMNLP), Doha, Qatar. 2014.
[0117]
[15] Shorten C, Khoshgoftaar T M. A survey on Image Data Augmentation for Deep Learning[J]. Journal of Big Data, 2019.
[0118]
[16] Cheplygina V, De Bruijne M, Pluim J P W. Not-so-supervised: A survey of semi-supervised, multi-instance, and transfer learning in medical image analysis[J]. Medical Image Analysis, 2019.
[0119]
[17] FOLMSBEE J, LIU X, BRANDWEIN-WEBER M, et al. Active deep learning: Improved training efficiency of convolutional neural networks for tissue classification in oral cavity cancer[C / OL] / / 2018 IEEE 15th International Symposium on Biomedical Imaging(ISBI 2018), Washington, DC. 2018.
[0120]
[18] SHABAN M, KHURRAM S A, FRAZ M M, et al. A Novel Digital Score for Abundance of Tumour Infiltrating Lymphocytes Predicts Disease Free Survival in Oral Squamous Cell Carcinoma.[J / OL]. Scientific Reports, 2019.
[0121]
[19] ARIJI Y, FUKUDA M, KISE Y, et al. Contrast-enhanced computed tomography image assessment of cervical lymph node metastasis in patients with oral cancer by using a deep learning system of artificial intelligence.[J / OL]. Oral Surgery, Oral Medicine, Oral Pathology and Oral Radiology, 2019.
[0122]
[20] HALICEK M, DORMER J D, LITTLE J V, et al. Hyperspectral Imaging of Head and Neck Squamous Cell Carcinoma for Cancer Margin Detection in Surgical Specimens from 102 Patients Using Deep Learning[J / OL]. Cancers, 2019.
[0123]
[21] HORIE Y, YOSHIO T, AOYAMA K, et al. Diagnostic outcomes of esophageal cancer by artificial intelligence using convolutional neural networks[J / OL]. Gastrointestinal Endoscopy, 2019.
[0124]
[22] Yang S Y, Li S H, Liu J L, et al. Histopathology-based diagnosis of oral squamous cell carcinoma using deep learning[J]. Journal of Dental Research, 2022.
[0125]
[23] MINAEE S, BOYKOV Y, PORIKLI F, et al. Image Segmentation Using Deep Learning: A Survey[J / OL]. IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE Transactions on Pattern Analysis and Machine Intell igence, 2020.
[0126]
[24] Acho S N, Rae W I. Interactive breast mass segmentation using a convex active contour model with optimal threshold values[J]. Physica Medica, 2016.
[0127]
[25] Mencattini A, Rabottino G, Salmeri M, et al. Breast mass segmentation in mammographic images by an effective region growing algorithm[C] / / International Conference on Advanced Concepts for Inelligent Vision Sysems..
[0128]
[26] Abbas Q, Celebi M E, Garci A I F. Breast mass segmentation using region-based and edge-based methods in a 4-stage multiscale sysem[J]. Biomedical Signal Processing & Control, 2013.
[0129]
[27] Gong Jinchang, Zhao Shangyi, Wang Yuanjun. Research progress of medical image segmentation based on deep learning[J]. Chinese Journal of Medical Physics, 2019.
[0130]
[28] Otsu N. A threshold selection method from gray-level histograms[. IEEE transactions on systems, man, and cybernetics, 1979.
[0131]
[29] Pun T. A new method for grey-level picture thresholding using the entropy of the histogram[J]. Signal processing, 1980.
[0132]
[30] Kittler J, Illingworth J. Minimum error thresholding. Pattern recognition, 1986.
[0133]
[31] Huang Jintao, Yang Shaoqing, Liu Songtao. Research on an adaptive multi-mode interactive image tracking algorithm [J]. Modern Defense Technology, 2015.
[0134]
[32] Ju Zhiyong, Zhai Chunyu, Zhang Wenxin. Color commodity label image segmentation method based on SVM and region growing [J]. Electronic Science and Technology, 2021.
[0135]
[33] Liu Yuan, Xia Chunlei. An algorithm for edge detection of strip surface defect images based on Sobel operator [J]. Electronic Measurement Technology, 2021.
[0136]
[34] Li Lei, Li Yingna, Zhao Zhengang. Edge detection of transmission line images based on improved Canny operator [J]. Electric Power Science and Engineering, 2021.
[0137]
[35] LONG J, SHELHAMER E, DARRELL T. Fully Convolutional Networks for Semantic Segmentation [C / OL] / / 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA. 2015.
[0138]
[36] ZENG X, LONG L. Generative Adversarial Networks [M / OL] / / Beginning Deep Learning with TensorFlow. 2022.
[0139]
[37] MARTINO F, BLOISID D, PENNISI A, et al. Deep Learning-Based Pixel-Wise Lesion Segmentation on Oral Squamous Cell Carcinoma Images [J / OL]. Applied Sciences, 2020.
[0140]
[38] BADRINARAYANAN V, KENDALL A, CIPOLLAR. SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation[J / OL]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.
[0141]
[39] RONNEBERGER O, FISCHER P, BROX T. U-Net: Convolutional Networks for Biomedical Image Segmentation[M / OL] / / Lecture Notes in Computer Science, Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015.
[0142]
[40] DOS SANTOSD F D, DE FARIAP R, BAN, et al. Automated Detection of Tumor Regions from Oral Histological Whole Slide Images using Fully Convolutional Neural Networks[J / OL]. Biomedical Signal Processing and Control, 2021.
[0143]
[41] ALBISHRI A, SHAH S, LEE Y, et al. OCU-Net: A Novel U-Net Architecture for Enhanced Oral Cancer Segmentation[J]. 2023.
[0144]
[42] Li Lianbing, Rui Yingying, Shang Jianwei, et al. Diagnosis and Segmentation Method of Oral Squamous Cell Carcinoma Based on Deep Learning[J]. Computer Applications and Software, 2021.
[0145]
[43] PENNISI A,BLOISID D,NARDID,et al.Multi-encoder U-Net for Oral Squamous Cell Carcinoma Image Segmentation[C / OL] / / 2022IEEE International Symposium on Medical Measurements and Applications(MeMeA),Messina,Italy.2022.
[0146]
[44] Wu Y,Koyuncu C F,Toro P,et al.A machine learning model for separating epithelial and stromal regions in oral cavity squamous cell carcinomas using H&E-stained histology images:a multi-center,retrospective study[J].Oral Oncology,2022.
[0147]
[45] LIU Y,BILODEAU E,POLLACK B,et al.Automated detection of premalignant oral lesions on whole slide images using convolutional neural networks[J].
[0148]
[46] MUSULIN J, D,ZULIJANI A,et al.An Enhanced Histopathology Analysis:An AI-Based System for Multiclass Grading of Oral Squamous Cell Carcinoma and Segmenting of Epithelial and Stromal Tissue.[J / OL].Cancers,2021.
[0149]
[47] FRAZ M M, KHURRAM S A, GRAHAM S, et al. FABnet: feature attention-based network for simultaneous segmentation of microvessels and nerves in routine histology images of oral cancer[J / OL]. Neural Computing and Applications, 2020.
[0150]
[48] Yan, W. C., Li, J. Z., Chen, M., & Lu, Y. H. (2021). Artificial intelligence assisting oral medicine: applications in the field of endodontics. International Journal of Stomatology.
[0151]
[49] Li, Y., Hao, Z., & Lei, H. Survey of convolutional neural network. J. Comput. Appl. 2016.
[0152]
[50] Luo, W., Li, Y., Urtasun, R., et al. Understanding the effective receptive field in deep convolutional neural networks[J]. Advances in neural information processing systems, 2016.
[0153]
[51] Liu, X. P., Luan, X. D., & Xie, Y. X. A review of transfer learning research and algorithms[J]. Journal of Changsha University, 2018.
[0154]
[52] Chaudhari, S., Polatkan, G., Ramanath, R., & Mithal, V. An Attentive Survey of Attention Models[J]. ACM Transactions on Intelligent Systems and Technology, 2021.
[0155]
[53] Jie H,Li S,Samuel A,Gang S,Enhua W.Squeeze-and-ExcitationNetworks[C].IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR),2017.
[0156]
[54] Hou Q,Zhou D,Feng J.Coordinate Attention for Efficient MobileNetwork Design[C].IEEE / CVF Conference on Computer Vision and PatternRecognition(CVPR),2021.
[0157]
[55] Woo S,Park J,Lee J-Y,et al.Cbam:Convolutional block attentionmodu le;proceedings of the Proceedings of the European conference on computervi sion(ECCV),F,2018.
[0158]
[56] CHEN L C,PAPANDREOU G,KOKKINOS I,et al.DeepLab:Semantic ImageSegmentation with Deep Convolutional Nets,Atrous Convolution,and FullyConnected CRFs[J / OL].IEEE Transactions on Pattern Analysis and Mac hineIntelligence,2018.
[0159]
[57] Chen L C,Zhu Y,Papandreou G,et al.Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation[C] / / European Conferen ceon Computer Vision.Springer,Cham,2018.
[0160]
[58] ZHAO H, SHI J, QI X, et al. Pyramid Scene Parsing Network[C / OL] / / 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI. 2017.
[0161]
[59] Zhou Z, Siddiquee M M R, Tajbakhsh N, et al. UNet++: A Nested U-Net Architecture for Medical Image Segmentation[J]. 2018..
[0162]
[60] TAN M, LE Quoc V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks[J]. 2019.
[0163]
[61] LI J, WEN Y, HE L. SCConv: Spatial and Channel Reconstruction Convolution for Feature Redundancy[J].
[0164] It should be understood that in the embodiments of the present invention, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the case of A existing alone, the case of A and B existing simultaneously, and the case of B existing alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0165] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other may be an indirect coupling or communication connection through some interfaces, devices or units, and may also be in the form of electrical, mechanical or other connections.
[0166] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0167] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0168] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for constructing an oral cancer diagnosis model, characterized in that, The diagnostic model is based on DeepLabv3+, with an auxiliary branch added to provide location information for the diagnostic model and SCConv attention added to achieve segmentation processing of oral cancer multi-tissue category pathological images.
2. The method according to claim 1, characterized in that The method includes obtaining sample set annotation data sets ORA-TCGA and ORA-Nine for training the diagnostic model, dividing the data set tissue categories into tumor (CTR), epithelium (ORER), stroma (SR), lymphatic vessels (LV) and skeletal muscle and vessels (MN), and then using OCMS-Net for diagnostic model training and prediction.
3. The method according to claim 2, characterized in that, The diagnostic model is based on OCMS-Net, including a main branch A and an auxiliary branch B. The input of main branch A is oral cancer pathology image, and the output is segmentation mask, which uses different colors to represent different tissue types; The input of auxiliary branch B is oral cancer pathology image, and the output is binary segmentation map of cell nuclei and nuclear pixel distance map.
4. The method according to claim 3, characterized in that, The main branch A is responsible for identifying and dividing different semantic information. Its structure includes: Backbone network - used to extract high-level semantic features of the input image. Pixel Decoder - Converts the features extracted by the backbone network into pixel-level predictions for generating the final segmentation mask.
5. The method according to claim 3, wherein Auxiliary branch B is used to extract the cell nucleus information of the labeled category to assist in tissue segmentation. Its structure includes: Backbone network - a backbone network shared with the main branch A, used to extract image features; Nuclear pixel classification branch: used to identify whether each pixel in the image belongs to the cell nucleus, distinguish nuclear pixels from background pixels through binary classification, and obtain a binary segmentation map of the cell nucleus; Nuclear pixel distance branch: used to predict the horizontal and vertical distances of each nuclear pixel to the centroid of the cell nucleus to which it belongs, and obtain a nuclear pixel distance map. The distance information is used to distinguish adjacent cell nuclei.
6. The method according to claim 3, wherein The workflow of the diagnostic model includes: (1) The input oral cancer pathology image is sent to the main branch and the auxiliary branch at the same time; (2) The main branch extracts high-level semantic features of the image through the backbone network and then generates a segmentation mask through a pixel decoder; (3) The auxiliary branch extracts image features through the backbone network, and then generates a binary segmentation map of the cell nucleus and a nuclear pixel distance map through the nuclear pixel classification branch and the nuclear pixel distance branch; (4) The outputs of the main branch and the auxiliary branch are combined for final tissue segmentation and diagnosis.
7. The method according to claim 4 or 5, characterized in that, The backbone network can be any pre-trained deep learning model, including ResNet and EfficientNet.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor runs the computer program to implement the method according to any one of claims 1 to 7.
9. A computer program product, comprising a computer program, characterized in that, The computer program is executed by a processor to implement the method according to any one of claims 1 to 7.
10. An oral cancer pathological diagnosis system, characterized in that, It includes tissue segmentation module and diagnosis module, among which, The tissue segmentation module has: Select Image Function - Allows users to select or upload oral cancer pathology images to be analyzed. Prediction function - perform tissue segmentation on selected oral cancer pathology images to identify and distinguish different tissue types; The diagnostic module has: Image selection function - a function for selecting from the oral cancer pathological images output by the tissue segmentation module, Tumor stroma ratio function - analyzing the ratio of tumor tissue to surrounding normal tissue, Vascular and perineural invasion function - detecting whether the tumor has invaded the surrounding blood vessels and nerves.