Hyperspectral data-based pharyngolaryngeal tumor semantic segmentation identification method
Through the semantic segmentation method combined with hyperspectral imaging technology and deep learning, the problem of inaccurate identification in the early diagnosis of throat tumors is solved, and the accurate identification of different lesion types and anatomical parts is achieved, objective basis for discrimination is provided, and the accuracy and consistency of the diagnosis is improved.
Patent Information
- Application Number
- CN202510412100.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to achieve objective, accurate, fast and robust identification of different lesion types and anatomical sites in the early diagnosis of throat tumors. The conventional laryngoscopy imaging resolution is low, the NBI spectrum selection is single, the CT/MRI cost is high and the examination time is long, the probability-based classification method lacks an intuitive judgment basis, the single pixel classification ignores spatial correlation, the instance segmentation network calculation cost is high, and the expert labeling error is large.
The semantic segmentation method based on hyperspectral imaging technology and deep learning is adopted to identify and segment the throat tumor through multispectral segment information and lesion spatial characteristics, combined with semantic segmentation neural network, and feature extraction and classification are used for the HybridSN model, and the spectral curve features are provided with a basis for discrimination.
It realizes accurate identification of throat tumors of different sizes, shapes and locations, provides objective basis for discrimination, improves the accuracy and consistency of diagnosis, and reduces calculation costs and labeling errors.
Smart Images

Figure CN120339615A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical imaging technology, and in particular to a method for semantic segmentation and recognition of laryngeal tumors based on hyperspectral data. Background Art
[0002] Laryngeal cancer, including nasopharyngeal cancer, oropharyngeal cancer and laryngeal cancer, is one of the most common malignant tumors of the head and neck. This type of tumor not only causes voice and eating disorders, but also because the early symptoms are not obvious, the prognosis is often poor, resulting in a high mortality rate. Early diagnosis of laryngeal cancer is of great significance in treatment and improving patient survival. Endoscopic examination is a commonly used method for early diagnosis of laryngeal cancer. Compared with other diagnostic methods such as imaging and biopsy, it has the advantages of being intuitive, time-saving and highly safe, and is suitable for periodic examinations required for early diagnosis. Among the various endoscopic imaging technologies, conventional laryngoscopes can generally only obtain visible light (RGB) images and cannot obtain lesion feature data. Early malignant tumors in the throat are superficial, and there are generally no obvious symptoms or atypical symptoms in the early stages of cancer. It is difficult to achieve high-precision identification using conventional laryngoscopes, which can easily lead to missed detections. Fluorescent staining imaging technology can obtain relatively rich feature data, but the throat tissue is special, and the pharyngeal reflex is enhanced to highly irritating dyes, and even mucosal damage is caused, so it cannot be applied to laryngeal endoscopy. Narrow Band Imaging (NBI) technology, based on conventional endoscopy, uses narrow-band spectra to develop microvascular structures, increases visual contrast, helps observe morphological changes in the mucosa, and plays an important role in the clinical diagnosis of laryngeal cancer. However, the existing NBI spectrum selection is relatively single, mainly focusing on two narrow-band lights with wavelengths of 415nm (blue light) and 540nm (green light), which have limited penetration ability into the diseased mucosa, low imaging resolution, and it is difficult to unify diagnostic standards between different anatomical sites and lesion types. In addition, for special cases such as white spots with "umbrella effect" and lesions after chemotherapy, the accuracy of interpretation is limited.
[0003] The Chinese patent document "Tumor Identification Method" with publication number CN107292312A and the Chinese patent document "A Method for Automatic Outlining of Nasopharyngeal Carcinoma Tumors Based on Deep Learning" with publication number CN115359881 perform organ and blood vessel segmentation on CT or MRI images, and train tumor classifiers after data preprocessing to identify or segment benign and malignant tumors in the images to be tested. The disadvantage is that this technology relies on CT or MRI images for training and classification, but conventional CT has low specificity or high missed diagnosis rate for certain types of diseases (such as laryngeal cancer with asymmetric chondrosclerosis), resulting in large differences in the diagnostic accuracy of different diseases. Although the combination of CT and MRI can improve the overall diagnostic accuracy, its high cost and long examination time make it difficult to meet the periodic screening needs required for early diagnosis.
[0004] The Chinese patent document "A method for identifying head and neck images" with the publication number CN116739969 uses artificial intelligence to identify features such as infection, bleeding, cauliflower shape, and irregular boundaries in white light endoscopy images of otolaryngology. At the same time, it identifies the spot and blood vessel shape features in NBI images, and combines with a trained classifier to determine whether there is carcinoma in situ in head and neck squamous cell carcinoma images. Its disadvantages are as follows: This method mainly relies on the spot and blood vessel shape features of NBI images to determine the lesion morphology, but lacks objective criteria such as spectral features and lesion size. Its classification method based on probability is difficult to provide an intuitive discrimination basis and lacks interpretability when the doctor's judgment is inconsistent with the classification result.
[0005] The US patent document "Method of non-invasive detection of tumour and / or healthy tissue and hyperspectral imaging apparatus" with the publication number US10964018 uses a hyperspectral imaging device to detect human tissues. After preprocessing the generated hyperspectral data, the pixels are clustered by using the spectral angle (SA), machine learning, or manual annotation method through the spectral values of each pixel, and then a pixel classifier is trained and generated by using the machine learning method to achieve pixel-level tumor classification. Its disadvantages are as follows: This method only classifies based on the spectral values of single pixels, ignoring the spatial correlation and morphological features between adjacent pixels, and is prone to be inflexible in the multi-scale and morphological differences of organs or lesions. Moreover, single pixel values are easily affected by noise pollution, and the prediction results are greatly affected by random factors. Although using a convolutional neural network (CNN) based on image patch classification can effectively improve these problems, it does not conform to the pixel-based classification method stated in this technology.
[0006] The Chinese patent document “A Microscopic Hyperspectral Leukocyte Detection Method Based on Improved Faster RCNN” with publication number CN114037671A establishes a leukocyte microscopic hyperspectral image dataset, combines pseudo-color images and spectral data to annotate different categories of leukocytes, and uses an improved Faster RCNN network to locate and classify leukocytes, thereby realizing automatic recognition of microscopic leukocyte images. However, this method uses Faster RCNN and other similar instance segmentation networks to locate and classify individual instances through bounding boxes, which is only applicable to microscopic images with clear individual boundaries such as leukocytes, and is not applicable to densely packed tumor cells or tumor epidermis that can only be identified based on continuous areas. Therefore, this method has the problems of difficulty in labeling and reduced recognition accuracy in such scenarios, which is fundamentally different from the semantic segmentation and classification method adopted by the present technical solution. In addition, the instance segmentation method has a high computational cost, and its practicality is lower than that of the semantic segmentation method in intraoperative detection applications with limited computing resources and real-time requirements.
[0007] The Chinese patent document "A Sen-UNet-based hyperspectral image segmentation method for dead wood with pine wilt disease" with publication number CN117593529A uses hyperspectral imaging equipment to collect hyperspectral images of forest areas, and after data preprocessing and expert annotation, uses a semantic segmentation network based on Sen-UNet to automatically identify pine wilt disease. This method is suitable for remote sensing ground object detection scenarios, because the outline of trees is relatively clear, and the proportion of the target to be measured in the image is limited, and the expert annotation error is small. However, in the detection of non-section specimens of throat tumors, experts usually perform pathological examinations on suspicious parts and make overall judgments on the entire specimen, rather than accurately calibrating the lesion area based on vision. This annotation method is easily affected by subjective factors, which may reduce the annotation quality and affect the model classification effect, especially when the sample size of the medical database is small.
[0008] Therefore, there is an urgent need for an objective, accurate, rapid, and robust detection method that can identify different lesion types and anatomical locations to promote the standardization and precision of early endoscopic diagnosis of laryngeal tumors. Summary of the invention
[0009] In view of the above problems, the present invention provides a semantic segmentation and recognition method for laryngeal tumors based on hyperspectral data. Based on hyperspectral imaging technology and deep learning, it can more effectively utilize the information abundance of multi-spectral bands of hyperspectral images and the spatial characteristics of lesions, so that it has a good recognition effect on laryngeal tumors of different sizes, shapes and locations, and can provide an objective judgment basis based on the spectral curve characteristics of different tumors.
[0010] A method for semantic segmentation and recognition of throat tumors based on hyperspectral data, comprising the following steps: S1, using a hyperspectral imaging device to collect hyperspectral data of throat tumor tissue samples or endoscopic images of throat tumor patients; S2, performing clipping preprocessing on the hyperspectral data obtained in step S1 to obtain preprocessed hyperspectral data; S3, performing black and white board correction on the preprocessed hyperspectral data in step S2 to obtain corrected hyperspectral data, and performing a filtering operation to reduce spatial or spectral noise; S4, synthesizing a pseudo-color image from the grayscale image corresponding to the hyperspectral data in step S1 at a specific spectral band, and performing annotation based on morphological features and spectral curves; S5, merging the corrected hyperspectral data in step S3 and the tumor labels in S4 in a one-to-one correspondence to establish a throat tumor sample data set; S6, using a semantic segmentation neural network based on image patches to preprocess and train the sample data set to obtain a throat tumor semantic segmentation and recognition model; S7, identifying throat tumors through the throat tumor semantic segmentation and recognition model; if there is a large difference between the recognition result and the label, analyze the corresponding spectral curve to determine whether the difference comes from recognition error or label error, and retrain the model after correcting the label in the case of label error.
[0011] Further, in step S3, the relative reflectance calculation formula for black and white board correction is: , where, represents the relative reflectance of the sample, represents the original spectral reflectance of the sample, represents the spectral reflectance of the black board, represents the spectral reflectance of the white board.
[0012] Further, in the said step S3, the filtering operation includes at least one of mean filtering, median filtering, Gaussian filtering, Savitzky-Golay filtering and bilateral smoothing filtering.
[0013] Further, in the said step S6, the preprocessing method includes at least one of standard normal transformation, logarithmic transformation, linear stretching, and median filtering; the semantic segmentation neural network classification model includes semantic segmentation neural networks such as FCN, U-Net, SegNet, and HybridSN.
[0014] Further, in the said step S4, the specific spectral bands include 461, 548, and 698 nm, and the labeled labels include a total of K label classifications including background, healthy tissue, and K - 2 different benign or malignant tumor categories.
[0015] Further, in step S5, the throat tumor sample dataset has a total of N sample data, the size of a single sample data is HxWxB, there are K types of labels, and the size of a single label is HxW. Corresponding samples and spectral data are selected according to the tumor type category.
[0016] Further, step S6 includes a data preprocessing step: The data is first preprocessed by standard normal transformation for normalization. The formula for standard normal transformation is: , where, is the original spectral value, is the mean of the spectrum, is the standard deviation of the spectrum, is the spectral value after standard normal transformation; The B dimension is compressed to the b dimension by principal component analysis, so that the size of a single sample data becomes HxWxb; then, in the spatial dimension of HxW, it moves at a fixed stride of k pixels in both the horizontal and vertical directions and selects the pixel center to generate image patch-label pairs for network training; an image patch of size hxw centered on each pixel is used as a preprocessed data, and the original label of this pixel is used as the preprocessed label. A total of n groups of image patch-label pairs are generated from this single sample, and the area beyond the boundary is filled with the value 0.
[0017] Further, step S6 includes a network model and parameter setting step: The Hybrid Spectral Network model is used for the segmentation and recognition of throat tumors. The segmentation network consists of a three-dimensional convolutional layer, a channel attention module, a spatial attention module, a two-dimensional convolutional layer, and a fully connected layer; the purpose of the convolutional layer is to achieve local perception and feature extraction, so that when distinguishing tumor categories, not only a single pixel point is considered, but also the spectral features and spatial features of several surrounding points; the purpose of the attention module is to perform dynamic weighting on different channels or spaces, emphasizing or ignoring special bands or spatial positions, and strengthening feature representation and model interpretability.
[0018] Further, step S6 includes a network training step: after generating image patch-label pairs for each sample, a part of them is set as the training set, and the rest is set as the test set. The training set is input into the segmentation network, and after several iterations of convergence, the training is ended to obtain a semantic segmentation and recognition model for throat tumors.
[0019] Beneficial effects brought by the technical solution of the present invention 1) Diagnosis is realized based on hyperspectral images with more spectral bands, the information is richer, and the accuracy is higher; 2) The semantic segmentation recognition results can accurately label the boundaries between normal tissues and different lesions in the image, more intuitively display the lesion boundaries, and support the inspection of spectral curves pixel by pixel to check for labeling errors or recognition errors, providing an objective basis for the recognition results; 3) Artificial neural networks, especially convolutional neural networks, can learn the global information and detailed information of the image, better distinguish the spatial features of the lesions, and thus improve the prediction accuracy and generalization ability of the semantic segmentation network model; 4) The unique spectral features of different tumor categories can be used as discriminant bases in addition to the lesion morphology, providing an objective and unified judgment standard for doctors. Brief Description of the Drawings
[0020] Figure 1 is a flowchart of the method for semantic segmentation and recognition of laryngeal tumors based on hyperspectral imaging technology and deep learning according to the present invention.
[0021] Figure 2 is a schematic structural diagram of the segmentation network according to an embodiment of the present invention.
[0022] Figure 3 is an example of the segmentation and recognition results of laryngeal tumors obtained by using the method of the present invention.
[0023] Figure 4 is an example of the comparison of the segmentation and recognition results of laryngeal tumors obtained by the single-pixel classification method in the US patent document "Method of non-invasive detection of tumour and / or healthy tissue and hyperspectral imaging apparatus" with the publication number US10964018 and the method based on image patch classification of the present invention.
[0024] Figure 5 is an example of the comparison of the instance segmentation method in the Chinese patent document "A method for microscopic hyperspectral white blood cell detection based on improved Faster RCNN" with the publication number CN114037671A and the semantic segmentation method of the present invention. Detailed Embodiments
[0025] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content recorded in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0026] The hyperspectral imaging mentioned in the present invention is a sophisticated technique that can capture and analyze the spectra at each point within a spatial region. Since it can detect the unique spectral "signatures" at different spatial positions of a single object, hyperspectral imaging technology can detect substances that are indistinguishable to the human eye. Different biological tissues, especially the physical and chemical properties of normal tissues and different types of diseased tumors, result in different absorption and reflection of wavelengths, thus leading to different spectral characteristics. Therefore, different throat tumors have their unique spectra. These spectral information that cannot be seen by the human eye can be distinguished through hyperspectral imaging technology, providing feasibility for computer intelligent visualization to identify throat tumors.
[0027] As Figure 1 shown, the method for semantic segmentation and recognition of throat tumors based on hyperspectral imaging technology and deep learning according to the present invention includes the following steps: S1. Collect throat tumor tissue samples or prepare throat tumor patients who have signed the informed consent form; the throat tumor tissue samples used in this embodiment are provided by Beijing Friendship Hospital and include 156 throat tumor samples extracted during surgery, including 14 types of lesion categories such as papillary thyroid carcinoma, squamous cell carcinoma of the tongue, squamous cell carcinoma of the larynx, lymph node metastasis, squamous cell carcinoma of the hypopharynx, supraglottic squamous cell carcinoma of the larynx, thyroid malignancy, malignancy of the root of the tongue, supraglottic malignancy, laryngeal mass, thyroid mass, recurrent carcinoma at the stoma, hypopharyngeal malignancy, and adenoid cystic carcinoma of the larynx, which were collected from partner hospitals between 2023 and 2024.
[0028] S2. Use a hyperspectral imaging device to collect the hyperspectral data of the throat tumor tissue samples or the endoscopic images of the throat tumor patients. The hyperspectral device can collect in the visible light range of several hundred nm, with a spectral resolution better than 10 nm, and can collect dozens or even hundreds of spectral channels. When the hyperspectral camera is collecting data, it generally needs to be externally connected to a stable broadband light source and illuminated at a certain angle, with uniform illumination and no strong reflection. The hyperspectral imaging device used in this embodiment is the Finnish Specim IQ handheld intelligent hyperspectral camera. The distance between the lens of the hyperspectral imager and the tissue is 20 - 30 cm, the spectral range of the spectrometer is 400 - 1000 nm, the band interval is 3 nm, and there are a total of 204 bands.
[0029] S3. Perform cropping preprocessing on the hyperspectral data described in S2 to obtain preprocessed hyperspectral data. Cropping preprocessing means removing the irrelevant regions and backgrounds in the image and retaining the key regions of interest, which can reduce the useless information and edge interference in the data, and improve the data quality and processing efficiency; the captured data is cropped according to the region of interest to a spatial size of pixel dimensions 512x512, with the center of the image being the throat tissue sample.
[0030] S4. Perform black-and-white calibration on the preprocessed hyperspectral data described in S3 to obtain calibrated hyperspectral data, and perform a filtering operation to reduce spatial or spectral noise. The purpose of black-and-white calibration is to eliminate noise and deviations in the hyperspectral imaging system, such as systematic errors or dark current noise introduced by imaging devices and environmental light sources, thereby improving the authenticity and consistency of the data. The filtering operation adopted in this embodiment is Savitzky-Golay (SG) filtering, with a window size of 11 and an order of 2, which is applicable to most spectral data and can retain the main features of the signal while smoothing the noise.
[0031] S5. Synthesize a pseudo-color image from the grayscale images corresponding to the hyperspectral data described in S2 at the wavelengths of 461, 548, and 698 nm, and have professional medical personnel annotate it based on morphological features and spectral curves. The labels are background, healthy tissue, and different benign / malignant tumor categories; in this embodiment, diseased tissues are used, combined with the background and healthy tissue categories, for a total of 3 label classifications, all of which are annotated by professional doctors from partner hospitals.
[0032] S6. Merge the calibrated hyperspectral data described in S4 with the tumor labels described in S5 to establish a pharyngeal tumor sample dataset. This dataset not only provides high-quality data support for subsequent model training and tumor detection algorithm development, but also facilitates the screening and analysis of spectral features based on different anatomical locations and tumor types; in this embodiment, the pharyngeal tumor sample dataset consists of a total of 156 sample data, with the size of a single sample being 512x512x204, 3 label categories, and the size of a single label being 512x512. The corresponding samples and spectral data can be screened according to the tumor type category.
[0033] S7. Use an artificial neural network to preprocess and train the sample dataset to obtain a pharyngeal tumor semantic segmentation recognition model; in this embodiment, S7 is mainly divided into three sub-steps: S71: Data preprocessing. The data is first preprocessed by normalizing it through standard normal transformation. The formula for standard normal transformation is , where is the original spectral value, is the mean of the spectrum, is the standard deviation of the spectrum, is the spectral value after standard normal transformation. Normalization helps to avoid the influence of data scale on the training process of the neural network model and helps to improve the convergence speed and accuracy of the model.
[0034] Compress the 204 dimensions to 30 dimensions through principal component analysis, so that the size of a single sample becomes 512x512x30. The main purpose is to reduce the amount of data while retaining the main information of the data.
[0035] Next, move and select pixel centers at a fixed stride of 10 pixels in both the horizontal and vertical directions in the 512x512 spatial dimension to generate image patch-label pairs for network training. Take a 25x25-sized image patch centered on each pixel as a preprocessed data, and use the original label of the pixel as the preprocessed label. A total of 2704 groups of image patch-label pairs can be generated from this single example sample. The area beyond the boundary is filled with the value 0.
[0036] S72: Network model and parameter settings. In this embodiment, an optimized HybridSN (Hybrid Spectral Network) model is used for the segmentation and recognition of throat tumors. The segmentation network consists of 3 three-dimensional convolutional layers, 1 channel attention module, 1 spatial attention module, 1 two-dimensional convolutional layer, and 3 fully connected layers. The main purpose of the convolutional layer is to achieve local perception and feature extraction, so that when distinguishing tumor categories, not only a single pixel point is considered, but also the spectral features and spatial features of several surrounding points. The main purpose of the attention module is to perform dynamic weighting processing on different channels or spaces, such as emphasizing or ignoring special bands or spatial positions, so as to strengthen feature representation and model interpretability. The network structure is shown in Figure 2 , and its working process is as follows: Select 3D image patches of size 25x25x30 from the training set data, and obtain the first output feature through a three-dimensional convolutional layer with a convolutional kernel size of 7x3x3 and 8 channels. The picture size is reduced to 23x23x24. The first output feature passes through a three-dimensional convolutional layer with a convolutional kernel size of 5x3x3 and 16 channels to obtain the second output feature, and the picture size is reduced to 21x21x20. The second output feature passes through a three-dimensional convolutional layer with a convolutional kernel size of 3x3x3 and 32 channels to obtain the third output feature, and the picture size is reduced to 19x19x18.
[0037] The third output feature, a total of 32 groups of 19x19x18 pictures, is transformed into a deformed third output feature of 19x19x576 by changing the shape, and sequentially passes through the channel attention module and the spatial attention module.
[0038] The channel attention module performs global max pooling and average pooling operations in the spatial dimension, that is, calculates the maximum value and average value of all spatial dimensions in the same channel, and generates pooling vectors of size 1x1x576 respectively. Through a multi-layer perceptron constructed by two fully connected layers with dimensions 36 and 576, the sum of the results of passing the two pooling vectors through the multi-layer perceptron is mapped through the Sigmoid function to obtain a channel attention coefficient of size 1x1x576. According to the coefficient, the channel dimension of the third output feature is modified to generate a fourth output feature of size 19x19x576. The formula is as follows: , wherein is the channel attention coefficient, is the deformed third output feature, is the multi-layer perceptron, is the average pooling, is the max pooling.
[0039] The spatial attention module generates a feature map of size 19x19x2 by performing max pooling and average pooling operations in the channel dimension. Then, through a two-dimensional convolutional layer with a 3x3 convolutional kernel and a Sigmoid activation function, a spatial attention coefficient of size 19x19x1 is generated. According to the coefficient, the spatial dimension of the fourth output feature is modified to generate the fifth output feature of size 19x19x576.
[0040] After the fifth output feature is processed by a two-dimensional convolutional layer with a 3x3 convolutional kernel and 64 channels, the sixth output feature is obtained, and the image size is reduced to 17x17x64. After deformation, a vector of size 18496 is generated, which is called the deformed sixth output feature.
[0041] The deformed sixth output feature passes through three fully connected layers with dimensions 256, 128, and 16 to obtain the output of the segmentation network. The output of the segmentation network is a 16-dimensional tensor data, where 16 represents the number of categories to be segmented, including tumors and the background. After Softmax processing, it corresponds to the probabilities corresponding to 16 labels.
[0042] All convolutional layers have a stride of 1, use the ReLU activation function, and use batch normalization (BN) to accelerate the training process to improve the stability and generalization ability of the model.
[0043] In the training parameters, the ADAM optimizer is used, the learning rate is set to 0.00005, gamma is 0.15, momentum is 0.9, and batch_size is 256. The loss function is the cross-entropy loss.
[0044] S73: Network training. After generating the image patch-label pairs for each sample according to S71, 90% is set as the training set and 10% is set as the test set. The training set is input into the segmentation network, and after 50 iterations of convergence, the training ends, and a semantic segmentation recognition model for throat tumors is obtained.
[0045] S8. The recognition of pharyngeal tumors is achieved through the above-mentioned pharyngeal tumor semantic segmentation recognition model. The criteria for judging the recognition accuracy are the overall discrimination accuracy rate and the discrimination accuracy rates of each label or each category of pharyngeal tumor types. When recognizing pharyngeal tumors in this embodiment, according to the data preprocessing method described in S71, the sliding step size is set to 1, image patches corresponding to each pixel of the image to be tested are generated, and the pharyngeal tumor semantic segmentation recognition model obtained in S73 is used for prediction. Finally, each pixel of the image to be tested is classified to obtain the recognition result of pharyngeal tumors. For the recognition result, see Figure 3 . By comparing the true labels with the predicted labels, the overall accuracy of the network model and the accuracies of the three classifications can be obtained. Among them, the accuracy of HybridSN has been greatly improved compared with the commonly used CNN2D and CNN3D networks. The overall accuracy rate reaches 97.72%. The recognition accuracy of the diseased tissue has been improved from about 60% of other models to 82.29%. The specific values are shown in Table 1, which can intuitively and accurately identify the location and type of pharyngeal tumors, providing scientific assistance for improving the accuracy and objectivity of early diagnosis of pharyngeal cancer.
[0046] Table 1
[0047] The above has described the embodiments of the present invention in detail with reference to the examples. However, the present invention is not limited to the above examples. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and these should also be regarded as the protection scope of the present invention.
Claims
1. A semantic segmentation and recognition method for throat tumors based on hyperspectral data, characterized in that It includes the following steps: S1. Use a hyperspectral imaging device to collect hyperspectral data of a throat tumor tissue sample or an endoscopic image of a throat cancer patient; S2. Perform cropping preprocessing on the hyperspectral data obtained in step S1 to obtain preprocessed hyperspectral data; S3. Perform black and white board correction on the preprocessed hyperspectral data in step S2 to obtain corrected hyperspectral data, and perform a filtering operation to reduce spatial or spectral noise; S4. Synthesize a pseudo-color image from the grayscale image corresponding to the hyperspectral data in step S1 at a specific spectral band, and perform annotation based on morphological features and spectral curves; S5. Merge the corrected hyperspectral data in step S3 and the tumor labels in step S4 in a one-to-one correspondence to establish a throat tumor sample data set; S6. Use a semantic segmentation neural network based on image patches to preprocess and train the sample data set to obtain a throat tumor semantic segmentation recognition model; S7. Identify throat tumors through the throat tumor semantic segmentation recognition model; if there is a large difference between the recognition result and the label, analyze the corresponding spectral curve to determine whether the difference is due to recognition error or label error, and retrain the model after correcting the label in the case of label error.
2. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 1, wherein, In step S3, the relative reflectance calculation formula for black and white board correction is: , Among them, represents the relative reflectance of the sample, represents the original spectral reflectance of the sample, represents the spectral reflectance of the blackboard, represents the spectral reflectance of the whiteboard.
3. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 1, wherein In the said step S3, the filtering operation includes at least one of mean filtering, median filtering, Gaussian filtering, Savitzky-Golay filtering and bilateral smoothing filtering.
4. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 1, wherein In the said step S6, the preprocessing method includes at least one of standard normal transformation, logarithmic transformation, linear stretching, and median filtering; the semantic segmentation neural network classification model includes semantic segmentation neural networks such as FCN, U-Net, SegNet, and HybridSN.
5. The throat tumor semantic segmentation and recognition method based on hyperspectral data according to claim 2, wherein In the said step S4, the specific spectra include 461, 548, and 698 nm, and the labeled tags include a total of K tag classifications of background, healthy tissue, and K-2 different benign or malignant tumor categories.
6. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 5, wherein In the said step S5, the throat tumor sample data set has a total of N sample data, the size of a single data is HxWxB, there are K tag categories, and the size of a single tag is HxW. Select the corresponding samples and spectral data according to the tumor type category.
7. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 6, wherein The said step S6 includes a data preprocessing step: The data is first normalized by standard normal transformation, and the formula for standard normal transformation is: , Among them, is the original spectral value, is the mean value of the spectrum, is the standard deviation of the spectrum, is the spectral value after standard normal transformation; Compress the B dimension to the b dimension through principal component analysis, so that the size of a single data becomes HxWxb; then move in the HxW spatial dimension at a fixed stride of k pixels in both the horizontal and vertical directions and select the pixel center to generate image patch-label pairs for network training; take an image patch of size hxw centered on each pixel as a preprocessed data, and the original label of this pixel as the preprocessed label, and generate a total of n groups of image patch-label pairs from this single example sample. The area beyond the boundary is filled with the value 0.
8. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 7, wherein, The said step S6 includes a network model and parameter setting step: Use the Hybrid Spectral Network model for the segmentation and recognition of throat tumors. The segmentation network consists of a three-dimensional convolutional layer, a channel attention module, a spatial attention module, a two-dimensional convolutional layer, and a fully connected layer. The purpose of the convolutional layer is to achieve local perception and feature extraction, so that when discriminating tumor categories, not only a single pixel point is considered, but also the spectral features and spatial features of several surrounding points are considered. The purpose of the attention module is to perform dynamic weighting on different channels or spaces, emphasizing or ignoring special bands or spatial positions, and strengthening feature representation and model interpretability.
9. The method for semantic segmentation and recognition of throat tumors based on hyperspectral data according to claim 8, wherein The step S6 includes a network training step: after generating image patch-label pairs for each sample, a part of them is set as the training set, and the rest is set as the test set. The training set is input into the segmentation network, and after several iterations of convergence, the training is ended to obtain a semantic segmentation and recognition model for throat tumors.
Citation Information
Patent Citations
Tumor identification method
CN107292312A
Improved Faster RCNN-based microscopic hyperspectral white blood cell detection method
CN114037671A
Bursaphelenchus xylophilus dead wood hyperspectral image segmentation method based on Sen-UNet
CN117593529A
Method of non-invasive detection of tumour and / or healthy tissue and hyperspectral imaging apparatus
US10964018B2