Accurate labeling method for microscopic hyperspectral pathological image
Through the K-means clustering algorithm of rough manual labeling and spectral characteristics, combined with filtering and manual adjustment, the accuracy and consistency of hyperspectral pathological image annotation is solved, efficient and accurate pathological image annotation is achieved, and the recognition ability of neural networks is improved.
Patent Information
- Application Number
- CN202510151687.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has high labor cost, low efficiency and low accuracy in hyperspectral pathological image annotation, making it difficult to cope with the problem of inconsistent staining time of pathological sections in different batches, resulting in inconsistent labeling quality and poor algorithm robustness.
Through rough manual labeling combined with spectral features, pixel-level adjustment is performed to generate high-quality data sets, and relative reflectance correction and multiple filtering processing are used, and parameters are adjusted in combination with manual evaluation to achieve accurate annotation.
It simplifies the operation process of pathologists, improves the accuracy and consistency of labeling, reduces the impact of different images, and enhances the training effect of neural network models.
Smart Images

Figure CN120299035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical imaging technology, and particularly to a precise annotation method for microscopic hyperspectral pathological images. Background Art
[0002] In the field of medical tumor recognition, pathological detection analyzes the microscopic structure and cell morphology of a patient's tissue sample to determine whether there is a lesion in the tissue and to gain an in-depth understanding of the nature, type, and development stage of the lesion, playing an important role in the early diagnosis of cancer and other major diseases. Traditional pathological judgment mainly relies on the subjective judgment of pathologists based on their understanding of tumor morphology, which is time-consuming and requires a great deal of doctor experience. Omissions are likely to occur in cases of complex pathology or heavy workload. With the development of assisted diagnosis methods such as artificial intelligence and image analysis technology, digital pathology mainly based on deep learning models can achieve more accurate and objective cancer detection and analysis, providing an auxiliary basis for pathologists' judgments. Based on RGB microscopic images, hyperspectral microscopic images can generate spectral fingerprints based on the physical and chemical properties of normal tissues and different types of diseased tumors, and can more precisely and objectively identify subtle tissue differences and the characteristics of different molecules based on cell morphology. However, intelligent pathological recognition algorithms based on microscopic hyperspectral images require a large number of high-quality datasets with accurate annotations. These datasets usually require professional pathologists to manually annotate RGB images, which is costly in terms of human resources and inefficient. Moreover, due to the difficulty in accurately judging the tumor edge, the annotation accuracy is further reduced, ultimately lowering the accuracy of the recognition model and the robustness of the algorithm.
[0003] The Chinese patent "Method and System for Generating a Benchmark Pathological Dataset Containing an Automatic Annotation File" with the publication number CN115274093B manually selects representative pixels and labels them as different categories, such as healthy tissues and diseased cell nuclei, and uses the LightBGM decision tree algorithm of machine learning for training to predict the pixels in the entire image, and selects the category with the highest probability in a single pixel to generate a label map. Its disadvantage is that this technology relies on manually selecting representative pixels as the main factor for fully automatic label generation, resulting in the generated label results only being able to screen out parts that are close enough to the manually selected pixels, and performing poorly on images with high noise, requiring additional filtering operations and subsequent label map processing. This method is difficult to adjust the point selection according to the quality of the generated results and also has high requirements for the consistency of the captured images, making it difficult to handle the differences in hyperspectral images in cases where the staining duration of different batches of pathological sections varies.
[0004] The Chinese patent "Automatic Generation Method of Precision-Labeled Digital Pathology Datasets Based on Hyperspectral Imaging" with the publication number CN114596298B performs segmentation processing on dual-stained hyperspectral images based on gradient boosting decision trees and graph cut algorithms, and filters out small noises by extracting the outer contour, thereby obtaining a binary image of the tumor region of interest. Dual-stained hyperspectral images refer to adding immunohistochemical marker staining on the basis of conventional hematoxylin-eosin staining. Dual-stained sections can better show the abnormal areas that react chemically with the markers, achieving precise labeling. However, its disadvantage is that this method generates labels relying on dual-stained hyperspectral images, while dual-stained labels do not belong to conventional pathological examination methods. The selection of immunohistochemical markers will also affect the generation quality, and the additional staining steps will be more time-consuming and increase the randomness of staining errors. This method is also difficult to apply to the already acquired hyperspectral pathological sections.
[0005] The Chinese patent "A Semi-Automatic Establishment Method of Hyperspectral Database Based on Deep Learning" with the publication number CN107704878B uses a semi-automatic method to manually label a part of hyperspectral pathological data to obtain the ground truth, and then trains a classifier through a neural network. Then, other parts of hyperspectral pathological data are automatically labeled through the classifier, thereby reducing the number of images that need to be manually labeled. Its disadvantage is that in this method, part of the manually labeled labels are needed to generate the remaining labels through training the classifier. The training of the classifier is based on hyperspectral data, while manual labeling is based on the observation of white light (RGB) images by the human eye and the understanding of the lesion morphology. Differences that are significant in the human eye may not hold true spectrally, and those with significant spectral differences may not be distinguished in manual labeling either. This inconsistent discrimination criterion brings uncertainty to the training of the algorithm model, increasing the difficulty and error of automatic label generation. Therefore, the accuracy of the generated labels depends on the accuracy of the manual labeling regarded as the ground truth, and it is difficult to achieve a high labeling accuracy in the case of difficult manual marking or the need for precise marking (such as precise to the cell nucleus), ultimately resulting in a poor overall quality of the automatically generated dataset. For example, it is easy to achieve high-precision manual labeling for the microscopic hyperspectral image of melanoma with obvious characteristics, but it is difficult to label in the microscopic hyperspectral image of lung cancer filled with numerous squamous and strip cells. In addition, when the total number of data is small, such as less than 100 cases, relying on the precise labeling of some samples to generate the precise labeling of all samples will have random deviations caused by sample selection. Similar to the prior art 1, when there are certain differences in the captured images, such as the inconsistency caused by different staining durations of pathological sections in different batches, the label marking based on the neural network will rely on the robustness of the algorithm and the selection of the training set.
[0006] In summary, the current hyperspectral-based pathological recognition is limited to cancer types with clear or highly distinguishable suspicious regions, making it difficult to apply to cancer types with complex tissue morphology. The lack of high-quality hyperspectral pathological datasets also restricts the development of hyperspectral-based pathological recognition algorithms and applications. Summary of the Invention
[0007] To address the above problems, the present invention provides a precise annotation method for microscopic hyperspectral pathological images. Based on rough manual marking, pixel-level adjustment is achieved through spectral features, and parameters are actively adjusted according to the generated results to obtain a high-quality dataset to ensure the accuracy of deep learning.
[0008] A precise annotation method for microscopic hyperspectral pathological images includes the following steps: 1) Obtain a microscopic hyperspectral pathological image and its preliminary rough marking; 2) Perform filtering on the microscopic hyperspectral pathological image described in step 1); 3) Select gray-scale images corresponding to three bands from the filtered microscopic hyperspectral pathological image obtained in step 2) to synthesize a pseudo-color image; 4) Divide the filtered microscopic hyperspectral pathological image obtained in step 2) into K classes using the K-means clustering algorithm based on the value vectors of pixels, where K is an adjustable constant, obtain the clustering result, and generate corresponding colors for each class to provide a visual clustering map for subsequent class selection; 5) Manually annotate the pseudo-color image obtained in step 3) and the clustering map obtained in step 4); 6) Compare the label selection result in step 5) with the preliminary rough marking in step 1), conduct manual evaluation, and take corresponding measures.
[0009] Further, the microscopic hyperspectral image in step 1) is corrected by black and white boards to obtain the relative reflectance, and the relative reflectance calculation formula is: , where represents the relative reflectance of the sample, represents the original spectral reflectance of the sample, represents the spectral reflectance of the black board, represents the spectral reflectance of the white board.
[0010] Further, the filtering in step 2) includes at least any one of mean filtering, median filtering, Gaussian filtering, Savitzky-Golay (SG) filtering, and mean smoothing filtering.
[0011] Further, in step 5), the types of labels marked include at least any one of background, healthy tissue, blood vessels, and diseased cell nuclei. Each type of label has zero or more clustering types, and each clustering type should be assigned to one and only one type of label.
[0012] Further, in step 6), if the manual evaluation passes or only fine-tuning is required, then proceed to the next process or generate labels for the next hyperspectral image; if the manual evaluation fails, return to step 2), and reselect the filtering operation in step 2) and K in step 4), or inquire whether the rough labels in step 1) are incorrect.
[0013] Beneficial effects brought by the technical solution of the present invention 1) Compared with the traditional manual pixel-level annotation, the pathologists of the present invention only need to outline the rough regions and edges on the H&E stained section, and then use the clustering algorithm and clustering selection to achieve label classification, with more concise operations.
[0014] 2) Compared with the traditional rough edge manual annotation based on vision or immunohistochemical markers, the spectral value vector clustering and grouping classification of the present invention can improve the spectral feature consistency between groups and the objectivity of classification in annotation, and reduce the negative impact on the training of subsequent neural network recognition models.
[0015] 3) Based on semi-automated classification, the present invention performs clustering analysis on a single image rather than all data, so the influence brought by the differences between different images can be reduced.
[0016] 4) The present invention can flexibly adjust the selection of the filtering and the number of clustering types K according to the classification results. During the process of observing the classification results and making adjustments, the centroid brightness and pixel number of each cluster can also be viewed, so as to better observe the data distribution and characteristics of the hyperspectral image. Description of the drawings
[0017] Figure 1 is a flowchart of a semi-automatic intelligent labeling method for a microscopic hyperspectral image according to the present invention.
[0018] Figure 2 is an example of a tool for label classification of clustering in an embodiment of the present invention.
[0019] Figure 3 is an example of a method for generating labels in an embodiment of the present invention.
[0020] Figure 4 is an example of the neural network structure of the recognition classification algorithm in an embodiment of the present invention Figure 5 is an example of the result of applying the recognition classification algorithm to the generated labels in an embodiment of the present invention. Detailed implementation manners
[0021] A semi-automatic intelligent labeling method for microscopic hyperspectral images according to the present invention includes the following steps: S1. Obtain a microscopic hyperspectral hematoxylin-eosin (H&E) image including the region of interest, and have a professional pathologist give region lesion labels; here, the region of interest refers to the region containing cell lesion characteristics or the healthy tissue region for comparison, and is cropped as needed. The hyperspectral device can collect visible light in the range of several hundred nm, with a spectral resolution better than 10 nm, and can collect dozens or even hundreds of spectral channels. When the hyperspectral camera is collecting, generally an external stable white light source needs to be tilted at a certain angle for illumination, with uniform light and no strong reflection. The obtained microscopic hyperspectral image should be a clear image obtained by the hyperspectral camera and has been corrected by a black and white board to obtain the relative reflectance. The purpose of black and white board correction is to eliminate the noise and deviation in the hyperspectral imaging system, such as the systematic error or dark current noise introduced by the imaging device and the environmental light source, so as to improve the authenticity and consistency of the data.
[0022] The formula for relative reflectance is: , where represents the relative reflectance of the sample, represents the original spectral reflectance of the sample, represents the spectral reflectance of the black board, represents the spectral reflectance of the white board.
[0023] S2. Perform filtering processing on the hyperspectral data described in S1 to reduce spatial or spectral noise; common filtering operations include mean filtering, median filtering, Gaussian filtering, Savitzky-Golay (SG) filtering, mean smoothing filtering, etc., without limitation.
[0024] S3. Synthesize a pseudo-color image from the gray-scale images corresponding to the hyperspectral data described in S2 at the 461, 548, and 698 nm bands, or the three bands generated by the band selection algorithm. S4. Divide the hyperspectral data described in S2 into K categories using the K-means clustering algorithm based on the value vectors of pixels, where K is an adjustable constant, obtain the clustering result, and generate corresponding colors for each category to provide a visual clustering map for subsequent category selection. S5. Manually label the pseudo-color image described in S3 and the clustering map described in S4, that is, select the corresponding clustering category for each label category based on the morphological characteristics of the pseudo-color image and the types of the clustering map; the label categories include background, healthy tissue, blood vessels, lesion cell nuclei, etc., and there can be more than one. Each label category can have zero or more clustering categories. Each clustering category should be assigned to one and only one label category.
[0025] S6. Compare the label selection result of S5 with the preliminary rough labeling described in S1, conduct manual evaluation and take corresponding measures; if the manual evaluation passes or only minor adjustments are required, proceed to the next step or generate the label for the next hyperspectral image. If the manual evaluation fails, return to S2, reselect the filtering operation of S2 and K of S4, or ask the pathologist whether the rough label given in S1 is incorrect.
[0026] The following further elaborates in detail on the precise labeling method for microscopic hyperspectral pathological images of the present invention in conjunction with the accompanying drawings and specific embodiments: Embodiment
[0027] As Figure 1 shown, a semi-automatic intelligent labeling method for microscopic hyperspectral images of the present invention includes the following steps: S1. Obtain a microscopic hyperspectral hematoxylin-eosin (H&E) image including the region of interest; the hyperspectral imaging device used in this embodiment is a self-developed hyperspectral microscope. The spectral range of the spectrometer is 420 - 750 nm, the spectral resolution is 5 nm, the pixel size is 3088x2064, the band interval is 5 nm, and there are a total of 67 bands. The pathological section is a lung squamous cell carcinoma pathological section from the National Key Laboratory of Respiratory Diseases, Guangzhou Medical University, and professional pathologists give regional lesion markings, including 2 types of labels for the lesion area and the non-lesion area, with a total of 62 hyperspectral images.
[0028] S2. Filter the hyperspectral data described in S1 to reduce spatial or spectral noise; the filtering operation used in this embodiment is SG filtering, with a window size of 11 and an order of 2, which is applicable to most spectral data and can retain the main features of the signal while smoothing the noise. Secondly, perform normalization preprocessing through standard normal transformation. The formula for standard normal transformation is , where is the original spectral value, is the mean of the spectrum, is the standard deviation of the spectrum, is the spectral value after standard normal transformation. Normalization helps to avoid the influence of data scale on the training process of the neural network model and helps to improve the convergence speed and accuracy of the model.
[0029] S3. Synthesize a pseudo-color image from the gray-scale images corresponding to the hyperspectral data described in S2 at the 461, 548, and 698 nm bands, or 3 bands generated by the band selection algorithm; the pseudo-color map bands used in this embodiment are approximately 460, 550, 700 nm.
[0030] S4. Use the K-means clustering algorithm to divide the pixel-based spectral value vectors of the hyperspectral data described in S2 into K categories, where K is an adjustable constant, to obtain the clustering result, and generate corresponding colors for each category to provide a visual clustering map for subsequent category selection. In this embodiment, K is 20, which can accurately select the nucleus regions in the hyperspectral image with a relatively small number, and initially distinguish between diseased and healthy nuclei.
[0031] S5. Manually annotate the pseudo-color map described in S3 and the clustering map described in S4, that is, select the corresponding clustering category for each label category based on the morphological features of the pseudo-color map and the types of the clustering map. In this embodiment, the label categories include 4 types: background, tissue, diseased nuclei, and healthy nuclei. Among them, the tissue and nuclei are manually annotated, and the distinction between diseased and healthy nuclei is determined by the regional lesion label of the doctor described in S1. In this embodiment, the self-developed graphical interface tool for selecting the corresponding clustering category of the selected label is as Figure 2 shown, and the clustering corresponding to the tissue and nucleus labels can be quickly manually selected based on the morphological features of the pseudo-color map and the spectral features of the clustering map. As Figure 3 shown, after generating the background and nucleus labels, other pixels are regarded as tissue labels, and the nucleus labels are further divided into diseased nuclei and healthy nuclei based on the regional lesion label given by the doctor, and the label maps are merged to generate the final label map required for the database and the recognition model.
[0032] S6. Manually evaluate the label selection result in S5 and take corresponding measures. The selection of SG filtering and clustering number K = 20 in this embodiment is the final result after experiments. The selection of mean filtering, median filtering, and different K values in the figure has poor effects. Since the difference between the generated labels of diseased and healthy nuclei is determined by the doctor's regional lesion marking, the generated labels in this embodiment are similar to the doctor's markings. The manual evaluation determines that the generated labels successfully extract the nucleus part of the image without detailed manual pixel-by-pixel annotation, achieving the purpose of rapid classification of this method.
[0033] To test the quality of the generated labels, in this embodiment, the HybridSN (Hybrid Spectral Network) network model is used to train and recognize the labels generated by this method and the regional lesion labels provided by the doctor respectively. The description of the network structure can be found in Figure 4 and the recognition results can be found in Figure 5 . The network trained according to the regional lesion label obviously has more noise and lower image interpretability. The network trained according to the labels generated by this method has an accuracy of 92.52%, which is much greater than the accuracy of 77.33% of the network trained according to the regional lesion label. Therefore, the purpose of accurate classification of this method is achieved.
[0034] The above embodiments have been described in detail with reference to the examples of the present invention. However, the present invention is not limited to the above examples. Various changes can be made within the scope of knowledge possessed by those of ordinary skill in the art without departing from the spirit of the present invention, and these should also be regarded as the protection scope of the present invention.
Claims
1. A precise annotation method for microscopic hyperspectral pathological images, characterized in that, It includes the following steps: 1) Obtain a microscopic hyperspectral pathological image and its preliminary rough label; 2) Perform filtering processing on the microscopic hyperspectral pathological image described in step 1); 3) Select the grayscale images corresponding to 3 bands from the filtered microscopic hyperspectral pathological image obtained in step 2) to synthesize a pseudo-color image; 4) Based on the value vector of pixels, use the K-means clustering algorithm to divide the filtered microscopic hyperspectral pathological image obtained in step 2) into K classes, where K is an adjustable constant, obtain the clustering result and generate corresponding colors for each class to provide a visual clustering map for subsequent class selection; 5) Manually annotate the pseudo-color image obtained in step 3) and the clustering map obtained in step 4); 6) Compare the label selection result in step 5) with the preliminary rough label in step 1), conduct manual evaluation and take corresponding measures.
2. The precise annotation method for microscopic hyperspectral pathological images according to claim 1, wherein The microscopic hyperspectral image in step 1) is corrected by a black and white board to obtain the relative reflectance, and the relative reflectance calculation formula is: ‘ wherein represents the relative reflectance of the sample, represents the original spectral reflectance of the sample, represents the spectral reflectance of the blackboard, represents the spectral reflectance of the whiteboard.
3. The precise annotation method for microscopic hyperspectral pathological images according to claim 2, wherein The filtering processing in step 2) includes at least any one of mean filtering, median filtering, Gaussian filtering, Savitzky-Golay (SG) filtering, and mean smoothing filtering.
4. The precise annotation method for microscopic hyperspectral pathological images according to claim 3, characterized in that In step 5), the types of labels for annotation include at least any one of background, healthy tissue, blood vessel, and diseased cell nucleus. Each label type has zero or more clustering types, and each clustering type should be assigned to one and only one label type.
5. The precise annotation method for microscopic hyperspectral pathological images according to claim 4, wherein In step 6), if the manual evaluation passes or only minor adjustments are required, proceed to the next step or generate the label for the next hyperspectral image; if the manual evaluation fails, return to step 2), and reselect the filtering operation in step 2) and K in step 4), or inquire whether the rough label in step 1) is incorrect.
Citation Information
Patent Citations
A semi-automatic method for building a hyperspectral database based on deep learning
CN107704878B
Automatic Generation Method of Highly Annotated Digital Pathology Datasets Based on Hyperspectral Imaging
CN114596298B
Method and system for generating benchmark pathology datasets containing automatically labeled files
CN115274093B
Cited By
Automatic lung slice cell nucleus segmentation method combining microscopic hyperspectrum and artificial intelligence
CN121213586A
Thyroid hyperspectral image risk prediction method based on quality control and uncertainty
CN122435417A