A method for processing immunohistochemical digital pathology images

By preprocessing, feature extraction and fusion of pathological images of immunohistochemical staining, and combining region growth and merging algorithms and K-means clustering algorithms, the shortcomings of staining intensity detection and cell region segmentation in the existing technology are solved, and high-precision staining region identification and classification analysis are achieved.

CN119181091BActive Publication Date: 2025-05-13QIAGEN SUZHOU TRANSLATIONAL MEDICINE CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411697291.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-13
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

The existing immunohistochemical image processing technology has shortcomings in staining intensity detection and precise segmentation of cell areas, making it difficult to deal with image noise and staining inhomogeneity, resulting in low recognition accuracy.

Method used

The pathological images of immunohistochemical staining were obtained for pre-processing, including Gaussian filtering, histogram equalization and median filtering; then the image was divided into regions, extracted and fused with multi-dimensional feature vectors, and optimized images using region growth and merging algorithms, and classified and quantitative analysis of stained regions were combined with K-means clustering algorithm.

Benefits of technology

Accurate segmentation and classification analysis of stained areas are realized, the recognition accuracy and coherence of stained areas are improved, and the automation and accuracy of the entire pathological image processing are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181091B_ABST
    Figure CN119181091B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for processing immunohistochemical digital pathological images, which relates to the technical field of image processing, including obtaining immunohistochemically stained pathological images and performing preprocessing; performing regional division on the preprocessed pathological images, extracting and fusing features for each sub-region, and generating a multidimensional feature vector; optimizing the pathological images based on the multidimensional feature vectors by using a regional growing and merging algorithm; performing staining intensity detection on the optimized pathological images, and classifying the stained regions by a K-means clustering algorithm; performing quantitative analysis and generating an immunohistochemical image analysis report based on the classification results. The present invention realizes accurate segmentation and classification analysis of immunohistochemically stained regions through preprocessing, feature extraction and fusion, and combines regional growing and merging algorithms, thereby improving the automation and recognition accuracy of image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of image processing, in particular to an immunohistochemical digital pathology image processing method. Background Art

[0002] Immunohistochemistry (IHC) is a practical biological experimental technique that uses the principle of antigen-antibody specific binding, enzymes, proteins or chemical dyes to display macromolecular proteins or gene expression in vivo on tissue sections or cell slices in situ. It is widely used in the pathological diagnosis of various diseases such as tumors and inflammation. With the development of digital pathology technology, traditional microscope observation has gradually been replaced by digital pathological images, and the demand for automated pathological image analysis is also growing. Existing immunohistochemical image processing technology mainly relies on manual annotation and semi-automatic analysis based on simple image processing methods. However, manual annotation is time-consuming and labor-intensive and easily affected by subjective factors, resulting in poor accuracy and consistency of analysis results. Automated immunohistochemical image analysis technology has begun to be used in clinical practice, but due to the complexity of images and the diversity of staining effects, existing automated methods still have many technical bottlenecks in dealing with image noise, accurately segmenting cell areas, and quantitative analysis of staining intensity.

[0003] Existing immunohistochemical image processing technologies have deficiencies in staining intensity detection and accurate segmentation of cell regions. On the one hand, existing methods often rely on simple threshold segmentation techniques, which are difficult to deal with uneven staining and noise effects, resulting in low recognition accuracy of stained areas; on the other hand, existing staining intensity analysis methods are mostly based on single-channel brightness or simple color feature extraction, ignoring the multidimensional information fusion of the staining area, and cannot fully reflect the true staining state of cells in pathological images. Therefore, in order to improve the accuracy and efficiency of immunohistochemical digital pathology image processing, an automated processing method that can comprehensively consider the multidimensional feature information of the image is urgently needed to achieve accurate segmentation of the staining area and quantitative analysis of the staining intensity. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an immunohistochemical digital pathology image processing method to solve the problems of low staining area recognition accuracy and inaccurate staining intensity analysis.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an immunohistochemical digital pathology image processing method, which includes obtaining an immunohistochemically stained pathology image and performing preprocessing; dividing the preprocessed pathology image into regions, extracting and fusing features of each sub-region, and generating a multidimensional feature vector; optimizing the pathology image based on the multidimensional feature vector by using a region growing and merging algorithm; detecting the staining intensity of the optimized pathology image, and classifying the staining regions by using a K-means clustering algorithm; and performing quantitative analysis based on the classification results and generating an immunohistochemical image analysis report.

[0008] As a preferred solution of the immunohistochemical digital pathological image processing method of the present invention, the specific steps of obtaining the immunohistochemical stained pathological image and preprocessing it are as follows:

[0009] The immunohistochemically stained pathological images were acquired through a high-resolution digital pathology scanner and converted into an analyzable digital image format;

[0010] Use a Gaussian filter to apply a weighted average convolution kernel to the image to reduce random noise in the image while preserving edge details;

[0011] Enhance the contrast of the image through histogram equalization;

[0012] Use a median filter to remove small spots and texture noise in the image, reduce false positives, and protect the edge information of the image.

[0013] As a preferred solution of the immunohistochemical digital pathological image processing method of the present invention, the specific steps of dividing the pre-processed pathological image into regions are as follows:

[0014] The preprocessed pathological image is recursively divided into several sub-regions of equal size using the quadtree segmentation algorithm. The non-uniformity index is calculated by linearly combining the variance of color and brightness and the texture metric of each sub-region. The expression is:

[0015] ;

[0016] in, represents the color variance of the ith sub-region, is the adjustment coefficient of the color variance of the ith sub-region, represents the texture metric of the ith sub-region, m is the adjustment coefficient of the texture metric of the ith sub-region, represents the brightness variance of the ith sub-region, n is the adjustment coefficient of the brightness variance of the ith sub-region, is the non-uniformity index of the ith sub-region;

[0017] Based on the heterogeneity of the divided areas in the historical segmentation results, a heterogeneity standard threshold is defined , when the inhomogeneity of a sub-region Reaching the non-uniformity threshold Stop dividing.

[0018] As a preferred solution of the immunohistochemical digital pathology image processing method of the present invention, wherein: the feature extraction and fusion of each sub-region to generate a multi-dimensional feature vector is performed, and the specific steps are as follows:

[0019] The Otsu algorithm was used to binarize the sub-regions of each pathological image to distinguish the stained area from the background, and the watershed algorithm was used to extract the morphological features in the pathological image;

[0020] Through the gray level co-occurrence matrix, the gray level co-occurrence relationship of pixels in the sub-region of each pathological image is identified, and the texture features in the pathological image are extracted;

[0021] Use histogram analysis to count the color distribution of images and extract color features in pathological images;

[0022] The morphological features, texture features and color features of each sub-region are nonlinearly fused to generate a multi-dimensional feature vector, which is expressed as:

[0023] ;

[0024] in, is the morphological feature vector of the ith sub-region, is the texture feature vector of the ith sub-region, is the color feature vector of the ith sub-region, is the multidimensional feature vector of the i-th sub-region.

[0025] As a preferred solution of the immunohistochemical digital pathology image processing method of the present invention, wherein: based on the multidimensional feature vector, the pathology image is optimized by using the region growing and merging algorithm, and the specific steps are as follows:

[0026] The similarity between each sub-region and its adjacent sub-region is calculated by nonlinear combination, and the expression is:

[0027] ;

[0028] in, is the color feature vector of the current growing area, is the color feature vector of the adjacent sub-region, is the texture feature vector of the current region, is the texture feature vector of the adjacent sub-region, is the morphological feature vector of the current growth area, is the morphological feature vector of the adjacent sub-region, Indicates the similarity between the current sub-region and the adjacent sub-region;

[0029] Based on the similarity of the regional features that were successfully segmented during the historical segmentation process, a similarity standard threshold τ is defined;

[0030] Select the similarity between the current sub-region and the adjacent sub-region The lowest sub-region is used as the seed point;

[0031] Starting from the selected seed point, based on the color features, texture features, and morphological features in each sub-region, an initial growth region is generated at the seed point, and the adjacent sub-regions around the seed point are checked step by step in an incremental manner, and the growth region is recursively expanded from the inside to the outside, and the similarity with the current growth region is calculated. Adjacent sub-regions with a similarity standard threshold τ are gradually merged into the current growing region;

[0032] Through region growing and region merging, the boundary clarity of the stained area and the background area in the pathological image is improved, the coherence of the stained area is enhanced, and finally an optimized pathological image is generated.

[0033] As a preferred solution of the immunohistochemical digital pathological image processing method of the present invention, wherein: the staining intensity detection of the optimized pathological image is performed, and the specific steps are as follows:

[0034] Convert the optimized pathological image from RGB color space to Lab color space;

[0035] Separate the color value distribution of a channel and b channel from the pathological image in Lab color space;

[0036] Analyze the pixels in the a channel and the b channel, and calculate the difference between each pixel and the reference color. The expression is:

[0037] ;

[0038] in, Represents pixels in Lab color space The difference from the reference color, Represents pixels The a channel value in the Lab color space, Represents pixels The b channel value in the Lab color space, Represents the mean value of the a channel of the reference color m, represents the mean value of the b channel of the reference color m, x represents the horizontal coordinate of the pixel in the pathological image, and y represents the vertical coordinate in the pathological image;

[0039] The reference color is defined based on known dyed area colors in historical data;

[0040] Based on color difference value , the optimized pathological image is initially binarized, and the optimal color difference value is automatically calculated based on Otsu's by maximizing the inter-class variance between foreground and background ;

[0041] when ≥ When , it is marked as the positive staining area;

[0042] when < When , it is marked as negative staining area.

[0043] As a preferred solution of the immunohistochemical digital pathology image processing method of the present invention, the staining area is classified by the K-means clustering algorithm, and the specific steps are as follows:

[0044] For each pixel in the positive staining area, extract the L channel, a channel, and b channel of each pixel in the positive staining area, and generate a three-dimensional feature vector based on the coordinates of each pixel in each positive staining area;

[0045] According to the total number of classification types, the number of clusters is set to k, and k centroids are randomly initialized;

[0046] Based on the three-dimensional feature vector, the Euclidean distance from each pixel to the centroid is calculated by the K-means clustering algorithm, and each pixel is assigned to the nearest centroid;

[0047] Based on the assignment results, the centroid position of each cluster was recalculated until the centroid position no longer changed, and finally the pixels in the positive staining area were divided into weak positive, moderate positive, and strong positive.

[0048] As a preferred solution of the immunohistochemical digital pathology image processing method of the present invention, wherein: based on the classification results, quantitative analysis is performed and an immunohistochemical image analysis report is generated. The specific steps are as follows:

[0049] Based on the classification results, the color deconvolution algorithm was used to analyze the area proportion and staining intensity of each category in the optimized pathological image according to the number of pixels in each positive staining category and the total number of pixels in the positive staining area;

[0050] Generate an immunohistochemistry image analysis report based on the quantitative analysis results.

[0051] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the immunohistochemical digital pathology image processing method as described in the first aspect of the present invention is implemented.

[0052] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the immunohistochemical digital pathology image processing method as described in the first aspect of the present invention is implemented.

[0053] The beneficial effects of the present invention are as follows: by preprocessing, feature extraction and fusion of immunohistochemically stained pathological images, and optimizing the images in combination with the region growing and merging algorithm, accurate segmentation and classification analysis of the stained regions are achieved. Among them, the multidimensional feature vector generated by feature fusion effectively improves the recognition accuracy of stained regions such as cell membranes and cell nuclei, while the region growing and merging algorithm enhances the coherence and boundary clarity of the stained regions. This makes the entire pathological image processing process more automated and accurate, meeting the clinical needs for efficient and accurate analysis of positive and negative cells in immunohistochemical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0055] Figure 1 This is a flow chart of the immunohistochemical digital pathology image processing method in Example 1.

[0056] Figure 2 This is a flow chart of classifying stained areas using the K-means clustering algorithm in Example 1. DETAILED DESCRIPTION

[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0059] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0060] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides an immunohistochemical digital pathology image processing method, comprising the following steps:

[0061] S1: Obtain pathological images of immunohistochemical staining and perform preprocessing.

[0062] S1.1: Obtain immunohistochemically stained pathological images using a high-resolution digital pathology scanner and convert them into analyzable digital image formats (such as TIFF, SVS, JPEG);

[0063] Furthermore, pathological images include images of tissue sections or cell samples acquired through a microscope or a digital scanner, which are used to diagnose diseases or study pathological characteristics.

[0064] S1.2: Apply a weighted average convolution kernel to the image using a Gaussian filter to reduce random noise in the image while preserving cell edge details;

[0065] Specifically, when the filter is applied to the cell region, it can effectively remove noise and spots in the background while keeping the stained nucleus and cell membrane outlines clear, thus ensuring that the cell region can be accurately identified during the subsequent segmentation and feature extraction process.

[0066] S1.3: Enhance the contrast of the image by histogram equalization;

[0067] Furthermore, the brightness difference between the stained area and the background is made more obvious, thereby improving the visibility of the image;

[0068] S1.4: Use a median filter to remove small spots and texture noise in the image, reduce misjudgment, and protect the edge information of the image.

[0069] For example, when processing immunohistochemically stained pathological images, the median filter can effectively remove small spots, texture noise, and staining background caused by uneven staining or sample preparation without blurring important cell boundaries.

[0070] S2: Divide the preprocessed pathological image into regions, extract and fuse features from each sub-region, and generate a multi-dimensional feature vector.

[0071] S2.1: The preprocessed pathological image is recursively divided into several sub-regions of equal size using the quadtree segmentation algorithm. The non-uniformity index is calculated by linearly combining the variance of color and brightness and the texture metric of each sub-region. The expression is:

[0072] ;

[0073] in, represents the color variance of the ith sub-region, is the adjustment coefficient of the color variance of the ith sub-region, represents the texture metric of the ith sub-region, m is the adjustment coefficient of the texture metric of the ith sub-region, represents the brightness variance of the ith sub-region, n is the adjustment coefficient of the brightness variance of the ith sub-region, is the non-uniformity index of the ith sub-region;

[0074] S2.1.1: Based on the heterogeneity of the divided areas in the historical segmentation results, define a heterogeneity standard threshold , when the inhomogeneity of a sub-region Reaching the non-uniformity threshold Stop dividing.

[0075] It should be noted that, first, the entire image is regarded as a large area. If the features (such as color, brightness, texture, etc.) in the area are not uniform enough, the area is divided into four sub-areas of equal size; then the same process is repeated for each sub-area to check its uniformity. If a sub-area still does not meet the uniformity standard, it is further divided into four smaller sub-areas, and so on recursively until the features of all sub-areas meet the preset uniformity conditions or reach the minimum allowable division size. Through this recursive division method, regions with different features in the image can be finely segmented, laying the foundation for subsequent feature extraction and analysis.

[0076] S2.2: Use the Otsu algorithm to binarize the sub-regions of each pathological image to distinguish the stained area (cells or tissues) from the background, and use the watershed algorithm to extract the morphological features in the pathological image;

[0077] Specifically, the Otsu algorithm is used to binarize the sub-regions of each pathological image, which can automatically calculate the global optimal threshold and accurately distinguish the stained area (cells or tissues) from the background. Subsequently, the morphological features in the pathological image are further extracted by the watershed algorithm. The watershed algorithm uses the binarization results to segment the image, which is particularly suitable for processing complex cell boundaries. By marking the foreground and background areas, it effectively separates adjacent cells or tissue structures, thereby accurately extracting the morphological contours of the stained area.

[0078] S2.3: Identify the gray level co-occurrence relationship of pixels in the sub-region of each pathological image through the gray level co-occurrence matrix, and extract the texture features in the pathological image;

[0079] For example, when analyzing immunohistochemically stained pathological images, the gray-level co-occurrence matrix can be used to calculate the texture features of the cell area. Assuming that the distance is 1 pixel and the direction is 0 degrees (horizontal) to construct the gray-level co-occurrence matrix, the gray value combination of adjacent pixels in the image can be counted, such as the frequency of co-occurrence of pixels with a gray value of 50 and pixels with a gray value of 100. The calculated gray-level co-occurrence matrix can be used to extract texture features such as contrast, homogeneity, and entropy.

[0080] Use histogram analysis to count the color distribution of images and extract color features in pathological images;

[0081] For example, when processing immunohistochemically stained pathological images, the distribution of pixels of different colors can be counted by analyzing the color histogram of the image. First, the image is converted from the RGB color space to a color space more suitable for analysis (such as Lab or HSV), and then a color histogram is generated on each channel (such as the L channel for brightness, the a channel for red and green distribution, and the b channel for blue and yellow distribution). Each bar of the histogram represents the number of pixels within a certain color range.

[0082] The morphological features, texture features and color features of each sub-region are nonlinearly fused to generate a multi-dimensional feature vector, which is expressed as:

[0083] ;

[0084] in, is the morphological feature vector of the ith sub-region, is the texture feature vector of the ith sub-region, is the color feature vector of the ith sub-region, is the multidimensional feature vector of the i-th sub-region.

[0085] It should be noted that the quadtree segmentation algorithm is used to recursively divide the stained images of pathological sections into sub-regions with more uniform features, ensuring the fine segmentation of cells in the stained area, which is a key innovation. Then, the texture features are extracted by gray-level co-occurrence matrix and the color features are extracted by color histogram analysis, and the multi-dimensional feature vector of specific cells is generated by combining the nonlinear fusion of morphological features. This multi-dimensional feature fusion strategy can fully capture the complex information of specific cell staining images, which not only improves the recognition accuracy of the staining area, but also provides a more accurate basis for the subsequent optimization and classification of specific cell staining images, significantly improving the reliability and automation of pathological image analysis.

[0086] S3: Based on the multidimensional feature vector, the pathological image is optimized using the region growing and merging algorithm.

[0087] S3.1: The similarity between each sub-region and its adjacent sub-regions is calculated by nonlinear combination, and the expression is:

[0088] ;

[0089] in, is the color feature vector of the current dyeing area, is the color feature vector of the adjacent sub-region, is the texture feature vector of the current region, is the texture feature vector of the adjacent sub-region, is the morphological feature vector of the current growth area, is the morphological feature vector of the adjacent sub-region, Indicates the similarity between the current sub-region and the adjacent sub-region;

[0090] It should be noted that the differences in color features, texture features, and morphological features between the cells in the current growing region and the adjacent sub-regions are calculated by nonlinear combination. Finally, these differences are combined into an overall similarity S, which is used to evaluate the similarity between the cells in the adjacent sub-regions and the current growing region, thereby guiding the regional cell growth and merging process. This process ensures that the weights of different features in the image are reasonably integrated, thereby improving the accuracy of segmentation.

[0091] S3.2: Based on the similarity of the regional features that were successfully segmented in the historical segmentation process, a similarity standard threshold τ is defined;

[0092] It should be noted that the feature similarity of the stained regions successfully segmented in the historical segmentation process refers to the statistical results of the feature similarity between the cell regions that have been correctly segmented in the previous stained image segmentation process. These successfully segmented cell regions usually have similar color, texture and morphological features. By recording the feature vectors of these cell regions and the similarities between them, an empirical similarity threshold can be summarized.

[0093] S3.3: Select the similarity between a current sub-region and an adjacent sub-region The lowest sub-region is used as the seed point;

[0094] It should be noted that the sub-region with the lowest similarity S is selected as the seed point in order to preferentially expand to the region closest to the characteristics of the current growing region, ensuring that the uniformity of the internal characteristics of the region can be maintained during the region growth process, thereby improving the accuracy of cell segmentation.

[0095] S3.4: Starting from the selected seed point, based on the color features, texture features, and morphological features in each sub-region, an initial growing region is generated at the seed point, and the adjacent sub-regions around the seed point are checked step by step in an incremental manner, and the growing region is recursively expanded from the inside to the outside, and the similarity with the current growing region is extended. Adjacent sub-regions with a similarity standard threshold τ are gradually merged into the current growing region;

[0096] For example, in pathological image segmentation, assuming that the seed point is located in a stained cell membrane, first check its four directly adjacent sub-regions (upper, lower, left, and right) and calculate their color, texture, and morphological similarity with the current growing region. If the similarity meets the threshold condition, these sub-regions will be merged into the growing region. Then, continue to check the adjacent regions of these newly merged sub-regions, expanding outward layer by layer, similar to the process of ripple diffusion, until there are no adjacent regions that meet the similarity criteria.

[0097] Through region growing and region merging, the boundary clarity between the cell staining area and the background area in the pathological image is improved, the coherence of the staining area is enhanced, and finally an optimized pathological image is generated.

[0098] S4: The optimized pathological image is subjected to staining intensity detection, and the cells in the stained area are classified using the K-means clustering algorithm.

[0099] S4.1: converting the optimized pathological image from RGB color space to Lab color space;

[0100] It should be noted that the reason for converting to the Lab color space is that this space can better separate the brightness information (L channel) and color information (a and b channels) of the image, thereby better analyzing the staining intensity;

[0101] L channel: represents the brightness of the image, ranging from black to white;

[0102] a channel: represents the color distribution from green to red, positive values ​​represent reddishness, and negative values ​​represent greenishness;

[0103] b channel: represents the color distribution from blue to yellow, positive values ​​represent yellowish color, and negative values ​​represent bluish color;

[0104] S4.2: Separate the color value distribution of a channel and b channel from the pathological image in Lab color space;

[0105] For example, the pathological image is first converted from the RGB color space to the Lab color space, where the L channel represents brightness, the a channel represents the color distribution from green to red, and the b channel represents the color distribution from blue to yellow. After separating the a channel, the distribution of green and red pixels in the image can be obtained, while after separating the b channel, the color distribution of blue and yellow can be seen.

[0106] S4.3: Analyze the pixels in the a channel and the b channel, and calculate the difference between each pixel and the reference color. The expression is:

[0107] ;

[0108] in, Represents pixels in Lab color space The difference from the reference color, Represents pixels The a channel value in the Lab color space, Represents pixels The b channel value in the Lab color space, Represents the mean value of the a channel of the reference color m, represents the mean value of the b channel of the reference color m, x represents the horizontal coordinate of the pixel in the pathological image, and y represents the vertical coordinate in the pathological image;

[0109] It should be noted that this expression is calculated by calculating the pixel The Euclidean distance between the a-channel and b-channel values ​​of the reference color and the mean of the a-channel and b-channel values ​​of the reference color is used to measure the similarity between the pixel color and the reference color. The smaller the difference value, the closer the pixel color is to the reference color.

[0110] S4.4: Reference colors are defined based on the specific colors of the cell membrane, cytoplasm, or nucleus of known stained regions in historical data;

[0111] Furthermore, the reference color is defined based on the specific color of the known stained areas (such as cell membrane, cytoplasm or nucleus) in the historical data. By analyzing the color characteristics of these known areas (such as the distribution in HSV or Lab color space), these reference colors can be used for subsequent cell detection using the OpenCV algorithm. At the same time, the OpenCV algorithm's color space conversion, threshold segmentation, and contour detection functions combined with the reference color can effectively identify and separate the cell structure in the stained area, ensuring that the detected cell area matches the staining in the historical data.

[0112] S4.5: Based on color difference value , the optimized pathological image is initially binarized, and the optimal color difference value is automatically calculated based on Otsu's by maximizing the inter-class variance between foreground and background ;

[0113] Furthermore, the optimal color difference value The expression is:

[0114] ;

[0115] in, is the optimal color difference value L, L represents the candidate color difference value, is the average color difference value of the negative area, is the average color difference value of the positive area;

[0116] S4.5.1: When ≥ When , it is marked as the positive staining area;

[0117] Furthermore, the positive staining area is a disease-related marker area, such as positive cells or lesion areas in immunohistochemical staining.

[0118] S4.5.2: When < When , it is marked as negative staining area.

[0119] Furthermore, the negatively stained area corresponds to normal cells or tissues that do not express the corresponding biomarkers, and is usually represented as background or non-lesion area in pathological images.

[0120] S4.6: for each pixel of the cells in the positive staining area, extract the L channel a channel and the L channel b channel of each pixel of the cells in the positive staining area, and generate a three-dimensional feature vector based on the coordinates of each pixel of the cells in each positive staining area;

[0121] Specifically, L channel: brightness, a channel: color values ​​from green to red, b channel: color values ​​from blue to yellow;

[0122] S4.7: According to the total number of classification types, set the number of clusters to k and randomly initialize k centroids;

[0123] S4.8: Based on the three-dimensional feature vector, the Euclidean distance from each pixel to the centroid is calculated by the K-means clustering algorithm, and each pixel is assigned to the nearest centroid;

[0124] It should be noted that the Euclidean distance from each pixel to the centroid is expressed as:

[0125] ;

[0126] Among them, represents the pixel The Euclidean distance to the centroid, is the horizontal coordinate of the center of mass, is the ordinate of the center of mass;

[0127] S4.9: Based on the allocation results, the centroid position of each cluster is recalculated until the centroid position no longer changes, and finally the pixels in the positive staining area are divided into weak positive, moderate positive and strong positive.

[0128] It should be noted that the centroid position of each cluster is recalculated based on the allocation result, which means that in each iteration, after the pixels are assigned to the closest centroid according to their Euclidean distance, the average position of all pixels in the cluster is calculated and the coordinates of the centroid are updated. This process continues until the centroid position no longer changes, indicating that the clustering converges. Finally, the cell pixels in the positive staining area are divided into weak positive, medium positive and strong positive according to their distance from different centroids (representing different positive intensities). This method ensures a more accurate grading of staining intensity and reflects the distribution of different staining intensities in pathological images.

[0129] S5: Based on the classification results, perform quantitative analysis and generate an immunohistochemistry image analysis report.

[0130] S5.1: Based on the classification results, the color deconvolution algorithm is used to analyze the area proportion and staining intensity of each category in the optimized pathological image according to the number of pixels in each positive staining category and the total number of pixels in the positive staining area;

[0131] Specifically, first, the different staining components are separated through the deconvolution algorithm, the number of pixels in each category (weakly positive, moderately positive, and strongly positive) is calculated, and compared with the total number of pixels in the positive staining area to obtain the area proportion of each category. At the same time, the color value or staining intensity of each type of pixel is counted to calculate the average staining intensity of each category. This analysis can intuitively show the distribution and proportion of different staining intensities in pathological images, which helps to assess the scope and severity of lesions and provide a more accurate basis for pathological diagnosis.

[0132] Furthermore, the quantitative analysis also includes evaluating the number of cells based on the area of ​​positive staining, which specifically includes the following steps:

[0133] Based on the segmented stained areas, the morphological characteristics of each independent cell (such as area, perimeter, circularity, aspect ratio, etc.) were extracted by performing connected domain analysis on each positive staining area. The area of ​​each cell was corrected according to the morphological feature vector to ensure that the morphological differences of cells can be reflected in the final cell number estimation.

[0134] Based on the area proportion of each positive staining area (weakly positive, moderately positive, and strongly positive) and the morphological feature vector of the cell, the number of cells in the positive staining area is predicted by combining the morphological correction factor and the staining intensity correction factor of each cell. The expression is:

[0135] ;

[0136] in, is the total number of cells in the positively stained area, is the actual area of ​​the jth positive staining region, is a known reference average cell area (based on historical data or observation under a microscope), is the morphological feature vector of cells in the jth stained area, is the morphology correction factor, is the staining intensity factor of the jth positive staining area, which is used to correct the influence of different staining intensities (weak positive, medium positive, strong positive) on the estimation of cell area, and z is the total number of positive staining areas;

[0137] It should be noted that It is a morphological correction factor based on the morphological feature vector of the cell. For example, the closer the morphology of a cell is to an ideal circle or a specific aspect ratio, the better the morphological correction factor. The closer it is to 1, if the morphology deviates from the ideal state, the correction factor will be adjusted accordingly;

[0138] Staining intensity correction factor It is obtained by comparing the staining intensity of each cell with a reference intensity, e.g. It can be obtained by calculating the distance between the staining intensity of each cell and the reference staining intensity (such as Euclidean distance). The smaller the distance, the closer the staining intensity is to the ideal value and the larger the correction factor.

[0139] S5.2: Generate an immunohistochemistry image analysis report based on the quantitative analysis results.

[0140] Furthermore, the immunohistochemistry image analysis report can accurately reflect the positive cells and negative cells and the lesion-free areas in the pathological images, helping pathologists to objectively evaluate the staining effect and avoid subjective errors. At the same time, the report provides visual charts and data to intuitively display the number and proportion of specific positive and negative cells, thereby improving the accuracy and efficiency of diagnosis and providing data support for the formulation of personalized treatment plans. In addition, this standardized reporting method is also helpful for long-term disease monitoring and treatment effect evaluation.

[0141] This embodiment also provides a computer device, which is suitable for the immunohistochemical digital pathology image processing method, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the immunohistochemical digital pathology image processing method proposed in the above embodiment.

[0142] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0143] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for processing immunohistochemical digital pathology images proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0144] In summary, the present invention achieves accurate segmentation and classification analysis of specific cells in the stained area by: preprocessing, feature extraction and fusion of immunohistochemical stained pathological images, and optimizing the images in combination with region growing and merging algorithms. Among them, the multidimensional feature vector generated by feature fusion effectively improves the recognition accuracy of specific cells in the stained area, while the region growing and merging algorithm enhances the coherence and boundary clarity of specific cells in the stained area. This makes the entire pathological image processing process more automated and accurate, meeting the clinical demand for efficient and accurate immunohistochemical image analysis.

[0145] Example 2, referring to Table 1, is the second example of the present invention. To further verify the technical solution of the present invention, experimental simulation data of the immunohistochemical digital pathology image processing method are provided.

[0146] The experimental preparation stage includes obtaining sections from pathological tissues and performing immunohistochemical staining, followed by obtaining immunohistochemically stained pathological images through a high-resolution digital pathology scanner and preprocessing them. First, the pathological images were converted into analyzable digital image formats. In order to reduce the noise in the image, the image was processed using a Gaussian filter, and a 5×5 convolution kernel was used to remove random noise in the image, but at the same time, the edge information in the image was retained to ensure that the stained cell nucleus or cell membrane outline was clearly visible. Then, the image contrast was further enhanced by histogram equalization, making the brightness difference between the stained area and the background more obvious, thereby enhancing the visibility of the image.

[0147] In order to better remove small speckles and texture noise generated by the sample preparation process, a median filter was used. The median filter uses a 3×3 window size to effectively remove small noise in the background while retaining the boundary information in the pathological image, making the outline of the cell nucleus or cell membrane clearer.

[0148] Next, the quadtree segmentation algorithm is used to partition the preprocessed image. The image is recursively divided into several sub-regions of equal size, and the non-uniformity index is calculated step by step based on the color variance, brightness variance and texture information of each sub-region. Non-uniformity threshold =0.85 is defined as the criterion for stopping division. When a sub-region reaches or exceeds this threshold, further division of the region is stopped.

[0149] Subsequently, the Otsu threshold method was used to binarize each sub-region to distinguish the stained area from the background area. After that, the morphological features of the pathological image were extracted by the Canny edge detection algorithm, and the texture features of each sub-region were extracted using the gray-level co-occurrence matrix. The color features in the pathological image were extracted by combining the color histogram analysis method. Finally, a multidimensional feature vector of the sub-region was generated based on these features. .

[0150] In the optimized pathological images, the positive staining areas were classified by the K-means clustering algorithm. The pathological images were first converted to the Lab color space, and the color values ​​of the a channel and the b channel were separated. Based on the difference value between each pixel and the reference color, the Euclidean distance was calculated, and the specific cell staining pixels in the pathological images were divided into three categories: weak positive, medium positive, and strong positive.

[0151] The experimental preparation stage includes obtaining sections from pathological tissues, performing immunohistochemical staining, and acquiring pathological images through a high-resolution digital pathology scanner. First, the images are converted into an analyzable digital format and Gaussian filtering is used to remove noise while retaining the outline of the cell nucleus or cell membrane. Then, histogram equalization enhances the image contrast, making the stained area more distinct from the background.

[0152] In order to further remove the spots and texture noise generated during the sample preparation process, a median filter was used to retain the edge information of the image. Next, the image was divided into regions using the quadtree segmentation algorithm, and the non-uniformity index was calculated based on the color variance, brightness variance, and texture features of each sub-region. When a sub-region reaches the preset non-uniformity threshold =0.85, stop dividing.

[0153] Subsequently, each sub-region was binarized using the Otsu threshold method to distinguish the stained area from the background area. The morphological features were extracted by Canny edge detection, and the texture and color features were extracted by combining the gray-level co-occurrence matrix and color histogram to generate a multi-dimensional feature vector for each sub-region.

[0154] Finally, the optimized images were classified using the K-means clustering algorithm. The images were converted to Lab color space, the Euclidean distance of the pixels was calculated, and they were divided into three categories: weak positive, moderate positive, and strong positive, completing the classification analysis of the positive staining area.

[0155] The details are shown in Table 1 below:

[0156] Table 1 Immunohistochemical staining image analysis table

[0157]

[0158] By analyzing the data in the above table, it can be clearly seen that there is a certain positive correlation between the total number of pixels in the positive staining area and the staining intensity. As the sum of the number of weakly positive, moderately positive, and strongly positive pixels increases, the staining intensity also increases accordingly. For example, image 5 has a total number of positive pixels of 3853 and a staining intensity of 0.772, which is the image with the highest staining intensity in the table; in contrast, image 1 has a total number of positive pixels of 2458 and a staining intensity of 0.647.

[0159] In addition, from the distribution of weakly positive, moderately positive, and strongly positive pixel numbers, it can be seen that weakly positive pixels usually occupy most of the positive area, while strongly positive pixels are relatively few. This trend indicates that the staining intensity of most stained areas is relatively moderate or weak, while the proportion of strongly stained areas is relatively low. This is consistent with the common phenomenon that weakly positive areas are more common and strongly positive areas are relatively rare in immunohistochemical staining.

[0160] Through the immunohistochemical digital pathology image processing method of the present invention, not only can the image noise be accurately removed, but also the positive staining area can be effectively classified, and weak positive, medium positive and strong positive areas can be distinguished. This reflects the efficiency and accuracy of the present invention in automated pathology image processing, and provides a more objective and reliable basis for pathological diagnosis.

[0161] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for processing immunohistochemical digital pathology images, characterized in that: include, Obtain pathological images of immunohistochemical staining and perform preprocessing; The preprocessed pathological image is divided into regions, and features are extracted and fused for each sub-region to generate a multi-dimensional feature vector; Based on multi-dimensional feature vectors, the pathological images are optimized using region growing and merging algorithms; The optimized pathological images were tested for staining intensity, and the staining areas were classified using the K-means clustering algorithm; Based on the classification results, quantitative analysis is performed and an immunohistochemistry image analysis report is generated; The specific steps of dividing the preprocessed pathological image into regions are as follows: The preprocessed pathological image is recursively divided into several sub-regions of equal size using the quadtree segmentation algorithm. The non-uniformity index is calculated by linearly combining the variance of color and brightness and the texture metric of each sub-region. The expression is: ; in, represents the color variance of the ith sub-region, is the adjustment coefficient of the color variance of the ith sub-region, represents the texture metric of the ith sub-region, m is the adjustment coefficient of the texture metric of the ith sub-region, represents the brightness variance of the ith sub-region, n is the adjustment coefficient of the brightness variance of the ith sub-region, is the non-uniformity index of the ith sub-region; Based on the heterogeneity of the divided areas in the historical segmentation results, define the heterogeneity standard threshold , when the inhomogeneity of a sub-region Reaching the non-uniformity threshold Stop dividing when The feature extraction and fusion of each sub-region to generate a multi-dimensional feature vector is performed in the following specific steps: The Otsu algorithm was used to binarize the sub-regions of each pathological image to distinguish the stained area from the background, and the watershed algorithm was used to extract the morphological features in the pathological image; Through the gray level co-occurrence matrix, the gray level co-occurrence relationship of pixels in the sub-region of each pathological image is identified, and the texture features in the pathological image are extracted; Use histogram analysis to count the color distribution of images and extract color features in pathological images; The morphological features, texture features and color features of each sub-region are nonlinearly fused to generate a multi-dimensional feature vector, which is expressed as: ; in, is the morphological feature vector of the ith sub-region, is the texture feature vector of the ith sub-region, is the color feature vector of the ith sub-region, is the multidimensional feature vector of the i-th sub-region; The pathological image is optimized by using a region growing and merging algorithm based on a multi-dimensional feature vector. The specific steps are as follows: The similarity between each sub-region and its adjacent sub-region is calculated by nonlinear combination, and the expression is: ; in, is the color feature vector of the current growing area, is the color feature vector of the adjacent sub-region, is the texture feature vector of the current region, is the texture feature vector of the adjacent sub-region, is the morphological feature vector of the current growth area, is the morphological feature vector of the adjacent sub-region, Indicates the similarity between the current sub-region and the adjacent sub-region; Based on the similarity of the regional features that were successfully segmented during the historical segmentation process, a similarity standard threshold τ is defined; Select the similarity between the current sub-region and the adjacent sub-region The lowest sub-region is used as the seed point; Starting from the selected seed point, based on the color features, texture features, and morphological features in each sub-region, an initial growth region is generated at the seed point, and the adjacent sub-regions around the seed point are checked step by step in an incremental manner, and the growth region is recursively expanded from the inside to the outside, and the similarity with the current growth region is expanded. Adjacent sub-regions with a similarity standard threshold τ are gradually merged into the current growing region; Through region growing and region merging, the boundary clarity of the stained area and the background area in the pathological image is improved, the coherence of the stained area is enhanced, and finally an optimized pathological image is generated.

2. The immunohistochemical digital pathology image processing method according to claim 1, characterized in that: The specific steps of obtaining the immunohistochemical staining pathological image and preprocessing are as follows: The immunohistochemically stained pathological images were acquired through a high-resolution digital pathology scanner and converted into an analyzable digital image format; Use a Gaussian filter to apply a weighted average convolution kernel on the image; Enhance the contrast of the image through histogram equalization; Use a median filter to remove small speckle and texture noise from the image.

3. The immunohistochemical digital pathology image processing method according to claim 2, characterized in that: The staining intensity detection of the optimized pathological image is performed in the following specific steps: Convert the optimized pathological image from RGB color space to Lab color space; Analyze the pixels in the a channel and the b channel, and calculate the difference between each pixel and the reference color. The expression is: ; in, Represents pixels in Lab color space The difference from the reference color, Represents pixels The a channel value in the Lab color space, Represents pixels The b channel value in the Lab color space, Represents the mean value of the a channel of the reference color m, represents the mean value of the b channel of the reference color m, x represents the horizontal coordinate of the pixel in the pathological image, and y represents the vertical coordinate in the pathological image; The reference color is defined based on known dyed area colors in historical data; Based on color difference value , the optimized pathological image is initially binarized, and the optimal color difference value is automatically calculated based on Otsu's by maximizing the inter-class variance between foreground and background ; when ≥ When , it is marked as the positive staining area; when < When , it is marked as negative staining area.

4. The immunohistochemical digital pathology image processing method according to claim 3, characterized in that: The stained areas are classified by K-means clustering algorithm. The specific steps are as follows: For each pixel in the positive staining area, extract the L channel, a channel, and b channel of each pixel in the positive staining area, and generate a three-dimensional feature vector based on the coordinates of each pixel in each positive staining area; According to the total number of classification types, the number of clusters is set to k, and k centroids are randomly initialized; Based on the three-dimensional feature vector, the Euclidean distance from each pixel to the centroid is calculated by the K-means clustering algorithm, and each pixel is assigned to the nearest centroid; Based on the assignment results, the centroid position of each cluster was recalculated until the centroid position no longer changed, and finally the pixels in the positive staining area were divided into weak positive, moderate positive, and strong positive.

5. The immunohistochemical digital pathology image processing method according to claim 4, characterized in that: Based on the classification results, quantitative analysis is performed and an immunohistochemical image analysis report is generated. The specific steps are as follows: Based on the classification results, the color deconvolution algorithm was used to analyze the area proportion and staining intensity of each category in the optimized pathological image according to the number of pixels in each positive staining category and the total number of pixels in the positive staining area; Generate an immunohistochemistry image analysis report based on the quantitative analysis results.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the immunohistochemical digital pathology image processing method according to any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the immunohistochemical digital pathology image processing method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and system for automatically analyzing panoramic image of digital pathological section

    CN105550651A

  • Lung extraction method and system for chest CT image

    CN117934534A

  • Intelligent control method and system of hair removal instrument

    CN118975849A