A method for pathological image preprocessing and feature extraction
By removing contaminants through RGB and HSV color spaces and combining open source tool segmentation and dicing processing, the standardization and feature extraction of pathological images are achieved, solving the problems of contaminant removal and staining inconsistency, and improving the accuracy and efficiency of pathological image analysis.
Patent Information
- Application Number
- CN202510036207.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Existing technologies fail to effectively remove contaminants from pathological images, and staining inconsistencies limit the accuracy and generalization of image analysis, and there is a lack of effective preprocessing and staining standardization methods.
RGB and HSV dual color spaces are used for pollutant removal, open source tools are used for pathological tissue and background separation and segmentation, and staining standardization is used to ensure image consistency. Feature extraction is performed in combination with a pre-trained visual converter model.
It significantly improves the quality and consistency of pathology images, enhances the accuracy of feature extraction and pathology analysis, and supports automated diagnosis and disease prediction.
Smart Images

Figure CN119964154B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pathological image processing, and in particular to a method for preprocessing and feature extraction of pathological images. Background Art
[0002] Pathology image analysis is a key area in modern medical diagnosis and research, playing a crucial role in cancer diagnosis, pathology research, and drug development. Pathology images, particularly whole slide images (WSIs), provide rich cellular and tissue structural information, which is crucial for understanding the nature and progression of disease. However, pathology image analysis faces multiple challenges. First, whole slide images are typically very large, often containing billions of pixels. This not only places high demands on storage devices but also significantly increases the computational burden of image processing and analysis. Furthermore, pathology images may be introduced during the scanning process with various contaminants, such as dust, scratches, and other impurities. These contaminants can obscure or distort critical tissue regions in the image, thus compromising diagnostic accuracy and reliability. Another significant issue is the inconsistency of the staining process. The staining process for pathology slides relies on a variety of chemical reagents and operating conditions, which can vary between laboratories or over time, leading to significant visual differences in images of the same tissue. This staining inconsistency not only affects human diagnostic accuracy but also obscures or distorts histological features, significantly interfering with the performance of automated image analysis algorithms, as most algorithms rely on color and texture consistency to identify and classify pathological features. To address these issues, preprocessing of pathology images before applying image analysis algorithms is essential. This preprocessing effectively removes contaminants from the image, followed by image segmentation to reduce storage requirements and computational complexity, facilitating subsequent in-depth analysis. Furthermore, standardization of staining is crucial for ensuring image analysis accuracy, reducing variability between images from different batches and sources and improving the algorithm's generalization capabilities.
[0003] Disadvantages of existing technologies: (1) Ignoring the phenomenon of glass slide contamination: Existing technologies often do not fully consider the possible contamination of glass slides in actual situations. Contaminants can cause obvious interference on the image, affecting the quality of the image and subsequent image analysis. If effective decontamination is not performed, these contaminants may cause the image recognition algorithm to mistakenly identify false positive or false negative results, thereby affecting the accuracy of diagnosis. (2) Staining inconsistency is not resolved: Different pathology departments may use different staining protocols, which will lead to staining inconsistency. Images under the same pathological conditions have significant differences in color; existing technologies often lack effective staining standardization methods, and cannot ensure that images from different sources have consistent staining levels when performing feature extraction and analysis, which limits the generalization ability and application scope of the algorithm. (3) Pathology images are not fully preprocessed, and the quality and availability of the extracted features are greatly limited, which will directly affect the feature learning and downstream task accuracy in the subsequent image analysis algorithm.
[0004] Prior art 1, application number CN202410700444.X, discloses a multimodal breast tumor risk prediction method and system based on pathological images. The method comprises acquiring pathological image data and gene sequence data; preprocessing the pathological image data and gene sequence data; constructing a multimodal breast tumor risk prediction model; and extracting features from the preprocessed data using the multimodal breast tumor risk prediction model, followed by feature fusion of the extracted features. While this method can efficiently integrate pathological images and genomic data to accurately predict the survival risk of breast cancer patients, its overreliance on data and lack of preprocessing of the source data have, to some extent, affected the accuracy of the model's output results.
[0005] Prior art 2, application number: CN202411697291.4 discloses a method for processing immunohistochemical digital pathology images, including obtaining immunohistochemically stained pathology images and performing preprocessing; dividing the preprocessed pathology images into regions, extracting and fusing features from each sub-region to generate a multidimensional feature vector; optimizing the pathology images using a region growing and merging algorithm based on the multidimensional feature vector; detecting the staining intensity of the optimized pathology images, and classifying the stained regions using a K-means clustering algorithm; and performing quantitative analysis based on the classification results and generating an immunohistochemical image analysis report. Although accurate segmentation and classification analysis of immunohistochemically stained regions are achieved through preprocessing, feature extraction and fusion, combined with the region growing and merging algorithm, improving the automation and recognition accuracy of image processing, the phenomenon of slide contamination is ignored, affecting the image quality and subsequent image analysis.
[0006] Prior art three, application number: CN 202411087165.7 discloses a pathology image imaging quality assessment method. Pathology images captured by an optical microscope at multiple defocused planes are used to construct a pathology image defocus dataset. The defocused pathology image dataset is preprocessed. A pathology image quality assessment network based on frequency domain features is established. A combined loss function is used to optimize and train a deep neural network model for pathology image quality assessment based on frequency domain features. The optimized pathology image quality assessment network is used to output the pathology image quality grade. Although an end-to-end deep learning model is used and a frequency convolution module is proposed to extract the frequency features of pathology images, improving the performance of pathology image quality assessment with good robustness and adaptability, the method lacks an effective staining standardization method, making it impossible to ensure that images from different sources have consistent staining levels during feature extraction and analysis. This limits the generalization ability and application scope of the algorithm.
[0007] Currently, the existing technologies 1, 2 and 3 have the problem of ignoring the phenomenon of glass slide contamination, not solving the problem of staining inconsistency, and not fully performing preprocessing operations on pathological images. Therefore, the present invention provides a method for preprocessing and feature extraction of pathological images. Summary of the Invention
[0008] In order to solve the above technical problems, the present invention provides a method for pathological image preprocessing and feature extraction, comprising the following steps:
[0009] Full-slice images were acquired for thumbnail extraction, and contaminants were removed using RGB and HSV dual color spaces;
[0010] Use open source tools to segment pathological images, separate pathological tissue from background, and perform block processing;
[0011] The pre-trained visual transformer model is used to extract features from the pre-processed whole-slide images to capture the complex features and subtle differences in the normalized whole-slide images.
[0012] Optionally, the process of pollutant removal using the dual color space of RGB and HSV includes the following steps:
[0013] In the RGB color space, the image is decomposed into three independent channels: red, green, and blue;
[0014] After preliminary filtering in the RGB domain, the image is converted to the HSV color space. In the HSV space, dynamic threshold ranges are set for hue, saturation, and brightness to locate and remove interference information in the complex background while retaining the detailed features of the pathological tissue.
[0015] According to the specific characteristics of the whole-slice image, the weight ratio of the RGB and HSV domains is automatically adjusted; by analyzing the global and local characteristics of the image, the threshold range of the RGB and HSV domains is automatically adjusted.
[0016] Optionally, for each red, green, and blue channel, a dynamic threshold range is set according to the characteristics of the full-slice image; by comparing the RGB value of each pixel with the preset threshold, pixels below the threshold are reset to white.
[0017] Optionally, after the threshold range is automatically adjusted, the image is decomposed into multiple scales, and the threshold range is set at different scales to adapt to pollutants of different sizes and types; after the pollutants are removed, the image is edge-smoothed.
[0018] Optionally, after the dicing process, an open source library is used for color standardization, and N×N image blocks are spliced to form a large image for batch processing.
[0019] Alternatively, the staining normalization process using an open-source library can include the following steps:
[0020] Extract staining distribution information from images through open source libraries, and statistically analyze staining features of multiple high-quality pathology images using open source libraries to establish staining reference templates.
[0021] Based on the staining reference template, staining mapping and correction are performed on the target pathology image to map the staining intensity of the target image to the standardized range of the reference template;
[0022] After the staining standardization is completed, statistical analysis methods are used to verify whether the distribution of staining in the image meets the standards of the reference template; feature matching technology is used to evaluate the consistency of staining intensity, uniformity and contrast.
[0023] Optionally, the process of creating a staining reference template includes the following steps:
[0024] Extract key parameters such as staining intensity, staining uniformity, and staining contrast from multiple high-quality pathology images using open-source libraries;
[0025] Based on the statistical analysis results of staining intensity, a standardized intensity range was constructed and used as the core parameter of the staining intensity template. Based on the statistical analysis results of staining uniformity, a uniformity reference standard was established and used as the core parameter of the staining uniformity template. Based on the statistical analysis results of staining contrast, a standardized contrast threshold was defined and used as the core parameter of the staining contrast template.
[0026] Optionally, the staining intensity is quantified by the distribution of pixel values, the staining uniformity is statistically analyzed by the spatial distribution of the staining in the tissue, and the staining contrast is calculated by the difference between the stained area and the background area.
[0027] Optionally, after the color intensity is mapped to the standardized range of the reference template, the uniformity of the staining in the tissue is improved by local contrast enhancement and staining distribution equalization; and the contrast between the staining and the background is optimized according to the contrast threshold of the reference template.
[0028] Optionally, parameters of the staining intensity, uniformity, and contrast template are adjusted through an iterative optimization method; and the applicability of the staining reference template in different images is verified.
[0029] The present invention uses full-slice image thumbnail extraction and contaminant removal. Thumbnail extraction reduces image resolution by generating thumbnails, facilitating quick preview and preliminary analysis, and reducing computational burden. RGB and HSV dual color space processing: In the RGB domain, color threshold filtering is used to reset pixels below a preset threshold to white, effectively removing obvious contaminants (such as dust, stains, etc.). In the HSV domain, by setting threshold ranges for hue, saturation, and brightness, different types of contaminants (such as unevenly stained areas) can be more accurately identified and removed. Pathological image segmentation, block cutting, and staining standardization: Image segmentation separates pathological tissue from the background, focusing on the target area and reducing interference from irrelevant information. Block cutting divides large-scale pathological images into N×N small blocks, facilitating batch processing and efficient training of deep learning models. Staining standardization uses an open source library to standardize image staining, eliminating color differences caused by different staining batches or equipment, ensuring image consistency. Image stitching reassembles the processed N×N image blocks into a large image for overall analysis and feature extraction. Visual Transformer model feature extraction, pre-trained Visual Transformer (ViT) model, uses the pre-trained ViT model to extract features from preprocessed images, which can capture complex features and subtle differences in full-slice images; global modeling capability ViT's self-attention mechanism can capture long-distance dependencies in images and extract richer contextual information; high-dimensional feature representation The extracted features can be used for subsequent classification, segmentation or diagnosis tasks, providing strong support for pathological analysis.
[0030] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0031] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0033] Figure 1 Flowchart of the method for pathological image preprocessing and feature extraction in Example 1 of the present invention;
[0034] Figure 2 Schematic diagram of the method for pathological image preprocessing and feature extraction in Example 1 of the present invention;
[0035] Figure 3 This is a diagram of the process of pollutant removal using the RGB and HSV dual color spaces in Example 2 of the present invention;
[0036] Figure 4 This is a process diagram for setting a dynamic threshold range based on the characteristics of a full-slice image in Example 3 of the present invention;
[0037] Figure 5 This is a process diagram for setting dynamic threshold ranges for hue, saturation, and brightness in HSV space in Example 4 of the present invention;
[0038] Figure 6 This is a diagram of the process of automatically adjusting the threshold ranges of RGB and HSV domains in Example 5 of the present invention;
[0039] Figure 7 This is a process diagram for dyeing standardization using an open source library in Example 6 of the present invention;
[0040] Figure 8 This is a diagram of the process of establishing a staining reference template in Example 7 of the present invention;
[0041] Figure 9 This is a process diagram for performing staining mapping and correction on a target pathological image in Example 8 of the present invention;
[0042] Figure 10 This is a diagram showing the process of feature extraction on a preprocessed full-slice image in Example 9 of the present invention;
[0043] Figure 11 This is a process diagram for calculating the correlation between image blocks on a global scale in embodiment 10 of the present invention. DETAILED DESCRIPTION
[0044] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0045] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present application. The singular forms "a", "the" and "the" used in the embodiments of the present application are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0046] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0047] Example 1: Figure 1 As shown, an embodiment of the present invention provides a method for pathological image preprocessing and feature extraction, comprising the following steps:
[0048] S100: Full-slice images are acquired for thumbnail extraction, and contaminants are removed using both RGB and HSV color spaces. In the RGB domain, pixels below a preset threshold are reset to white through color threshold filtering to remove obvious contaminants. In the HSV domain, different types of contaminants are identified and removed by setting threshold ranges for hue, saturation, and brightness.
[0049] S200: Use open source tools to segment pathology images, separate pathological tissue from background, and perform block processing. Use open source libraries to standardize staining and splice N×N image blocks to form large images for batch processing.
[0050] S300: Use the pre-trained visual transformer model to extract features from the pre-processed whole-slide images to capture the complex features and subtle differences in the normalized whole-slide images.
[0051] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the whole-slice image is first obtained for thumbnail extraction, and the RGB and HSV dual color spaces are used for pollutant removal; in the RGB domain, pixels below the preset threshold are reset to white through color threshold filtering to remove obvious pollutants; in the HSV domain, different types of pollutants are identified and removed by setting the threshold range of hue, saturation and brightness; secondly, the pathological image is segmented using open source tools to separate the pathological tissue from the background and perform block processing; the open source library is used for staining standardization, and N×N image blocks are spliced to form a large image for batch processing; finally, the pre-trained visual converter model is used to extract features from the pre-processed whole-slice image to capture the complex features and subtle differences in the standardized whole-slice image (for specific principles, please refer to the attached Figure 2). Step S100 of the above scheme involves full-slice image thumbnail extraction and contaminant removal. Thumbnail extraction generates thumbnails, reduces image resolution, facilitates quick preview and preliminary analysis, and reduces computational burden. RGB and HSV dual color space processing: In the RGB domain, color threshold filtering is used to reset pixels below the preset threshold to white, effectively removing obvious contaminants (such as dust, stains, etc.). In the HSV domain, by setting the threshold range for hue, saturation, and brightness, different types of contaminants (such as unevenly stained areas) can be more accurately identified and removed. Significance: This method significantly improves image quality, provides clean and reliable input data for analysis, and reduces the interference of noise on feature extraction and diagnostic results. Step S200 involves pathology image segmentation, tiling, and staining standardization. Image segmentation separates pathological tissue from background, focusing on the target area and reducing interference from irrelevant information. Tile tiling divides large pathology images into N×N blocks, facilitating batch processing and efficient training of deep learning models. Staining standardization uses an open-source library to standardize image staining, eliminating color variations caused by different staining batches or equipment and ensuring image consistency. Image stitching reassembles the processed N×N image blocks into a large image, facilitating overall analysis and feature extraction. Significance: This achieves standardized and modular processing of pathology images, providing a high-quality, consistent data foundation for feature extraction and model training, significantly improving the accuracy and repeatability of analysis. Step S300: Visual Transformer Model Feature Extraction. A pre-trained Visual Transformer (ViT) model is used to extract features from pre-processed images. This model can capture complex features and subtle differences in full-slice images. ViT's self-attention mechanism, with its global modeling capabilities, can capture long-range dependencies in images and extract richer contextual information. The extracted features, represented by high-dimensional features, can be used for subsequent classification, segmentation, or diagnostic tasks, providing strong support for pathology analysis. Significance: This approach fully leverages the power of deep learning models to extract key features from complex pathology images, providing technical support for accurate diagnosis and pathology research. Furthermore, the introduction of ViT has brought new possibilities to pathology image analysis and promoted technological advancement in this field.
[0052] In summary, the pathology image preprocessing and feature extraction of this embodiment achieves full-process optimization, from raw data to high-quality features. This not only improves the efficiency and accuracy of pathology analysis, but also provides reliable technical support for automated diagnosis, disease prediction, and personalized treatment. Furthermore, the standardized and modular design of this process lays the foundation for large-scale pathology image analysis and promotes the development of medical artificial intelligence. This embodiment effectively addresses the aforementioned issues by implementing image decontamination, image slicing, staining standardization, and feature extraction using advanced deep learning models. This improves the efficiency and accuracy of pathology image analysis and provides reliable technical support for automated pathology image analysis. This embodiment provides a systematic preprocessing process, including image decontamination, image slicing, and staining standardization. This ensures high quality and consistency of the images input to the feature extraction model, laying a solid foundation for analysis. This embodiment enhances feature extraction accuracy by significantly improving feature extraction accuracy by extracting features from high-quality preprocessed images. Combined with a pretrained deep learning model, this process can capture more subtle and complex pathology features, enhancing the accuracy and reliability of image analysis and facilitating subsequent downstream tasks such as disease classification and diagnosis.
[0053] The improved pathology image quality in this embodiment addresses the issue of improving image quality during pathology image analysis, specifically by removing contaminants from slides and resolving staining inconsistencies. These improvements ensure image clarity and color consistency, thereby increasing analysis reliability. Feature extraction from preprocessed, high-quality images addresses the issue of feature extraction accuracy. Leveraging advanced pretrained deep learning models, more precise and detailed features can be extracted from pathology images, supporting more accurate pathology analysis and diagnosis. This not only improves the efficiency and accuracy of pathology image analysis but also provides strong technical support for automated pathology analysis.
[0054] Example 2: Figure 3 As shown, based on Example 1, the process of removing pollutants using the dual color space of RGB and HSV provided by the embodiment of the present invention includes the following steps:
[0055] S101: In the RGB color space, the image is decomposed into three independent channels: red, green, and blue. For each channel, a dynamic threshold range is set based on the characteristics of the full-slice image. The RGB value of each pixel is compared with the preset threshold, and pixels below the threshold are reset to white.
[0056] S102: After preliminary filtering in the RGB domain, the image is converted to the HSV color space. In the HSV space, dynamic threshold ranges are set for hue, saturation, and brightness to locate and remove interference information in the complex background while retaining detailed features of the pathological tissue;
[0057] S103: Automatically adjust the weight ratio of the RGB and HSV domains based on the specific characteristics of the full-slice image; automatically adjust the threshold range of the RGB and HSV domains by analyzing the global and local characteristics of the image; perform multi-scale decomposition on the image and set the threshold range at different scales to accommodate pollutants of different sizes and types; after removing the pollutants, perform edge smoothing on the image.
[0058] The working principle and beneficial effects of the above technical solution are as follows: this embodiment first decomposes the image into three independent channels of red, green and blue in the RGB color space; for each channel, a dynamic threshold range is set according to the characteristics of the full-slice image; by comparing the RGB value of each pixel with the preset threshold, the pixels below the threshold are reset to white; secondly, after preliminary filtering in the RGB domain, the image is converted to the HSV color space. In the HSV space, dynamic threshold ranges are set for hue, saturation and brightness respectively to locate and remove interference information in the complex background while retaining the detailed features of the pathological tissue; finally, according to the specific characteristics of the full-slice image, the weight ratio of the RGB and HSV domains is automatically adjusted; by analyzing the global and local features of the image, the threshold range of the RGB and HSV domains is automatically adjusted; the image is decomposed at multiple scales, and the threshold range is set at different scales to adapt to pollutants of different sizes and types; after the pollutants are removed, the image is edge smoothed. In step S101 of the above scheme, preliminary filtering, channel separation, and dynamic thresholding in the RGB color space decompose the image into three independent channels: red, green, and blue. This allows for more detailed analysis of the pixel distribution in each channel. By setting a dynamic threshold range, pathological tissue can be effectively distinguished from background noise. Pixel reset to white resets pixels below the threshold to white, preliminarily removing low-luminance contaminants from the image while preserving high-luminance pathological tissue information. Significance: Preliminary noise reduction, through preliminary filtering in the RGB space, can quickly remove significant low-luminance interference, laying the foundation for refined processing. Setting a detail-preserving dynamic threshold avoids over-filtering, ensuring that detailed features of pathological tissue are not destroyed. Step S102, refined processing in the HSV color space, involves color space conversion from RGB to HSV, which better separates hue, saturation, and brightness information, facilitating more precise processing of complex backgrounds. Dynamic thresholding and noise removal, by setting dynamic thresholds for hue, saturation, and brightness in HSV space, effectively locates and removes interference from complex backgrounds. Detail preservation, through fine-tuning the threshold, ensures that detailed features of pathological tissue are preserved. Significance: Complex background processing HSV space is more suitable for processing complex backgrounds with large changes in color and brightness, and can more accurately remove interference; Enhanced image quality Through fine processing, further improves the clarity and readability of the image, providing high-quality data for analysis. Step S103 Weight adjustment and multi-scale processing, weight ratio adjustment automatically adjusts the weight ratio of RGB and HSV domains according to the specific characteristics of the image to ensure that the processing effects of the two color spaces are optimally balanced; Multi-scale decomposition and threshold setting perform multi-scale decomposition of the image, setting threshold ranges at different scales, and can adapt to pollutants of different sizes and types; Edge smoothing After the pollutants are removed, the image is smoothed to eliminate jagged or uneven edges caused by filtering.Significance: Adaptability and robustness: Through multi-scale processing and weight adjustment, the algorithm can adapt to the characteristics of different images and pollutant types, improving the robustness of the algorithm; image optimization and edge smoothing further enhance the visual effect of the image, making it more in line with the needs of medical analysis.
[0059] In summary, this embodiment, through the collaborative processing of the RGB and HSV dual color spaces, can effectively remove contaminants from images and improve image quality. Dynamic thresholding and multi-scale processing ensure that the detailed features of pathological tissue are not destroyed, providing reliable data support for medical diagnosis. Automatic adjustment of weights and threshold ranges reduces the need for manual intervention and improves processing efficiency and accuracy. Through the above steps, the contaminant removal method for the RGB and HSV dual color spaces not only effectively improves image quality but also provides a solid foundation for further analysis and diagnosis of medical images.
[0060] Example 3: Figure 4 As shown, based on Example 2, the process of setting the dynamic threshold range according to the characteristics of the full-slice image provided by the embodiment of the present invention includes the following steps:
[0061] S1011: Perform global pixel value distribution analysis on the red, green, and blue channels to generate histograms. Calculate the pixel value histogram for each channel to identify the brightness distribution characteristics of the main areas. Combined with the pathological characteristics of the full-slice image, determine the upper and lower limits of the threshold range and exclude extreme values through statistical analysis. For example, in the red channel, pathological tissue may exhibit higher pixel values, while contaminants may exhibit lower pixel values.
[0062] S1012: Performing a local scan of the full-slice image based on a sliding window, dividing the image into multiple local windows, and performing statistical analysis on the pixel values within each window; for each window, calculating statistics such as the mean and variance of the pixel values, and dynamically adjusting the local threshold range based on the global threshold range;
[0063] S1013: Decompose the full-slice image into multiple scales by performing multi-scale decomposition. Analyze the pixel value distribution at each scale to determine the threshold range of the scale. Combine the threshold ranges of each scale and determine the final dynamic threshold range by weighted averaging.
[0064] Among them, S1011 global pixel value distribution analysis and threshold interval determination:
[0065]
[0066] Where, Represents pixel value; Indicates channels (R, G, B); represents the channel weight, which indicates the contribution of each channel to the overall distribution; represents the number of Gaussian distributions; Indicates the The mixing coefficient of the kth Gaussian distribution in the channel satisfies ; Indicates the The mean of the k-th Gaussian distribution in the channel; Indicates the The variance of the k-th Gaussian distribution in the channel;
[0067] Step S1012 represents local scanning and dynamic threshold adjustment:
[0068]
[0069] Where, Represents a local window Dynamic threshold of represents the global threshold; Represents a local window The mean pixel value of ; Represents the global pixel value mean; Represents a local window The pixel value variance of Represents the global pixel value variance; , Represents the weight coefficient, which is used to balance the influence of mean and variance on the local threshold;
[0070] Step S1013 represents multi-scale decomposition and comprehensive threshold determination:
[0071]
[0072]
[0073] Where, represents the final dynamic threshold; Indicates the total number of scales; Represents the weight of the s-th scale, satisfying ; represents the threshold of the sth scale; Indicates the number of pixel blocks at the sth scale; represents the mean value of the i-th pixel block in the s-th scale; represents the variance of the i-th pixel block in the s-th scale; represents the variance weight coefficient. The above algorithm and formula combine complex mathematical methods such as probability statistics, regression analysis, and multiscale decomposition. They effectively adapt to the complex characteristics of whole-slice images and enable precise setting of dynamic thresholds. They not only consider global and local pixel value distributions but also capture detailed image information through multiscale analysis, providing more reliable technical support for pathological diagnosis.
[0074] The working principle and beneficial effects of the above technical solution are as follows: this embodiment first performs a global pixel value distribution analysis on the red, green and blue channels respectively to generate a histogram; calculates the pixel value histogram of each channel to identify the brightness distribution characteristics of the main area; combines the pathological characteristics of the full-slice image to determine the upper and lower limits of the threshold range, and excludes extreme values through statistical analysis; for example, in the red channel, pathological tissue may present a higher pixel value, while pollutants may present a lower pixel value; secondly, based on a sliding window, the full-slice image is locally scanned, the image is divided into multiple local windows, and statistical analysis is performed on the pixel values in each window; for each window, the statistical quantities such as the mean and variance of its pixel values are calculated, and the local threshold range is dynamically adjusted in combination with the global threshold range; finally, by performing multi-scale decomposition on the full-slice image, the image is decomposed into multiple scales, and pixel value distribution analysis is performed on each scale to determine the threshold range of the scale; the threshold ranges of each scale are comprehensively combined and the final dynamic threshold range is determined by weighted average. In step S1011 of the above scheme, global pixel value distribution analysis and threshold interval determination are performed. By analyzing the pixel value histograms of the red, green, and blue channels, the overall brightness distribution characteristics of the image can be fully understood. Combined with the pathological characteristics, the pixel value differences between pathological tissue and background or contaminants can be identified, thereby determining the reasonable upper and lower limits of the threshold interval. Significance: It provides a global reference for local threshold adjustment, ensuring that the threshold range can adapt to the overall characteristics of the image, while eliminating the interference of extreme values and improving the accuracy of the analysis. Step S1012 local scanning and dynamic threshold adjustment performs local scanning of the image through a sliding window, which can capture the brightness changes in different areas of the image. Statistical analysis is performed on the pixel values within each window, and combined with the global threshold interval, the local threshold range is dynamically adjusted so that the threshold can adapt to the local characteristics of the image. Significance: Local dynamic adjustment can effectively solve the problem of uneven brightness in the image and avoid misjudgment caused by the global threshold. Especially in full-slice images, the distribution of pathological tissue may be uneven, which can significantly improve the accuracy of the analysis. Step S1013 involves multi-scale decomposition and comprehensive threshold determination. Through multi-scale decomposition, the image is broken down into multiple scales. Pixel value distribution analysis is performed on each scale to determine the threshold range for each scale. The threshold ranges for each scale are then combined and weighted averaged to determine the final dynamic threshold range, ensuring that the threshold takes into account both global and local characteristics of the image. Significance: Multi-scale analysis captures detailed information at different levels within the image, ensuring that the threshold range setting neither misses subtle features nor is affected by noise. This further improves the robustness and adaptability of threshold setting.
[0075] In summary, this embodiment gradually refines the threshold range setting from global to local, and from a single scale to multiple scales. This not only adapts to the complex characteristics of whole-slice images, but also effectively eliminates interfering factors, improving the accuracy and reliability of pathological analysis. This dynamic threshold setting method has important application value in the field of medical image processing, providing more accurate data support for pathological diagnosis.
[0076] Example 4: Figure 5 As shown, based on Example 2, the process of setting the dynamic threshold range for hue, saturation and brightness in the HSV space provided by the embodiment of the present invention includes the following steps:
[0077] S1021: Based on the global hue distribution of the whole-slice image, multi-peak Gaussian fitting is used to identify the main peak areas of the hue distribution;
[0078] Among them, the dynamic threshold range of hue is determined by calculating the skewness and kurtosis of the hue distribution; the upper and lower thresholds of hue are set in combination with the global features of the whole slice image to exclude extreme values; the global saturation distribution is analyzed, and the main distribution area of saturation is identified by adaptive quantile regression; the dynamic threshold range of saturation is determined by calculating the cumulative distribution function (CDF) of saturation; for example, pathological tissue usually has a higher saturation, while the background or pollutants may have a lower saturation; the upper and lower thresholds of saturation are set in combination with global features; the global brightness distribution is analyzed, and the main distribution area of brightness is identified by local weighted density estimation; the dynamic threshold range of brightness is determined by calculating the probability density function (PDF) of brightness; for example, pathological tissue usually has a medium brightness, while pollutants may have extremely high or extremely low brightness. In combination with global features, the upper and lower thresholds of brightness are set;
[0079] The cumulative distribution function equation is expressed as:
[0080]
[0081] Where, Indicates that the random variable X is less than or equal to The cumulative probability of; X represents a random variable, representing the saturation value; represents a specific saturation value; f(t) represents the probability density function (PDF), which describes the distribution of saturation values; t represents the integral variable, which represents the value of saturation; represents the mean of the random variable X; represents the standard deviation of the random variable X; represents the exponential function; represents the normalization constant, ensuring that the integral of the probability density function is 1;
[0082] The probability density function equation is expressed as:
[0083]
[0084] Where, Indicates that the random variable X is The probability density at ; represents the cumulative distribution function; Indicates a specific brightness value; represents the mean of the random variable X; represents the standard deviation of the random variable X; represents the exponential function; represents the normalization constant, ensuring that the integral of the probability density function is 1;
[0085] S1022: Divide the image into multiple local windows and perform statistical analysis on the hue distribution within each window; dynamically adjust the local hue threshold range by calculating the mean and variance of the local hue distribution and combining it with the global hue threshold; and use nonlinear mapping technology to correct abnormal hue values in the local windows;
[0086] Among them, local saturation threshold adjustment: statistical analysis is performed on the saturation distribution in each local window to calculate the mean and variance of the local saturation; combined with the global saturation threshold, the local saturation threshold range is dynamically adjusted through weighted regression technology; for abnormal saturation values in the local window, adaptive filtering is used to correct them;
[0087] Local brightness threshold adjustment: Statistical analysis is performed on the brightness distribution within each local window to calculate the mean and variance of the local brightness. Combined with the global brightness threshold, the local brightness threshold range is dynamically adjusted through gradient optimization technology. Local contrast enhancement technology is used to correct abnormal brightness values in local windows.
[0088] S1023: Perform multi-scale decomposition on the full-slice image to obtain hue distributions at different scales; determine the final threshold value comprehensively through weighted averaging technology; calculate the dynamic threshold range of hue at each scale, and determine the final hue threshold comprehensively through weighted averaging technology; set multi-scale saturation threshold: perform multi-scale decomposition on the full-slice image to obtain saturation distributions at different scales; calculate the dynamic threshold range of saturation at each scale, and determine the final saturation threshold comprehensively through weighted averaging technology; set multi-scale brightness threshold: perform multi-scale decomposition on the full-slice image to obtain brightness distributions at different scales; calculate the dynamic threshold range of brightness at each scale, and determine the final brightness threshold comprehensively through weighted averaging technology.
[0089] The working principle and beneficial effects of the above technical solution are as follows: This embodiment first uses multi-peak Gaussian fitting to identify the main peak area of the hue distribution based on the global hue distribution of the whole slice image; wherein, the dynamic threshold range of the hue is determined by calculating the skewness and kurtosis of the hue distribution; the upper and lower thresholds of the hue are set in combination with the global features of the whole slice image to exclude extreme values; the global saturation distribution is analyzed, and the main distribution area of saturation is identified by adaptive quantile regression; the dynamic threshold range of saturation is determined by calculating the cumulative distribution function (CDF) of saturation; for example, pathological tissue usually has a higher saturation, while background or pollutants may have a lower saturation; combined with the global features, the saturation distribution is gradually reduced. Based on the characteristics, the upper and lower saturation thresholds are set; the global brightness distribution is analyzed, and the main distribution area of brightness is identified by local weighted density estimation; the dynamic threshold range of brightness is determined by calculating the probability density function (PDF) of brightness; for example, pathological tissue usually has medium brightness, while pollutants may have extremely high or low brightness. Combining the global characteristics, the upper and lower thresholds of brightness are set; secondly, the image is divided into multiple local windows, and the hue distribution in each window is statistically analyzed; by calculating the mean and variance of the local hue distribution, combined with the global hue threshold, the local hue threshold range is dynamically adjusted; for abnormal hue values in the local window, nonlinear mapping technology is used to correct them; among them, Adjustment of local saturation threshold: Statistical analysis is performed on the saturation distribution in each local window, and the mean and variance of the local saturation are calculated; combined with the global saturation threshold, the local saturation threshold range is dynamically adjusted through weighted regression technology; for abnormal saturation values in the local window, adaptive filtering is used for correction; Adjustment of local brightness threshold: Statistical analysis is performed on the brightness distribution in each local window, and the mean and variance of the local brightness are calculated; combined with the global brightness threshold, the local brightness threshold range is dynamically adjusted through gradient optimization technology; for abnormal brightness values in the local window, local contrast enhancement technology is used for correction; finally, multi-scale decomposition is performed on the full slice image to obtain color images of different scales. Phase distribution; the final threshold is determined comprehensively by weighted averaging technology; wherein, at each scale, the dynamic threshold range of hue is calculated separately, and the final hue threshold is determined comprehensively by weighted averaging technology; multi-scale saturation threshold setting: multi-scale decomposition of the full slice image is performed to obtain saturation distributions of different scales; at each scale, the dynamic threshold range of saturation is calculated separately, and the final saturation threshold is determined comprehensively by weighted averaging technology; multi-scale brightness threshold setting: multi-scale decomposition of the full slice image is performed to obtain brightness distributions of different scales; at each scale, the dynamic threshold range of brightness is calculated separately, and the final brightness threshold is determined comprehensively by weighted averaging technology. Step S1021 of the above scheme sets the global threshold, and the hue dynamic threshold setting identifies the main peak area of the hue distribution through multi-peak Gaussian fitting, determines the dynamic threshold interval in combination with skewness and kurtosis, and excludes extreme values.Significance: Ensure that the hue threshold setting can accurately reflect the distribution characteristics of the main colors in the image, avoid threshold deviation caused by interference from extreme values, and improve the accuracy of analysis. The saturation dynamic threshold setting uses adaptive quantile regression to identify the main distribution area of saturation, and determines the dynamic threshold interval through the cumulative distribution function (CDF). Significance: Based on the saturation difference between pathological tissue and background or pollutants, a reasonable threshold range is set to effectively distinguish target areas from non-target areas, thereby improving the accuracy of image segmentation; the brightness dynamic threshold setting uses local weighted density estimation to identify the main distribution area of brightness, and determines the dynamic threshold interval through the probability density function (PDF). Significance: Based on the brightness difference between pathological tissue and pollutants, a reasonable threshold range is set to avoid misjudgment due to brightness abnormalities and ensure the reliability of image analysis. Step S1022: Local threshold adjustment dynamically adjusts the local threshold by calculating the mean and variance of the local hue distribution and combining it with the global threshold. Nonlinear mapping is used to correct outliers. Significance: It adapts to hue variations in local image regions, improves threshold setting flexibility, and ensures more accurate hue analysis in local regions. Local saturation threshold adjustment dynamically adjusts the local saturation threshold using weighted regression techniques and uses adaptive filtering to correct outliers. Significance: It optimizes threshold settings based on saturation variations in local regions, reduces noise interference, and improves local consistency in image segmentation. Local brightness threshold adjustment dynamically adjusts the local brightness threshold using gradient optimization techniques and corrects outliers using local contrast enhancement techniques. Significance: It adapts to local brightness variations, optimizes threshold settings, enhances the visibility of image details, and improves the accuracy of image analysis. Step S1023: Multi-scale threshold synthesis uses multi-scale decomposition and weighted averaging techniques to comprehensively determine the final hue threshold. Significance: It comprehensively considers hue distribution characteristics at different scales, ensures consistency in threshold setting across different resolutions, and improves the robustness of image analysis. Multi-scale saturation threshold setting uses multi-scale decomposition and weighted averaging techniques to comprehensively determine the final saturation threshold. Significance: Comprehensively consider the saturation distribution characteristics of different scales, optimize the threshold setting, and improve the global consistency of image segmentation; multi-scale brightness threshold setting uses multi-scale decomposition and weighted averaging technology to comprehensively determine the final brightness threshold. Significance: Comprehensively consider the brightness distribution characteristics of different scales, optimize the threshold setting, and ensure the accuracy and stability of image analysis.
[0090] In summary, the dynamic threshold setting method of this embodiment can adapt to changes in global and local image features, effectively eliminate interference from noise and outliers, and improve the accuracy and robustness of image segmentation and analysis. In particular, in pathological image analysis, this method can accurately distinguish pathological tissue from background or contaminants, providing reliable data support for medical diagnosis and research. Furthermore, the introduction of multi-scale analysis technology further enhances the method's adaptability at different resolutions, enabling it to perform even better in complex scenarios.
[0091] Example 5: Figure 6 As shown, based on Example 2, the process of automatically adjusting the threshold range of RGB and HSV domains provided by the embodiment of the present invention includes the following steps:
[0092] S1031: Perform global feature analysis on the full-slice image, including the overall brightness distribution, color distribution, and texture characteristics of the image; calculate statistical quantities such as the image mean, variance, and histogram to preliminarily determine the distribution range and intensity of pollutants in the image;
[0093] S1032: Divide the whole slice image into several local regions, and calculate the brightness, color and texture characteristics of each region respectively;
[0094] S1033: Dynamically assigning weight ratios of the RGB and HSV domains based on the analysis results of the global and local features; fusing the processing results of the RGB and HSV domains through a weighted average or adaptive fusion algorithm;
[0095] RGB domain weighting: If the contaminants in the full-slice image are mainly manifested as intensity distribution (such as brightness differences), increase the weight of the RGB domain and use the intensity information in the RGB space for filtering;
[0096] HSV domain weight: If the contaminants in the full-slice image are mainly manifested as color distribution (such as hue or saturation differences), increase the weight of the HSV domain and use the color information in the HSV space for filtering.
[0097] The working principle and beneficial effects of the above technical solution are as follows: This embodiment first performs a global feature analysis on the full-slice image, including the image's overall brightness distribution, color distribution, and texture features. By calculating statistics such as the image's mean, variance, and histogram, the distribution range and intensity of pollutants in the image are preliminarily determined. Secondly, the full-slice image is divided into several local regions, and the brightness, color, and texture features of each region are calculated separately. Finally, based on the analysis results of the global and local features, the weight ratios of the RGB and HSV domains are dynamically assigned. The processed results of the RGB and HSV domains are fused using a weighted average or adaptive fusion algorithm. Step S1031 of the above solution, global feature analysis, analyzes the overall brightness distribution, color distribution, and texture features of the full-slice image to comprehensively understand the basic characteristics of the image and the distribution of pollutants. By calculating statistics such as the image's mean, variance, and histogram, the distribution range and intensity of pollutants can be preliminarily determined, providing a global reference for processing. Significance: Global feature analysis provides basic data for local feature extraction and weight assignment, ensuring that the processing process covers the entire range of the image. Through global analysis, large-scale pollutant areas can be quickly located, improving processing efficiency. Step S1032, local feature extraction, divides the full-slice image into several local regions and calculates the brightness, color, and texture features of each region, capturing detailed information within the image. Local feature analysis allows for more accurate identification of the distribution of small-scale or detailed contaminants. Significance: Local feature extraction compensates for the shortcomings of global analysis, ensuring that the processing process covers detailed areas within the image. Local analysis allows for more accurate identification and processing of small-scale contaminants, improving the sophistication of the processing results. Step S1033, dynamic weight allocation and fusion, dynamically allocates weight ratios for the RGB and HSV domains based on the analysis results of global and local features, enabling flexible response to different types of contaminants. The processing results from the RGB and HSV domains are fused using a weighted average or adaptive fusion algorithm to ensure comprehensive and accurate contaminant removal. Significance: Dynamic weight allocation allows for flexible adjustment of processing strategies based on the specific characteristics of the image, ensuring the adaptability and robustness of the processing results. Fusion of the RGB and HSV domains fully utilizes information in both color and intensity spaces, improving the comprehensiveness and accuracy of contaminant removal.
[0098] In summary, this embodiment, through global feature analysis, local feature extraction, and dynamic weight allocation and fusion, comprehensively covers the distribution of pollutants in an image, ensuring the accuracy and robustness of the processing results. It can flexibly respond to different types of pollutants and adapt to the processing needs of different scenarios. It significantly improves the efficiency and accuracy of pollutant removal, providing an innovative solution for the field of image processing. By fully leveraging the multi-scale and color space characteristics of images, this method has broad application prospects, particularly in medical image processing, industrial inspection, and other fields.
[0099] Example 6: Figure 7 As shown, based on Example 1, the process of using an open source library to perform dyeing standardization provided in the embodiment of the present invention includes the following steps:
[0100] S201: Extract staining distribution information from images using open-source libraries, including key parameters such as staining intensity, staining uniformity, and staining contrast. Statistically analyze staining characteristics of multiple high-quality pathology images using open-source libraries to establish a staining reference template.
[0101] Among them, the staining intensity template: constructs a standardized intensity range based on the distribution law of staining intensity; the staining uniformity template: establishes a uniformity reference standard by analyzing the uniformity of staining distribution in the tissue; the staining contrast template: defines a standardized contrast threshold based on the contrast relationship between staining and background;
[0102] S202: Based on the staining reference template, stain mapping and correction are performed on the target pathology image to map the staining intensity of the target image to the standardized range of the reference template; local contrast enhancement and staining distribution equalization are performed to improve the uniformity of staining in the tissue; and the contrast between the staining and the background is optimized according to the contrast threshold of the reference template.
[0103] S203: After the staining standardization is completed, the distribution of the staining in the image is verified to be consistent with the standard of the reference template through statistical analysis methods; and the consistency of the staining intensity, uniformity and contrast is evaluated through feature matching technology.
[0104] The working principle and beneficial effects of the above technical solution are as follows: This embodiment first extracts the staining distribution information in the image through the open source library, including key parameters such as staining intensity, staining uniformity and staining contrast; uses the open source library to perform statistical analysis of the staining features of multiple high-quality pathological images to establish a staining reference template; among them, the staining intensity template: constructs a standardized intensity range based on the distribution law of staining intensity; the staining uniformity template: establishes a uniformity reference standard by analyzing the distribution uniformity of staining in the tissue; the staining contrast template: defines a standardized contrast threshold based on the contrast relationship between staining and background; secondly, based on the staining reference template, the target pathological image is stained and corrected, and the staining intensity of the target image is mapped to the standardized range of the reference template; the uniformity of staining in the tissue is improved through local contrast enhancement and staining distribution equalization; the contrast between staining and background is optimized according to the contrast threshold of the reference template; finally, after the staining standardization is completed, the statistical analysis method is used to verify whether the distribution of staining in the image meets the standards of the reference template; and the consistency of staining intensity, uniformity and contrast is evaluated through feature matching technology. Step S201 of the above scheme extracts staining distribution information and establishes a staining reference template. Parameters such as staining intensity, uniformity, and contrast are extracted from the image using an open-source library, providing a data foundation for template establishment. Significance: These parameters are core indicators of staining standardization, comprehensively reflecting the staining characteristics of the image and providing a scientific basis for standardization. Establishing a staining reference template establishes a standardized intensity range based on the distribution of staining intensity. This provides a unified standard for staining intensity, avoiding misdiagnosis due to excessively dark or light staining. Significance: This ensures comparability of staining intensity across different images, improving the accuracy of subsequent analysis. By analyzing the uniformity of staining distribution within the tissue, a uniformity reference standard is established. This provides a quantitative standard for staining uniformity, reducing the impact of uneven staining on image analysis. Significance: This improves overall image quality, ensuring more consistent staining distribution and facilitating subsequent pathological diagnosis. Based on the contrast relationship between stain and background, a standardized contrast threshold is defined, providing a clear standard for stain-background contrast and enhancing image clarity. Significance: This improves image readability, enhances the prominence of pathological features, and reduces the possibility of misdiagnosis. Step S202 performs stain mapping and correction based on the staining reference template, adjusting the staining intensity of the target image to within the standardized range of the reference template. Significance: Eliminates differences in staining intensity between images, ensuring image intensity consistency; improves the uniformity of staining in tissue through local contrast enhancement and distribution equalization. Significance: Improves the visual effect of the image, making the staining distribution more uniform and facilitating the identification of pathological features; optimizes the contrast between the staining and the background based on the contrast threshold of the reference template. Significance: Enhances image clarity and readability, making pathological features more apparent and improving diagnostic accuracy.Step S203: Staining Standardization Verification and Consistency Assessment. Statistical analysis verifies that the staining distribution conforms to the reference template. Significance: Ensures the effectiveness of staining standardization and provides reliable data support for image analysis. Evaluates consistency in staining intensity, uniformity, and contrast. Significance: Confirms the successful implementation of staining standardization and provides a high-quality image foundation for pathological diagnosis.
[0105] In summary, the staining standardization process in this example achieves uniformity in staining intensity, uniformity, and contrast in pathology images through staining feature extraction, reference template establishment, mapping correction, and verification and evaluation. This not only improves image quality and consistency but also provides reliable technical support for pathology diagnosis and image analysis, with significant clinical and scientific significance.
[0106] Example 7: Figure 8 As shown, based on Example 6, the process of establishing a staining reference template provided by the embodiment of the present invention includes the following steps:
[0107] S2011: We extracted key parameters such as staining intensity, staining uniformity, and staining contrast from multiple high-quality pathology images using an open-source library. Staining intensity was quantified based on the distribution of pixel values, while staining uniformity was statistically analyzed based on the spatial distribution of staining within the tissue. Staining contrast was calculated based on the difference between the stained and background regions.
[0108] S2012: Based on the statistical analysis results of staining intensity, a standardized intensity range was constructed and used as the core parameter of the staining intensity template. Based on the statistical analysis results of staining uniformity, a uniformity reference standard was established and used as the core parameter of the staining uniformity template. Based on the statistical analysis results of staining contrast, a standardized contrast threshold was defined and used as the core parameter of the staining contrast template.
[0109] S2013: Adjust the parameters of the staining intensity, uniformity, and contrast templates through iterative optimization methods; verify the applicability of the staining reference template in different images.
[0110] The working principle and beneficial effects of the above technical solution are as follows: This embodiment first extracts key parameters such as staining intensity, staining uniformity and staining contrast from multiple high-quality pathological images using an open source library; the staining intensity is quantified by the distribution law of pixel values, and the staining uniformity is statistically analyzed by the spatial distribution of staining in the tissue; the staining contrast is calculated by the difference between the stained area and the background area; secondly, based on the statistical analysis results of the staining intensity, a standardized intensity range is constructed and used as the core parameter of the staining intensity template; based on the statistical analysis results of the staining uniformity, a uniformity reference standard is established and used as the core parameter of the staining uniformity template; based on the statistical analysis results of the staining contrast, a standardized contrast threshold is defined and used as the core parameter of the staining contrast template; finally, through an iterative optimization method, the parameters of the staining intensity, uniformity and contrast templates are adjusted; and the applicability of the staining reference template in different images is verified. Step S2011 of the above scheme, staining feature extraction, quantifies staining intensity through the distribution of pixel values, accurately reflecting the depth of staining in the image. This provides a data basis for standardization of staining intensity. By statistically analyzing the spatial distribution of staining within the tissue, the uniformity of staining within the tissue can be quantified, providing spatial distribution features for standardization of staining uniformity. By calculating the difference between the stained and background regions, the contrast relationship between staining and background can be clarified, providing a quantitative basis for standardization of staining contrast. Significance: This serves as the foundation for establishing the entire staining reference template. By extracting key parameters such as staining intensity, uniformity, and contrast, it provides high-quality data input for subsequent statistical analysis. Abstracting and quantifying staining features from the image provides a scientific basis for standardization. Step S2012 constructs a staining reference template. Based on the statistical analysis results of staining intensity, a standardized intensity range is constructed to provide a unified reference standard for staining intensity, ensuring consistency of staining intensity across different images. Based on the statistical analysis results of staining uniformity, a uniformity reference standard is established to provide a quantitative reference for staining uniformity, ensuring uniform distribution of staining within the tissue. Based on the statistical analysis results of staining contrast, a standardized contrast threshold is defined to provide a clear reference standard for stain-background contrast, ensuring consistency of the contrast relationship between staining and background. Significance: Through statistical analysis, staining features are converted into standardized reference templates, providing specific technical support for staining standardization. Through template-based methods, staining features are elevated from the data level to the standardization level, providing a basis for staining correction and optimization. Step S2013 optimizes and verifies the staining reference template. Through iterative optimization methods, the parameters of the staining intensity, uniformity, and contrast templates are adjusted to improve the template's applicability and accuracy, ensuring the template's universality across different images. Statistical analysis is used to verify the applicability of the staining reference template across different images, ensuring the template's reliability and providing a basis for the scientific and practical application of the template.Significance: Through optimization and verification, the staining reference template is ensured to be universal and reliable in practical applications; through scientific verification, the accuracy and consistency of staining standardization are improved, providing technical support for the subsequent analysis and diagnosis of pathological images.
[0111] In summary, this embodiment provides a data basis for staining standardization through staining feature extraction; provides specific technical support for staining standardization through the construction of staining reference templates; ensures the accuracy and consistency of staining standardization through template optimization and verification; and constructs a universal and standardized staining reference template through a hierarchical and systematic technical process, providing a scientific basis and technical support for the staining standardization of pathological images, thereby improving the accuracy and reliability of pathological image analysis.
[0112] Example 8: Figure 9 As shown, based on Example 6, the process of performing staining mapping and correction on the target pathological image provided by the embodiment of the present invention includes the following steps:
[0113] S2021: Obtain the staining intensity value of each pixel in the image, compare the staining intensity distribution of the target image with the standardized intensity range of the reference template, and identify the offset or deviation of the staining intensity in the target image; based on the intensity range matching result, design a staining intensity mapping function to linearly or nonlinearly map the staining intensity value of the target image to the standardized range of the reference template;
[0114] S2022: Applying a mapping function to correct the staining intensity of the target image, eliminating the offset or deviation of the staining intensity so that it conforms to the standardized range of the reference template; performing local intensity optimization for areas in the target image where the staining intensity distribution is uneven;
[0115] S2023: Verify whether the distribution of the target image staining intensity after mapping and correction conforms to the standardized range of the reference template; evaluate the consistency of the target image staining intensity and the reference template; dynamically adjust the standardized intensity range of the reference template based on the mapping and correction results of the target image staining intensity.
[0116] The working principle and beneficial effects of the above technical solution are as follows: this embodiment first obtains the staining intensity value of each pixel in the image, compares the staining intensity distribution of the target image with the standardized intensity range of the reference template, and identifies the offset or deviation of the staining intensity in the target image; based on the intensity range matching result, a staining intensity mapping function is designed to linearly or nonlinearly map the staining intensity value of the target image to the standardized range of the reference template; secondly, the mapping function is applied to correct the staining intensity of the target image to eliminate the offset or deviation of the staining intensity so that it conforms to the standardized range of the reference template; local intensity optimization is performed on areas in the target image with uneven staining intensity distribution; finally, it is verified whether the distribution of the target image staining intensity after mapping and correction conforms to the standardized range of the reference template; the consistency of the target image staining intensity and the reference template is evaluated; based on the mapping and correction results of the target image staining intensity, the standardized intensity range of the reference template is dynamically adjusted. In step S2021 of the above scheme, a staining intensity mapping relationship is established. By obtaining the staining intensity value of each pixel in the image and comparing it with the standardized intensity range of the reference template, the offset or deviation of the staining intensity in the target image is identified, providing a data basis for mapping and correction. Based on the intensity range matching results, a staining intensity mapping function is designed to linearly or nonlinearly map the staining intensity value of the target image to the standardized range of the reference template to ensure the alignment and standardization of the staining intensity. Significance: It provides a scientific basis for staining mapping and correction, ensuring the accuracy and reliability of the operation. Through the design of the mapping function, the staining intensity of the target image can be aligned with the standardized range of the reference template, laying the foundation for staining standardization. Step S2022: Correction and optimization of staining intensity. The mapping function is applied to correct the staining intensity of the target image, eliminating the offset or deviation of the staining intensity, so that it conforms to the standardized range of the reference template and improves the overall consistency of the staining intensity. Local intensity optimization is performed on areas of uneven staining intensity distribution in the target image to ensure a more balanced distribution of staining intensity in the tissue and improve the local quality of the image. Significance: Through intensity correction, the staining intensity of the target image is made to meet the standardization requirements of the reference template as a whole, improving the usability and analyzability of the image; through local optimization, the problem of uneven staining intensity distribution is solved, ensuring that the image is clearer and more accurate in detail areas. Step S2023 Verification of the staining intensity mapping results and template update: Through statistical analysis methods, verify whether the distribution of the target image staining intensity after mapping and correction meets the standardization range of the reference template to ensure the accuracy of the correction results; through feature matching technology, evaluate the consistency of the target image staining intensity with the reference template to ensure the reliability and scientific nature of the mapping and correction; based on the mapping and correction results of the target image staining intensity, dynamically adjust the standardization intensity range of the reference template to ensure the universality and applicability of the template.Significance: Through verification and evaluation, the scientificity and practicality of the staining intensity mapping and correction results are ensured, providing a reliable basis for analysis; through dynamic adjustment of the reference template, the continuous improvement and long-term applicability of the staining standardization technology are ensured, and the universality and stability of the technology are enhanced.
[0117] In summary, this embodiment, through hierarchical technical operations, achieves the establishment of a staining intensity mapping relationship, correction and optimization of staining intensity, verification of mapping results, and template updating. Ultimately, the staining intensity of the target pathology image conforms to the standardized range of the reference template, improving the image's consistency and analyzability. Through local optimization and global correction, the image quality is improved in both overall and local areas, ensuring the uniformity of staining distribution and the rationality of contrast. Through dynamic adjustment of the reference template, the continuous improvement and long-term applicability of the staining standardization technology are ensured, providing a scientific basis and technical support for the standardized analysis of pathology images. This method is not only technologically innovative and practical, but also provides a systematic solution for the staining standardization of pathology images, with important scientific significance and application value.
[0118] Example 9: Figure 10 As shown, based on Example 1, the process of extracting features from the preprocessed full-slice image provided by the embodiment of the present invention includes the following steps:
[0119] S301: The Visual Transformer model divides the input full-slice image into multiple fixed-size image patches. Each patch is mapped to a high-dimensional feature vector. The image patches capture local features of the pathology image at different scales, such as cell nucleus morphology and staining intensity distribution. The Visual Transformer model uses a self-attention mechanism to globally calculate the correlation between image patches.
[0120] S302: In the multi-layer structure of the visual transformer model, each layer gradually refines features through a self-attention mechanism and a feedforward neural network. The lower layers extract local details (such as cell nucleus boundaries and staining particles), while the higher layers fuse local features to generate higher-level semantic features (such as tissue type and lesion area).
[0121] S303: The visual converter model uses position encoding to embed the spatial position information of the image block into the feature vector; through the feature vector of the last layer, a high-dimensional feature representation of the pathological image is generated, which contains local detail information and global semantic information.
[0122] The working principle and beneficial effects of the above technical solution are as follows: First, in this embodiment, the visual converter model divides the input full-slice image into multiple image blocks of fixed size, and each image block is mapped to a high-dimensional feature vector. The image block captures local features of the pathological image at different scales, such as the morphology of the cell nucleus, the distribution of staining intensity, etc.; the visual converter model uses the self-attention mechanism to calculate the correlation between image blocks on a global scale; secondly, in the multi-layer structure of the visual converter model, each layer gradually refines the features through the self-attention mechanism and the feedforward neural network; the low-level network extracts local detail features (such as cell nucleus boundaries, staining particles, etc.), while the high-level network fuses local features to generate higher-level semantic features (such as tissue type, lesion area, etc.); finally, the visual converter model uses position encoding to embed the spatial position information of the image block into the feature vector; through the feature vector of the last layer, a high-dimensional feature representation of the pathological image is generated, which contains local detail information and global semantic information. Step S301 of the above scheme involves image block partitioning and global context modeling. The full slide image is divided into fixed-size image blocks, each of which is mapped into a high-dimensional feature vector. This allows for the capture of local features of the pathology image at different scales, such as cell nucleus morphology and staining intensity distribution. This partitioning ensures the accurate extraction of local details. A self-attention mechanism is used to globally calculate correlations between image blocks, dynamically capturing long-range dependencies between different regions in the pathology image, such as the interaction between the tumor region and surrounding tissue, and the spatial distribution of staining intensity. Significance: Through image block partitioning and the self-attention mechanism, the model can simultaneously focus on local details and global structure, avoiding the limitations of single-scale feature extraction in traditional methods. Global context modeling enables the model to more comprehensively understand the overall semantic information of the image, providing richer feature input for subsequent tasks. Step S302 involves hierarchical feature extraction and semantic enhancement. Within the multi-layered structure of the visual transformer model, each layer gradually refines features through a self-attention mechanism and a feedforward neural network. Low-level networks extract local, detailed features (such as cell nuclear boundaries and staining particles), while high-level networks fuse these local features to generate higher-level semantic features (such as tissue type and lesion area). Through hierarchical feature extraction, the model can achieve semantic understanding from local to global, generating more discriminative feature representations. Significance: Hierarchical feature extraction enables the model to gradually extract higher-level semantic information from local details, enhancing the depth and richness of feature expression. Through semantic enhancement, the model can more accurately identify key areas and lesion features in pathology images, providing more reliable support for diagnosis and treatment.In step S303, position encoding and feature expression, the visual transformer model uses position encoding to embed the spatial position information of image blocks into feature vectors, enhancing the model's understanding of the spatial structure in pathological images, such as the morphological distribution of tumors and regional differences in staining intensity. Through the feature vectors of the final layer, a high-dimensional feature representation of the pathological image is generated, containing both local detail information and global semantic information, providing high-quality feature input for subsequent tasks. Significance: Position encoding enables the model to perceive the spatial relationship between image blocks, improving its understanding of spatial structure in pathological images. Through feature expression, the high-dimensional feature representation generated by the model not only contains local detail information but also incorporates global semantic information, enabling a comprehensive description of the content of the pathological image and providing more accurate feature input for subsequent tasks.
[0123] In summary, this embodiment achieves the integration of local details and global features through image block segmentation and global context modeling, improving the comprehensiveness of feature expression; through hierarchical feature extraction and semantic enhancement, it achieves semantic understanding from local to global, enhancing the model's discriminative ability; through position encoding and feature expression, it enhances the model's perception of spatial structure and generates high-quality feature output. These steps work together to enable the visual converter model to efficiently capture the complex features and subtle differences in pathological images, providing strong technical support for tasks such as pathological diagnosis and lesion detection, and has important clinical significance and application value.
[0124] Example 10: Figure 11 As shown, based on Example 9, the process of calculating the correlation between image blocks in a global scope provided by the embodiment of the present invention includes the following steps:
[0125] S3011: For each image patch's feature vector, the model generates query, key, and value vectors. The query vector represents the current image patch's focus, the key vector represents the potential relevance of other image patches, and the value vector conveys specific information.
[0126] S3012: Calculate the correlation weights between the current image block and other image blocks by performing a dot product operation on the query vector and the key vector. The weights reflect the dynamic correlation between the image blocks, such as the similarity of staining intensity and the spatial distribution of cell nuclear morphology.
[0127] S3013: Perform weighted summation on the value vector using the correlation weight to generate a global context feature of the current image block.
[0128] The working principle and beneficial effects of the above technical solution are as follows: First, for each image block's feature vector, the model generates query, key, and value vectors. The query vector represents the current image block's attention needs, the key vector represents the potential relevance to other image blocks, and the value vector conveys specific information. Secondly, the correlation weights between the current image block and other image blocks are calculated by performing a dot product operation on the query vector and the key vector. The weights reflect the dynamic relevance between the image blocks, such as similarity in staining intensity or the spatial distribution of cell nuclear morphology. Finally, the value vectors are weighted and summed using the correlation weights to generate the global context feature of the current image block. Step S3011 of the above solution generates query, key, and value vectors, decomposing the feature vector of each image block into query, key, and value vectors, each of which performs different functions: the query vector represents the current image block's attention needs, the key vector represents the potential relevance to other image blocks, and the value vector conveys specific information. Through this decomposition, the model can flexibly adjust the role of each image block in the global context, providing a basis for correlation calculation. Significance: The generation of query, key, and value vectors enables the model to dynamically focus on the relationship between image blocks, rather than relying solely on fixed local features. This mechanism provides the possibility for global feature interaction. By decomposing the feature vector, the model can describe the content of the image block in more detail, providing richer information for global correlation calculation. Step S3012 calculates the correlation weight. By performing a dot product operation on the query vector and the key vector, the model can quantify the dynamic correlation between the current image block and other image blocks. The weight reflects the similarity, spatial distribution patterns, etc. between image blocks. The calculation of the correlation weight enables the model to understand the relationship between image blocks from a global perspective, such as the similarity of staining intensity, the spatial distribution of cell nuclear morphology, etc. Significance: The calculation of the correlation weight enables the model to capture the relationship between image blocks on a global scale, breaking through the limitations of traditional methods that are limited to local features. By dynamically adjusting the correlation weight, the model can optimize feature expression according to task requirements, such as enhancing the contrast features between lesion areas and non-lesion areas in lesion area detection tasks. Step S3013 generates global context features. The value vectors are weighted and summed using correlation weights to generate global context features for the current image block. This ensures that the features of each image block not only contain its own local information but also incorporate the global information of other image blocks. The generation of global context features enables the model to comprehensively describe the content of the image block, providing high-quality feature input for tasks such as lesion detection and tissue classification. Significance: The generation of global context features achieves a deep fusion of local details and global semantics, enabling the model to more comprehensively understand the content of pathological images. By integrating global context information, the model can more accurately identify lesion areas and classify tissue types, significantly improving task performance.
[0129] In summary, this embodiment provides a basis for dynamic attention mechanisms and feature interactions by generating query, key, and value vectors; achieves dynamic feature interactions on a global scale by calculating correlation weights; and achieves a deep fusion of local details and global semantics by generating global context features. These steps together constitute a hierarchical and dynamic feature extraction process, providing a new technical path for the intelligent analysis of pathological images. Its significance lies in breaking through the limitations of traditional methods, achieving a comprehensive understanding and efficient use of pathological image features, and providing strong technical support for medical diagnosis and research.
[0130] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention's equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for pathological image preprocessing and feature extraction, characterized in that: The following steps are involved: Full-slice images were acquired for thumbnail extraction, and contaminants were removed using RGB and HSV dual color spaces; Use open-source tools to segment the whole-slide images after removing contaminants, separate pathological tissue from the background, and perform block processing; After tiling, an open-source library is used for staining standardization, and the image blocks are spliced together to form a large image for batch processing. A pre-trained visual transformer model is used to extract features from the pre-processed full-slide image, capturing the complex features and subtle differences in the standardized full-slide image. The preprocessing process includes image decontamination, tiling, staining standardization, and splicing of image blocks. The process of feature extraction for the preprocessed full-slice image includes the following steps: The visual transformer model divides the input full-slice image into multiple fixed-size image patches. Each patch is mapped to a high-dimensional feature vector. The image patches capture local features of the pathology image at different scales. The visual transformer model uses a self-attention mechanism to calculate the correlation between image patches globally. In the multi-layer structure of the visual transformer model, each layer gradually refines features through a self-attention mechanism and a feedforward neural network. The low-level network extracts local detail features, while the high-level network fuses local features to generate higher-level semantic features. The visual transformer model uses position encoding to embed the spatial position information of image patches into feature vectors. The feature vectors in the last layer are used to generate high-dimensional feature representations of pathological images, which contain both local details and global semantic information. The process of pollutant removal using the dual color space of RGB and HSV includes the following steps: In the RGB color space, the full-slice image after thumbnail extraction is decomposed into three independent channels: red, green, and blue; After preliminary filtering in the RGB domain, the whole-slice image is converted to the HSV color space. In the HSV space, dynamic threshold ranges are set for hue, saturation, and brightness to locate and remove interference information in the complex background while retaining the detailed features of the pathological tissue. Automatically adjust the weight ratio of RGB and HSV domains according to the specific characteristics of the whole slice image; automatically adjust the threshold range of RGB and HSV domains by analyzing the global and local characteristics of the image; The process of automatically adjusting the threshold range of RGB and HSV domains includes the following steps: Perform global feature analysis on the whole-slice image, including the overall brightness distribution, color distribution, and texture characteristics of the image; calculate the mean, variance, and histogram statistics of the image to preliminarily determine the distribution range and intensity of the pollutants in the image; The whole slice image is divided into several local areas, and the brightness, color and texture characteristics of each area are calculated respectively; Dynamically allocate weights of RGB and HSV domains based on the analysis results of global and local features; fuse the processing results of RGB and HSV domains through weighted averaging or adaptive fusion algorithm; If the contaminants in the whole-slice image are mainly manifested as intensity distribution, the weight of the RGB domain is increased and the intensity information of the RGB space is used for filtering; If the contaminants in the whole-slice image are mainly manifested as color distribution, the weight of the HSV domain is increased and the color information of the HSV space is used for filtering.
2. The method for pathological image preprocessing and feature extraction according to claim 1, wherein: For each red, green, and blue channel, a dynamic threshold range is set according to the characteristics of the full-slice image; by comparing the RGB value of each pixel with the preset threshold, pixels below the threshold are reset to white.
3. The method for pathological image preprocessing and feature extraction according to claim 1, wherein: After the threshold range is automatically adjusted, the image is decomposed into multiple scales, and the threshold range is set at different scales to adapt to pollutants of different sizes and types; after the pollutants are removed, the image is edge smoothed.
4. The method for pathological image preprocessing and feature extraction according to claim 1, wherein: The process of staining standardization using open source libraries includes the following steps: Extract staining distribution information from images through open source libraries, and statistically analyze staining features of multiple high-quality pathology images using open source libraries to establish staining reference templates. Based on the staining reference template, staining mapping and correction are performed on the target pathology image to map the staining intensity of the target image to the standardized range of the reference template; After the staining standardization is completed, statistical analysis methods are used to verify whether the distribution of staining in the image meets the standards of the reference template; feature matching technology is used to evaluate the consistency of staining intensity, uniformity and contrast.
5. The method for pathological image preprocessing and feature extraction according to claim 4, wherein: The process of establishing a staining reference template includes the following steps: Extract key parameters such as staining intensity, staining uniformity, and staining contrast from multiple high-quality pathology images using open-source libraries; Based on the statistical analysis results of staining intensity, a standardized intensity range was constructed and used as the core parameter of the staining intensity template. Based on the statistical analysis results of staining uniformity, a uniformity reference standard was established and used as the core parameter of the staining uniformity template. Based on the statistical analysis results of staining contrast, a standardized contrast threshold was defined and used as the core parameter of the staining contrast template.
6. The method for pathological image preprocessing and feature extraction according to claim 5, wherein: The staining intensity was quantified by the distribution of pixel values, the staining uniformity was statistically analyzed by the spatial distribution of staining in the tissue, and the staining contrast was calculated by the difference between the stained area and the background area.
7. The method for pathological image preprocessing and feature extraction according to claim 5, wherein: After the staining intensity is mapped to the standardized range of the reference template, the uniformity of staining in the tissue is improved through local contrast enhancement and staining distribution equalization; the contrast between staining and background is optimized according to the contrast threshold of the reference template.
8. The method for pathological image preprocessing and feature extraction according to claim 5, wherein: Through an iterative optimization method, the parameters of the staining intensity, uniformity, and contrast templates were adjusted; and the applicability of the staining reference templates in different images was verified.
Citation Information
Patent Citations
Multi-modal breast tumor risk prediction method and system based on pathological image
CN118675618A
Pathological image imaging quality assessment method
CN119048446A
Immunohistochemical digital pathological image processing method
CN119181091A
Image defogging method with high fidelity
CN106023110A
Lung adenocarcinoma HE staining pathological image tumor area multi-scale feature extraction and prognosis analysis method and device
CN114841947A