Systems and methods for preprocessing immunocytochemistry images for machine learning image-to-image translation
Patent Information
- Application Number
- CN202311545595.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-04
- Filing Date
- 2023-11-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-11-20
AI Technical Summary
然而,大尺寸的配对图像(染色和未染色的图像,每个图像的大小约为五千兆字节(~5GB))所获得的格式无法简单地使用机器学习和深度学习算法进行分析
Smart Images

Figure CN119808715B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to machine learning image-to-image conversion, and more specifically, to systems and methods for preprocessing immunocytochemical images to improve the accuracy of machine learning image-to-image conversion. Background Technology
[0002] Canine fine-needle aspiration (FNA) based on immunocytochemistry is commonly used to anatomically observe the localization of specific proteins or antibodies in cells by using specific primary antibodies called biomarkers that bind to specific proteins or antigens. However, the format of the large paired images (stained and unstained images, each approximately five gigabytes (~5 GB)) cannot be easily analyzed using machine learning and deep learning algorithms.
[0003] Currently, deep learning algorithms require significantly longer training times and are highly complex when dealing with paired images; furthermore, localization and labeling take far too long to track and annotate each cell. For meaningful machine learning, transforming image pixels used to depict cell features and morphology into features for analysis requires well-defined preprocessing before machine learning. This preprocessing involves converting the raw features into data that machine learning algorithms can understand and learn.
[0004] Therefore, there is a need for methods and systems for preprocessing immunocytochemical images for machine learning image-to-image conversion, which overcome the shortcomings of existing systems and methods and provide simpler machine learning with higher accuracy and reduced computation time. Furthermore, other desirable features and characteristics will become apparent from the following detailed description and appended claims, taken in conjunction with the accompanying drawings and the background of this disclosure. Summary of the Invention
[0005] According to at least one aspect of this embodiment, a method for preprocessing immunocytochemical images for machine learning is provided. The preprocessing method includes: receiving paired positive and negative polyprotein images; and staining and segmenting the paired polyprotein images. The method further includes: labeling each image pixel corresponding to a cell in the paired polyprotein image; and converting the image pixel information of the labeled image pixels into a tabular form to generate a table of cell coordinates and geometric features for each paired polyprotein image. The method further includes: generating two cell pairing tables for each paired polyprotein image based on Euclidean distance pairings before inputting them for the machine learning; wherein the Euclidean distance pairings are based on cell coordinates in the table of cell coordinates and geometric features for each paired polyprotein image.
[0006] According to another aspect of this embodiment, a system for preprocessing immunocytochemical images for machine learning is provided. The system includes an image receiving module, an image segmentation module, a cell labeling module, a table generation module, and a data output module. The image receiving module is configured to receive paired positive and negative polyprotein images, and the image segmentation module is configured to perform staining segmentation on the paired polyprotein images. The cell labeling module is configured to label each image pixel of the paired polyprotein images corresponding to cells in the paired polyprotein images, and the table generation module is configured to convert the image pixel information of the labeled image pixels into a tabular form to generate a table of cell coordinates and geometric features for each paired polyprotein image. The data output module is configured to, when input for the machine learning, generate two cell pairing tables for each paired polyprotein image based on Euclidean distance-based pairing, wherein the Euclidean distance-based pairing is based on cell coordinates in the table of cell coordinates and geometric features for each paired polyprotein image. Attached Figure Description
[0007] The accompanying drawings are used to illustrate various embodiments and explain the various principles and advantages according to the embodiments. Similar reference numerals in the drawings refer to the same or functionally similar elements in the various views. The drawings, together with the following detailed description, are incorporated into and form a part of the specification.
[0008] Figure 1 A framework diagram of the image preprocessing system according to this embodiment is shown.
[0009] Figure 2 includes Figure 2A , Figure 2B and Figure 2C The illustration shows the separation of hematoxylin-stained and eosin-stained images according to this embodiment, wherein... Figure 2A The images stained with hematoxylin show good separation. Figure 2B This demonstrates hue, saturation, value (HSV) separation and red-green-blue (RGB) separation. Figure 2C It shows the use of according to Figure 2B The developed filter effectively separates eosin-stained images.
[0010] Figure 3 includes Figure 3A and Figure 3B This illustrates instance segmentation of a cropped WG image according to this embodiment, thereby generating normalized images and labeled images at different scaling levels, wherein... Figure 3A Normalized WG-H images at eight different scaling levels are shown, while Figure 3B Eight labeled images at corresponding zoom levels are shown.
[0011] Figure 4 includes Figures 4A to 4D The image shown is a cropped and labeled image with cell atlas and cell counts according to this embodiment, wherein... Figure 4A Images of WG markers and their corresponding cell markers are shown. Figure 4B Images of PAX5 markers and their corresponding cell markers are shown. Figure 4C The cellular atlases of both are shown. Figure 4D A bar graph showing cell counts is displayed.
[0012] Figure 5 A scatter plot showing the normalized, scaled, and rounded values of WG immune cells and CD3 immune-positive cells according to this embodiment is shown.
[0013] Those skilled in the art will understand that the elements in the figure are shown for simplicity and clarity and are not necessarily depicted to scale, and that the numbers in the figure may have been normalized for simplicity and clarity. Detailed Implementation
[0014] The following detailed description is merely exemplary in nature and is not intended to limit the invention or its application and uses. Furthermore, it is not intended to be limited by the foregoing background art or any theory set forth in the following detailed description. The object of this embodiment is to provide a unique system and method for preprocessing immunocytochemical images for machine learning image-to-image conversion. The preprocessing system and method according to this embodiment provides improved robustness and accuracy compared to conventional methods and includes preprocessing protein images (e.g., Wright Giemsa(WG)-CD3 / PAX5 images) by converting and simplifying pixel information for each cell (even if the cell count is equal to or greater than 4 million) into a tabular form with geometric features. Furthermore, the preprocessing system and method according to this embodiment converts and simplifies pixel information for each cell in images with cell arrangements in paired image datasets with a prediction accuracy greater than 90% and up to 92%, requiring only a very low minimum learning time for machine learning on deep learning algorithms, for example, approximately 10 seconds or at most a few minutes compared to conventional preprocessing times of several days. Therefore, the system and method according to this embodiment provides a preprocessing scheme that improves accuracy, reduces computation time, and simplifies and makes machine learning understandable.
[0015] According to this embodiment, image preprocessing obtains paired WG-CD3 / PAX5 positive or negative images, and the staining used to label each cell is normalized and separated based on a Non-Maximum Separation (NMS) convolutional neural network (CNN) using StartDist. The pixel information (i.e., cell information) of the labeled images, along with coordinates and geometric features, is converted into a tabular form for pairing CD3 / PAX5 positive or negative cells based on minimum distance intervals. According to this embodiment, the preprocessing generates two cell pairing tables for each pair of WG-CD3 / PAX5 images before the input for machine learning.
[0016] refer to Figure 1Figure 100 illustrates the framework of the image preprocessing system according to this embodiment. The framework according to this embodiment receives paired whole-side images (WSI) of WG-stained cells in an image as input, the images being CD3 positive or negative (+ / -) and PAX5 positive or positive (+ / -). Then, in cropping and alignment steps 102a, 102b, the processing includes cropping and image registration. The registered pairs are separated into WG-CD3 / PAX5 positive or negative pairs. Then, the registered pairs are normalized and filtered 104a, 104b to separate hematoxylin and eosin staining into a normalized hematoxylin WG image 106, a normalized hematoxylin CD3 / PAX5 image 108, and a normalized eosin CD3 / PAX5 image 110. StarDist is a deep learning-based 2D and 3D cell nucleus detection method. It uses StarDist's NMS convolutional neural network 112 to perform image segmentation 106, 108, and 110, with cells labeled as WG cells 114a, CD3 / PAX5 immunopositive cells 114b, and CD3 / PAX5 immunonegative cells 114c. With the help of the Regionprops tool in Python, cell counts, coordinates, and their geometric features (e.g., area, orientation, principal axis length, diameter, orientation, solidity, and mean intensity) are generated for paired WG-CD3 / PAX5 labeled images in tabular form (including overlapping tables). Each table is sorted based on area, then normalized, scaled, and sorted 118 times. Once each table for paired WG-CD3 / PAX5 cells has been sorted 118 times, all cells have been paired from WG to their corresponding stained cells in CD3 / PAX5 by a minimum Euclidean interval 120, indicated by the overlapping tables. According to this embodiment, the pairing of cells from paired WG-CD3 / PAX5 images advantageously presents a simplified form of possible tabular data for machine learning, based on only nine parameters (x-coordinate, y-coordinate, area, principal axis length, secondary axis length, orientation, equivalent diameter, mean intensity, and solidity) for each counted cell to be analyzed. Therefore, the framework shown in Figure 100 preprocesses image information for machine learning image-to-image conversion and advantageously makes learning, prediction, and classification easier to use for machine learning, with sufficient accuracy in a shorter learning time.
[0017] In cropping steps 102a and 102b, the acquired data, including the full-side images (WSIs) of the WGs paired with a coarsely stained CD3 / PAX5 image of approximately 5 GB, are cropped to a lower size, such that each WG image is cropped 102a to approximately 200 MB to 250 MB in size. Each WSI image in 102a and 102b is cropped according to the normalization and labeling constraints of the StartDist algorithm by keeping the WSI size (in GB) divided by the number of cropped images less than or equal to approximately 200 MB to 250 MB. For example, if the WG image size is 2 GB and it is paired with a 1.5 GB CD3 / PAX5 image, the WG remains standard, and the paired images are cropped into 20 images in PNG format. QuPath software is used to read the WSIs from the original NDPI scan.
[0018] As described above, cropping steps 102a and 102b also include cell alignment. Cell alignment is used for image registration, and the first pass of image registration is adjusted using PTGui software and Python to pair the cropped WG image with the stained CD3 / PAX5 cropped image using a scale-invariant feature transform (SIFT) algorithm. To achieve this, fifteen to twenty-five cell coordinates are selected from each paired image, these coordinates having properties such as edge, scaling, rotation, and translation.
[0019] Before image annotation using StartDist 112, each image pair was advantageously normalized and filtered 104a, 104b prior to color separation for cell nucleus segmentation to achieve improved quantitative analysis. Reference Figure 2A Figure 200 shows a hematoxylin (H) staining image filtered 104a from a normalized WG / CD3 / PAX5 image 210 using conventional methods to derive a well-separated image 220, thereby deriving a normalized hematoxylin WG image 106 and a normalized hematoxylin CD3 / PAX5 image 108. It should be noted that hematoxylin staining helps to separate cyanocytes from the WG image.
[0020] However, as Figure 2B As shown in Figure 230, the eosin (E) tinted image is separated from the CD3 / PAX5 image 240 by using Hue, Saturation, Value (HSV) segmentation 250 and Red, Green, Blue (RGB) segmentation 255. Figure 230 shows the extracted HSV 252 segmentation values and the extracted RGB 257 segmentation values. Based on the HSV 252 and RGB 257 segmentation values, a filter is designed to fully separate the eosin tinted image 270 from the CD3 / PAX5 image 260, as shown below. Figure 2C As shown in Figure 260.
[0021] According to the framework of this embodiment, StartDist is used to perform instance segmentation 112 for cell detection. StartDist employs trained per-pixel cell segmentation followed by subsequent pixel grouping with shape refinement. To avoid any errors at cell intersection boundaries or diffusion segments, the framework 100 of this embodiment uses a non-maximum separation (NMS) convolutional neural network (CNN) of StartDist to locate cell nuclei via star-shaped convex polygons. This CNN predicts a polygon for each pixel representing a cell instance with a reasonable shape at that location. The NMS-CNN then labels and counts conflicting objects, regardless of color, and separates them based on a threshold. By separating the color information of the image, cells can be counted and labeled, and various quantitative analyses can be performed. Figure 2A Image 220 in the image shows an instance segmentation of one of the cropped WG images according to this embodiment.
[0022] After instance segmentation 112 of the image, the image is labeled as 114a, 114b, 114c. Figure 1 Then, Regionprops is used to extract information about the markers. Regionprops is an image processing function used to calculate various properties of pixel regions in an image, thereby measuring a set of properties for each marked region in a marker matrix or image. (Reference) Figure 3A and 3B , Figure 3A Normalized WG-H images at eight different scaling levels are shown, and Figure 3B Eight labeled images are shown at their respective zoom levels. Figure 3A In the first panel, one can notice that there are "n" regions whose centroid, coordinates, area, diameter, principal axis length, secondary axis length, average pixel intensity, and solidity can be easily measured. Therefore, a list or table with such information can be generated and extracted. According to this embodiment, information on cell coordinates and geometric features, such as the x-coordinate, y-coordinate, area, principal axis length, secondary axis length, orientation, equivalent diameter, average intensity, and solidity of each labeled cell or region in the paired WG-CD3 / PAX5 images used for registration, is generated and extracted into Table 116. Figure 1 Table 1 is... Figure 3A Here is an example of a table of cell coordinates and geometric features generated for a WG image.
[0023] 1 1572.241915 4293.429495 2319 54.338223 116.724450 0.979307 2 3206.334807 9831.170016 2488 56.283390 127.234727 0.977603 3 9738.976842 10023.875370 2375 54.990398 98.668632 0.979381 4 983.006436 3173.770919 3418 65.969180 118.709479 0.983597 5 1963.540326 798.415373 2641 57.988151 103.094283 0.982881
[0024] Table 1
[0025] After generating the labeling information 114a, 114b, and 114c, cell counting and various analyses can be performed. (Reference) Figures 4A to 4D The image shown is a cropped and labeled image with cell diagrams and cell counts according to this embodiment. Figure 4A Images of WG markers with their corresponding cell markers are shown, while Figure 4B An image of the PAX5 marker with its corresponding cell marker is shown. Figure 4C The cell diagrams of both are shown, while Figure 4D A bar graph of cell counts is shown, where bar 410 represents the cell counts in the WG-labeled image, and bar 420 represents the cell counts in the PAX5-labeled image. Therefore, after generating labeling information 114a, 114b, and 114c, as... Figure 4C and 4D As shown, cell maps (x, y) and cell counts can be generated. These cells in WG-labeled and PAX5-labeled images can then be classified into immunopositive and immunonegative cells.
[0026] When stained with CD3 / PAX5, each WG image is stained with information on immunopositive and immunonegative cells, and when segmented 112, immunopositive cells are labeled 114b and immunonegative cells are labeled 114c. Cells 114b and 114c in the CD3 / PAX5 image belong to the original WG image.
[0027] As part of normalization, scaling, and sorting 118, the features of the markers are normalized because the geometric attributes of each cell in the table 116 generated from Regionprops are normalized in their respective columns / matrices using minimum / maximum normalization. For example, scaling can scale the values to six-digit integer values before rounding. Rounding and sorting are performed using a Python script before pairing using the corresponding coordinates (x, y) 120 of immunopositive and immunonegative cells. Figure 5 The normalized, scaled, and rounded scatter plots show black WG immune cells and gray CD3 immune-positive cells.
[0028] A sorting matrix with a sorting unit size of 50 units was defined for each (x, y) coordinate of WG-CD3 / PAX5 paired immunopositive cells and WG-CD3 / PAX5 paired immunonegative cells. Each sorting matrix was then labeled as a sorting label matrix.
[0029] After rounding and sorting 118, cells from WG images paired with stained CD3 / PAX5 images were paired by finding the center coordinates with the smallest Euclidean interval in the paired images. Prior to pairing, WG cells and immunopositive / negative cells were sorted based on the x-values of the cut CD3 / PAX5 images.
[0030] After sorting, a script is formulated to find cells based on the proximity (x1, y1) of cells in the WG image to their corresponding immunopositive / negative cells in the CD3 / PAX5 image at approximately the same location, using Euclidean interval separation for pairing. For example, the first row (x1, y1) of the sorted WG image is compared with all rows of the sorted (x2, y2) CD3 / PAX5 image. This creates an n×n-dimensional distance matrix with distance information from each cell in the WG image to each cell in the CD3 / PAX5 image. From this matrix, WG cells are paired with CD3 / PAX5 cells based on the minimum distance between cells. Finally, a table is generated with the WG cells paired with CD3 / PAX5 cells and labeled with binary tags (1 or 0). Immunonegative cells are labeled with 0, while immunopositive cells are labeled with 1.
[0031] According to this embodiment, the preprocessing of immunocytochemical staining for improving the accuracy of machine learning involves normalization, staining separation, cell labeling based on nonmaximal separation CNN, and segmentation methods, and can be performed in AI-assisted pathology. For example, the preprocessing according to this embodiment can be advantageously applied to AI-assisted pathology for the early diagnosis of suspected cancer, where cancer cells are typically identified prior to machine learning by biomarkers such as hematoxylin and eosin-stained cells. More broadly, the preprocessing according to this embodiment can be advantageously used in cytopathology, immunocytochemistry, staining separation by digital filtering, and cell localization.
[0032] One application of the pretreatment system and method according to this embodiment that will benefit the global veterinary community is the immunocytochemical image analysis of canine lymphoma. This will assist the fine-needle aspiration (FNA) procedure by reducing diagnostic time and converting image data into tabular form for clinical advice. Furthermore, by using the system and method according to this embodiment, medical, veterinary, scientific, and engineering personnel will benefit from training in AI research and development for their respective goals and analyses. In the future, the system and method according to this embodiment can assist in the diagnosis of lymphoma, which can be used to help diagnostic laboratories, thereby improving diagnostic accuracy and the quality of life for cancer patients.
[0033] Therefore, it can be seen that the method and system according to this embodiment provide a novel and effective image information preprocessing for machine learning image-to-image conversion, making machine learning simpler, faster, and more efficient. The method and system according to this embodiment improve the accuracy of machine learning in converted images, reduce machine learning computation time, and make machine learning simpler and more understandable. Currently, deep learning algorithms require a long training time and are highly complex when applied to the localization and labeling of paired images of stained cells, requiring too much time to track and annotate each cell. For meaningful machine learning, converting image pixel information of cell features and morphology into features requires well-defined preprocessing before machine learning. Preprocessing transforms raw features into data that machine learning algorithms can understand and learn. This preprocessing not only simplifies learning but also achieves better accuracy and less computation time. According to this implementation, multi-protein images (e.g., WG-CD3 / PAX5 images) are preprocessed to convert and simplify the pixel information of each cell (which can be a cell count of up to 4 million cells) in tabular form. This pixel information has geometric features of cell arrangement in the paired image dataset, with a prediction accuracy of 92%. The minimum learning time for machine learning is about 10 seconds to a few minutes, while deep learning algorithms using traditional preprocessing require several days.
[0034] Although exemplary embodiments have been given in the foregoing detailed description of these embodiments, it should be understood that numerous variations exist. It should also be understood that the exemplary embodiments are merely examples and are not intended to limit the scope, applicability, operation, or configuration of the invention in any way. The foregoing detailed description will provide those skilled in the art with a convenient roadmap for implementing the exemplary embodiments of the invention, and it should be understood that various changes can be made to the functionality and arrangement of the steps and methods of operation described in the exemplary embodiments without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A method for preprocessing immunocytochemical images for machine learning, the method comprising: Receive paired positive and negative polyprotein images; Staining and segmenting paired multi-protein images; Each image pixel is labeled with the cell corresponding to the paired multiprotein image; The image pixel information of the labeled image pixels is converted into a table format to generate a table of cell coordinates and geometric features for each of the paired multi-protein images; as well as Before inputting into the machine learning, two cell pairing tables are generated for each paired multiprotein image based on Euclidean distance pairing, wherein the Euclidean distance pairing is based on cell coordinates in a table of cell coordinates and geometric features for each paired multiprotein image, wherein the geometric features include at least one of area, principal axis length, secondary axis length, orientation, equivalent diameter, average strength, and solidity.
2. The method according to claim 1, wherein, The paired multiprotein images include paired WG-CD3 / PAX5 images.
3. The method of claim 1, further comprising staining the received paired multiprotein image with hematoxylin and / or eosin prior to segmentation.
4. The method of claim 3, further comprising normalizing the staining of the paired multiprotein images prior to segmentation.
5. The method according to claim 1, wherein, The segmentation includes: performing instance segmentation by staining the paired multiprotein images using cell nucleus detection.
6. The method according to claim 5, wherein, The instance segmentation of staining in the paired multi-protein images using cell nucleus detection includes: using a StarDist nonmaximal separation convolutional neural network to separate the staining.
7. The method of claim 1, further comprising sorting cells of the paired multiprotein images based on cell coordinates in a table of cell coordinates and geometric features for each of the paired multiprotein images.
8. The method according to claim 7, wherein, Cell sorting based on the cell coordinates of the paired multi-protein image includes: Based on the cell coordinates, a sorting matrix comprising multiple sorting units is defined; and In response to the cell coordinates, each cell is sorted into one of the plurality of sorting units.
9. The method according to claim 1, wherein, The Euclidean distance-based pairing includes: pairing corresponding cells in each of the paired multi-protein images based on the minimum Euclidean interval.
10. The method according to claim 1, wherein, Receiving the paired positive and negative polyprotein images includes: receiving a full-view image of multiple protein cells and cropping the full-view image to generate the paired positive and negative polyprotein images.
11. A system for preprocessing immunocytochemical images for machine learning, the system comprising: The image receiving module is configured to receive paired positive and negative polyprotein images; The image segmentation module is configured to perform staining segmentation on paired multi-protein images; A cell labeling module is configured to label each image pixel of the paired multiprotein image corresponding to a cell in the paired multiprotein image; A table generation module is configured to convert the image pixel information of the labeled image pixels into a table format to generate a table of cell coordinates and geometric features for each of the paired multi-protein images; as well as The data output module is configured to generate two cell pairing tables for each paired multiprotein image when the input is used for the machine learning, based on Euclidean distance-based pairing, wherein the Euclidean distance-based pairing is based on cell coordinates in a table of cell coordinates and geometric features for each paired multiprotein image, wherein the geometric features include at least one of area, principal axis length, secondary axis length, orientation, equivalent diameter, average strength, and solidity.
12. The system according to claim 11, wherein, The paired multiprotein images include paired WG-CD3 / PAX5 images.
13. The system of claim 11, further comprising a cell staining module configured to stain the received paired multiprotein image with hematoxylin and / or eosin, wherein, The cell staining module is connected to the image segmentation module to provide it with stained images.
14. The system according to claim 13, wherein, The cell staining module is further configured to normalize the staining of the paired multiprotein images before providing the stained images to the image segmentation module.
15. The system according to claim 11, wherein, The image segmentation module performs instance segmentation by using cell nucleus detection to stain the paired multi-protein images.
16. The system according to claim 15, wherein, The image segmentation module uses the StarDist nonmaximal separation convolutional neural network to separate staining to perform instance segmentation of the staining in the paired multiprotein image using cell nucleus detection.
17. The system of claim 11, further comprising a sorting module configured to sort cells of the paired multiprotein images based on cell coordinates in a table of cell coordinates and geometric features for each of the paired multiprotein images.
18. The system according to claim 17, wherein, The sorting module is configured to sort cells in the paired multiprotein image based on the cell coordinates by defining a sorting matrix comprising multiple sorting units based on the cell coordinates, and sorting each cell into one of the multiple sorting units in response to the cell coordinates.
19. The system according to claim 11, wherein, The Euclidean distance-based pairing includes: pairing corresponding cells in each of the paired multi-protein images based on the minimum Euclidean interval.
20. The system according to claim 11, wherein, The image receiving module is configured to receive full-side images of multiple protein cells and crop the full-side images to generate the paired positive and negative multiprotein images.
Citation Information
Patent Citations
Detecting cells of interest in large image datasets using artificial intelligence
CN113678143A
Liquid-based cell pathology image generation method based on weak supervised learning
CN113903030A