Blood smear image staining standardization method, device and equipment and storage medium
By adopting a decoupled deep attention model in the normalization of blood smear image staining, the decomposition task is to enhance the low-quality image and unify the staining style. The Laplace pyramid generative network using the attention mechanism has solved multiple problems in the existing technology and achieved better dyeing standardization effect and image quality.
Patent Information
- Application Number
- CN202510099269.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has several problems with standardization of blood smear image staining, including limited generalization ability due to statistical characteristics, strong dependence on domain labels, inability to deal with unknown domains, and performance bottlenecks caused by degradation and staining as a whole.
A decoupled deep attention model is proposed to decompose the standardization task of blood smear image staining into two stages: low-quality image enhancement and staining style unification. By constructing a generative network of Laplace pyramids based on attention mechanisms, independent representations of cell morphology and staining profiles were extracted and staining standardization was achieved without changing the underlying cell structure.
It achieves better dyeing standardization effect, improves image quality, overcomes the dependence on domain labels in the prior art, and improves the generalization ability of the model.
Smart Images

Figure CN120147222A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of pathological image staining standardization, and particularly to a method, device, equipment and storage medium for blood smear image staining standardization. Background Technique
[0002] In recent years, the development of generative adversarial networks (GANs) has inspired the emergence of a variety of GAN-based H&E staining image transfer methods. By using conditional generators, many existing frameworks are able to transfer images from one or more staining domains to a given target staining domain. In addition, the success of CycleGAN in achieving unsupervised domain transfer has prompted the development of one-to-one and many-to-one staining frameworks using cycle consistency. The literature "BenTaieb A, Hamarneh G. Adversarial stain transfer for histopathology image analysis[J]. IEEE transactions on medical imaging, 2017, 37(3): 792-802." designs a generative network to learn dataset-specific staining attributes and image-specific color transformations to achieve stain transfer. Inspired by CycleGAN, the literature "M.T. Shaban, C. Baur, N. Navab, and S. Albarqouni, “Staingan: Stain style transfer for digital histological images,” in 2019 Ieee 16th international symposium on biomedical imaging(Isbi 2019). IEEE, 2019, pp. 953–956." proposes an end-to-end deep learning solution to transfer the stain of an image from the source domain to the target domain in an end-to-end manner. CAGAN introduces a color-adaptive generative network to achieve normalization, which combines supervised learning in the target domain and unsupervised learning in the source domain to enhance the performance of the model. The literature "H. Liang, K.N. Plataniotis, and X. Li, “Stain style transfer of histopathology images via structure-preserved generative learning,” in International Workshop on Machine Learning for Medical Image Re-construction. Springer, 2020, pp. 153–162." proposes two generative adversarial network-based stain style transfer models, SSIM-GAN and DSCSI-GAN. By combining structure-preserving metrics and the feedback of an auxiliary diagnostic network during the learning process, medical-related information, including image texture, structure, and chromatic contrast features, can be retained in the color-normalized images, and DSCSI-GAN is used to improve the normalization.The literature "J. Vasiljevic, F. Feuerhake, C. Wemmert, and T. Lampert, 'Towards histopathological stain invariance by unsupervised domain augmentation using generative adversarial networks,' Neurocomputing, vol. 460, pp. 277–291, 2021." proposed an unsupervised augmentation method based on adversarial image-to-image translation to achieve realistic transformation between the source domain and the target domain. The literature "Wagner S J, Khalili N, Sharma R, et al. Structure-preserving multi-domain stain color augmentation using style-transfer with disentangled representations[C] / / Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part VIII 24. Springer International Publishing, 2021: 257-266." proposed a novel color augmentation technique to simulate various realistic histological stains, enabling the neural network to be stain-invariant when applied during training. The literature "D. Mahapatra, B. Bozorgtabar, J.-P. Thiran, and L. Shao, 'Structure preserving stain normalization of histopathology images using self supervised semantic guidance,' in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2020, pp. 309–319." proposed a self-supervised method to incorporate semantic information into a GAN-based stain normalization framework to preserve the structural information of tissues.These GAN-based methods require defining multiple different staining domains according to laboratory sources and cannot handle unknown domains. Therefore, there is an urgent need to develop an algorithm that can treat the entire set of real staining domains as coming from a single domain, thereby breaking its dependence on domain labels, making the model have better generalization ability and being more suitable for practical applications.
[0003] The Laplacian pyramid method decomposes an image into band-pass images at octave intervals and a low-resolution residual image, and performs reconstruction through the low-frequency residual. This method has been successfully applied to various deep learning tasks. The literature "Denton E L, Chintala S, Fergus R. Deep generative image models using a laplacian pyramid of adversarial networks[J]. Advances in neural information processing systems, 2015, 28." uses a cascaded convolutional network within the Laplacian pyramid framework to generate high-quality natural image samples in a coarse-to-fine manner, and trains a separate generative model at each level of the pyramid using the generative adversarial network method to capture the image structure at a specific scale. The literature "Lin T, Ma Z, Li F, et al. Drafting and revision: Laplacian pyramid network for fast high-quality artistic style transfer[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021: 5141-5150." introduces a Laplacian pyramid feed-forward method. It first transfers the global style pattern at low resolution through a sketch network, then modifies the local details at high resolution through a refinement network, and finally obtains the stylized image by aggregating the outputs of all pyramid levels. The literature "Lai W S, Huang J B, Ahuja N, et al. Deep laplacian pyramid networks for fast and accurate super-resolution[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 624-632." proposes a Laplacian pyramid super-resolution network for gradually reconstructing the sub-band residuals of a high-resolution image. At each pyramid level, the model takes a coarse-resolution feature map as input, predicts the high-frequency residuals, and uses transposed convolution for upsampling to a finer level.The literature "Liang J, Zeng H, Zhang L. High-resolution photorealistic image translation in real-time: A laplacian pyramid translation network[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:9392-9400." reveals that attribute transformations (such as illumination and color operations) are more related to low-frequency components, while content details can be adaptively refined on high-frequency components. Based on this, a laplacian pyramid transformation network is proposed to perform these two tasks simultaneously, and a progressive masking strategy is adopted to effectively refine high-frequency components. The literature "Li F, Hu Z, Chen W, et al. A laplacian pyramid based generative h&e stain augmentation network[J]. IEEE Transactions on Medical Imaging, 2023." proposes a generative stain augmentation network to simulate real stain changes. This network adopts a laplacian pyramid-based generator architecture and can separate stains from cell morphology. However, this network lacks more context information to associate spatial features. In addition, it treats the degradation and normalization problems as a whole to handle, and will face performance bottlenecks because even the best network is difficult to take into account all problems at the same time and can only sacrifice performance to seek a balance between the two.
[0004] These methods can achieve stain normalization to a certain extent, but they have the following problems:
[0005] (1) Traditional methods use the statistical characteristics of images for stain normalization, which highly depends on the statistical similarity between the source image and the target image, resulting in limited generalization ability;
[0006] (2) GAN-based methods need to define multiple different stain domains according to laboratory sources, have a strong dependence on domain labels, and cannot solve unknown domains;
[0007] (3) Previous studies treat degradation and staining as a whole to handle, facing performance bottlenecks, resulting in poor processing effects;
[0008] (4) These methods are all for H&E images, and the research on stain normalization algorithms for blood smear images is severely lacking. Summary of the Invention
[0009] The present application provides a method, apparatus, device and storage medium for standardizing the staining of blood smear images to solve the above technical problems.
[0010] In a first aspect, the present application provides a method for standardizing the staining of blood smear images, including:
[0011] Obtain a blood smear image and a light prior map, merge the light prior map into the blood smear image, simulate the interaction between regions under different lighting conditions, extract a light feature map, and aggregate the extracted light feature map to generate a light map, and use the light map to enhance the blood smear image to obtain a first image;
[0012] Obtain an enhanced image according to the first image and the light feature map;
[0013] Decompose the enhanced image into a group of band-pass images with octave intervals and a low-resolution layer residual image;
[0014] Extract a first feature map from the low-resolution layer residual image, extract a first morphological representation and a staining contour from the first feature map, and combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image;
[0015] Extract a second morphological representation from the group of band-pass images with octave intervals, and combine the second morphological representation with the staining contour to obtain a standardized band-pass image.
[0016] In a possible design, obtaining a blood smear image and a light prior map, merging the light prior map into the blood smear image, simulating the interaction between regions under different lighting conditions, extracting a light feature map, and aggregating the extracted light feature map to generate a light map, and using the light map to enhance the blood smear image to obtain a first image, includes:
[0017] Merge the blood smear image with the light prior map, fuse the light prior map into the blood smear image using unit convolution, and use depthwise separable convolution to simulate the interaction between regions under different lighting conditions, and further extract features to generate a light feature map F lu ;
[0018] Aggregate the light feature map F using unit convolution lu to generate a light map;
[0019] Decompose the blood smear image into a reflectance component and a lighting component through the following formula:
[0020]
[0021] Wherein, S represents the blood smear image, R represents the reflectance component, and I represents the illumination component. represents element-wise multiplication;
[0022] Introduce perturbations of the reflectance component and the illumination component into the blood smear image to obtain an intermediate image, denoted as:
[0023]
[0024] Wherein, S1 represents the intermediate image. represents the perturbation of the reflectance component. represents the perturbation of the illumination component;
[0025] Multiply the intermediate image by the illumination map to obtain:
[0026]
[0027] Wherein, represents the illumination map. represents the noise and artifacts amplified during the lighting process of the low-exposure image. represents underexposure, overexposure, and / or color distortion caused during the lighting process;
[0028] Simplify formula (3) to obtain:
[0029]
[0030] Wherein, S lu represents the first image.
[0031] In a possible design, according to the first image and the illumination feature map, calculate the enhanced image through the following formula:
[0032] I en = g(S lu , F lu ) (5)
[0033] Wherein, g represents a damage repairer, I en represents the enhanced image, S lu represents the first image, F lu represents the illumination feature map.
[0034] In a possible design, the damage repairer includes an encoder and a decoder; wherein, the encoder is used to downsample the input S lu to match F luThe dimension is reduced by the downsampling process while deep feature extraction is performed. The downsampling process is divided into the first level and the second level. Both the first level and the second level are composed of an IFMT module and a conv 4×4 module. After the downsampling processing of the first level and the second level, an intermediate feature map is obtained.
[0035] The decoder is used to upsample the intermediate feature map. The upsampling is divided into the third level and the fourth level. Both the third level and the fourth level are composed of a dconv 2×2 module, a conv 1×1 module, and an IFMT module. The output of the dconv2×2 module is concatenated with the output of the IFMT module in the corresponding level of the downsampling process to alleviate the loss of image information during the downsampling process. After the upsampling processing of the third level and the fourth level, a second feature map is obtained, and the dimension of the second feature map is adjusted to restore it to the RGB three channels to obtain a restored image. The restored image is connected with the first image by residual connection to obtain an enhanced image.
[0036] In a possible design, the IFMT module includes a first multi-head self-attention mechanism, a second multi-head self-attention mechanism, layer normalization, a feed-forward network, and a residual connection.
[0037] The data processing process of the first multi-head self-attention mechanism is as follows:
[0038] And the input feature F in is reshaped into the feature X;
[0039] The feature X is divided into k heads by the following formula:
[0040] X = [X 1 , X 2 , …, X i , …, X k (6)
[0041] In the formula, X i represents the feature of the i-th head, represents the set of real numbers, H represents the height of the image, W represents the width of the image, d k represents the feature dimension of each head, and C represents the number of channels;
[0042] For each head i , three bias-free fully connected layers are used to linearly project X i to
[0043]
[0044] In the formula, are the learnable parameters of the fully connected layer, T is the matrix transpose operation, and Q i is the query, K i is the key, and V i is the value;
[0045] Using the illumination feature map F lu as a guide for the calculation of the self-attention mechanism to enhance the interaction between regions of different exposure levels, reshape F lu into features and divide the feature Y into k heads:
[0046] Y = [Y 1 , Y 2 , …, Y i , …, Y k (8)
[0047] where Y i represents the feature of the i-th head,
[0048] Express the self-attention of each head i as:
[0049]
[0050] where are learnable parameters used for adaptive scaling of matrix multiplication, Attention is the data processing process of the multi-head self-attention mechanism, and softmax is the softmax function;
[0051] Connect the k heads and send them into the fully connected layer, and add the position encoding to obtain the output token
[0052] Reshape X out to obtain the output feature
[0053] The data processing process of the second multi-head self-attention mechanism is as follows:
[0054] Process the input feature through Linear and depthwise separable convolution, and after passing through the SiLU activation function, expand it from the four directions of up, down, left, and right of the image using 2D scanning. The height and width after the image is flattened are merged and converted into the token length, and each sequence in the scanning is input into the feature extraction module for global feature extraction. The calculation process is expressed as:
[0055]
[0056] where x t is the input feature, and yt as the output feature, and are both learnable parameters, h t is the state vector at time step t;
[0057] Sum and combine the outputs in four different directions of the extracted features y 1 , y 2 , y 3 and y 4 and adjust the size of the output feature to be the same as that of the input feature;
[0058] Set the number of hidden layers in the 2D scan to double as the level of the IFMT deepens, so as to be able to extract deeper features layer by layer from the vector containing the integrated illumination features.
[0059] In a possible design, extract features from the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation and a staining contour from the first feature map, and combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image, including;
[0060] Construct a low-resolution layer residual network and a style network; wherein, the low-resolution layer residual network is used to extract features from the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation from the first feature map, and combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image; the style network is used to extract the staining contour from the first feature map;
[0061] The low-resolution layer residual network obtains the standardized low-resolution layer residual image through the following method:
[0062] Obtain the low-resolution layer residual image; use two conv 3×3 units and LeakyReLU units to extract features, perform deeper feature extraction through two residual self-attention mechanism modules, and finally obtain the first feature map through a conv 3×3 unit;
[0063] Use a conv 3×3 unit, an InstanceNorm2d unit, a LinearBlock unit and a LeakyReLU unit to extract the first morphological representation from the first feature map and combine it with the staining contour extracted by the style network, and send it into the generator. First, pass through two residual self-attention mechanism modules, and then pass through a conv 3×3 unit, a LeakyReLU unit and two conv 3×3 units to obtain the standardized low-resolution layer residual image
[0064] The style network includes an encoding module and a decoding module. The style network extracts the staining contour from the first feature map through the following method:
[0065] In the encoding module, a Color Transform module is designed based on conv 1×1, BatchNorm2d, and ReLU to transform the first feature map, improve the color consistency of the input features, and reduce the influence of non-target style features on the generation of the style vector;
[0066] AdaptiveAvgPool2d is used for feature compression in the spatial dimension to obtain a global feature representation; during the generation of the global feature representation, a channel attention mechanism is introduced to adaptively adjust the importance of channels, enabling the network to focus on features that significantly contribute to the style information; the output of the encoding module is the mean and variance, and the mean and variance are used to generate the style feature vector through VAE;
[0067] In the decoding module, the style feature vector is reconstructed through a fully connected layer and further non-linearly transformed through an activation function to generate the adaptive instance normalization parameters (α i , β i ) to represent the staining contour.
[0068] In a second aspect, the present application provides a blood smear image staining standardization device, including:
[0069] An image enhancement module, configured to obtain a blood smear image and a lighting prior map, merge the lighting prior map into the blood smear image, simulate the interaction between regions under different lighting conditions, extract a lighting feature map, aggregate the extracted lighting feature map to generate a lighting map, and use the lighting map to enhance the blood smear image to obtain a first image; obtain an enhanced image based on the first image and the lighting feature map;
[0070] A staining style unification module, configured to decompose the enhanced image into a set of band-pass images with octave intervals and a low-resolution layer residual image; extract a first feature map from the low-resolution layer residual image, extract a first morphological representation and a staining contour from the first feature map, combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image; extract a second morphological representation from the set of band-pass images with octave intervals, and combine the second morphological representation with the staining contour to obtain a standardized band-pass image.
[0071] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one processor and a memory; the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the blood smear image staining standardization method described in the first aspect above and various possible designs of the first aspect.
[0072] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the blood smear image staining standardization method described in the first aspect above and various possible designs of the first aspect are implemented.
[0073] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the blood smear image staining standardization method described in the first aspect above and various possible designs of the first aspect are implemented.
[0074] The blood smear image staining standardization method, device, equipment and storage medium provided by the present application are used to realize the standardization of blood smear image staining. On the one hand, the present application develops a decoupled deep attention model, which decouples the standardization task into two stages: low-quality image enhancement and staining style unification, breaking through the performance bottleneck faced by existing standardization algorithms in treating degradation and staining as a whole. On the other hand, the present application constructs a Laplacian pyramid generative network based on the attention mechanism, and extracts two independent representations from the input image: cell morphology as content and staining contour as style. By combining the staining representation extracted from the reference image with the morphological representation of the input image, staining standardization is achieved without changing the underlying cell structure. In addition, a low-resolution layer residual network based on the attention mechanism and a style network based on the channel attention mechanism are designed, which avoid the forgetting problem of deep networks and realize global attention and adaptive adjustment of weights. The present application regards the whole set of real staining appearances as coming from a single domain, providing a new idea for breaking the dependence of generative adversarial networks on domain labels during standardization. Description of the Drawings
[0075] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application and used together with the description to explain the principles of the present application.
[0076] Figure 1 It is a flowchart of the blood smear image staining standardization method provided by an embodiment of the present application;
[0077] Figure 2 It is an overall network framework diagram for implementing the blood smear image staining standardization method provided by an embodiment of the present application;
[0078] Figure 3 The network structure diagram of Net1 for realizing low-quality image enhancement provided by the embodiments of the present application;
[0079] Figure 4 The schematic diagram of the design details of IFMT provided by the embodiments of the present application;
[0080] Figure 5 The network structure diagram of Net2 for realizing unified staining style provided by the embodiments of the present application;
[0081] Figure 6 The analysis diagram of each layer in the Laplacian pyramid provided by the embodiments of the present application;
[0082] Figure 7 The structure diagram of the residual self-attention mechanism module provided by the embodiments of the present application;
[0083] Figure 8 The visual comparison between the method of the present application and other methods on the SegPC2021 dataset provided by the embodiments of the present application Figure 1 ;
[0084] Figure 9 The visual comparison between the method of the present application and other methods on the SegPC2021 dataset provided by the embodiments of the present application Figure 2 ;
[0085] Figure 10 The visual comparison between the method of the present application and other methods in cell segmentation provided by the embodiments of the present application Figure 1 ;
[0086] Figure 11 The visual comparison between the method of the present application and other methods in cell segmentation provided by the embodiments of the present application Figure 2 ;
[0087] Figure 12 The visual comparison diagram between the method of the present application and other methods in the segmentation of myeloma plasma cells provided by the embodiments of the present application;
[0088] Figure 13 The structural schematic diagram of the blood smear image staining normalization device provided by the embodiments of the present application;
[0089] Figure 14 The structural schematic diagram of the electronic device provided by the embodiments of the present application.
[0090] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Implementation Modes
[0091] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation modes described in the following exemplary embodiments do not represent all implementation modes consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0092] In the technical solution of the present application, the collection, storage, use, processing, transmission, provision, and disclosure of information such as financial data, user data, or medical image data all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0093] It should be noted that in the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0094] The following uses specific embodiments to describe in detail the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application in conjunction with the accompanying drawings.
[0095] An embodiment of the present application provides a method for staining a blood smear image. Figure 1 is a flowchart of the method for staining a blood smear image provided by the embodiment of the present application. As Figure 1 shown, the method for staining a blood smear image can be implemented based on an electronic terminal, and the method for staining a blood smear image includes the following steps S10 to S50.
[0096] S10. Obtain a blood smear image and a lighting prior map, merge the lighting prior map into the blood smear image, simulate the interaction between regions under different lighting conditions, extract a lighting feature map, and aggregate the extracted lighting feature map to generate a lighting map, and use the lighting map to enhance the blood smear image to obtain a first image;
[0097] S20. Obtain an enhanced image according to the first image and the lighting feature map;
[0098] S30. Decompose the enhanced image into a group of band-pass images with octave intervals and a low-resolution layer residual image;
[0099] S40. Extract features from the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation and a staining contour from the first feature map, and combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image;
[0100] S50. Extract features from a group of band-pass images with an octave interval to obtain a second morphological representation, and combine the second morphological representation with the staining contour to obtain a standardized band-pass image.
[0101] In this embodiment, the purpose of the blood smear image staining standardization method is to achieve blood smear image staining standardization. This embodiment designs a network framework to implement the above steps S10 to S50 respectively. The network framework is as Figure 2 shown. In this network framework, the blood smear image staining standardization task is decoupled into two stages: low-quality image enhancement and staining style unification. Among them, the low-quality image enhancement stage corresponds to the above steps S10 - S20, and the staining style unification stage corresponds to the above steps S30 - S50.
[0102] To achieve blood smear image staining standardization, this embodiment designs two subnets, Net1 and Net2, to model these two stages respectively. Among them, the task of Net1 is to repair images with uneven quality, and the task of Net2 is to unify images with different staining styles. In addition, to verify the impact of the standardization algorithm on downstream tasks, myeloma plasma cell segmentation will be used to evaluate the proposed method later. Two branches, Branch1 and Branch2, are used to segment cells and cell nuclei respectively, and finally the segmented cells and cell nuclei are added to obtain the final segmentation result.
[0103] Net1 combines Transformer and Mamba to develop a structure called MambaFormer, and embeds it into the Retinex model to construct a low-quality image enhancement framework called RetinexMambaFormer. This framework not only retains the long-range dependence and non-local interaction modeling ability of Transformer, but also utilizes the selection mechanism of Mamba to achieve linear-level global and local attention. Net2 combines the Laplacian pyramid and the attention mechanism to achieve staining style unification in a generative manner. It is worth mentioning that in this stage, a low-resolution layer residual network based on the residual attention mechanism is designed to avoid the forgetting problem of deep networks and achieve adaptive adjustment of weights; a style network based on the channel attention mechanism is also designed to reallocate weights using channel attention information, enabling it to better extract staining contours and thus achieve better standardization.
[0104] The overall structure of Net1 is as Figure 3As shown, the Net1 network consists of a light estimator (a) and a damage repairer (b). The light estimator is designed based on the idea of the Retinex model, and the damage repairer is constructed based on Illumination - Guided Mamba and Transformer (IGMT). As Figure 3 shown in (b) therein, the core component of IGMT is Illumination Fused Mamba and Transformer (IFMT), which consists of a light - guided multi - head self - attention mechanism, layer normalization, Vision Mamba, and a feed - forward network. The design details of IFMT are as Figure 4 shown.
[0105] The traditional Retinex model decomposes the observed image into a reflectance component and an illumination component where the reflectance component is an inherent property of the object and does not change with external factors, and the brightness of the image is presented by the illumination component:
[0106]
[0107] where represents element - wise multiplication. However, this theory ignores the uneven light distribution in real - world exposure scenarios or the noise and artifacts generated in low - exposure environments, and these factors are usually infinitely amplified during the enhancement process. Therefore, inspired by the Retinex model, this embodiment uses the perturbation terms proposed by Retinexformer for modeling, and introduces perturbation terms for the reflectance component and the illumination component in formula (1):
[0108]
[0109] where is the perturbation of the reflectance component, is the perturbation of the illumination component. Regarding R as a well - exposed image, in order to enhance S1, multiply both sides of formula (2) element - wise by the illumination map where
[0110]
[0111] where represents the noise and artifacts amplified during the lighting process of the low - exposure image, represents the underexposure / overexposure and color distortion caused during the lighting process. Based on this, formula (3) can be simplified:
[0112]
[0113] wherein is the first image, and the first image is a Lit-up image, is all the damages mentioned above. Therefore, the data processing process of network Net1 can be expressed as:
[0114] (S lu , F lu ) = f(S, I p )
[0115] I en = g(S lu , F lu ) (5)
[0116] where f(·) represents a light estimator, and g(·) represents a damage repairer. f(·) takes the low-quality image S and the light prior as inputs, and the Lit-up image S lu and the light feature map as outputs. Wherein I p = mean c (S), and mean c (·) is the average value of each pixel point of the image over all color channels. Then S lu and F lu are input into g(·) to repair the damages magnified during the lighting process of the image and generate the final enhanced image
[0117] Light estimator. The design details of the light estimator are as shown in Figure 2 (a) below. First, the low-quality image S is merged with the light prior I p , and I p is fused into S using conv 1×1. Then, a 5×5 depthwise separable convolution is used to simulate the interaction between regions under different lighting conditions, and further extract features to generate the light feature map F lu . Finally, another conv 1×1 is used to aggregate F lu to generate the light map Set as a three-channel RGB tensor to improve its representation ability in simulating the nonlinear relationship between RGB channels, thereby enhancing the color. In addition, in order to enable each convolutional layer to extract useful features, maintain the stability of training, and accelerate the convergence speed, Batchnorm layers are added after the first conv 1×1 and the 5×5 depthwise separable convolution respectively, and ReLU is used immediately afterwards to improve the nonlinear expression ability of the network.
[0118] Damage repairer. The design details of the damage repairer are as follows Figure 2 As shown in (b), it consists of an encoder and a decoder based on Illumination-Guided Mamba and Transformer. The encoder and decoder represent the downsampling and upsampling processes, respectively. They are symmetrical and divided into two levels. First, S is obtained by conv 3×3. lu Downsample to match F lu dimension. Next, downsampling is used to reduce the computational complexity and perform deep feature extraction. The downsampling process is divided into two levels, each level consists of an Illumination-Guided Mamba and Transformer (IFMT, i.e., the MambaFormer structure proposed in this embodiment) and a conv 4×4 with a step size of 2. After each convolution operation, the size of the image is halved, while the feature dimension is doubled. Therefore, after two levels of downsampling, the dimension of the obtained feature map is 4C. After the downsampling operation is completed, upsampling is performed to restore the image. Similar to downsampling, upsampling is also divided into two levels, each level consists of dconv 2×2, conv 1×1, and IFMT with a step size of 2. After each dconv 2×2 operation is performed, the width and height of the image are doubled, while the feature dimension is halved. Then, the output of dconv is concat-operated with the output of the downsampled IFMT at the corresponding level to alleviate the loss of image information during the downsampling process. Finally, conv 3×3 is used to reduce the dimension of the feature and restore it to the RGB three-channel to obtain the repaired image S re , and with S lu Perform residual connection to obtain the final enhanced image S en .
[0119] The CNN-based method has shown great advantages in performance compared with the traditional method, but it is insufficient in capturing long-range dependencies. Due to the great computational complexity of the Transformer-based method, some works based on the hybrid of CNN and Transformer only use the Transformer layer in the bottleneck layer of U-Net, so the potential of Transformer has yet to be explored. In order to fully explore the potential of Transformer, this embodiment designs an Illumination Fused Mamba and Transformer (IFMT) structure. The design details of IFMT are as follows: Figure 4As shown. The classic Transformer consists of two multi - head self - attention mechanisms, layer normalization, a feed - forward network, and residual connections. To reduce the computational complexity of the Transformer in capturing long - range dependencies, in this embodiment, a light - guided multi - head self - attention mechanism and Vision Mamba (i.e., the first multi - head self - attention mechanism and the second multi - head self - attention mechanism) are respectively designed to replace the two multi - head self - attention mechanisms in the classic Transformer, so that it can not only maintain the advantages of the Transformer structure but also enjoy the global and local attention at the linear level of Mamba.
[0120] Light - guided multi - head self - attention mechanism. As Figure 4 shown, the F obtained by f(·) lu is fed into the light - guided multi - head self - attention mechanism. To solve the huge computational complexity of the multi - head attention mechanism in the Transformer, the light - guided multi - head self - attention mechanism regards a single - channel feature map as a token to calculate self - attention. First, the input feature is reshaped into the feature Then X is divided into k heads:
[0121] X = [X 1 , X 2 , …, X i , …, X k (6)
[0122] where X i represents the feature of the i - th head, Figure 3 only shows the case of k = 1 and some details are omitted for simplicity. For each head i , three unbiased fully - connected layers are used to linearly project X i to
[0123]
[0124] where Q i is the query, K i is the key, V i is the value, where are the learnable parameters of the fully - connected layer, and T is the matrix transpose operation. It can be understood that the illumination information of the image is not evenly distributed in each region, but shows different intensities in different regions, and the damage is more serious in the under - exposed regions, and its recovery is relatively difficult. The well - exposed regions often contain the semantic context representation of the image, and these semantic information can assist in the recovery of the under - exposed regions. Therefore, in this embodiment, F luAs a guidance for self-attention mechanism calculation to enhance the interaction between regions with different exposure levels. To make F lu consistent with the shape of X, F lu is reshaped into features and divided into k heads:
[0125] Y = [Y 1 , Y 2 , …, Y i , …, Y k (8)
[0126] In the formula, Y i represents the features of the i-th head,
[0127] Then, the self-attention of each head i is expressed as:
[0128]
[0129] In the formula, is a learnable parameter for adaptively scaling matrix multiplication. Then, the k heads are concatenated and fed into a fully connected layer, and the position encoding (learnable parameter) is added to obtain the output token Finally, X out is reshaped to obtain the output feature The computational complexity of this design mainly comes from k calculations of two matrix multiplications in formula (10), that is and Its time complexity can be expressed as:
[0130]
[0131] Therefore, the computational complexity of the light-guided multi-head self-attention mechanism designed in this embodiment is linearly related to the spatial size, greatly reducing the computational complexity of the Transformer architecture.
[0132] Vision Mamba. Inspired by the selective state space, this embodiment integrates Vision Mamba into the framework and replaces the second multi-head self-attention mechanism in the classical Transformer with it. Vision Mamba consists of Linear, depthwise separable convolution, SiLU, 2D scanning, normalization, and residual connection. First, the input features are processed through Linear and depthwise separable convolution, and after passing through the SiLU activation function, 2D scanning is used to expand from the four directions of up, down, left, and right of the image. After the image is flattened, its height (H) and width (W) are merged and converted into the token length (L). Then, each sequence in the scanning is input into the feature extraction module for global feature extraction. The operation of this step can be expressed by the following formula:
[0133]
[0134] In the formula, x t is the input feature, y t is the output feature, and are both learnable parameters, and h t is the state vector at time step t.
[0135] Subsequently, the features y 1 、y 2 、y 3 、y 4 output in four different directions are summed and merged, and the size of the output feature is adjusted to be the same as that of the input feature. In addition, in order to extract deeper latent features, the number of hidden layers in the 2D scanning is set to double as the level of IFMT deepens. Only as an example, the initial number of hidden layers d state in the 2D scanning is set to 16, and as the sampling level deepens, d state also doubles and reaches 64 at the deepest sampling. This setting allows deeper features to be extracted layer by layer from the vector containing integrated illumination features.
[0136] Net1 enhances the quality of the input image, but due to the decoupled idea, Net1 only considers the degradation problem and does not consider any staining-related problems. Therefore, on this basis, this embodiment designs Net2 to unify the staining style of the image, thereby achieving complete staining standardization. The overall structure of Net2 is as Figure 5As shown in the figure, it is a generative network that combines the attention mechanism and the Laplacian pyramid. Net2 includes a Laplacian pyramid, a band-pass network, a low-resolution layer residual network, and a style network. Among them, the band-pass network and the low-resolution layer residual network are encoder-generator structures. The encoder of the low-resolution layer residual network extracts the cell morphology as the content, and the style network extracts the staining contour as the style. Then, the staining representation extracted from the reference is combined with the morphological representation of the input image to achieve staining normalization without changing the underlying cell structure.
[0137] Laplacian pyramid. The Laplacian pyramid (LP) decomposes an image into a set of band-pass images with octave intervals and a low-resolution layer residual image. The histograms are used to analyze the features of these two layers respectively, and the analysis results are as Figure 6 shown. It can be seen from the results that the band-pass images mainly reflect the high-frequency details of the image, and these layers are not sensitive to the color information of the image. The staining and content information of the image are mainly presented by the low-resolution layer residual image. Motivated by this result, only a lightweight network needs to be designed to fine-tune the band-pass images, and more detailed processing of the image focuses on the low-resolution layer residual network, which makes the network design clearer and the tasks of each module more explicit.
[0138] Low-resolution layer residual network. As can be seen from the above description, in order to better separate the content and style of the image, the design of the low-resolution layer residual network is particularly important. To solve the problem of information forgetting caused by the deepening of the network layer, this embodiment designs a residual module, and on this basis, designs a self-attention mechanism to pay attention to the global information to achieve adaptive adjustment of weights. Finally, the self-attention mechanism is embedded into the residual module to form a residual self-attention mechanism module for deep feature processing and adaptive adjustment of weight parameters, as Figure 7 shown. In this embodiment, the low-resolution layer residual image I 3 obtained by decomposing LP is input into the low-resolution layer residual network. First, two conv 3×3 and LeakyReLU are used to extract features, then two residual self-attention mechanism modules are passed through for deeper feature extraction, and finally LR Image Encoder Feature is obtained through conv 3×3. To better achieve staining normalization, the morphological representation is combined with the staining contour extracted by the style network using conv 3×3, InstanceNorm2d, LinearBlock, and LeakyReLU and sent into the generator. First, two residual self-attention mechanism modules are passed through, and then conv 3×3, LeakyReLU, and two conv 3×3 are passed through to obtain the stained low-resolution layer residual image.
[0139] Style network. In order to separate the cell morphology and staining outline, a style network is designed to process the LR ImageEncoder Feature. The style network consists of an encoding module and a decoding module, which are used to realize style information extraction and parameter generation respectively. In the encoding module, the ColorTransform module is first designed based on conv 1×1, BatchNorm2d, and ReLU to transform the LR Image Encoder Feature to improve the color consistency of the input features and reduce the influence of non-target style features on the generation of style vectors. Next, AddaptiveAvgPool2d is used to compress the features in the spatial dimension to obtain a global feature representation. Subsequently, in the process of generating the style vector, the channel attention mechanism is introduced to adaptively adjust the importance of the channel, so that the network can better focus on the features that contribute significantly to the style information and avoid excessive attention to irrelevant or minor features. The output of the encoding stage is the mean and variance, which are used to generate the style feature vector through VAE. In the decoding module, the style feature vector is first reconstructed through a fully connected layer and further nonlinearly transformed through an activation function to generate the adaptive instance normalization (AdaIN) parameters (α i ,β i ) to characterize the staining profile.
[0140] Bandpass network. The bandpass image mainly reflects the high-frequency details of the image. These layers are not sensitive to the color information of the image, so it is only necessary to design a lightweight network to fine-tune it. The input of the bandpass network is the bandpass image h of different levels. 0 、h 1 、h 2 Similar to the low-resolution residual network, the input is mapped to the encoder to extract the morphological representation, which is then compared with the AdaIN parameters (α i ,β i ) to achieve color transfer, and finally get the output through the generator On this basis, this embodiment introduces a pyramid feature fusion mechanism. For each layer of bandpass image, the bandpass features of the current layer are added to the pyramid features of the previous layer. This design can introduce contextual information in the bandpass image adjustment process, supplement the feature dependencies between different resolution layers, and enhance the global structure perception ability of the high-frequency layer while maintaining the detail texture. In addition, since the bandpass image mainly reflects high-frequency details, its mean is usually 0, and its standard deviation is much smaller than that of the low-resolution layer residual image. Therefore, this embodiment introduces a non-learnable parameter θ i and δ i The input and output of the bandpass network are scaled separately and constrained to (-1,1), which is more conducive to the learning of the bandpass network.
[0141] Next, the embodiments of the present application will conduct a qualitative evaluation of the method proposed in the present application and a downstream task - myeloma plasma cell segmentation to prove the feasibility and progressiveness of the present application.
[0142] 1) Qualitative evaluation
[0143] Figure 8 and Figure 9 show the visual comparison of the method proposed in the present application with other methods in staining standardization. Among them, ParamNet and RandStainNA are algorithms developed for H&E staining standardization, GCTI:SN is an algorithm for staining standardization of blood smears, and gsan is a general staining architecture. It can be seen from the results that ParamNet and RandStainNA failed to achieve the goal of standardization. As mentioned above, the vast majority of existing staining standardization algorithms are for H&E images and cannot be directly applied to blood smears, and there is a serious shortage of standardization algorithms in the field of blood smears. Figure 8 and Figure 9 The results presented in (b) and (c) in again prove this view. GCTI:SN corrects the illumination change during the staining standardization process and considers both, resulting in room for further improvement in the visual effect after standardization. In addition, the quality of the reference itself will directly affect the standardization result. If low-quality images are used as references for staining, it is difficult to achieve this goal even if the brightness is corrected during the standardization process, as shown in (d) in Figure 8 and Figure 9 Benefiting from the principle of separating content and style, gsan can be used as a general framework for H&E and blood smears, but it treats standardization and degradation as a whole, which will sacrifice the performance of the network to a certain extent. Because even the best network is difficult to handle all problems at the same time and can only sacrifice network performance to seek a balance between the two. In addition, it ignores the importance of the attention mechanism in the Laplacian pyramid. Especially when extracting cell morphology and staining contours, the attention mechanism can provide more accurate feature expression and information capture ability for it, thus effectively improving the performance of the model in detail processing and global consistency, which is crucial for the extraction of cell morphology and staining contours. Our method uses the decoupled idea to decompose the standardization task into two stages: low-quality image enhancement and staining style unification, decomposing the complex task into two simpler subtasks. Each subtask only needs to focus on its specific goal, simplifying the standardization difficulty. We first repair the images with uneven quality before unifying the staining style, providing feasible high-quality image data for the staining style unification, and developing an attention module in the staining style unification stage to enhance the adaptive adjustment ability of feature extraction. As shown in Figure 8 and Figure 9As shown in (f), the method proposed in this application can not only well achieve staining standardization, but also greatly improve the quality of the standardized images.
[0144] 2) Downstream task - myeloma plasma cell segmentation
[0145] The performance of the method proposed in this application has been fully verified in the qualitative evaluation. To further evaluate the effectiveness of the method proposed in this application from a quantitative perspective, this embodiment selects myeloma plasma cell segmentation as the downstream task for in-depth analysis. As mentioned before, only two methods, GCTI:SN and gsan, can be applied to blood smears. Therefore, only the comparison results with these two methods are shown in the downstream task.
[0146] This embodiment uses the mean Intersection over Union (mIoU) and the Aggregated Jaccard Index (AJI) as the evaluation indicators for the segmentation performance. mIoU is mainly used to measure the overlapping degree between the model prediction result and the true label, while AJI comprehensively evaluates the performance of the model in the instance segmentation task from two dimensions: the segmentation accuracy of the target instance and the regional coverage rate. By combining these two indicators, the performance and robustness of the algorithm can be comprehensively evaluated from different levels of semantic segmentation and instance segmentation. In this embodiment, Branch1 and Branch2 are used to segment cells and cell nuclei respectively, and finally the segmented cells and cell nuclei are added together to obtain the final segmentation result.
[0147] Table 1 shows the comparison of the indicators of the method proposed in this application with GCTI:SN and gsan in the segmentation of cells, cell nuclei, and myeloma plasma cells.
[0148] Table 1 Comparison results of the indicators for the downstream segmentation task
[0149]
[0150] The experimental results show that the method proposed in this application exhibits better performance and is superior to other advanced methods in both mIoU and AJI. Compared with the input, the accuracy of gsan has been improved, which indicates that standardization has a significant impact on the performance of the downstream task. However, due to its neglect of the key role of attention information in the extraction of cell morphology and staining contours, the performance improvement is limited. The accuracy of GCTI:SN has been greatly improved, which benefits from its correction of brightness during the standardization process. However, treating brightness correction and standardization as a whole restricts the further improvement of its performance. In addition, using low-quality images as a reference for brightness correction will lead to a performance decline. The method of this application effectively overcomes these problems, and both mIoU and AJI reach the optimal values.
[0151] Figure 10 , Figure 11 , Figure 12 Visualization displays of the segmentation results of cells, cell nuclei, and myeloma plasma cells were respectively performed. As Figure 12 shown, the accuracy of (c2) is higher than that of (b2), but the accuracy of (c3) is lower than that of (b3). This indicates that in the brightness correction process of GCTI:SN, the quality of the reference image has an important impact on the correction effect: when the quality of the reference image is higher than that of the input image, brightness correction can improve performance; while when the quality of the reference image is lower than that of the input image, brightness correction may lead to performance degradation. The accuracy of (d2) is lower than that of (b2), but the accuracy of (d3) is higher than that of (b3). This shows that gsan exhibits good performance when processing images with higher quality, while its performance significantly degrades when processing low-quality images. After being standardized by the method of this application, the segmentation performance has been significantly improved, further verifying the effectiveness of the decoupling strategy. Especially before the staining style is unified, designing a low-quality image enhancement network to repair the image quality plays an indispensable role and makes up for the performance bottlenecks faced by GCTI:SN and gsan.
[0152] In summary, the embodiment of this application proposes a decoupled deep attention model for blood smear image staining standardization. This model can achieve blood smear image staining standardization. The key technical points of this application are reflected in the following three aspects:
[0153] First, this application proposes a decoupled deep attention model for blood smear image staining standardization. The present invention decouples the blood smear image staining standardization task into two stages: low-quality image enhancement and staining style unification, and establishes two subnets Net1 and Net2 for these two stages respectively to model them. This decoupled idea decomposes the standardization task into two simpler subtasks, greatly reducing the difficulty of standardization. In addition, the present invention overcomes the performance bottlenecks faced by existing standardization algorithms in treating degradation and standardization as a whole. With the help of this decoupled idea, the present invention shows excellent results in achieving standardization.
[0154] Second, Net1 of this application integrates the Retinex model, Transformer, and Mamba to construct a framework named RetinexMambaFormer. This framework not only retains the long-range dependence and non-local interaction modeling capabilities of Transformer, but also uses the selection mechanism of Mamba to achieve linear-level global and local attention. This is the first attempt to combine Transformer and Mamba to solve the low-quality image enhancement problem.
[0155] Third, Net2 of the present application constructs a Laplacian pyramid generative network based on the attention mechanism, which overcomes the problem of the dependence on domain labels in the existing generative adversarial network when implementing standardization. It is worth mentioning that in this stage, a low-resolution layer residual network based on the residual attention mechanism is designed to avoid the forgetting problem of deep networks and achieve adaptive adjustment of weights; a style network based on the channel attention mechanism is also designed to reallocate weights using channel attention information so that it can better extract the staining contour, thereby achieving better standardization.
[0156] Figure 13 It is a schematic structural diagram of a blood smear image staining standardization device provided by an embodiment of the present application. An embodiment of the present application also provides a blood smear image staining standardization device. As Figure 13 shown, the blood smear image staining standardization device includes:
[0157] An image enhancement module 1301, configured to obtain a blood smear image and a light prior map, merge the light prior map into the blood smear image, simulate the interaction between regions under different lighting conditions, extract a light feature map, and aggregate the extracted light feature map to generate a light map, and use the light map to enhance the blood smear image to obtain a first image; obtain an enhanced image according to the first image and the light feature map;
[0158] A staining style unification module 1302, configured to decompose the enhanced image into a group of band-pass images with octave intervals and a low-resolution layer residual image; extract features from the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation and a staining contour according to the first feature map, and combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image; extract features from the group of band-pass images with octave intervals to obtain a second morphological representation, and combine the second morphological representation with the staining contour to obtain a standardized band-pass image.
[0159] In some embodiments, the image enhancement module is further configured to:
[0160] Merge the blood smear image with the light prior map, fuse the light prior map into the blood smear image using unit convolution, and use depthwise separable convolution to simulate the interaction between regions under different lighting conditions, and further extract features to generate a light feature map F lu ;
[0161] Aggregate the light feature map F lu to generate a light map;
[0162] Decompose the blood smear image into a reflectance component and a light component through the following formula:
[0163]
[0164] In the formula, S represents the blood smear image, R represents the reflectance component, and I represents the illumination component. represents element-wise multiplication;
[0165] Perturbations of the reflectance component and the illumination component are introduced into the blood smear image to obtain an intermediate image, denoted as:
[0166]
[0167] In the formula, S1 represents the intermediate image. represents the perturbation of the reflectance component. represents the perturbation of the illumination component;
[0168] Multiply the intermediate image by the illumination map to obtain:
[0169]
[0170] In the formula, represents the illumination map. represents the noise and artifacts amplified during the lighting process of the low-exposure image. represents the underexposure, overexposure, and / or color distortion caused during the lighting process;
[0171] Simplify formula (3) to obtain:
[0172]
[0173] In the formula, S lu represents the first image.
[0174] In some embodiments, the image enhancement module is further configured to:
[0175] I en = g(S lu , F lu ) (5)
[0176] In the formula, g represents a damage repairer, I en represents the enhanced image, S lu represents the first image, and F lu represents the illumination feature map.
[0177] In some embodiments, the damage repairer includes an encoder and a decoder; wherein, the encoder is used to downsample the input S lu to match F luThe dimension is reduced through a downsampling process while deep feature extraction is performed. The downsampling process is divided into a first level and a second level. Both the first level and the second level are composed of an IFMT module and a conv 4×4 module. After the downsampling processing of the first level and the second level, an intermediate feature map is obtained.
[0178] The decoder is used to upsample the intermediate feature map. The upsampling is divided into a third level and a fourth level. Both the third level and the fourth level are composed of a dconv 2×2 module, a conv 1×1 module, and an IFMT module. The output of the dconv 2×2 module is concatenated with the output of the IFMT module in the corresponding level of the downsampling process to alleviate the loss of image information during the downsampling process. After the upsampling processing of the third level and the fourth level, a second feature map is obtained, and the dimension of the second feature map is adjusted to restore it to the RGB three channels to obtain a restored image. The restored image is subjected to residual connection with the first image to obtain an enhanced image.
[0179] In some embodiments, the IFMT module includes a first multi-head self-attention mechanism, a second multi-head self-attention mechanism, layer normalization, a feed-forward network, and a residual connection.
[0180] The data processing process of the first multi-head self-attention mechanism is as follows:
[0181] And the input feature F in is reshaped into the feature X.
[0182] The feature X is divided into k heads through the following formula:
[0183] X = [X 1 , X 2 , …, X i , …, X k (6)
[0184] In the formula, X i represents the feature of the i-th head, represents the set of real numbers, H represents the height of the image, W represents the width of the image, d k represents the feature dimension of each head, and C represents the number of channels.
[0185] For each head i , three bias-free fully connected layers are used to linearly project X i onto
[0186]
[0187] In the formula, are the learnable parameters of the fully connected layer, T is the matrix transpose operation, and Q i is the query, K i is the key, V i is the value.
[0188] Utilize the illumination feature map F lu as the guidance for the self-attention mechanism calculation to enhance the interaction between regions with different exposure levels, reshape F lu into features and divide the feature Y into k heads:
[0189] Y = [Y 1 , Y 2 , …, Y i , …, Y k ] (8)
[0190] In the formula, Y i represents the feature of the i-th head,
[0191] Represent the self-attention of each head i as:
[0192]
[0193] In the formula, are learnable parameters used for adaptive scaling of matrix multiplication, Attention is the data processing process of the multi-head self-attention mechanism, and softmax is the softmax function;
[0194] Connect the k heads and send them into the fully connected layer, and add the position encoding to obtain the output token
[0195] Reshape X out to obtain the output feature
[0196] The data processing process of the second multi-head self-attention mechanism is:
[0197] Process the input feature through Linear and depthwise separable convolution, and after passing through the SiLU activation function, expand it from the four directions of up, down, left, and right of the image using 2D scanning. The height and width after the image is flattened are merged and converted into the token length, and each sequence in the scanning is input into the feature extraction module for global feature extraction. The calculation process is expressed as:
[0198]
[0199] In the formula, x t is the input feature, y tas the output feature, and are both learnable parameters, h t is the state vector at time step t;
[0200] Sum and combine the outputs in four different directions of the extracted features y 1 , y 2 , y 3 and y 4 , and adjust the size of the output feature to be the same as that of the input feature;
[0201] Set the number of hidden layers in the 2D scan to double as the level of the IFMT deepens, so as to be able to extract deeper features layer by layer from the vector containing the integrated illumination features.
[0202] In some embodiments, the staining style unification module is further configured to:
[0203] Construct a low-resolution layer residual network and a style network; wherein, the low-resolution layer residual network is used to extract features from the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation from the first feature map, and combine the first morphological representation with the staining contour to obtain a standardized low-resolution layer residual image; the style network is used to extract the staining contour according to the first feature map;
[0204] The low-resolution layer residual network obtains the standardized low-resolution layer residual image through the following method:
[0205] Obtain the low-resolution layer residual image; use two conv 3×3 units and LeakyReLU units to extract features, perform deeper feature extraction through two residual self-attention mechanism modules, and finally obtain the first feature map through a conv 3×3 unit;
[0206] Use a conv 3×3 unit, an InstanceNorm2d unit, a LinearBlock unit, and a LeakyReLU unit to extract the first morphological representation from the first feature map and combine it with the staining contour extracted by the style network, and send it into the generator. First, pass through two residual self-attention mechanism modules, then through a conv 3×3 unit, a LeakyReLU unit, and two conv 3×3 units to obtain the standardized low-resolution layer residual image
[0207] The style network includes an encoding module and a decoding module. The style network extracts the staining contour according to the first feature map through the following method:
[0208] In the encoding module, a Color Transform module is designed based on conv 1×1, BatchNorm2d, and ReLU to transform the first feature map, improve the color consistency of the input features, and reduce the influence of non-target style features on the generation of the style vector.
[0209] AdaptiveAvgPool2d is used to compress the features in the spatial dimension to obtain a global feature representation. During the generation of the global feature representation, a channel attention mechanism is introduced to adaptively adjust the importance of channels, enabling the network to focus on the features that significantly contribute to the style information. The output of the encoding module is the mean and variance, and the mean and variance are used to generate the style feature vector through VAE.
[0210] In the decoding module, the style feature vector is reconstructed through a fully connected layer and further non-linearly transformed through an activation function to generate the adaptive instance normalization parameters (α i , β i ) to represent the staining contour.
[0211] The blood smear image staining normalization device provided by the embodiments of the present application can be used to execute the technical solutions of the blood smear image staining normalization method in the above embodiments. The implementation principle and technical effects are similar and will not be elaborated here.
[0212] Figure 14 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 14 shown, the electronic device may include: a processor 141 and a memory 142. Among them, the processor 141 and the memory 142 can communicate. Exemplarily, the processor 141 and the memory 142 communicate through a communication bus 143.
[0213] The processor 141 executes the computer execution instructions stored in the memory 142, enabling the processor 141 to execute the solutions in the above embodiments. The processor 141 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit ASIC, a field-programmable gate array FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0214] The communication bus 143 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The transceiver is used to implement the communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM), and may also include non-volatile memory.
[0215] The electronic device provided by the embodiment of the present application can be the terminal device in the above embodiment.
[0216] The embodiment of the present application also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer, the computer is enabled to execute the technical solution of the blood smear image staining standardization method in the above embodiment.
[0217] The embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, the technical solution of the blood smear image staining standardization method in the above embodiment can be implemented.
[0218] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in an electrical, mechanical or other form.
[0219] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0220] In addition, in each embodiment of the present application, each functional module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0221] The integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module is stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in each embodiment of the present application.
[0222] It should be understood that the above processor can be a Central Processing Unit (CPU for short), or other general-purpose processors, Digital Signal Processors (DSP for short), Application Specific Integrated Circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.
[0223] The memory may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0224] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0225] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0226] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a master control device.
[0227] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical disks.
[0228] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A blood smear image staining standardization method, characterized in that: include: Acquire a blood smear image and a prior illumination map, merge the prior illumination map into the blood smear image, simulate the interaction between regions under different illumination conditions, extract an illumination feature map, aggregate the extracted illumination feature map to generate an illumination map, and enhance the blood smear image using the illumination map to obtain a first image; Obtaining an enhanced image according to the first image and the illumination feature map; Decomposing the enhanced image into a set of octave-spaced bandpass images and a low-resolution layer residual image; Performing feature extraction on the low-resolution layer residual image to obtain a first feature map, extracting a first morphological representation and a dyeing contour according to the first feature map, and combining the first morphological representation with the dyeing contour to obtain a standardized low-resolution layer residual image; Feature extraction is performed on the set of bandpass images spaced at octave intervals to obtain a second morphological representation, and the second morphological representation is combined with the staining profile to obtain a standardized bandpass image.
2. The method according to claim 1, characterized in that Acquiring a blood smear image and a priori illumination map, merging the priori illumination map into the blood smear image, simulating the interaction between regions under different illumination conditions, extracting an illumination feature map, aggregating the extracted illumination feature map to generate an illumination map, and enhancing the blood smear image using the illumination map to obtain a first image, including: The blood smear image is combined with the illumination prior map, and the illumination prior map is fused into the blood smear image using unit convolution. The interaction between regions under different illumination conditions is simulated using depthwise separable convolution, and features are further extracted to generate an illumination feature map F. lu ; Aggregate the illumination feature map F using unit convolution lu To generate the light map; The blood smear image is decomposed into reflectance component and illumination component by the following formula: In the formula, S represents the blood smear image, R represents the reflectivity component, and I represents the illumination component. Represents element-wise multiplication; The disturbance of the reflectivity component and the illumination component is introduced into the blood smear image to obtain an intermediate image, which is expressed as: Where S1 represents the intermediate image, represents the perturbation of the reflectivity component, represents the disturbance of the illumination component; Multiplying the intermediate image by the light map gives: In the formula, represents the light map, It indicates the noise and artifacts of low-exposure images that are amplified during the lighting process. Indicates underexposure, overexposure and / or color distortion caused by the lighting process; Simplifying formula (3) yields: In the formula, S lu Represents the first image.
3. The method according to claim 1, characterized in that According to the first image and the illumination feature map, the enhanced image is calculated by the following formula: I en =g(S lu ,F lu ) (5) In the formula, g represents the damage repairer, I en represents the enhanced image, S lu represents the first image, F lu Represents the illumination feature map.
4. The method according to claim 3, characterized in that The damage repairer includes an encoder and a decoder; wherein the encoder is used to convert the input S lu Downsample to match F lu The downsampling process is used to reduce the computational complexity and perform deep feature extraction. The downsampling process is divided into the first level and the second level. The first level and the second level are both composed of an IFMT module and a conv4×4 module. After the first level and the second level downsampling process, an intermediate feature map is obtained. The decoder is used to upsample the intermediate feature map, and the upsampling is divided into a third level and a fourth level. The third level and the fourth level are both composed of a dconv2×2 module, a conv1×1 module and an IFMT module. The output of the dconv2×2 module is concat-operated with the output of the IFMT module in the downsampling process of the corresponding level to alleviate the loss of image information in the downsampling process. After the upsampling process of the third level and the fourth level, a second feature map is obtained, and the dimension of the second feature map is adjusted and restored to RGB three channels to obtain a repaired image. The repaired image is residually connected with the first image to obtain an enhanced image.
5. The method according to claim 4, characterized in that The IFMT module includes a first multi-head self-attention mechanism, a second multi-head self-attention mechanism, layer normalization, a feedforward network and a residual connection; The data processing process of the first multi-head self-attention mechanism is: And input feature F in Reshape into feature X; The feature X is divided into k heads by the following formula: X=[X1,X2,…,X i ,…,X k ] (6) Where, X i represents the features of the i-th head, i=1,2,…,k, represents a real number set, H represents the height of the image, W represents the width of the image, d k represents the feature dimension of each head, and C represents the number of channels; For each head i , use three unbiased fully connected layers to transform X i Linear projection to In the formula, is the learnable parameter of the fully connected layer, T is the matrix transpose operation, Q is the query, K i is the key, V i is the value; Using the illumination feature map F lu As a guide for the calculation of the self-attention mechanism to enhance the interaction between regions with different exposure levels, F lu Reshape into features And divide the feature Y into k heads: Y=[Y1,Y2,…,Y i ,…,Y k ] (8) Where Y i represents the features of the i-th head, i=1,2,…,k; Each head i The self-attention of is expressed as: In the formula, is a learnable parameter used to adaptively scale matrix multiplication, Attention is the data processing process of the multi-head self-attention mechanism, and softmax is the softmax function; Connect the k heads and send them to the fully connected layer, and add position encoding Get output token X out Reshape to get output features The data processing process of the second multi-head self-attention mechanism is: The input features are processed by Linear and depth-wise separable convolution, and after SiLU activation function, they are expanded from the top, bottom, left, and right directions of the image using 2D scanning. The height and width of the flattened image are merged and converted into token length. Each sequence in the scan is input into the feature extraction module for global feature extraction. The calculation process is expressed as: In the formula, x t is the input feature, y t is the output feature, and are all learnable parameters, h t is the state vector at time step t; The outputs of the extracted features y1, y2, y3 and y4 in four different directions are summed and merged, and the size of the output features is adjusted to be consistent with the size of the input features; The number of hidden layers in the 2D scan is set to double as the level of IFMT goes deeper, so that deeper features can be extracted layer by layer from the vector containing the integrated illumination features.
6. The method according to claim 1, characterized in that Extracting features from the low-resolution layer residual image to obtain a first feature map, extracting a first morphological representation and a dyeing contour according to the first feature map, and combining the first morphological representation with the dyeing contour to obtain a standardized low-resolution layer residual image, including: Constructing a low-resolution layer residual network and a style network; wherein the low-resolution layer residual network is used to perform feature extraction on the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation based on the first feature map, and combine the first morphological representation with the dyeing contour to obtain a standardized low-resolution layer residual image; the style network is used to extract the dyeing contour based on the first feature map; The low-resolution layer residual network obtains a standardized low-resolution layer residual image by the following method: Get the low-resolution layer residual image; use two conv3×3 units and LeakyReLU units to extract features, and perform deeper feature extraction through two residual self-attention mechanism modules, and finally get the first feature map through the conv3×3 unit; The first morphological representation is extracted from the first feature map using conv3×3 units, InstanceNorm2d units, LinearBlock units, and LeakyReLU units, combined with the coloring contour extracted by the style network, and sent to the generator. It first passes through two residual self-attention mechanism modules, and then passes through conv3×3 units, LeakyReLU units, and two conv3×3 units to obtain a standardized low-resolution layer residual image. The style network includes an encoding module and a decoding module. The style network extracts the coloring contour according to the first feature map by the following method: In the encoding module, the Color Transform module is designed based on conv1×1, BatchNorm2d, and ReLU to transform the first feature map, improve the color consistency of the input features, and reduce the impact of non-target style features on style vector generation; AdaptiveAvgPool2d is used to compress the features in the spatial dimension to obtain the global feature representation. In the process of generating the global feature representation, the channel attention mechanism is introduced to adaptively adjust the importance of the channel so that the network focuses on the features that contribute significantly to the style information. The output of the encoding module is the mean and variance, which are used to generate the style feature vector through VAE. In the decoding module, the style feature vector is reconstructed through a fully connected layer and further nonlinearly transformed through an activation function to generate an adaptive instance normalization parameter (α i ,β i ) to characterize the staining profile.
7. A blood smear image staining standardization device, characterized in that: include: an image enhancement module, configured to acquire a blood smear image and an illumination prior map, merge the illumination prior map into the blood smear image, simulate the interaction between regions under different illumination conditions, extract an illumination feature map, aggregate the extracted illumination feature map to generate an illumination map, enhance the blood smear image using the illumination map to obtain a first image; and obtain an enhanced image based on the first image and the illumination feature map; The coloring style unification module is configured to decompose the enhanced image into a group of bandpass images with an interval of octave and a low-resolution layer residual image; perform feature extraction on the low-resolution layer residual image to obtain a first feature map, extract a first morphological representation and a coloring contour according to the first feature map, and combine the first morphological representation with the coloring contour to obtain a standardized low-resolution layer residual image; perform feature extraction on the group of bandpass images with an interval of octave to obtain a second morphological representation, and combine the second morphological representation with the coloring contour to obtain a standardized bandpass image.
8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.