Two-stage image generation learning method for cell detection and analysis based on molecular expression prediction

Through the two-stage image generation learning method, combined with cell segmentation and molecular expression prediction, the problem of cell subtype classification and molecular expression level acquisition in pathological images is solved, achieving higher accuracy and application value.

CN119942216AActive Publication Date: 2025-05-06ANHUI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510099375.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is difficult to use pathological images to accurately classify cell subtypes, and it is difficult to quickly obtain the expression levels of cell molecular markers.

Method used

Using a two-stage image generation learning method based on molecular expression prediction, the cell segmentation model and molecular expression prediction model are constructed to achieve the prediction of cell binary segmentation, cell semantic segmentation and molecular expression level.

Benefits of technology

It improves the accuracy of cell segmentation and molecular expression prediction, enhances the accuracy of cell semantic segmentation, and provides a convenient and accurate tool for cell-level cancer diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942216A_ABST
    Figure CN119942216A_ABST
Patent Text Reader

Abstract

The invention relates to a two-stage image generation learning method for cell detection and analysis based on molecular expression prediction. The two-stage image generation learning method comprises the following steps: acquiring a Hamp; forming a data set by the E data set and the paired mIHC data set, and preprocessing the data set; constructing and training a cell segmentation model; a cell binary segmentation result is obtained and visually displayed; constructing and training a molecular expression prediction model; a molecular expression prediction result is obtained and visually displayed; and obtaining a cell semantic segmentation result for visual display. In the first stage, a generative cell segmentation model is used from Hamp to Hamp; e, extracting cell characteristic information from the image, generating a more accurate cell binary segmentation result, and improving the cell segmentation accuracy of the model; in the second stage, the generative molecular expression prediction model is used for extracting pathological information features and generating a virtual mIHC image, and meanwhile, the feature extraction capability of the model is improved by using the pathological basic model, so that the acquisition efficiency of the mIHC image is improved, and the accuracy of the generated content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular to a two-stage image generation learning method for cell detection and analysis based on molecular expression prediction. Background Art

[0002] WSI (Whole Slide Image) uses a digital scanner to scan traditional pathological sections to collect high-resolution tissue section images, which contain a large amount of pathological information and are an important tool in cancer research. H&E (Hematoxylin and Eosin) stained images are a commonly used digital pathology image that can be used to help pathologists analyze cancer at the cellular level. Analyzing WSI at the cellular level is an important research topic, which includes accurately identifying the location of cells, distinguishing cell subtypes, and analyzing the expression levels of related molecules. These tasks rely on cell segmentation, cell subtype identification, and obtaining the expression levels of related molecules. However, it is difficult or even impossible for pathologists to complete these tasks by naked eye observation alone. Therefore, it is necessary to use H&E images to simultaneously achieve cell binary segmentation, cell semantic segmentation, and molecular expression prediction.

[0003] Deep learning technology has achieved a series of successful applications in the analysis of tissue pathology. Through artificial neural networks, the model can automatically extract and encode pathological features, and use these features to efficiently and accurately implement pathological tasks. Unsupervised learning technology can extract effective information from data that does not rely on manual annotation, and use this information to perform related pathological analysis and diagnosis tasks. Using unsupervised learning technology for cell segmentation can overcome the problem of requiring a large amount of manually labeled data in general supervised learning tasks and improve the robustness of the model. However, unsupervised learning also has certain difficulties in application, mainly because the information that can be directly used from data that has not been manually labeled is limited, and the design requirements of the algorithm are high. Therefore, when performing cell segmentation through unsupervised learning, it is necessary to make full use of the prior information in the existing data, select appropriate technical strategies based on actual conditions, and fully extract and encode effective information in unlabeled data to improve the accuracy and application value of cell segmentation.

[0004] The expression level of related molecules in cells plays an important role in the identification of cell subtypes and other related pathological studies. Although mIHC (Multiplex Immunohistochemistry) images can be used to analyze the expression levels of related molecules, it is usually very difficult to obtain mIHC images. The preparation of mIHC images takes a lot of time and is expensive. At the same time, since the staining process will damage the tissue sections, it is impossible to repeatedly stain the same section, so the acquired H&E images and mIHC images are generally difficult to align. Summary of the invention

[0005] In order to solve the current problems of difficulty in accurately classifying cell subtypes using pathological images and difficulty in quickly obtaining the expression levels of cell molecular markers, the purpose of the present invention is to provide a two-stage image generation learning method for cell detection and analysis based on molecular expression prediction, which increases the accuracy of generated content, classifies cell subtypes at the molecular expression level, and increases the accuracy of cell semantic segmentation.

[0006] To achieve the above object, the present invention adopts the following technical solution: a two-stage image generation learning method for cell detection and analysis based on molecular expression prediction, the method comprising the following steps in order:

[0007] (1) Obtaining the H&E dataset and the paired mIHC dataset to form a dataset and preprocessing it, and dividing the preprocessed dataset into a training set and a test set;

[0008] (2) constructing a cell segmentation model, wherein the cell segmentation model is composed of a pre-trained cell segmentation network, a cell segmentation generator network, an information integrator, and a discriminator network; the pre-trained cell segmentation network includes a pre-trained Cellpose network, a pre-trained Hover-Net network, and a pre-trained CellViT network;

[0009] (3) training the cell segmentation model to obtain a trained cell segmentation model;

[0010] (4) Input the H&E images in the test set into the trained cell segmentation model to obtain the cell binary segmentation results, and visualize the cell binary segmentation results;

[0011] (5) constructing a molecular expression prediction model, which consists of a molecular expression prediction generator network and a molecular expression prediction discriminator network;

[0012] (6) training the molecular expression prediction model to obtain a trained molecular expression prediction model;

[0013] (7) Inputting the H&E images in the test set into the trained molecular expression prediction model to obtain the molecular expression prediction results, and visually displaying the molecular expression prediction results;

[0014] (8) The cell binary segmentation results and the molecular expression prediction results are mapped one-to-one at the pixel level in space to obtain the cell semantic segmentation results, and the cell semantic segmentation results are visualized.

[0015] In step (1), the preprocessing specifically refers to: using the valis method to align the WSI data of the H&E image in the H&E dataset and the WSI data of the mIHC image in the mIHC dataset, and then downsampling the WSI data of the H&E image and the WSI data of the mIHC image by four times, and then cutting them into image pairs of size 128×128.

[0016] In step (2), the cell segmentation generator network adopts the Attention U-Net architecture. The depth of Attention U-Net is 5, the input feature channel and the output feature channel are both 3, and the feature channels are set to 64, 128, 256, 512, and 1024 respectively during the feature extraction process; downsampling is achieved through maximum pooling, and upsampling is achieved through bilinear interpolation.

[0017] In step (2), the Cellpose network is pre-trained on the Cellpose official dataset to obtain a pre-trained Cellpose network, and the pre-trained weights are loaded with the nuclei weight parameters; the Hover-Net network is pre-trained on the PanNuke dataset to obtain a pre-trained Hover-Net network, and the pre-trained weights are loaded with the hovernet_fast_pannuke weight parameters; the CellViT network is pre-trained on the PanNuke dataset to obtain a pre-trained CellViT network, and the pre-trained weights are loaded with the CellViT-SAM-H-x40 weight parameters; the nuclei weight parameters, hovernet_fast_pannuke weight parameters, and CellViT-SAM-H-x40 weight parameters remain frozen during the cell segmentation model training process.

[0018] In step (2), the information integrator and discriminator network are constructed based on the PatchGAN network. An upsampling module is added before the output layer of the PatchGAN network, and a feature matrix of size (1, 128, 128) is obtained by bilinear interpolation upsampling.

[0019] Step (3) specifically includes the following steps in order:

[0020] (3a) Input the H&E images in the training set into the pre-trained cell segmentation network with frozen weights to obtain the feature matrix F;

[0021] (3b) Input the H&E images in the training set into the cell segmentation generator network to obtain the feature matrix P;

[0022] (3c) Input the feature matrix F into the information integrator and the discriminator network to obtain the feature matrix A; input the feature matrix P into the information integrator and the discriminator network to obtain the feature matrix B;

[0023] (3d) A loss function is constructed based on the feature matrix A and the feature matrix B. The loss function is a multivariate discriminant loss, and its expression is:

[0024]

[0025] Among them, G s represents the cell segmentation generator network, D s represents the information integrator and discriminator network, x represents the input H&E image, z represents random noise, y cepo ,y hove and cevi Respectively represent the cell binary segmentation results obtained from the weight-frozen pre-trained Cellpose network, the pre-trained Hover-Net network, and the pre-trained CellViT network;

[0026] (3e) Based on the loss function, all non-frozen parameters of the cell segmentation model are updated through the back-propagation mechanism;

[0027] (3f) Repeat steps (3a) to (3e) until the training is completed, and save the final model parameter weights;

[0028] (3g) Load the final model parameter weights and use the test set to test the network performance of the cell segmentation generator.

[0029] In step (5), the molecular expression prediction generator network includes:

[0030] The input layer uses a feature matrix with a scale of (3, 128, 128);

[0031] The pathological basis model uses the PLIP pathological basis model image encoder to extract and encode the features of the input layer through a multi-head self-attention module to obtain a feature matrix with a scale of (50, 768);

[0032] In the reorganization layer, the first feature representing the category code in the first dimension of the feature matrix obtained by the pathological basic model is first removed to obtain a feature matrix with a scale of (49, 768). Then, the 49 features of the first dimension are reorganized according to the spatial position to obtain a feature matrix with a size of (7, 7, 768). Then, the third dimension is adjusted to the front of the first dimension to obtain a feature matrix with a size of (768, 7, 7).

[0033] In the projection layer, the feature matrix obtained by the reorganization layer is convolved to obtain a feature matrix with a scale of (512, 16, 16);

[0034] In the eleventh activation layer, the feature matrix of the projection layer is transformed into a feature matrix with a scale of (512, 16, 16) by nonlinear transformation ReLU;

[0035] The first convolutional layer passes the feature matrix of the input layer through the convolutional layer to obtain a feature matrix of scale (64, 128, 128);

[0036] The first activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (64, 128, 128) on the output feature matrix of the first convolutional layer.

[0037] The first downsampling layer downsamples the feature matrix obtained by the first activation layer through maximum pooling to obtain a feature matrix with a scale of (64, 64, 64);

[0038] In the second convolution layer, the feature matrix obtained by the first downsampling layer is passed through the convolution layer to obtain a feature matrix with a scale of (128, 64, 64);

[0039] The second activation layer uses the output feature matrix of the second convolutional layer through the nonlinear transformation ReLU to obtain a feature matrix with a scale of (128, 64, 64);

[0040] The second downsampling layer downsamples the feature matrix obtained by the second activation layer through maximum pooling to obtain a feature matrix with a scale of (128, 32, 32);

[0041] The third convolutional layer passes the feature matrix obtained by the second downsampling layer through the convolutional layer to obtain a feature matrix with a scale of (256, 32, 32);

[0042] The third activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (256, 32, 32) for the output feature matrix of the third convolutional layer;

[0043] The third downsampling layer downsamples the feature matrix obtained by the third activation layer through maximum pooling to obtain a feature matrix with a scale of (256, 16, 16);

[0044] The fourth convolutional layer passes the feature matrix obtained by the third downsampling layer through the convolutional layer to obtain a feature matrix with a scale of (512, 16, 16);

[0045] The fourth activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (512, 16, 16) on the output feature matrix of the fourth convolutional layer.

[0046] The fourth downsampling layer downsamples the feature matrix obtained by the fourth activation layer through maximum pooling to obtain a feature matrix with a scale of (512, 8, 8);

[0047] The fifth convolutional layer passes the feature matrix obtained by the fourth downsampling layer through the convolutional layer to obtain a feature matrix with a scale of (512, 8, 8);

[0048] The fifth activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (512, 8, 8) on the output feature matrix of the fifth convolutional layer.

[0049] In the first upsampling layer, the feature matrix obtained in the fifth activation layer is upsampled by bilinear interpolation to obtain a feature matrix with a scale of (512, 16, 16);

[0050] The first attention gate operates the feature matrix obtained by the first upsampling layer and the feature matrix obtained by the fourth activation layer through the attention gate to obtain a feature matrix of (512, 16, 16);

[0051] The first concatenation layer concatenates the feature matrix obtained by the first attention gate and the feature matrix obtained by the eleventh activation layer in the first dimension to obtain a feature matrix with a scale of (1024, 16, 16);

[0052] The sixth convolutional layer convolves the feature matrix obtained from the first concatenation layer to obtain a feature matrix with a scale of (512, 16, 16);

[0053] The sixth activation layer uses the nonlinear transformation ReLU to obtain a feature matrix of size (512, 16, 16) on the output matrix of the sixth convolutional layer.

[0054] In the second upsampling layer, the feature matrix obtained in the sixth activation layer is upsampled by bilinear interpolation to obtain a feature matrix with a scale of (512, 32, 32);

[0055] The second attention gate operates the feature matrix obtained by the second upsampling layer and the feature matrix obtained by the third activation layer through the attention gate to obtain a feature matrix of (256, 32, 32);

[0056] The second concatenation layer concatenates the feature matrix obtained by the second attention gate and the feature matrix obtained by the second upsampling layer in the first dimension to obtain a feature matrix with a scale of (768, 32, 32);

[0057] In the seventh convolutional layer, the feature matrix obtained in the second concatenation layer is convolved to obtain a feature matrix with a scale of (256, 32, 32);

[0058] The seventh activation layer, the output matrix of the seventh convolutional layer is transformed by nonlinear ReLU to obtain a feature matrix of size (256, 32, 32);

[0059] In the third upsampling layer, the feature matrix obtained in the seventh activation layer is upsampled by bilinear interpolation to obtain a feature matrix with a scale of (256, 64, 64);

[0060] The third attention gate operates the feature matrix obtained by the third upsampling layer and the feature matrix obtained by the second activation layer through the attention gate to obtain a feature matrix of (128, 64, 64);

[0061] The third concatenation layer concatenates the feature matrix obtained by the third attention gate and the feature matrix obtained by the third upsampling layer in the first dimension to obtain a feature matrix with a scale of (384, 64, 64);

[0062] In the eighth convolutional layer, the feature matrix obtained in the third concatenation layer is convolved to obtain a feature matrix with a scale of (128, 64, 64);

[0063] The eighth activation layer, the output matrix of the eighth convolutional layer is transformed by nonlinear ReLU to obtain a feature matrix of size (128, 64, 64);

[0064] The fourth upsampling layer upsamples the feature matrix obtained in the eighth activation layer by bilinear interpolation to obtain a feature matrix with a scale of (128, 128, 128);

[0065] The fourth attention gate operates the feature matrix obtained by the fourth upsampling layer and the feature matrix obtained by the first activation layer through the attention gate to obtain a feature matrix of (64, 128, 128);

[0066] The fourth concatenation layer concatenates the feature matrix obtained by the fourth attention gate and the feature matrix obtained by the fourth upsampling layer in the first dimension to obtain a feature matrix with a scale of (192, 128, 128);

[0067] The ninth convolutional layer convolves the feature matrix obtained from the fourth concatenation layer to obtain a feature matrix with a scale of (64, 128, 128);

[0068] The ninth activation layer, the output matrix of the ninth convolutional layer is transformed into a feature matrix of size (64, 128, 128) through the nonlinear transformation ReLU;

[0069] The tenth convolutional layer convolves the feature matrix obtained in the ninth activation layer to obtain a feature matrix of scale (3, 128, 128);

[0070] The tenth activation layer, the output matrix of the tenth convolutional layer is transformed into a feature matrix of size (3, 128, 128) through the nonlinear transformation ReLU;

[0071] Output layer, outputs the feature matrix obtained by the tenth activation layer.

[0072] In step (5), the molecular expression prediction discriminator network is constructed based on the PatchGAN network, and an upsampling module is added before the output layer of the PatchGAN network. By bilinear interpolation upsampling, a feature matrix of size (1, 128, 128) is obtained.

[0073] Step (6) specifically includes the following steps in order:

[0074] (6a) Input the H&E images in the training set into the molecular expression prediction generator network to obtain the feature matrix Q;

[0075] (6b) Inputting the mIHC images in the training set into the molecular expression prediction discriminator network to obtain the feature matrix T; inputting the feature matrix Q into the molecular expression prediction discriminator network to obtain the feature matrix S;

[0076] (6c) A total loss function is constructed based on the mIHC image, feature matrix Q, feature matrix T, and feature matrix S. The formula of the total loss function is:

[0077] L Total (G,D)=L AD (G,D)+L F (G)+L Mul (G)

[0078] Among them, G represents the molecular expression prediction generator network, D represents the molecular expression prediction discriminator network; L AD represents the discrimination loss; L F Indicates frequency loss; L Mul represents multi-scale loss;

[0079] The expression of discriminant loss is:

[0080] L AD (G,D)=E x,y [log D(x,y)]+E x,e[log(1-D(x,G(x,e)))]

[0081] Where x, y, and e represent the input H&E image, the real mIHC image, and random noise, respectively; E x,y Represents the mathematical expectation of the joint distribution of x and y; E x,e Represents the mathematical expectation of the joint distribution of x and e;

[0082] The expression for frequency loss is:

[0083] L F (G) = p·E x,y,e [||F HP (y)-F HP (G(x,e))‖1]+q·E x,y,e [||F LP (y)-F L P(G(x,e))||1]

[0084] Among them, F HP (·)=IFFT[HP[FFT(·)]],F LP (·)=IFFT[LP[FFT(·)]], where FFT and IFFT represent Fourier transform and inverse Fourier transform, respectively, HP and LP represent high-pass filter and low-pass filter, respectively, p and q represent high-frequency weight and low-frequency weight, respectively, p is set to 1, and q is set to 2; E x,y,e Represents the mathematical expectation of the joint distribution of x, y, and e;

[0085] The expression of multi-scale loss is:

[0086]

[0087] Among them, λ i Represents the weight of the i-th downsampling loss; DS i (·) indicates that the real molecular expression image or the generated molecular expression prediction image has undergone i scale transformations, and each scale transformation includes four Gaussian filtering and one downsampling using a Gaussian kernel with a mean of 0 and a variance of 1;

[0088] (6d) updating all parameters of the molecular expression prediction model according to the total loss function through the back propagation mechanism;

[0089] (6e) Repeat steps (6a) to (6d) until the training is completed, and save the final model parameter weights;

[0090] (6) Load the final model parameter weights and use the test set to test the performance of the molecular expression prediction generator network.

[0091] It can be seen from the above technical scheme that the beneficial effects of the present invention are as follows: first, the present invention realizes end-to-end cell binary segmentation, cell semantic segmentation and prediction of molecular expression levels through a two-stage generative adversarial learning framework; second, in the first stage of the present invention, a generative cell segmentation model is used to extract cell feature information from H&E images, and the cell binary segmentation results obtained from the existing pre-trained cell segmentation network are integrated as prior information, which can generate more accurate cell binary segmentation results and improve the accuracy of the model for cell segmentation; third, in the second stage of the present invention, a generative molecular expression prediction model is used to extract pathological information features from H&E images and generate virtual mIHC images reflecting molecular expression levels. At the same time, the use of the pathological basic model improves the feature extraction ability of the model, thereby improving the efficiency of acquiring mIHC images and increasing the accuracy of the generated content; fourth, the present invention can classify cell subtypes at the molecular expression level through the mapping of cell binary segmentation results and virtual mIHC images at the spatial pixel level, which increases the accuracy of cell semantic segmentation, thereby providing a convenient and accurate tool for cancer diagnosis at the cellular level. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 is a flow chart of the method of the present invention;

[0093] Figure 2 It is a schematic diagram of the structure of the cell segmentation model and the molecular expression prediction model in the present invention;

[0094] Figure 3 for Figure 2 Schematic diagram of the structure of the cell segmentation model;

[0095] Figure 4 for Figure 3 Schematic diagram of the structure of the pre-trained cell segmentation network;

[0096] Figure 5 for Figure 3 Schematic diagram of the structure of the information integrator and discriminator network;

[0097] Figure 6 for Figure 2 Schematic diagram of the structure of the molecular expression prediction model;

[0098] Figure 7 for Figure 6 Schematic diagram of the structure of the molecular expression prediction generator network;

[0099] Figure 8 A flow chart of cell detection and analysis based on molecular expression prediction in the present invention;

[0100] Fig. 9 This is an example diagram of the visualization of the cell binary segmentation result of the present invention;

[0101] Fig.10 This is an example diagram of the visualization of the molecular expression prediction results of the present invention;

[0102] Fig.11 This is an example diagram of the visualization of the cell semantic segmentation results of the present invention;

[0103] Fig.12 It is the evaluation index of the present invention in the tasks of cell binary segmentation, cell semantic segmentation and molecular expression prediction. DETAILED DESCRIPTION

[0104] like Figure 1 , Figure 2 A two-stage image generation learning method for cell detection and analysis based on molecular expression prediction is shown, the method comprising the following steps in order:

[0105] (1) Obtaining the H&E dataset and the paired mIHC dataset to form a dataset and preprocessing it, and dividing the preprocessed dataset into a training set and a test set;

[0106] (2) Construct a cell segmentation model, such as Figure 3 As shown, the cell segmentation model consists of a pre-trained cell segmentation network, a cell segmentation generator network, an information integrator and a discriminator network; the pre-trained cell segmentation network includes a pre-trained Cellpose network, a pre-trained Hover-Net network and a pre-trained CellViT network, as shown in Figure 4 As shown;

[0107] (3) training the cell segmentation model to obtain a trained cell segmentation model;

[0108] (4) Input the H&E images in the test set into the trained cell segmentation model to obtain the cell binary segmentation results, such as Fig. 9 As shown, the results of cell binary segmentation are visualized;

[0109] (5) Constructing molecular expression prediction models, such as Figure 6 As shown, the molecular expression prediction model consists of a molecular expression prediction generator network and a molecular expression prediction discriminator network;

[0110] (6) training the molecular expression prediction model to obtain a trained molecular expression prediction model;

[0111] (7) Input the H&E images in the test set into the trained molecular expression prediction model to obtain the molecular expression prediction results, such as Fig.10 As shown, the molecular expression prediction results are visualized;

[0112] (8) The cell binary segmentation results and the molecular expression prediction results are mapped one-to-one at the pixel level in space to obtain the cell semantic segmentation results, such as Fig.11 As shown in the figure, the results of cell semantic segmentation are visualized. By mapping the cell binary segmentation results and the virtual mIHC image at the spatial pixel level, cell subtypes can be classified at the molecular expression level, which increases the accuracy of cell semantic segmentation and provides a convenient and accurate tool for cancer diagnosis at the cellular level.

[0113] In step (1), the preprocessing specifically refers to: using the valis method to align the WSI data of the H&E image in the H&E dataset and the WSI data of the mIHC image in the mIHC dataset, and then downsampling the WSI data of the H&E image and the WSI data of the mIHC image by four times, and then cutting them into image pairs of size 128×128.

[0114] In step (2), the cell segmentation generator network adopts the Attention U-Net architecture. The depth of Attention U-Net is 5, the input feature channel and the output feature channel are both 3, and the feature channels are set to 64, 128, 256, 512, and 1024 respectively during the feature extraction process; downsampling is achieved through maximum pooling, and upsampling is achieved through bilinear interpolation. The cell segmentation model extracts cell feature information from H&E images and integrates the cell binary segmentation results obtained from the existing pre-trained cell segmentation network as prior information, which can generate more accurate cell binary segmentation results and improve the accuracy of the model for cell segmentation.

[0115] In step (2), the Cellpose network is pre-trained on the Cellpose official dataset to obtain a pre-trained Cellpose network, and the pre-trained weights are loaded with the nuclei weight parameters; the Hover-Net network is pre-trained on the PanNuke dataset to obtain a pre-trained Hover-Net network, and the pre-trained weights are loaded with the hovernet_fast_pannuke weight parameters; the CellViT network is pre-trained on the PanNuke dataset to obtain a pre-trained CellViT network, and the pre-trained weights are loaded with the CellViT-SAM-H-x40 weight parameters; the nuclei weight parameters, hovernet_fast_pannuke weight parameters, and CellViT-SAM-H-x40 weight parameters remain frozen during the cell segmentation model training process.

[0116] In step (2), if Figure 5As shown, the information integrator and discriminator network are constructed based on the PatchGAN network. An upsampling module is added before the output layer of the PatchGAN network, and a feature matrix of size (1, 128, 128) is obtained by bilinear interpolation upsampling.

[0117] Step (3) specifically includes the following steps in order:

[0118] (3a) Input the H&E images in the training set into the pre-trained cell segmentation network with frozen weights to obtain the feature matrix F;

[0119] (3b) Input the H&E images in the training set into the cell segmentation generator network to obtain the feature matrix P;

[0120] (3c) Input the feature matrix F into the information integrator and the discriminator network to obtain the feature matrix A; input the feature matrix P into the information integrator and the discriminator network to obtain the feature matrix B;

[0121] (3d) A loss function is constructed based on the feature matrix A and the feature matrix B. The loss function is a multivariate discriminant loss, and its expression is:

[0122]

[0123] Among them, G s represents the cell segmentation generator network, D s represents the information integrator and discriminator network, x represents the input H&E image, z represents random noise, y cepo ,y hove and cevi Respectively represent the cell binary segmentation results obtained from the weight-frozen pre-trained Cellpose network, the pre-trained Hover-Net network, and the pre-trained CellViT network;

[0124] (3e) Based on the loss function, all non-frozen parameters of the cell segmentation model are updated through the back-propagation mechanism;

[0125] (3f) Repeat steps (3a) to (3e) until the training is completed, and save the final model parameter weights;

[0126] (3g) Load the final model parameter weights and use the test set to test the network performance of the cell segmentation generator.

[0127] In step (5), if Figure 7 As shown, the molecular expression prediction generator network includes:

[0128] The input layer uses a feature matrix with a scale of (3, 128, 128);

[0129] The pathological basis model uses the PLIP pathological basis model image encoder to extract and encode the features of the input layer through a multi-head self-attention module to obtain a feature matrix with a scale of (50, 768);

[0130] The reorganization layer is Reshape. First, the first feature representing the category code in the first dimension of the feature matrix obtained by the pathological basic model is removed to obtain a feature matrix with a scale of (49, 768). Then, the 49 features of the first dimension are reorganized according to the spatial position to obtain a feature matrix of (7, 7, 768). Then, the third dimension is adjusted to the front of the first dimension to obtain a feature matrix of (768, 7, 7).

[0131] Projection layer: The feature matrix obtained by the reorganization layer is passed through the eleventh convolutional layer Conv11 to obtain a feature matrix with a scale of (512, 16, 16);

[0132] The eleventh activation layer, the feature matrix of the projection layer is transformed through the nonlinear ReLU, i.e., ReLU11, to obtain a feature matrix with a scale of (512, 16, 16);

[0133] The first convolutional layer Conv1 convolves the feature matrix of the input layer to obtain a feature matrix with a scale of (64, 128, 128);

[0134] The first activation layer, the output feature matrix of the first convolutional layer Conv1 is transformed into a feature matrix with a scale of (64, 128, 128) by nonlinear transformation ReLU, i.e. ReLU1;

[0135] The first downsampling layer Down1 downsamples the feature matrix obtained by the first activation layer through maximum pooling to obtain a feature matrix with a scale of (64, 64, 64);

[0136] The second convolutional layer Conv2 convolves the feature matrix obtained by the first downsampling layer Down1 to obtain a feature matrix with a scale of (128, 64, 64);

[0137] The second activation layer, the output feature matrix of the second convolutional layer Conv2 is transformed into a feature matrix with a scale of (128, 64, 64) by nonlinear transformation ReLU, i.e. ReLU2;

[0138] The second downsampling layer Down2 downsamples the feature matrix obtained by the second activation layer through maximum pooling to obtain a feature matrix with a scale of (128, 32, 32);

[0139] The third convolutional layer Conv3 convolves the feature matrix obtained by the second downsampling layer Down2 to obtain a feature matrix with a scale of (256, 32, 32);

[0140] The third activation layer, the output feature matrix of the third convolutional layer Conv3 is transformed into a feature matrix with a scale of (256, 32, 32) by nonlinear transformation ReLU, i.e. ReLU3;

[0141] The third downsampling layer Down3 downsamples the feature matrix obtained by the third activation layer through maximum pooling to obtain a feature matrix with a scale of (256, 16, 16);

[0142] The fourth convolutional layer Conv4 performs a convolution operation on the feature matrix obtained by the third downsampling layer Down3 to obtain a feature matrix with a scale of (512, 16, 16);

[0143] The fourth activation layer, the output feature matrix of the fourth convolutional layer Conv4 is transformed into a feature matrix with a scale of (512, 16, 16) by nonlinear transformation ReLU, i.e. ReLU4;

[0144] The fourth downsampling layer Down4 downsamples the feature matrix obtained by the fourth activation layer through maximum pooling to obtain a feature matrix with a scale of (512, 8, 8);

[0145] The fifth convolutional layer Conv5 performs a convolution operation on the feature matrix obtained by the fourth downsampling layer Down4 to obtain a feature matrix with a scale of (512, 8, 8);

[0146] The fifth activation layer, the output feature matrix of the fifth convolutional layer Conv5 is transformed into a feature matrix with a scale of (512, 8, 8) by nonlinear transformation ReLU, i.e. ReLU5;

[0147] The first upsampling layer Up1 upsamples the feature matrix obtained from the fifth activation layer by bilinear interpolation to obtain a feature matrix with a scale of (512, 16, 16);

[0148] The first attention gate A1 operates the feature matrix obtained by the first upsampling layer Up1 and the feature matrix obtained by the fourth activation layer through the attention gate to obtain the feature matrix of (512, 16, 16);

[0149] The first concatenation layer C1 concatenates the feature matrix obtained by the first attention gate A1 and the feature matrix obtained by the eleventh activation layer in the first dimension to obtain a feature matrix with a scale of (1024, 16, 16);

[0150] The sixth convolutional layer Conv6 performs a convolution operation on the feature matrix obtained by the first concatenation layer C1 to obtain a feature matrix with a scale of (512, 16, 16);

[0151] The sixth activation layer, the output matrix of the sixth convolutional layer Conv6 is transformed into a feature matrix of size (512, 16, 16) by nonlinear transformation ReLU, i.e. ReLU6;

[0152] The second upsampling layer Up2 upsamples the feature matrix obtained by the sixth activation layer through bilinear interpolation to obtain a feature matrix with a scale of (512, 32, 32);

[0153] The second attention gate A2 operates the feature matrix obtained by the second upsampling layer Up2 and the feature matrix obtained by the third activation layer through the attention gate to obtain a feature matrix of (256, 32, 32);

[0154] The second concatenation layer C2 concatenates the feature matrix obtained by the second attention gate A2 and the feature matrix obtained by the second upsampling layer Up2 in the first dimension to obtain a feature matrix with a scale of (768, 32, 32);

[0155] The seventh convolutional layer Conv7 performs a convolution operation on the feature matrix obtained by the second concatenation layer C2 to obtain a feature matrix with a scale of (256, 32, 32);

[0156] The seventh activation layer, the output matrix of the seventh convolutional layer Conv7 is transformed into a feature matrix of size (256, 32, 32) by nonlinear transformation ReLU, i.e. ReLU7;

[0157] The third upsampling layer Up3 upsamples the feature matrix obtained by the seventh activation layer by bilinear interpolation to obtain a feature matrix with a scale of (256, 64, 64);

[0158] The third attention gate A3 operates the feature matrix obtained by the third upsampling layer Up3 and the feature matrix obtained by the second activation layer through the attention gate to obtain a feature matrix of (128, 64, 64);

[0159] The third concatenation layer C3 concatenates the feature matrix obtained by the third attention gate A3 and the feature matrix obtained by the third upsampling layer Up3 in the first dimension to obtain a feature matrix with a scale of (384, 64, 64);

[0160] The eighth convolutional layer Conv8 convolves the feature matrix obtained by the third concatenation layer C3 to obtain a feature matrix with a scale of (128, 64, 64);

[0161] The eighth activation layer, the output matrix of the eighth convolutional layer Conv8 is transformed into a feature matrix of size (128, 64, 64) by nonlinear transformation ReLU, i.e. ReLU8;

[0162] The fourth upsampling layer Up4 upsamples the feature matrix obtained by the eighth activation layer by bilinear interpolation to obtain a feature matrix with a scale of (128, 128, 128);

[0163] The fourth attention gate A4 operates the feature matrix obtained by the fourth upsampling layer Up4 and the feature matrix obtained by the first activation layer through the attention gate to obtain the feature matrix of (64, 128, 128);

[0164] The fourth concatenation layer C4 concatenates the feature matrix obtained by the fourth attention gate A4 and the feature matrix obtained by the fourth upsampling layer Up4 in the first dimension to obtain a feature matrix with a scale of (192, 128, 128);

[0165] The ninth convolutional layer Conv9 performs a convolution operation on the feature matrix obtained by the fourth concatenation layer C4 to obtain a feature matrix with a scale of (64, 128, 128);

[0166] The ninth activation layer, the output matrix of the ninth convolutional layer Conv9 is transformed into a feature matrix of size (64, 128, 128) by nonlinear transformation ReLU, i.e. ReLU9;

[0167] The tenth convolutional layer Conv10 convolves the feature matrix obtained by the ninth activation layer to obtain a feature matrix of scale (3, 128, 128);

[0168] The tenth activation layer, the output matrix of the tenth convolutional layer Conv10 is transformed into a feature matrix of size (3, 128, 128) by nonlinear transformation ReLU, i.e. ReLU10;

[0169] Output layer, outputs the feature matrix obtained by the tenth activation layer.

[0170] In step (5), the molecular expression prediction discriminator network is constructed based on the PatchGAN network, and an upsampling module is added before the output layer of the PatchGAN network. By upsampling by bilinear interpolation, a feature matrix of size (1, 128, 128) is obtained. The molecular expression prediction model extracts pathological information features from the H&E image and generates a virtual mIHC image reflecting the molecular expression level. At the same time, the use of the pathological basis model improves the feature extraction ability of the model, thereby improving the efficiency of acquiring mIHC images and increasing the accuracy of the generated content.

[0171] Step (6) specifically includes the following steps in order:

[0172] (6a) Input the H&E images in the training set into the molecular expression prediction generator network to obtain the feature matrix Q;

[0173] (6b) Inputting the mIHC images in the training set into the molecular expression prediction discriminator network to obtain the feature matrix T; inputting the feature matrix Q into the molecular expression prediction discriminator network to obtain the feature matrix S;

[0174] (6c) A total loss function is constructed based on the mIHC image, feature matrix Q, feature matrix T, and feature matrix S. The formula of the total loss function is:

[0175] L Total (G,D)=L AD (G,D)+L F (G)+L Mul (G)

[0176] Among them, G represents the molecular expression prediction generator network, D represents the molecular expression prediction discriminator network; L AD represents the discrimination loss; L F Indicates frequency loss; L Mul represents multi-scale loss;

[0177] The expression of discriminant loss is:

[0178] L AD (G,D)=E x,y [log D(x,y)]+E x,e [log(1-D(x,G(x,e)))]

[0179] Where x, y, and e represent the input H&E image, the real mIHC image, and random noise, respectively; E x,y Represents the mathematical expectation of the joint distribution of x and y; E x,e Represents the mathematical expectation of the joint distribution of x and e;

[0180] The expression for frequency loss is:

[0181] L F (G) = p·E x,y,e [||F HP (y)-F HP (G(x,e))||1]+q·E x,y,e [||F LP (y)-F LP (G(x,e))||1]

[0182] Among them, F HP (·)=IFFT[HP[FFT(·)]],FLP (·)=IFFT[LP[FFT(·)]], where FFT and IFFT represent Fourier transform and inverse Fourier transform, respectively, HP and LP represent high-pass filter and low-pass filter, respectively, p and q represent high-frequency weight and low-frequency weight, respectively, p is set to 1, and q is set to 2; E x,y,e Represents the mathematical expectation of the joint distribution of x, y, and e;

[0183] The expression of multi-scale loss is:

[0184]

[0185] Among them, λ i Represents the weight of the i-th downsampling loss; DS i (·) indicates that the real molecular expression image or the generated molecular expression prediction image has undergone i scale transformations, and each scale transformation includes four Gaussian filtering and one downsampling using a Gaussian kernel with a mean of 0 and a variance of 1;

[0186] (6d) updating all parameters of the molecular expression prediction model according to the total loss function through the back propagation mechanism;

[0187] (6e) Repeat steps (6a) to (6d) until the training is completed, and save the final model parameter weights;

[0188] (6) Load the final model parameter weights and use the test set to test the performance of the molecular expression prediction generator network.

[0189] like Figure 8 As shown, the process of cell detection and analysis based on molecular expression prediction is as follows:

[0190] Step 1: Acquire H&E images;

[0191] Step 2: Input the H&E image into the trained cell segmentation generator network to perform cell binary segmentation;

[0192] Step 3: Input the H&E image into the trained molecular expression prediction generator network to perform molecular expression prediction;

[0193] Step 4: Map the cell binary segmentation results and the molecular expression prediction results at the pixel level in space to obtain the cell semantic segmentation results;

[0194] Step 5: Save the cell binary segmentation results, molecular expression prediction results, and cell semantic segmentation results.

[0195] The cell binary segmentation results are visualized. Fig. 9As shown, (a) is the H&E image, (b) is the real cell binary segmentation result, and (c) is the generated cell binary segmentation result.

[0196] The molecular expression prediction results are visualized. Fig.10 As shown, (a) is the H&E image, (b) is the actual molecular expression result, and (c) is the predicted molecular expression result.

[0197] The cell semantic segmentation results are visualized. Fig.11 As shown, (a) is the H&E image, (b) is the real cell semantic segmentation result, and (c) is the generated cell semantic segmentation result.

[0198] like Fig.12 As shown, the network for cell detection and analysis based on molecular expression prediction can perform accurate cell segmentation and molecular expression prediction in cell binary segmentation, cell semantic segmentation, and molecular expression prediction tasks, and has high evaluation indicators. Fig.12 In the figure, ACC indicates the consistency between the cell segmentation result and the true result; Dice indicates the ratio of the intersection area size of the segmentation result and the true result multiplied by 2 to the sum of the segmentation result area size and the true result area size; IoU indicates the ratio of the intersection area size of the binary segmentation result and the true result to the area size of the union of the two; MIoU indicates the ratio of the intersection area size of the semantic segmentation result and the true result to the area size of the union of the two. ACC, Dice, IoU and MIoU are all important indicators for measuring segmentation accuracy. The larger the indicator value, the better the segmentation accuracy. PSNR reflects the degree of loss of the original signal; SNR reflects the ratio between signal power and noise power; SSIM is based on visual perception and comprehensively evaluates the image from the aspects of brightness, contrast and structure. PSNR, SNR and SSIM are all important indicators for measuring images. The larger the indicator value, the better the image quality.

[0199] In summary, the present invention realizes end-to-end cell binary segmentation, cell semantic segmentation and prediction of molecular expression levels through a two-stage generative adversarial learning framework; in the first stage of the present invention, a generative cell segmentation model is used to extract cell feature information from H&E images, and the cell binary segmentation results obtained from the existing pre-trained cell segmentation network are integrated as prior information, which can generate more accurate cell binary segmentation results and improve the accuracy of the model for cell segmentation; in the second stage of the present invention, a generative molecular expression prediction model is used to extract pathological information features from H&E images and generate virtual mIHC images reflecting molecular expression levels. At the same time, the use of the pathological basic model improves the feature extraction ability of the model, thereby improving the efficiency of acquiring mIHC images and increasing the accuracy of the generated content; by mapping the cell binary segmentation results and the virtual mIHC images at the spatial pixel level, the present invention can classify cell subtypes at the molecular expression level, increase the accuracy of cell semantic segmentation, and thus provide a convenient and accurate tool for cancer diagnosis at the cellular level.

Claims

1. A two-stage image generation learning method for cell detection and analysis based on molecular expression prediction, characterized by: The method comprises the following steps in order: (1) Obtaining the H&E dataset and the paired mIHC dataset to form a dataset and preprocessing it, and dividing the preprocessed dataset into a training set and a test set; (2) constructing a cell segmentation model, wherein the cell segmentation model is composed of a pre-trained cell segmentation network, a cell segmentation generator network, an information integrator, and a discriminator network; the pre-trained cell segmentation network includes a pre-trained Cellpose network, a pre-trained Hover-Net network, and a pre-trained CellViT network; (3) training the cell segmentation model to obtain a trained cell segmentation model; (4) Input the H&E images in the test set into the trained cell segmentation model to obtain the cell binary segmentation results, and visualize the cell binary segmentation results; (5) constructing a molecular expression prediction model, which consists of a molecular expression prediction generator network and a molecular expression prediction discriminator network; (6) training the molecular expression prediction model to obtain a trained molecular expression prediction model; (7) Inputting the H&E images in the test set into the trained molecular expression prediction model to obtain the molecular expression prediction results, and visually displaying the molecular expression prediction results; (8) The cell binary segmentation results and the molecular expression prediction results are mapped one-to-one at the pixel level in space to obtain the cell semantic segmentation results, and the cell semantic segmentation results are visualized.

2. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: In step (1), the preprocessing specifically refers to: using the valis method to align the WSI data of the H&E image in the H&E dataset and the WSI data of the mIHC image in the mIHC dataset, and then downsampling the WSI data of the H&E image and the WSI data of the mIHC image by four times, and then cutting them into image pairs of size 128×128.

3. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: In step (2), the cell segmentation generator network adopts the Attention U-Net architecture. The depth of Attention U-Net is 5, the input feature channel and the output feature channel are both 3, and the feature channels are set to 64, 128, 256, 512, and 1024 respectively during the feature extraction process; downsampling is achieved through maximum pooling, and upsampling is achieved through bilinear interpolation.

4. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: In step (2), the Cellpose network is pre-trained on the Cellpose official dataset to obtain a pre-trained Cellpose network, and the pre-trained weights are loaded with the nuclei weight parameters; the Hover-Net network is pre-trained on the PanNuke dataset to obtain a pre-trained Hover-Net network, and the pre-trained weights are loaded with the hovernet_fast_pannuke weight parameters; the CellViT network is pre-trained on the PanNuke dataset to obtain a pre-trained CellViT network, and the pre-trained weights are loaded with the CellViT-SAM-H-x40 weight parameters; the nuclei weight parameters, hovernet_fast_pannuke weight parameters, and CellViT-SAM-H-x40 weight parameters remain frozen during the cell segmentation model training process.

5. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: In step (2), the information integrator and discriminator network are constructed based on the PatchGAN network. An upsampling module is added before the output layer of the PatchGAN network, and a feature matrix of size (1, 128, 128) is obtained by bilinear interpolation upsampling.

6. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: Step (3) specifically includes the following steps in order: (3a) Input the H&E images in the training set into the pre-trained cell segmentation network with frozen weights to obtain the feature matrix F; (3b) Input the H&E images in the training set into the cell segmentation generator network to obtain the feature matrix P; (3c) Input the feature matrix F into the information integrator and discriminator network to obtain the feature matrix A; Input the feature matrix P into the information integrator and discriminator network to obtain the feature matrix B; (3d) A loss function is constructed based on the feature matrix A and the feature matrix B. The loss function is a multivariate discriminant loss, and its expression is: Among them, G s represents the cell segmentation generator network, D s represents the information integrator and discriminator network, x represents the input H&E image, z represents random noise, y cepo ,y hove and cevi Respectively represent the cell binary segmentation results obtained from the weight-frozen pre-trained Cellpose network, the pre-trained Hover-Net network, and the pre-trained CellViT network; (3e) Based on the loss function, all non-frozen parameters of the cell segmentation model are updated through the back-propagation mechanism; (3f) Repeat steps (3a) to (3e) until the training is completed, and save the final model parameter weights; (3g) Load the final model parameter weights and use the test set to test the network performance of the cell segmentation generator.

7. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: In step (5), the molecular expression prediction generator network includes: The input layer uses a feature matrix with a scale of (3, 128, 128); The pathological basis model uses the PLIP pathological basis model image encoder to extract and encode the features of the input layer through a multi-head self-attention module to obtain a feature matrix with a scale of (50, 768); In the reorganization layer, the first feature representing the category code in the first dimension of the feature matrix obtained by the pathological basic model is first removed to obtain a feature matrix with a scale of (49, 768). Then, the 49 features of the first dimension are reorganized according to the spatial position to obtain a feature matrix with a size of (7, 7, 768). Then, the third dimension is adjusted to the front of the first dimension to obtain a feature matrix with a size of (768, 7, 7). In the projection layer, the feature matrix obtained by the reorganization layer is convolved to obtain a feature matrix with a scale of (512, 16, 16); In the eleventh activation layer, the feature matrix of the projection layer is transformed into a feature matrix with a scale of (512, 16, 16) by nonlinear transformation ReLU; The first convolutional layer convolves the feature matrix of the input layer to obtain a feature matrix of scale (64, 128, 128); The first activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (64, 128, 128) on the output feature matrix of the first convolutional layer. The first downsampling layer downsamples the feature matrix obtained by the first activation layer through maximum pooling to obtain a feature matrix with a scale of (64, 64, 64); In the second convolutional layer, the feature matrix obtained by the first downsampling layer is convolved to obtain a feature matrix with a scale of (128, 64, 64); The second activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (128, 64, 64). The second downsampling layer downsamples the feature matrix obtained by the second activation layer through maximum pooling to obtain a feature matrix with a scale of (128, 32, 32); In the third convolutional layer, the feature matrix obtained in the second downsampling layer is convolved to obtain a feature matrix with a scale of (256, 32, 32); The third activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (256, 32, 32) for the output feature matrix of the third convolutional layer; The third downsampling layer downsamples the feature matrix obtained by the third activation layer through maximum pooling to obtain a feature matrix with a scale of (256, 16, 16); The fourth convolutional layer convolves the feature matrix obtained in the third downsampling layer to obtain a feature matrix with a scale of (512, 16, 16); The fourth activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (512, 16, 16) on the output feature matrix of the fourth convolutional layer. The fourth downsampling layer downsamples the feature matrix obtained by the fourth activation layer through maximum pooling to obtain a feature matrix with a scale of (512, 8, 8); The fifth convolutional layer convolves the feature matrix obtained from the fourth downsampling layer to obtain a feature matrix with a scale of (512, 8, 8); The fifth activation layer uses the nonlinear transformation ReLU to obtain a feature matrix with a scale of (512, 8, 8) on the output feature matrix of the fifth convolutional layer. In the first upsampling layer, the feature matrix obtained in the fifth activation layer is upsampled by bilinear interpolation to obtain a feature matrix with a scale of (512, 16, 16); The first attention gate operates the feature matrix obtained by the first upsampling layer and the feature matrix obtained by the fourth activation layer through the attention gate to obtain a feature matrix of (512, 16, 16); The first concatenation layer concatenates the feature matrix obtained by the first attention gate and the feature matrix obtained by the eleventh activation layer in the first dimension to obtain a feature matrix with a scale of (1024, 16, 16); The sixth convolutional layer convolves the feature matrix obtained from the first concatenation layer to obtain a feature matrix with a scale of (512, 16, 16); The sixth activation layer uses the nonlinear transformation ReLU to obtain a feature matrix of size (512, 16, 16) on the output matrix of the sixth convolutional layer. In the second upsampling layer, the feature matrix obtained in the sixth activation layer is upsampled by bilinear interpolation to obtain a feature matrix with a scale of (512, 32, 32); The second attention gate operates the feature matrix obtained by the second upsampling layer and the feature matrix obtained by the third activation layer through the attention gate to obtain a feature matrix of (256, 32, 32); The second concatenation layer concatenates the feature matrix obtained by the second attention gate and the feature matrix obtained by the second upsampling layer in the first dimension to obtain a feature matrix with a scale of (768, 32, 32); In the seventh convolutional layer, the feature matrix obtained in the second concatenation layer is convolved to obtain a feature matrix with a scale of (256, 32, 32); The seventh activation layer, the output matrix of the seventh convolutional layer is transformed by nonlinear ReLU to obtain a feature matrix of size (256, 32, 32); In the third upsampling layer, the feature matrix obtained in the seventh activation layer is upsampled by bilinear interpolation to obtain a feature matrix with a scale of (256, 64, 64); The third attention gate operates the feature matrix obtained by the third upsampling layer and the feature matrix obtained by the second activation layer through the attention gate to obtain a feature matrix of (128, 64, 64); The third concatenation layer concatenates the feature matrix obtained by the third attention gate and the feature matrix obtained by the third upsampling layer in the first dimension to obtain a feature matrix with a scale of (384, 64, 64); In the eighth convolutional layer, the feature matrix obtained in the third concatenation layer is convolved to obtain a feature matrix with a scale of (128, 64, 64); The eighth activation layer, the output matrix of the eighth convolutional layer is transformed by nonlinear ReLU to obtain a feature matrix of size (128, 64, 64); The fourth upsampling layer upsamples the feature matrix obtained in the eighth activation layer by bilinear interpolation to obtain a feature matrix with a scale of (128, 128, 128); The fourth attention gate operates the feature matrix obtained by the fourth upsampling layer and the feature matrix obtained by the first activation layer through the attention gate to obtain a feature matrix of (64, 128, 128); The fourth concatenation layer concatenates the feature matrix obtained by the fourth attention gate and the feature matrix obtained by the fourth upsampling layer in the first dimension to obtain a feature matrix with a scale of (192, 128, 128); The ninth convolutional layer convolves the feature matrix obtained from the fourth concatenation layer to obtain a feature matrix with a scale of (64, 128, 128); The ninth activation layer, the output matrix of the ninth convolutional layer is transformed into a feature matrix of size (64, 128, 128) through the nonlinear transformation ReLU; The tenth convolutional layer convolves the feature matrix obtained in the ninth activation layer to obtain a feature matrix with a scale of (3, 128, 128); The tenth activation layer, the output matrix of the tenth convolutional layer is transformed into a feature matrix of size (3, 128, 128) through the nonlinear transformation ReLU; Output layer, outputs the feature matrix obtained by the tenth activation layer.

8. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: In step (5), the molecular expression prediction discriminator network is constructed based on the PatchGAN network, and an upsampling module is added before the output layer of the PatchGAN network. By bilinear interpolation upsampling, a feature matrix of size (1, 128, 128) is obtained.

9. The two-stage image generation learning method for cell detection and analysis based on molecular expression prediction according to claim 1, characterized in that: Step (6) specifically includes the following steps in order: (6a) Input the H&E images in the training set into the molecular expression prediction generator network to obtain the feature matrix Q; (6b) Inputting the mIHC images in the training set into the molecular expression prediction discriminator network to obtain the feature matrix T; Input the feature matrix Q into the molecular expression prediction discriminator network to obtain the feature matrix S; (6c) A total loss function is constructed based on the mIHC image, feature matrix Q, feature matrix T, and feature matrix S. The formula of the total loss function is: L Total (G,D)=L AD (G,D)+L F (G)+L Mul (G) Among them, G represents the molecular expression prediction generator network, D represents the molecular expression prediction discriminator network; L AD represents the discrimination loss; L F Indicates frequency loss; L Mul represents multi-scale loss; The expression of discriminant loss is: L AD (G,D)=E x,y [logD(x,y)]+E x,e [log(1-D(x,G(x,e)))] Where x, y, and e represent the input H&E image, the real mIHC image, and random noise, respectively; E x,y Represents the mathematical expectation of the joint distribution of x and y; E x,e Represents the mathematical expectation of the joint distribution of x and e; The expression for frequency loss is: L F (G)=p·E x,y,e [||F H P(y)-F H P(G(x,e))||1]+q·E x,y,e [||F L P(y)-F LP (G(x,e))||1] Among them, F HP (·)=IFFT[HP[FFT(·)]],F LP (·)=IFFT[LP[FFT(·)]], where FFT and IFFT represent Fourier transform and inverse Fourier transform, respectively, HP and LP represent high-pass filter and low-pass filter, respectively; p and q represent high-frequency weight and low-frequency weight, respectively, p is set to 1, and q is set to 2; E x,y,e Represents the mathematical expectation of the joint distribution of x, y, and e; The expression of multi-scale loss is: Among them, λ i Represents the weight of the i-th downsampling loss; DS i (·) indicates that the real molecular expression image or the generated molecular expression prediction image has undergone i scale transformations, and each scale transformation includes four Gaussian filtering and one downsampling using a Gaussian kernel with a mean of 0 and a variance of 1; (6d) updating all parameters of the molecular expression prediction model according to the total loss function through the back propagation mechanism; (6e) Repeat steps (6a) to (6d) until the training is completed, and save the final model parameter weights; (6) Load the final model parameter weights and use the test set to test the performance of the molecular expression prediction generator network.

Citation Information

Patent Citations

  • IHC nuclear expression pathological image cell classification device and method

    CN115862008A

  • Non-small cell lung cancer digital pathological image tissue intelligent segmentation system and method

    CN118365890A

  • Analysis method of multiple immunotissue fluorescent staining images and related equipment

    CN119169012A

  • Analysis of histopathology samples

    GB202106397D0