Method for generating HER2-targeted drug efficacy prediction model based on pathological images

Through the generation method of HER2-targeted drug efficacy prediction model based on pathological images, the pathological image data of HER2-positive breast cancer patients is extracted and fused using convolutional neural network and multi-scale feature fusion module, which solves the problem of difficulty in accurately predicting the response of HER2-positive breast cancer patients to targeted treatment in the prior art, and achieves higher prediction accuracy and the formulation of personalized treatment plans.

CN118967598BActive Publication Date: 2025-06-17CHONGQING UNIV CANCER HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410996852.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-06-17
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the response of HER2-positive breast cancer patients to targeted treatment before treatment, resulting in poor treatment results and serious drug resistance problems.

Method used

Using the method of predictive model generation of HER2-targeted drug efficacy based on pathological images, the HER2-positive pathological image data is featured and fused through convolutional neural network and multi-scale feature fusion module to generate a processing framework that is adapted to multi-scale pathological images.

Benefits of technology

Improve the accuracy of the prediction of the treatment response to HER2-positive breast cancer, help develop personalized treatment plans, optimize patient treatment results and improve quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967598B_ABST
    Figure CN118967598B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a HER2-targeted drug efficacy prediction model based on pathological images, comprising: obtaining a plurality of HER2-positive pathological image data and corresponding data in the TCGA database to construct a training set and a test set; defining the data form of the training set input into the model, and extracting features from the data input into the training set through each residual block set in the convolutional neural network to obtain basic image features; fusing the basic image features through a multi-fold magnification feature fusion module and a multi-scale feature fusion module to obtain fused image features; obtaining the fused image features obtained from all branches for model training in the initial model, and ending the training to generate a HER2-targeted drug efficacy prediction model when the test set training conditions are met; wherein, during training, the cross-entropy loss function is used as the optimization objective to calculate the comprehensive training loss of the model; the present invention is beneficial to improving the accuracy of predicting the efficacy response of HER2-positive breast cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical auxiliary diagnosis, and particularly relates to a method for generating a HER2-targeted drug efficacy prediction model based on pathological images. Background Art

[0002] HER2-positive breast cancer is a subtype of breast cancer with a high risk of recurrence and metastasis. Its chemotherapy effect is poor, resulting in a lower survival rate than other subtypes. Although the emergence of HER2-targeted drugs has significantly improved the treatment effect, the problem of drug resistance in patients remains significant. It is reported that 16%-22% of patients with early breast cancer may develop drug resistance to treatment, and studies have pointed out that more than half of HER2-positive breast cancer patients develop drug resistance after one year of using trastuzumab.

[0003] Therefore, this uncertainty urgently requires a method that can accurately predict the response of HER2-positive breast cancer patients to targeted therapy before treatment, to assist in optimizing the treatment outcomes of patients and improving their quality of life. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method for generating a HER2-targeted drug efficacy prediction model based on pathological images to solve the above technical problems.

[0005] To achieve the above purpose, a method for generating a HER2-targeted drug efficacy prediction model based on pathological images includes:

[0006] Obtain a number of HER2-positive pathological image data and corresponding data in the TCGA database to construct a training set and a test set;

[0007] Define the data form of the training set input to the model, and perform feature extraction on the data input to the training set through each residual block set in the convolutional neural network to obtain basic image features;

[0008] Fuse the basic image features through a multi-fold magnification feature fusion module and a multi-scale feature fusion module to obtain fused image features;

[0009] Obtain the fused image features obtained from all branches and perform model training on the initial model. When the test set training conditions are met, the training ends to generate a HER2-targeted drug efficacy prediction model; among them, the cross-entropy loss function is used as the optimization objective during training to calculate the comprehensive training loss of the model.

[0010] Further, obtaining a number of HER2-positive pathological image data and corresponding data in the TCGA database to construct a training set and a test set includes:

[0011] Obtain a number of HER2-positive pathological image data for ROI region annotation to obtain annotation information;

[0012] Perform position - corresponding grid processing on HER2 - positive pathological image data according to the annotation information to obtain tiles at two magnification levels of 10× and 40×;

[0013] Perform color normalization on all tiles according to the structure - preserving color normalization method, and perform data augmentation on the tiles after color normalization to obtain target tiles;

[0014] Divide all target tiles into a training set and a test set according to the stratified sampling of sensitive and drug - resistant results to obtain the CUCH training set and the CUCH test set;

[0015] Obtain corresponding data in the TCGA database to construct an external validation set to obtain the TCGA test set.

[0016] Furthermore, define the data form of the training set input to the model, and extract features from the data input to the training set through each residual block set in the convolutional neural network to obtain basic image features, including:

[0017] Define the data form of the training set input to the model {X i , T i}; where, i ∈ N, N is the total number of data pairs in the training set, represents the input of the 40 - fold image, X 10 represents the input of the 10 - fold image, X 10_dowwn represents the image input of the 10 - fold image after downsampling operation, T i is the corresponding binary label indicating the classification of the data pair, T i = 1 represents that the classification of X i is drug - resistant, T i = 0 represents that the classification of X i is sensitive to the drug;

[0018] Adjust the image resolution of the 40 - fold image and the 10 - fold image to 256×256 pixels, and extract features from the image bag composed of images through each block in the ResNet34 network to obtain basic image features Φ m (·) is the output function of the final residual block of the preset group of blocks in the ResNet34 network, I 40 is an image bag composed of 16 tiles, R represents the feature matrix, including the following dimensions: B = 16, C m is the feature dimension extracted by each block, H m ×Wm is the size of the feature image with an input image of 256 pixels;

[0019] Feature extraction is performed on X through each block in the ResNet34 network 10 to obtain the basic image features

[0020] Feature extraction is performed on X through each block in the ResNet34 network 10_down to obtain the basic image features where X 10_down has an image resolution of 128×128 pixels, is the size of the feature image with an input image of 128 pixels.

[0021] Furthermore, the basic image features are fused through the multi-fold magnification feature fusion module and the multi-scale feature fusion module to obtain the fused image features, including:

[0022] Information fusion is performed on the basic image features of the 10-fold image through the multi-fold magnification feature fusion module to obtain the first fused image feature where M is the MMFF operation, m ∈ {1, 2, 3, 4}, f represents a convolutional block containing a 3×3 convolutional layer with the same input and output dimensions of the convolutional layer, followed by a batch normalization function and a tanh activation function, and θ is the upsampling method of bilinear interpolation;

[0023] The first fused image feature and the basic image features of the 40-fold image are jointly processed through the multi-scale feature fusion module to obtain the fused image feature Z m = S(Y m , f(Φ m (I 40 ))); where S is the MSFF operation.

[0024] Furthermore, the MMFF operation includes:

[0025] Performing a multiplication operation and a convolutional block operation on the feature matrices and to obtain the reweighted feature where * represents element-wise multiplication between matrix elements;

[0026] Obtaining the feature Y′ m and adding it to the basic image features to obtain where Represents a function for aggregating different feature information, characterizing the enhancement of feature expression within adjacent feature spaces.

[0027] Furthermore, the MSFF operation includes:

[0028] After extracting features through each block in the ResNet34 network, perform convolutional block processing on the 40 - fold basic image features Φ m (I 40 ) and Y m Perform an aggregation operation to obtain comprehensive fused image features

[0029] Furthermore, obtain the fused image features obtained from all branches and perform model training on the initial model, including:

[0030] Obtain the fused image features Z obtained from the four branches m , predict the classification result through a single max - pooling and multi - layer perceptron operation to obtain the classification result

[0031] Obtain the preset parameter W corresponding to the contribution degree of each operation stage m =[a, b, c, d]; where a + b + c + d = 1;

[0032] According to W m Perform an element - wise multiplication operation with the classification result of each corresponding operation stage to obtain the target prediction result

[0033] In the post - processing stage, determine the prediction output of the HER2 - positive pathological image data by averaging and aggregating the target prediction results of all data pairs in each HER2 - positive pathological image data; among them, the Adam optimizer is used during the model training process, the initial learning rate is 0.0001, and a dynamic learning rate adjustment mechanism is cited.

[0034] Furthermore, during training, the cross - entropy loss function is used as the optimization objective to calculate the comprehensive training loss of the model, including:

[0035]

[0036] Among them, N is the number of samples in the training set, O is the true label of sample i, is the target prediction result, that is, the probability that the model predicts sample i as the positive class.

[0037] The beneficial effects of the present invention are as follows:

[0038] The present invention provides a method for generating a HER2-targeted drug efficacy prediction model based on pathological images, and generates a processing framework suitable for multi-scale pathological images to improve the accuracy of predicting the treatment response of HER2-positive breast cancer.

[0039] Other advantages, objectives and features of the present invention will be described in the following specification, and to some extent will be obvious to those skilled in the art, or those skilled in the art can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the accompanying drawings.

[0040] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0041] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0042] Figure 1 is a flowchart of a method for generating a HER2-targeted drug efficacy prediction model based on pathological images in an embodiment of the present invention;

[0043] Figure 2 is a schematic structural diagram of the DFSNet network model of a method for generating a HER2-targeted drug efficacy prediction model based on pathological images in an embodiment of the present invention;

[0044] Figure 3 is a comparison chart of the performance of DFSNet in the CUCH training set of a method for generating a HER2-targeted drug efficacy prediction model based on pathological images in an embodiment of the present invention. Detailed Embodiments

[0045] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0046] Please refer to Figure 1 , the present invention provides a method for generating a HER2-targeted drug efficacy prediction model based on pathological images, including:

[0047] S101. Obtain a number of HER2-positive pathological image data and corresponding data in the TCGA database to construct a training set and a test set;

[0048] S102. Define the data form of the training set input to the model, and extract features from the data input to the training set through each residual block set in the convolutional neural network to obtain basic image features;

[0049] S103. Fuse the basic image features through a multi-fold magnification feature fusion module and a multi-scale feature fusion module to obtain fused image features;

[0050] S104. Obtain the fused image features obtained from all branches and perform model training on the initial model. When the test set training conditions are met, end the training to generate a HER2 targeted drug efficacy prediction model; among them, the cross-entropy loss function is used as the optimization objective during training, and the comprehensive training loss of the model is calculated;

[0051] The working principle and beneficial effects of the above technical solution are as follows: The application of deep learning algorithms in analyzing clinical pathological tissue section images of tumors shows great potential. It can deeply mine and analyze a large amount of labeled data and raw data, and transform them into deep digital signals that help understand tumor development. In particular, a multi-scale prediction model constructed by integrating multi-scale image data can provide a more comprehensive and accurate prediction of treatment effects, providing key guidance for formulating personalized treatment plans; in response to the challenge of predicting the effects of HER2-positive patients after targeted therapy in the current medical field, the present invention develops a processing framework DFSNet (DualFusionScaleNet) suitable for multi-scale pathological images. Among them, the schematic diagram of the DFSNet network model structure can be referred to Figure 2 , to improve the accuracy of predicting the treatment response of HER2-positive breast cancer; the process of constructing a HER2-positive breast cancer targeted therapy efficacy prediction model for pathological images includes: first, annotating the ROI regions of the original image data; subsequently, the ROI regions of these images are cut into small tiles according to multi-scale position information, and preprocessing steps such as color normalization and data augmentation are performed on these small tiles; then, the Resnet34 network is used to extract features through the difference in feature expressions between different convolutional layers. These features are then sent to a multi-magnification and multi-scale feature fusion module to fuse the image features and enhance the prediction performance of the model; finally, a vote is performed through trainable weights to generate a prediction result;

[0052] When the method is actually applied, it preferably includes the following steps:

[0053] S1: Obtain clinically significant pathological images and clinical information according to strict inclusion and exclusion criteria;

[0054] S2: Perform grid processing for position correspondence of multi-layer pathological images to obtain tiles at two magnification levels of 10× and 40×;

[0055] S3: Normalize the colors of all tiles by a structure-preserving color normalization method;

[0056] S4: Divide 341 WSI with clinical significance into the CUCH training set and the CUCH test set through stratified sampling according to sensitive and drug-resistant results. The data in the TCGA database is used as an external validation set, denoted as the TCGA test set;

[0057] S5: Define the data form of the model input {X i , T i}. For each data pair i in the dataset, there is where i ∈ N, and N represents the total number of data pairs in the dataset, represents the input of the 40x image, X 10 represents the input of the 10x image, X 10_down represents the image input after downsampling the 10x image, T i is the corresponding binary label, indicating the classification of the data pair. T i = 1 represents that the classification of X i is drug-resistant, and T i = 0 represents that the classification of X i is sensitive to the drug;

[0058] S6: Adjust the image resolutions of 40x and 10x to 256×256 pixels to ensure the consistency of the input data and the standardization of the processing flow. Then, process through each set of residual blocks (blocks) of the convolutional neural network. The feature representations extracted by each block of the ResNet34 network for the image bag composed of 16 40x images are , m ∈ {1, 2, 3, 4}, Φ m (·) is the output function of the final residual block of the preset group of blocks in the ResNet34 network, I 40 is the image bag composed of 16 Tiles, and R represents the feature matrix, including the following dimensions: B = 16, C m is the feature dimension extracted by each block, H m ×W m is the size of the feature image with a 256-pixel input image;

[0059] Extract the feature of the X 10 image through each block in the ResNet34 network to obtain the basic image feature

[0060] Extract the feature of the X 10_down image through each block in the ResNet34 network to obtain the basic image feature where X10_down The image resolution is 128×128 pixels, which is the size of the feature image with 128 pixels as the input image;

[0061] S7: Perform information fusion between features at different magnification levels through the Multi - magnification Feature Fusion (MMFF) module:

[0062]

[0063] where Y m is the feature after fusing features at multiple magnification levels, M is the MMFF operation, which can be specifically obtained from S8 and S9, is the feature after convolution block processing after extracting features from each block of the 10 - fold image through the ResNet34 network, is the 10 - fold magnified image after downsampling. Similarly, features are extracted from each block of the ResNet34 network, and then after feature alignment and convolution block processing, the obtained feature is specifically expressed as: f represents a convolution block containing a 3×3 convolutional layer, where the input and output dimensions of the convolutional layer remain unchanged, followed by a batch normalization function and a tanh activation function, and θ is the upsampling method of bilinear interpolation;

[0064] S8: First, perform a multiplication operation and a convolution block operation between the feature matrices and obtained in S7 to get the re - weighted feature Y′ m , expressed as where * represents element - by - element multiplication between matrix elements;

[0065] S9: Then, by adding the feature Y′ m to the original image feature of the 10 - fold tile to generate Y m , expressed as: where represents a function for aggregating different feature information, aiming to strengthen the feature expression within the adjacent feature space;

[0066] S10: Subsequently, further jointly process the fused feature Y m with the feature set I 40 of the 40 - fold image through the Multi - scale Feature Fusion (MSFF) module to form the final feature representation Z m = S(Y m , f(Φ m (I 40))) where f is the same convolution block operation, S is the multi-scale feature fusion MSFF, and m ∈ {1, 2, 3, 4};

[0067] S11: The S module in S10 can be obtained by getting the 40-fold feature representation obtained by performing convolution block processing on the features extracted from each block of the ResNet34 network and aggregating it with Y m to obtain the comprehensive fused feature representation Z m ;

[0068] S12: The feature Z obtained from the four branches m , after a single max pooling and multi-layer perceptron (MLP) operation to predict the result, to obtain the classification result

[0069] To finely adjust the contribution of each stage, we set the trainable parameter W m = [a, b, c, d]; where a + b + c + d = 1;

[0070] These parameters generate a more accurate and reliable prediction result through an element-wise multiplication operation with the classification result of each stage

[0071] S13: In this study, the cross-entropy loss function is used as the optimization objective to calculate the comprehensive training loss of the model. The loss function is as follows:

[0072]

[0073] where N is the number of samples in the training set, O is the true label of sample i, is the target prediction result, that is, the probability that the model predicts sample i as the positive class;

[0074] S14: In the post-processing stage, the prediction output of the WSI is determined by averaging and aggregating the prediction results of all data pairs in each whole-slide image (WSI);

[0075] S15: In the model training of this study, an Omnisky server configured with PyTorch (version information: Python 3.9.7: PyTorch 1.11.0: Skleam 0.24.2) was selected for model training; the optimizer used for training was the Adam optimizer, with its initial learning rate set to 0.0001, and a dynamic learning rate adjustment mechanism was introduced: when the loss on the validation set increased twice in a row, a strategy of reducing the learning rate by 10% was adopted to stabilize the training process; the batch size was set to 64, and a total of 20 epochs were iterated; among them, the performance of DFSNet on the CUCH training set can be referred to Figure 3, where A is the ROC curve and B is the confusion matrix.

[0076] In one embodiment, obtaining a number of HER2-positive pathological image data and corresponding data in the TCGA database to construct a training set and a test set includes:

[0077] Obtaining a number of HER2-positive pathological image data for ROI region annotation to obtain annotation information;

[0078] Performing position-corresponding grid processing on the HER2-positive pathological image data according to the annotation information to obtain tiles at two magnification levels of 10× and 40×;

[0079] Performing color normalization processing on all the tiles according to the structure-preserving color normalization method, and performing data augmentation processing on the tiles after color normalization processing to obtain target tiles;

[0080] Dividing all the target tiles into a training set and a test set according to the stratified sampling of sensitive and drug-resistant results to obtain a CUCH training set and a CUCH test set;

[0081] Obtaining corresponding data in the TCGA database to construct an external validation set to obtain a TCGA test set;

[0082] The working principle and beneficial effects of the above technical solution are: obtaining clinically significant and high-quality pathological images according to strict inclusion and exclusion criteria, and then performing digital processing on the H&E-stained tissue pathological sections through a KF-PRO-005-HI digital slide scanner to generate whole slide imaging (WSI) with high clarity and accuracy; in the training set and test set finally constructed in this solution, the data included is WSI;

[0083] Using the OpenSlide library in Python to obtain the detailed layers of the WSI at magnifications of 10 times and 40 times; then for the images at 10 times magnification, dividing the tumor region into small tiles of 512×512 pixel size; then, converting these tiles into binary images to evaluate the proportion of blank areas within each tile, and screening out those tiles with blank areas exceeding 50% to ensure the quality of the data; then, accurately obtaining the corresponding tiles at 40 times magnification based on the coordinate position information of these tiles;

[0084] The structure-preserving color normalization (SPCN) method is used to perform color normalization on each image block. Specifically, 200 tiles are randomly extracted from each source domain image folder, and then the stain_separate function in the Vahadane algorithm is used to perform color separation on each extracted image block to obtain the color matrix W of each sample image block. Next, the average value of the W matrices of these sample image blocks is calculated to obtain an average w matrix that reflects the color characteristics of the entire image folder. Subsequently, a representative standard image block and the average w matrix calculated previously are selected as the normalization standard, and the image color matrix is ​​carefully adjusted and normalized by applying the SPCN_2 method of the Vahadane algorithm. The integrity of subsequent feature extraction can also be improved through data enhancement.

[0085] In this scheme, we prefer to adopt the preset data partitioning strategy, and stratify the 341 clinically significant pathological WSIs into CUCH training set and CUCH test set according to the ratio of 3:1. The 110 pathological slides of TCGA (The cancer genome atlas) database are used as an external validation set, recorded as TCGA test set.

[0086] In one embodiment, the data form of the training set in the model input is defined, and the data of the training set input is subjected to feature extraction through each residual block set in the convolutional neural network to obtain basic image features, including:

[0087] Define the training set in the model input data form {X i , T i};in, N is the total number of data pairs in the training set, represents the input of 40 times the image, X 10 represents the input of 10 times the image, X 10_down represents the image input after 10 times image downsampling operation, T i is the corresponding binary label, indicating the classification of the data pair, T i =1 represents X i The classification is resistant, T i =0 represents X i The classification is sensitive to the drug;

[0088] The image resolution of the 40x and 10x images is adjusted to 256×256 pixels, and each block in the ResNet34 network Feature extraction is performed on the image bag composed of images to obtain the basic image features Φ m (·) is the output function of the final residual block of the preset group of blocks of the ResNet34 network, and I 40 is an image bag composed of 16 Tiles, B = 16, C m is the feature dimension extracted for each block, H m ×W m is the size of the feature image with 256 pixels as the input image;

[0089] The preset group of blocks of the network is preferably 4 groups, and m is the number of groups of 4 groups of blocks;

[0090] Feature extraction is performed on the X 10 image through each block in the ResNet34 network to obtain the basic image features

[0091] Feature extraction is performed on the X 10_down image through each block in the ResNet34 network to obtain the basic image features Among them, X 10_down has an image resolution of 128×128 pixels, is the size of the feature image with 128 pixels as the input image.

[0092] In one embodiment, the basic image features are fused through a multi-fold magnification feature fusion module and a multi-scale feature fusion module to obtain the fused image features, including:

[0093] Information fusion is performed on the basic image features of the 10-fold image through the multi-fold magnification feature fusion module, aiming to promote and optimize the information fusion between features at different magnification levels to obtain the first fused image feature Among them, M is the MMFF operation, which is the key to realizing the fusion of image features at different magnification multiples. m ∈ {1, 2, 3, 4}, f represents a convolutional block containing a 3×3 convolutional layer, the input and output dimensions of the convolutional layer remain unchanged, followed by a batch normalization function and a tanh activation function, and θ is the upsampling method of bilinear interpolation, aiming to align the feature space; specifically, when a pair of features from different magnification multiples are aligned in the feature space, they are respectively input into the convolutional block operation to fuse the interpolated features;

[0094] The first fused image feature and the basic image features of the 40-fold image are jointly processed through the multi-scale feature fusion module to obtain the fused image feature Zm = S(Y m , f(Φ m (I 40 ))); where S is the MSFF operation, emphasizing the importance of different scale features in improving the performance of the prediction model.

[0095] In one embodiment, the MMFF operation includes:

[0096] Performing a multiplication operation and a convolutional block operation on the feature matrix and to obtain the reweighted feature where * represents element-wise multiplication between matrix elements, realizing the reweighting of features and thus strengthening the interaction between them;

[0097] Obtaining the feature Y' m and the basic image feature performing an addition operation, and performing a skip connection similar to the residual structure of the ResNet34 network to generate to reuse old features while learning new features; where represents a function for aggregating different feature information, characterizing the enhancement of feature expressions within adjacent feature spaces.

[0098] In one embodiment, the MSFF operation includes:

[0099] The 40-fold basic image feature Φ m (I 40 ) after convolution block processing by extracting features through each block in the ResNet34 network is aggregated with Y m to obtain the comprehensive fused image feature

[0100] Specifically: The 40-fold features are also processed through the convolutional block operation to ensure that the numerical orders of magnitude of the features are consistent. Then, the 40-fold feature representation after convolution block processing by extracting features through each block of the ResNet34 network is aggregated with Y m to obtain the comprehensive fused feature expression Z m , realizing the in-depth parsing and understanding of complex image content.

[0101] In one embodiment, obtaining the fused image features obtained from all branches for model training in the initial model includes:

[0102] Obtaining the fused image feature Z from four branches m, the classification result is predicted through a single max pooling and multi-layer perceptron operations. The MLP structures of the four branches are: block1(64-8-2), block2(128-16-2), block3(256-32-2), block4(512-64-2), and the classification result is obtained

[0103] Obtain the preset parameter W corresponding to the contribution degree of each operation stage m = [a, b, c, d]; where a + b + c + d = 1;

[0104] Preset parameter W m is trainable. The purpose of introducing the preset parameter W m is to finely adjust the contribution degree of each stage and further optimize the prediction performance of the model; These parameters achieve precise weighting of the outputs of each stage through element-wise multiplication operations with the classification results of each stage;

[0105] According to W m Perform element-wise multiplication operations with the classification results of each corresponding operation stage. This weighting mechanism allows the model to obtain more accurate and reliable target prediction results on the basis of comprehensively considering the feature importance of each stage

[0106] In the post-processing stage, the prediction output of the HER2-positive pathological image data (WSI) is determined by averaging and aggregating the target prediction results of all data pairs in each HER2-positive pathological image data (WSI); Among them, the Adam optimizer is used during the model training process, the initial learning rate is 0.0001, and a dynamic learning rate adjustment mechanism is introduced;

[0107] Specifically, in the model training of this study, an Omnisky server configured with PyTorch (version information: Python3.9.7, PyTorch 1.11.0, Skleam 0.24.2) was selected for model training; we connected to this server via SSH. This server is equipped with two Intel(R) Xeon(R) Gold 6348 CPUs @ 2.60GHz, with a total of 56 cores and 112 threads; the GPU used is NVIDIA A800 80GB PCIe, and this GPU is installed with a driver with version number 535.104.05 and supports CUDA 12.2; in addition, the server also has 515.6GB of memory and 1.8TB of hard disk storage space; the optimizer used is the Adam optimizer, with its initial learning rate set to 0.0001, and a dynamic learning rate adjustment mechanism is introduced: when the loss of the validation set rises twice in a row, a strategy of reducing the learning rate by 10% is adopted to stabilize the training process; the batch size is set to 64, and a total of 20 epochs are iterated;

[0108] For the ablation study of the model, we constructed a model SFSNet (SingleFusionScaleNet) that only uses the multi-magnification feature fusion (MMFF) module to explore the independent and cooperative effects of MMFF and MSFF; at the same time, we compared the performance of SFSNet with single-scale (10x and 40x) models and the DFSNet model in various average performance metrics. Please refer to Tables 1 and 2:

[0109] Table 1

[0110] Model AUC ACC Sen Spe PPV NPV 10× 0.795 0.802 0.631 0.850 0.545 0.890 40× 0.682 0.813 0.526 0.895 0.588 0.869 SFSNet 0.805 0.857 0.614 0.925 0.717 0.895 DFSNet 0.838 0.860 0.666 0.915 0.698 0.906

[0111] Table 2

[0112] Model AUC ACC Sen Spe PPV NPV 10× 0.748 0.700 0.666 0.701 0.114 0.973 40× 0.687 0.545 0.666 0.538 0.076 0.965 SFSNet 0.717 0.794 0.556 0.808 0.146 0.969 DFSNet 0.749 0.896 0.556 0.916 0.303 0.972

[0113] Among them, Table 1 shows the performance of DFSNet and its ablation models on the CUCH test set, and Table 2 shows the performance of DFSNet and its ablation models on the TCGA test set; this comparison particularly focuses on the single-scale baseline models at a 256-pixel resolution, defined as 10x and 40x, to highlight the superiority of the multi-scale model.

[0114] In one embodiment, the cross-entropy loss function is used as the optimization objective during training to calculate the comprehensive training loss of the model, including:

[0115]

[0116] Among them, N is the number of samples in the training set, O is the true label of sample i, is the target prediction result, that is, the probability that the model predicts sample i as the positive class.

[0117] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in terms of form and details without departing from the scope defined by the claims of the present invention.

Claims

1. A method for generating a HER2 targeted drug efficacy prediction model based on pathological images, characterized in that: include: Obtain several HER2-positive pathological image data and corresponding data in the TCGA database to construct training sets and test sets; Define the data format of the training set in the model input, extract features of the training set input data through each residual block set in the convolutional neural network, and obtain basic image features; The basic image features are fused through the multi-magnification feature fusion module and the multi-scale feature fusion module to obtain the fused image features; The fused image features obtained from all branches are obtained to train the initial model. When the test set training conditions are met, the training is terminated to generate a HER2 targeted drug efficacy prediction model. The cross entropy loss function is used as the optimization target during training to calculate the comprehensive training loss of the model. Define the data format of the training set in the model input, extract features of the training set input data through each residual block set in the convolutional neural network, and obtain basic image features, including: Define the data format of the training set in the model input ;in, , , N is the total number of data pairs in the training set, represents 40 times the input of the image, represents the input of 10 times the image, represents the image input after 10 times image downsampling operation, is the corresponding binary label, indicating the classification of the data pair, represent The classification is resistant, represent 's classification as sensitive to the drug; Adjust the image resolution of 40x and 10x images to 256 256 pixels, through each block in the ResNet34 network The image bag composed of images is used for feature extraction to obtain basic image features , , The output function of the final residual block of the preset group blocks for the ResNet34 network, is an image bag consisting of 16 Tiles. R represents the feature matrix, which contains the following dimensions: B=16, The feature dimensions extracted for each block are is the size of the feature image with 256 pixels as the input image; Through each block in the ResNet34 network Image feature extraction to obtain basic image features , ; Through each block in the ResNet34 network Image feature extraction to obtain basic image features , ;in, The image resolution is 128 128 pixels, is the size of the feature image with 128 pixels as the input image; The basic image features are fused through the multi-magnification feature fusion module and the multi-scale feature fusion module to obtain the fused image features, including: The first fused image feature is obtained by fusing the basic image features of the 10x image through the multi-magnification feature fusion module. ;in, , , M is the MMFF operation, , represents a convolutional block consisting of a 3×3 convolutional layer with unchanged input and output dimensions, followed by a batch normalization function and a tanh activation function. It is the upsampling method of bilinear interpolation; The first fused image features and the basic image features of the 40-fold image are jointly processed by the multi-scale feature fusion module to obtain the fused image features. ; Where S is the MSFF operation.

2. The method for generating a HER2 targeted drug efficacy prediction model based on pathological images according to claim 1, characterized in that: Several HER2-positive pathological image data and corresponding data from the TCGA database were obtained to construct training and test sets, including: Obtaining several HER2-positive pathological image data to mark the ROI region and obtain the marking information; According to the annotation information, the HER2-positive pathological image data is gridded corresponding to the position of the multi-layer pathological image to obtain tiles at two magnifications of 10× and 40×; All tiles are color normalized according to the structure-preserving color normalization method, and data enhancement is performed on the color-normalized tiles to obtain the target tiles; All target tiles are divided into training sets and test sets according to stratified sampling of sensitive and resistant results to obtain CUCH training sets and CUCH test sets; The corresponding data in the TCGA database were obtained to construct an external validation set and obtain the TCGA test set.

3. The method for generating a HER2 targeted drug efficacy prediction model based on pathological images according to claim 1, characterized in that: MMFF operations include: For the feature matrix and Perform multiplication and convolution block operations to obtain re-weighted features ; Where * represents element-by-element multiplication between matrix elements; The characteristics and Perform the addition operation and get ;in, Represents a function used to aggregate different feature information, representing and strengthening the feature expression in the adjacent feature space.

4. The method for generating a HER2 targeted drug efficacy prediction model based on pathological images according to claim 1, characterized in that: MSFF operations include: 40 times the basic image features after extracting features from each block in the ResNet34 network and processing them with convolution blocks and Perform aggregation operation to obtain comprehensive fused image features .

5. The method for generating a HER2 targeted drug efficacy prediction model based on pathological images according to claim 1, characterized in that: The fused image features obtained from all branches are obtained to train the model in the initial model, including: Get the fused image features from the four branches , predict the classification result through a single maximum pooling and multi-layer perceptron operation to obtain the classification result ; Among them, the MLP structures of the four branches are: block1 (64-8-2), block2 (128-16-2), block3 (256-32-2), block4 (512-64-2); Get the preset parameters corresponding to the contribution of each operation stage ;in, ; according to Perform element-wise multiplication with the classification results of each corresponding operation stage to obtain the target prediction result ; In the post-processing stage, the predicted output of the HER2-positive pathological image data is determined by averaging and aggregating the target prediction results of all data pairs in each HER2-positive pathological image data; the Adam optimizer is used in the model training process, the initial learning rate is 0.0001, and the dynamic learning rate adjustment mechanism is used.

6. The method for generating a HER2 targeted drug efficacy prediction model based on pathological images according to claim 1, characterized in that: During training, the cross entropy loss function is used as the optimization target to calculate the comprehensive training loss of the model, including: Where N is the number of samples in the training set, O is the true label of sample i, is the target prediction result, that is, the probability that the model predicts that sample i is a positive class.

Citation Information

Patent Citations

  • Neural network enhancement method and system and application thereof

    CN112416293A

  • Breast pathology image classification method based on DenseNet and conditional random field

    CN117152520A