Method for predicting status of human epidermal growth factor receptor-2 and related apparatus
Patent Information
- Application Number
- CN202610775281.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-28
AI Technical Summary
然而,现有方法的全视野数字病理切片与磁共振图像属于差异显著的异构模态,两者在成像视角、信息密度以及数据维度等方面存在较大差异,其显式对应关系通常较为模糊
[0042] The advantages and beneficial effects of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application:
Smart Images

Figure CN122657572A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and related equipment for predicting the state of human epidermal growth factor receptor-2. Background Technology
[0002] Whole-slide image (WSI) and magnetic resonance imaging (MRI) are two commonly used medical imaging modalities. Among related technologies, the fusion of WSI and MRI images can simultaneously utilize both microscopic and macroscopic information about tumors. WSI reflects cell morphology, tissue structure, and the tumor microenvironment at the microscopic level, while MRI provides information on the overall structure and functional status of the tumor at the macroscopic level. The two technologies are highly complementary in their information expression. Combining WSI and MRI for breast cancer-related tasks helps to more comprehensively characterize tumor biological features, thus providing richer information support for molecular subtyping and treatment evaluation. Human epidermal growth factor receptor-2 (HGF-2) status is an important biomarker in breast cancer subtyping and treatment decisions, closely related to the selection of targeted therapy regimens and patient prognosis. Therefore, predicting HGF-2 status based on multimodal fusion of WSI and MRI images has significant application value in precision diagnosis and treatment of breast cancer.
[0003] Existing multimodal fusion methods typically achieve joint modeling of information from different modalities through attention mechanisms, contrastive learning, and modality weighting. This involves constructing interaction relationships between cross-modal features, mapping features from different modalities to a unified representation space to enhance semantic consistency, or weighting and integrating features based on their contribution to the task results. This allows for the synergistic utilization of multi-source information and has demonstrated application value in various medical AI tasks. However, existing methods treat full-view digital pathology slides and magnetic resonance imaging (MRI) images as significantly different modalities, exhibiting substantial differences in imaging perspective, information density, and data dimensionality, often resulting in ambiguous explicit correspondences. Existing methods often directly interact features between the two modalities, failing to adequately consider this significant modal domain difference. Furthermore, the information content of full-view digital pathology slides and MRI images is often unbalanced during fusion. Compared to MRI images, full-view digital pathology slides typically contain richer information and are more likely to dominate the model training process, leading to over-reliance on the dominant modality and insufficient learning of non-dominant modalities. This further reduces the model's accuracy in predicting the state of human epidermal growth factor receptor-2. Therefore, there are still technical problems that need to be solved in the relevant technologies. Summary of the Invention
[0004] The purpose of this application is to at least partially solve one of the technical problems existing in the prior art.
[0005] Therefore, one objective of this application is to provide a method and related equipment for predicting the state of human epidermal growth factor receptor-2, which can improve the accuracy of predicting the state of human epidermal growth factor receptor-2.
[0006] To achieve the above-mentioned technical objectives, the technical solution adopted in this application includes: a method for predicting the state of human epidermal growth factor receptor-2, comprising the following steps: acquiring a full-view digital pathological slide, magnetic resonance imaging (MRI), and a first state label of human epidermal growth factor receptor-2 (HGF-2) from a breast cancer patient; inputting the full-view digital pathological slide and the MRI into a modal branch network to obtain a first modal feature representation corresponding to the full-view digital pathological slide and a second modal feature representation corresponding to the MRI; mapping the first modal feature representation and the second modal feature representation to probability distribution representations in a shared latent space, and aligning the first modal feature representation and the second modal feature representation based on the probability distribution representations to obtain the first state of the human epidermal growth factor receptor-2 (HGF-2) after alignment in the latent space. The system employs a first modality feature representation and a second modality feature representation. The first state label, the aligned first modality feature representation, and the second modality feature representation are enhanced and fused, and then input into a shared classifier to obtain a first human epidermal growth factor receptor-2 (HFR-2) state prediction result corresponding to the first modality feature representation and a second HFR-2 state prediction result corresponding to the second modality feature representation. Based on the first and second HFR-2 state prediction results, the parameters of the state prediction model are dynamically modulated to obtain a trained first state prediction model. The full-view digital pathological slides and magnetic resonance imaging of the patient to be tested are input into the trained first state prediction model to obtain the state of human epidermal growth factor receptor-2.
[0007] In addition, the method for predicting the state of human epidermal growth factor receptor-2 according to the above embodiments of the present invention may also have the following additional technical features:
[0008] Further, in this embodiment of the application, the step of mapping the first modal feature representation and the second modal feature representation to probability distribution representations in a shared latent space, and aligning the first modal feature representation and the second modal feature representation based on the probability distribution representations to obtain the aligned first modal feature representation and the second modal feature representation in the latent space, specifically includes:
[0009] Mapping said first modality feature representation and said second modality feature representation to a mean vector and a standard deviation vector respectively;
[0010] Converting said mean vector and said standard deviation vector into a Gaussian distribution;
[0011] Aligning said first modality feature representation and said second modality feature representation based on a distribution constraint loss and said Gaussian distribution representation, to obtain said aligned first modality feature representation and said aligned second modality feature representation in a latent space.
[0012] Further, in the embodiments of the present application, said distribution constraint loss is:
[0013]
[0014] wherein, represents the distribution constraint loss; and respectively represent the probability distribution representations of the magnetic resonance imaging modality and the whole slide imaging modality of the i-th sample in a shared latent space; KL(·||·) represents Kullback-Leibler divergence; λ represents a regularization weight; m∈{WSI, MRI} represents a modality index; represents a multivariate standard normal distribution with a mean of 0 and a covariance matrix of identity matrix I.
[0015] Further, in the embodiments of the present application, said dynamically modulating parameters of a state prediction model based on said first human epidermal growth factor receptor-2 status prediction result and said second human epidermal growth factor receptor-2 status prediction result to obtain a trained first state prediction model specifically comprises:
[0016] Calculating a contribution coefficient ratio of said first human epidermal growth factor receptor-2 status prediction result and said second human epidermal growth factor receptor-2 status prediction result;
[0017] Constructing a gradient modulation coefficient based on said contribution coefficient ratio;
[0018] Determining a total model loss based on a classification loss and the distribution constraint loss;
[0019] Dynamically modulating parameters of the state prediction model based on the total model loss to obtain a trained first state prediction model.
[0020] Further, in the embodiments of the present application, said calculating a contribution coefficient ratio of said first human epidermal growth factor receptor-2 status prediction result and said second human epidermal growth factor receptor-2 status prediction result specifically comprises:
[0021] Determine the sum of all first-person epidermal growth factor receptor-2 state predictions and the sum of all second-person epidermal growth factor receptor-2 state predictions in a training batch;
[0022] The ratio of the sum of all first-person epidermal growth factor receptor-2 status prediction results to the sum of all second-person epidermal growth factor receptor-2 status prediction results is used as the contribution coefficient ratio.
[0023] Furthermore, in this embodiment, the model loss value is jointly optimized using classification loss and distribution constraint loss, and its total loss function is:
[0024]
[0025] in, This represents the total loss of the model. Represents classification loss. α represents the distribution constraint loss, and α represents the weight parameter for balancing the classification objective and the distribution alignment objective.
[0026] Furthermore, in this embodiment of the application, the step of constructing gradient modulation coefficients based on the contribution coefficient ratio includes:
[0027] The gradient modulation coefficient is:
[0028]
[0029] in, This represents the gradient modulation coefficient of mode m in the j-th training batch. β represents the contribution coefficient ratio of mode m in the j-th training batch, and β represents the modulation intensity hyperparameter. When it is greater than 1, the gradient update magnitude of this mode is reduced; when it is less than or equal to 1, the gradient update magnitude of this mode remains unchanged.
[0030] On the other hand, embodiments of this application also provide a human epidermal growth factor receptor-2 status prediction system, comprising:
[0031] The first processing unit is used to acquire full-view digital pathological slides, magnetic resonance imaging, and the first state label of human epidermal growth factor receptor-2 in breast cancer patients.
[0032] The second processing unit is used to obtain the first modal feature representation corresponding to the full-view digital pathological slice and the second modal feature representation corresponding to the magnetic resonance imaging from the full-view digital pathological slice and the magnetic resonance imaging input modal branch network.
[0033] The third processing unit is configured to map the first modal feature representation and the second modal feature representation to probability distribution representations in a shared latent space, and align the first modal feature representation and the second modal feature representation based on the probability distribution representations, to obtain the aligned first modal feature representation and the second modal feature representation in the latent space.
[0034] The fourth processing unit is used to supervise the training of the state prediction model based on the first state label, enhance and fuse the aligned first modal feature representation and the second modal feature representation and input them into the shared classifier to obtain the fused state prediction result and the state prediction results corresponding to the first modal feature representation and the second modal feature representation respectively.
[0035] The fifth processing unit is used to dynamically modulate the parameters of the state prediction model based on the state prediction results of the first human epidermal growth factor receptor-2 and the state prediction results of the second human epidermal growth factor receptor-2 to obtain the trained first state prediction model.
[0036] The sixth processing unit is used to input the full-view digital pathological slices and magnetic resonance imaging of the patient to be tested into the trained first-state prediction model to obtain the state of human epidermal growth factor receptor-2.
[0037] On the other hand, this application also provides a human epidermal growth factor receptor-2 status prediction device, comprising:
[0038] At least one processor;
[0039] At least one memory for storing at least one program;
[0040] When the at least one program is executed by the at least one processor, the at least one processor implements a method for predicting the state of human epidermal growth factor receptor-2 as described in any one of the inventions.
[0041] In addition, this application also provides a computer-readable storage medium storing processor-executable instructions, which, when executed by a processor, are used to perform a human epidermal growth factor receptor-2 state prediction method as described in any of the preceding claims.
[0042] The advantages and beneficial effects of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application:
[0043] This application can obtain full-view digital pathological slides, magnetic resonance imaging (MRI) images, and first state labels of human epidermal growth factor receptor-2 (HGF-2) from breast cancer patients; input the full-view digital pathological slides and MRI images into a modal branching network to obtain the first modal feature representation corresponding to the full-view digital pathological slides and the second modal feature representation corresponding to the MRI images; map the first and second modal feature representations to probability distribution representations in a shared latent space, and align the first and second modal feature representations based on the probability distribution representations to obtain aligned first and second modal feature representations in the latent space; calculate the classification loss based on the first state labels and supervise the training of the model; enhance and fuse the aligned first and second modal feature representations and input them into a shared classifier to obtain the first HGF-2 state prediction result corresponding to the first modal feature representation and the second modal feature representation. The first and second human epidermal growth factor receptor-2 (HGF-2) state prediction results are represented by the first and second HGF-2 state prediction results. Based on these results, the parameters of the state prediction model are dynamically modulated to obtain a trained first state prediction model. The full-view digital pathological slides and magnetic resonance imaging (MRI) images of the patient are input into the trained first state prediction model to obtain the HGF-2 state. This application can achieve more consistent semantic representations between the full-view digital pathological slides and MRI images in a shared latent space through distribution representation alignment, reducing the domain gap between heterogeneous modalities and thus more accurately characterizing tumor biological features related to the HGF-2 state. Furthermore, dynamic gradient modulation alleviates the modality imbalance problem in multimodal fusion, enhancing the model's ability to learn information from non-dominant modalities, thereby improving the accuracy and stability of HGF-2 state prediction. Attached Figure Description
[0044] Figure 1 This is a schematic diagram illustrating the steps of a method for predicting the state of human epidermal growth factor receptor-2 in a specific embodiment of the present invention;
[0045] Figure 2 This is a flowchart illustrating a method for predicting the state of human epidermal growth factor receptor-2 in a specific embodiment of the present invention.
[0046] Figure 3 This is a schematic diagram of the structure of the human epidermal growth factor receptor-2 state prediction system in another specific embodiment of the present invention;
[0047] Figure 4 This is a schematic diagram of the structure of a human epidermal growth factor receptor-2 state prediction device in a specific embodiment of the present invention;
[0048] Figure 5 This is a graph showing the change in the modal contribution ratio of full-view digital pathology slides with and without dynamic gradient modulation in a specific embodiment of the present invention as a function of the number of training rounds.
[0049] Figure 6 This is a schematic diagram illustrating the classification performance of different single-modal and multi-modal fusion methods on internal validation sets and external test sets in a specific embodiment of the present invention;
[0050] Figure 7 This is the ablation experiment result of each component module in a specific embodiment of the present invention. Detailed Implementation
[0051] The following detailed description of the embodiments of the present invention, with reference to the accompanying drawings, illustrates the principles and processes of the human epidermal growth factor receptor-2 state prediction method and related equipment in the embodiments of the present invention.
[0052] Reference Figure 1 This application discloses a method for predicting the state of human epidermal growth factor receptor-2, comprising the following steps:
[0053] S101. Obtain full-view digital pathological sections, magnetic resonance imaging, and first-state labels of human epidermal growth factor receptor-2 in breast cancer patients.
[0054] S102. Input the full-view digital pathological slides and magnetic resonance imaging into the modal branch network to obtain the first modal feature representation corresponding to the full-view digital pathological slides and the second modal feature representation corresponding to the magnetic resonance imaging.
[0055] S103. Map the first modal feature representation and the second modal feature representation to probability distribution representations in the shared latent space, and align the first modal feature representation and the second modal feature representation based on the probability distribution representations to obtain the aligned first modal feature representation and second modal feature representation in the latent space.
[0056] S104. Enhance and fuse the aligned first modality feature representation and second modality feature representation and input them into the shared classifier to obtain the first human epidermal growth factor receptor-2 state prediction result corresponding to the first modality feature representation and the second human epidermal growth factor receptor-2 state prediction result corresponding to the second modality feature representation.
[0057] S105. Based on the state prediction results of the first human epidermal growth factor receptor-2 and the state prediction results of the second human epidermal growth factor receptor-2, the parameters of the state prediction model are dynamically modulated to obtain the trained first state prediction model.
[0058] S106: Input the whole-slide digital pathological section and magnetic resonance imaging of the patient to be tested into the trained first state prediction model to obtain the state of human epidermal growth factor receptor 2.
[0059] Further, in some feasible embodiments of the present application, mapping the first modality feature representation and the second modality feature representation respectively into probability distribution representations in a shared latent space, and aligning the first modality feature representation and the second modality feature representation based on the probability distribution representations to obtain the aligned first modality feature representation and second modality feature representation in the latent space specifically includes:
[0060] S201: Map the first modality feature representation and the second modality feature representation into a mean vector and a standard deviation vector respectively.
[0061] S202: Convert the mean vector and the standard deviation vector into a Gaussian distribution.
[0062] S203: Align the first modality feature representation and the second modality feature representation based on the distribution constraint loss and the Gaussian distribution representation to obtain the aligned first modality feature representation and second modality feature representation in the latent space.
[0063] Further, in some feasible embodiments of the present application, the distribution constraint loss is:
[0064]
[0065] Wherein, represents the distribution constraint loss; and respectively represent the probability distribution representations of the magnetic resonance imaging modality and the whole-slide digital pathological section modality of the i-th sample in the shared latent space; KL(·||·) represents Kullback-Leibler divergence; λ represents the regularization weight; m∈{WSI, MRI} represents the modality index; represents a multivariate standard normal distribution with a mean of 0 and a covariance matrix of identity matrix I.
[0066] Further, in some feasible embodiments of the present application, dynamically modulating the parameters of the state prediction model based on the first human epidermal growth factor receptor 2 state prediction result and the second human epidermal growth factor receptor 2 state prediction result to obtain the trained first state prediction model specifically includes:
[0067] S301: Calculate the contribution coefficient ratio of the first human epidermal growth factor receptor 2 state prediction result and the second human epidermal growth factor receptor 2 state prediction result.
[0068] S302: Construct a gradient modulation coefficient based on the contribution coefficient ratio.
[0069] S303. Determine the total model loss based on classification loss and distribution constraint loss.
[0070] S304. Based on the total loss of the model, the parameters of the state prediction model are dynamically modulated to obtain the trained first state prediction model.
[0071] Furthermore, in some feasible embodiments of this application, the contribution ratio of the first human epidermal growth factor receptor-2 state prediction result and the second human epidermal growth factor receptor-2 state prediction result is calculated, specifically including:
[0072] S401. Determine the sum of all first-person epidermal growth factor receptor-2 state prediction results and the sum of all second-person epidermal growth factor receptor-2 state prediction results in a training batch.
[0073] S402. The ratio of the sum of all first-person epidermal growth factor receptor-2 status prediction results to the sum of all second-person epidermal growth factor receptor-2 status prediction results is used as the contribution coefficient ratio.
[0074] Specifically, for the i-th sample of the method in this application, let the modality indices of the full-view digital pathological slide modality and the magnetic resonance image modality be denoted as:
[0075]
[0076] Modal feature representations are obtained after each encoder:
[0077]
[0078] To reduce the domain gap between two heterogeneous modes, modal features are... Mapped to the mean vector in the shared latent space respectively and standard deviation vector :
[0079]
[0080] in, and Let these represent the mean mapping function and the standard deviation mapping function corresponding to mode m, respectively. This represents the mean vector of the i-th sample in mode m. Let represent the standard deviation vector of the i-th sample in mode m. Further, the features of mode m are represented as a Gaussian distribution in the latent space:
[0081]
[0082] in, represents the latent probability distribution of the i-th sample in modality m, and Z represents a random variable in the latent space, (·) represents the Gaussian distribution, and diag(·) represents the diagonal matrix constructed from vectors. Based on said Gaussian distribution representation, the original deterministic features are converted into a probability distribution representation to model the semantic consistency between the two modalities. To facilitate distribution alignment of the whole slide image modality and the magnetic resonance imaging modality in the shared latent space, this embodiment adopts the following distribution constraint loss:
[0083]
[0084] wherein, represents the distribution constraint loss; and respectively represent the probability distribution representations of the magnetic resonance imaging modality and the whole slide image modality of the i-th sample in the shared latent space; KL(·||·) represents Kullback-Leibler divergence; λ represents the regularization weight; m∈{WSI, MRI} represents the modality index; represents the multivariate standard normal distribution with a mean of 0 and a covariance matrix being the identity matrix I.
[0085] The aforementioned first term is used to implement cross-modal distribution alignment, and the second term is used to constrain the latent distribution near the standard Gaussian prior to keep the structure of the latent space stable.
[0086] On this basis, the mean vector of each modality is further enhanced to obtain a modality representation, and the whole slide image modality representation and the magnetic resonance imaging modality representation are concatenated and then input into a shared classifier, which outputs the final state prediction result of human epidermal growth factor receptor-2. At the same time, the classifier also respectively outputs the independent prediction results and of each modality, which are used to evaluate the contribution degree of each modality to the current batch prediction task.
[0087] To alleviate the modality imbalance problem during training, for the j-th training batch B j , the contribution ratio of the whole slide image modality relative to the magnetic resonance imaging modality is defined as:
[0088]
[0089] and a gradient modulation coefficient is constructed:
[0090]
[0091] wherein, represents the gradient modulation coefficient of modality m in the j-th training batch, Let represent the contribution ratio of mode m in the j-th training batch, and β represent the modulation intensity hyperparameter. When β is greater than 1, the gradient update magnitude of that mode is reduced; when β is less than or equal to 1, the gradient update magnitude of that mode remains unchanged. When a mode's contribution is too high, its gradient update magnitude is reduced; when a mode's contribution is low, its gradient update magnitude is maintained, thereby alleviating the problem of dominant modes dominating training. Accordingly, the mode parameters are updated as follows:
[0092]
[0093] in, This represents the learnable parameters of the branch corresponding to mode m. Indicates the learning rate. This represents the gradient of the total loss with respect to . In this embodiment, the entire model employs joint optimization using classification loss and distribution constraint loss, and its total loss function is:
[0094]
[0095] in, This represents the total loss of the model. Represents classification loss. α represents the distribution constraint loss, and α represents the weight parameter for balancing the classification objective and the distribution alignment objective.
[0096] To verify the performance of the proposed method in the human epidermal growth factor receptor-2 (HGF-2) state prediction task, experiments were conducted on data from 866 breast cancer patients from three independent medical centers. Each patient's data included full-field digital pathological sections obtained from biopsy, dynamic contrast-enhanced magnetic resonance imaging (MRI) images in the third phase, and corresponding HGF-2 state labels. The HGF-2 state labels were binary labels, including positive and negative labels, which were encoded as 1 and 0 respectively during model training. Data from two cohorts were used for training and internal validation, while data from another independent cohort was used for external testing. Model performance was evaluated using AUC (Area Under the Curve), ACC (Accuracy), F1 (F1-score), SEN (Sensitivity), and SPE (Specificity).
[0097] In terms of experimental setup, this embodiment uses the Adam optimizer, with an initial learning rate η set to 1×, a batch size of 2, a maximum number of training epochs of 100, and relevant hyperparameters α, β, and λ set to 0.2, 0.1, and 2×, respectively. The classification performance of this application is as follows: Figure 6As shown, Figure 6 This section presents the classification performance of different unimodal and multimodal fusion methods on internal validation and external test sets. From... Figure 6 Experimental results show that the proposed method achieves AUCs of 81.7% and 86.1% on the internal validation set and external test set, respectively, which are superior to many existing single-modal and multi-modal fusion methods. This indicates that the proposed method can effectively reduce the gap between heterogeneous modal domains and alleviate the modal imbalance problem, thereby improving the accuracy and generalization performance of human epidermal growth factor receptor-2 state prediction.
[0098] Furthermore, to verify the contribution of each component module in this embodiment to the state prediction performance of human epidermal growth factor receptor-2, ablation experiments were conducted on the distribution characterization alignment module and the dynamic gradient modulation module, respectively. The experimental results are referenced from... Figure 7 , Figure 7 These are the ablation experiment results for each component module in this embodiment. Figure 7 In the ablation experiments, the performance changes of the models using only point representation classification, adding distribution representation alignment to point representation, and further adding dynamic gradient modulation to distribution representation alignment were compared. Experimental results show that, compared to the point representation classification model, adding distribution representation alignment improved the AUC of 1.1% and 3.4% on the internal validation set and external test set, respectively, indicating that this module can effectively reduce the heterogeneous modal differences between full-view digital pathology slides and magnetic resonance images. Further adding dynamic gradient modulation further improved the AUC of 1.2% and 3.1% on the internal validation set and external test set, indicating that this module can effectively alleviate the modal imbalance problem in the multimodal fusion process.
[0099] In addition, refer to Figure 3 ,and Figure 1Corresponding to the method described above, embodiments of this application also provide a human epidermal growth factor receptor-2 (HGF-2) state prediction system. This system may include a first processing unit 1001, a second processing unit 1002, a third processing unit 1003, a fourth processing unit 1004, a fifth processing unit 1005, and a sixth processing unit 1006. The first processing unit 1001 is used to acquire full-view digital pathological slides, magnetic resonance imaging (MRI), and a first state label of HGF-2 from breast cancer patients. The second processing unit 1002 is used to input the full-view digital pathological slides and MRI into a modal branching network to obtain a first modal feature representation corresponding to the full-view digital pathological slides and a second modal feature representation corresponding to the MRI. The third processing unit 1003 is used to map the first modal feature representation and the second modal feature representation to probability distribution representations in a shared latent space, and align the first modal feature representation and the second modal feature representation based on the probability distribution representations to obtain aligned first modal feature representations and second modal feature representations in the latent space. The fourth processing unit 1004 enhances and fuses the aligned first modality feature representation and the second modality feature representation, and inputs them into a shared classifier to obtain the first human epidermal growth factor receptor-2 state prediction result corresponding to the first modality feature representation and the second human epidermal growth factor receptor-2 state prediction result corresponding to the second modality feature representation. The fifth processing unit 1005 dynamically modulates the parameters of the state prediction model based on the first and second human epidermal growth factor receptor-2 state prediction results to obtain a trained first state prediction model. The sixth processing unit 1006 inputs the full-field digital pathological slides and magnetic resonance imaging of the patient to be tested into the trained first state prediction model to obtain the state of human epidermal growth factor receptor-2.
[0100] It should be noted that the first processing unit can be any integrated circuit unit or microprocessor unit obtained by integrating a chip with processing functions and its peripheral circuits using existing integration technology. The first processing unit and the second processing unit can also be any integrated circuit module or microprocessor module obtained by integrating a chip with processing functions and its peripheral circuits using existing integration technology. Furthermore, the first processing unit and the second processing unit may include one or more memories.
[0101] It should be noted that the content of the above-mentioned human epidermal growth factor receptor-2 state prediction method embodiments is applicable to my own epidermal growth factor receptor-2 state prediction system embodiments. The specific functions implemented by my own epidermal growth factor receptor-2 state prediction system embodiments are the same as those of the above-mentioned human epidermal growth factor receptor-2 state prediction method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-mentioned human epidermal growth factor receptor-2 state prediction method embodiments.
[0102] and Figure 1 Corresponding to the method described herein, embodiments of this application also provide a human epidermal growth factor receptor-2 state prediction device, the specific structure of which can be referred to Figure 4 ,include:
[0103] At least one processor 1011;
[0104] At least one memory 1012 is used to store at least one program;
[0105] When the at least one program is executed by the at least one processor, the at least one processor implements the human epidermal growth factor receptor-2 state prediction method.
[0106] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0107] and Figure 1 Corresponding to the method described above, embodiments of this application also provide a computer-readable storage medium storing processor-executable instructions, which, when executed by a processor, are used to perform the human epidermal growth factor receptor-2 state prediction method.
[0108] The contents of the above-described human epidermal growth factor receptor-2 state prediction method embodiments are all applicable to this storage medium embodiment. The specific functions implemented by this storage medium embodiment are the same as those of the above-described human epidermal growth factor receptor-2 state prediction method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-described human epidermal growth factor receptor-2 state prediction method embodiments.
[0109] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this application are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0110] Furthermore, although this application is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding this application. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional technology for an engineer. Therefore, those skilled in the art can implement the application set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of this application, which is determined by the full scope of the appended claims and their equivalents.
[0111] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable programs for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, a program execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can retrieve and execute a program from or in conjunction with such a program execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit a program for use by or in conjunction with a program execution system, apparatus, or device.
[0113] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory, read-only memory, erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0114] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0115] In the foregoing description of this specification, the references to terms such as "one embodiment," "another embodiment," or "some embodiments," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0116] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
[0117] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for predicting the state of human epidermal growth factor receptor-2, characterized in that, Includes the following steps: Obtain full-view digital pathological slides, magnetic resonance imaging, and first-state labels of human epidermal growth factor receptor-2 in breast cancer patients; The full-view digital pathological slice and the magnetic resonance imaging input modal branch network are used to obtain the first modal feature representation corresponding to the full-view digital pathological slice and the second modal feature representation corresponding to the magnetic resonance imaging. The first modal feature representation and the second modal feature representation are respectively mapped to probability distribution representations in a shared latent space, and the first modal feature representation and the second modal feature representation are aligned based on the probability distribution representations to obtain the first modal feature representation and the second modal feature representation aligned in the latent space; The state prediction model is trained under supervision based on the first state label; the aligned first modality feature representation and the second modality feature representation are enhanced and fused and input into a shared classifier to obtain the fused state prediction result, as well as the state prediction results corresponding to the first modality feature representation and the second modality feature representation respectively. Based on the first human epidermal growth factor receptor-2 state prediction results and the second human epidermal growth factor receptor-2 state prediction results, the parameters of the state prediction model are dynamically modulated to obtain the trained first state prediction model. By inputting the full-view digital pathological slides and magnetic resonance imaging of the patient under test into the trained first-state prediction model, the state of human epidermal growth factor receptor-2 is obtained.
2. The method for predicting the state of human epidermal growth factor receptor-2 according to claim 1, characterized in that, The step of mapping the first modal feature representation and the second modal feature representation to probability distribution representations in a shared latent space, and aligning the first modal feature representation and the second modal feature representation based on the probability distribution representations to obtain aligned first modal feature representations and second modal feature representations in the latent space, specifically includes: The first modal feature representation and the second modal feature representation are mapped to mean vector and standard deviation vector, respectively; Transform the mean vector and the standard deviation vector into Gaussian distribution representations; Based on the distribution constraint loss and the Gaussian distribution representation, the first modal feature representation and the second modal feature representation are aligned to obtain the aligned first modal feature representation and the second modal feature representation in the latent space.
3. The method for predicting the state of human epidermal growth factor receptor-2 according to claim 2, characterized in that, The distribution constraint loss is: wherein, represents the distribution constraint loss; and respectively represent the probability distribution representations of the magnetic resonance imaging modality and the whole slide imaging modality of the i-th sample in the shared latent space; KL(·丨丨·) represents the Kullback-Leibler divergence; λ represents the regularization weight; m∈{WSI, MRI} represents the modality index; represents the multivariate standard normal distribution with mean 0 and covariance matrix being the identity matrix I.
4. The method for predicting the state of human epidermal growth factor receptor-2 according to claim 1, characterized in that, The step of dynamically modulating the parameters of the state prediction model based on the first and second human epidermal growth factor receptor-2 state prediction results to obtain a trained first state prediction model specifically includes: Calculate the contribution ratio of the first person's epidermal growth factor receptor-2 status prediction result and the second person's epidermal growth factor receptor-2 status prediction result; Based on the aforementioned contribution coefficient ratio, gradient modulation coefficients are constructed; The total model loss is determined based on classification loss and distribution constraint loss. Based on the total model loss, the parameters of the state prediction model are dynamically modulated to obtain the trained first state prediction model.
5. The method for predicting the state of human epidermal growth factor receptor-2 according to claim 4, characterized in that, The calculation of the contribution coefficient ratio between the first human epidermal growth factor receptor-2 state prediction result and the second human epidermal growth factor receptor-2 state prediction result specifically includes: Determine the sum of all first-person epidermal growth factor receptor-2 state predictions and the sum of all second-person epidermal growth factor receptor-2 state predictions in a training batch; The ratio of the sum of all first-person epidermal growth factor receptor-2 status prediction results to the sum of all second-person epidermal growth factor receptor-2 status prediction results is used as the contribution coefficient ratio.
6. The method for predicting the state of human epidermal growth factor receptor-2 according to claim 4, characterized in that, The model loss value is jointly optimized using classification loss and distribution constraint loss, and its total loss function is: in, This represents the total loss of the model. Represents classification loss. α represents the distribution constraint loss, and α represents the weight parameter for balancing the classification objective and the distribution alignment objective.
7. The method for predicting the state of human epidermal growth factor receptor-2 according to claim 4, characterized in that, The construction of gradient modulation coefficients based on the contribution coefficient ratio includes: The gradient modulation coefficient is: in, This represents the gradient modulation coefficient of mode m in the j-th training batch. β represents the contribution coefficient ratio of mode m in the j-th training batch, and β represents the modulation intensity hyperparameter. When it is greater than 1, the gradient update magnitude of this mode is reduced; when it is less than or equal to 1, the gradient update magnitude of this mode remains unchanged.
8. A human epidermal growth factor receptor-2 status prediction system, characterized in that, include: The first processing unit is used to acquire full-view digital pathological slides, magnetic resonance imaging, and the first state label of human epidermal growth factor receptor-2 in breast cancer patients. The second processing unit is used to obtain the first modal feature representation corresponding to the full-view digital pathological slice and the second modal feature representation corresponding to the magnetic resonance imaging from the full-view digital pathological slice and the magnetic resonance imaging input modal branch network. The third processing unit is configured to map the first modal feature representation and the second modal feature representation to probability distribution representations in a shared latent space, and align the first modal feature representation and the second modal feature representation based on the probability distribution representations, to obtain the aligned first modal feature representation and the second modal feature representation in the latent space. The fourth processing unit performs supervised training on the state prediction model based on the first state label, enhances and fuses the aligned first modal feature representation and the second modal feature representation, and inputs them into the shared classifier to obtain the fused state prediction result, as well as the state prediction results corresponding to the first modal feature representation and the second modal feature representation respectively. The fifth processing unit is used to dynamically modulate the parameters of the state prediction model based on the state prediction results of the first human epidermal growth factor receptor-2 and the state prediction results of the second human epidermal growth factor receptor-2 to obtain the trained first state prediction model. The sixth processing unit is used to input the full-view digital pathological slices and magnetic resonance imaging of the patient to be tested into the trained first-state prediction model to obtain the state of human epidermal growth factor receptor-2.
9. A device for predicting the state of human epidermal growth factor receptor-2, characterized in that... include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the human epidermal growth factor receptor-2 state prediction method as described in any one of claims 1-7.
10. A computer-readable storage medium storing processor-executable instructions, characterized in that, The processor-executable instructions, when executed by the processor, are used to perform a method for predicting the state of human epidermal growth factor receptor-2 as described in any one of claims 1-7.