AMD curative effect prediction method, system and equipment based on feature fusion and medium
By combining 3DResNet, radiomics, and demographic features, and utilizing a feature fusion method with pre-interactive LSTM and Transformer modules, the problem of insufficient accuracy in predicting the efficacy of AMD in existing technologies is solved, and efficient prediction of the therapeutic effect of anti-vascular endothelial growth factor is achieved.
Patent Information
- Application Number
- CN202511419470.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing methods for predicting the efficacy of AMD treatment are not accurate enough under small data conditions and are difficult to effectively integrate radiomics features, deep learning features and demographic features, resulting in inaccurate prediction of the efficacy of anti-vascular endothelial growth factor therapy.
We employ a three-dimensional residual network (3DResNet) and a radiomics feature extraction network, combined with a pre-interactive long short-term memory (LSTM) network and a Transformer module, to integrate deep learning, radiomics, and demographic features. Through a time-series analysis network, we obtain the relationship between image features of patients at different time points and achieve efficacy prediction.
It improves the accuracy and reliability of predicting the treatment efficacy of AMD, enhances the automated processing capability of optical coherence tomography images, and provides technical support for the research and clinical practice of age-related macular degeneration.
Smart Images

Figure CN120932807A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of efficacy prediction technology, and specifically to a method, system, device, and medium for predicting the efficacy of AMD based on feature fusion. Background Technology
[0002] Age-related macular degeneration (AMD) is a common retinal disease, mainly divided into two types: dry (non-neovascular) and wet (neovascular). Wet AMD accounts for about 10% to 20% of all AMD cases, but the visual impairment caused by wet AMD is more severe than that caused by dry AMD. The clinical features of wet AMD are rapid appearance of distorted vision and loss of central vision within weeks to months, which can lead to blindness in severe cases. The main feature of wet AMD is the formation of choroidal neovascularization (CNV).
[0003] Choroidal neovascularization refers to the abnormal new blood vessels growing beneath the retina, also known as subretinal neovascularization. These are proliferating vessels originating from choroidal capillaries, expanding through tears in Bruch's membrane. They proliferate and form between Bruch's membrane and the retinal pigment epithelium, between the neuroretina and the retinal pigment epithelium, or between the retinal pigment epithelium and the choroid. They are most commonly found in the macula, thus severely impairing central vision. In optical coherence tomography (OCT), choroidal neovascularization typically appears as hyperreflective lesions located beneath the retinal pigment epithelium or neuroepithelium. Early detection and timely treatment of choroidal neovascularization are of great clinical significance.
[0004] In recent years, anti-vascular endothelial growth factor (VEGF) therapy has become one of the main methods for treating AMD. Multiple clinical studies have confirmed that anti-VEGF therapy can improve vision and ocular lesions in patients with wet AMD, slow disease progression, and improve patients' quality of life. VEGF is a protein that promotes angiogenesis and leakage. By inhibiting the action of VEGF, leakage and exudation in ocular blood vessels can be reduced, thereby alleviating the symptoms of AMD patients.
[0005] In predicting the overall efficacy of anti-vascular endothelial growth factor therapy in AMD patients, no outstanding deep learning network has yet achieved very good predictive results with a small amount of data. In addition, existing disease prediction methods usually use radiomics features or deep learning features alone, and few methods can effectively integrate radiomics features, deep learning features and other clinical information. Summary of the Invention
[0006] In view of the above-mentioned problems, the present invention is proposed.
[0007] Therefore, the technical problem solved by this invention is: how to obtain high-quality image features by using a feature extraction network based on a three-dimensional residual network (3DResNet) and radiomics, effectively fuse deep learning, radiomics features, and demographic features obtained by feature processing and selection networks, and deeply acquire and integrate the image feature relationships of patients at different time points through a pre-interactive long short-term memory (LSTM) network and a transformer subtraction module, thereby enhancing the accuracy of classification and prediction and achieving accurate classification and prediction of the effect of anti-vascular endothelial growth factor therapy on patients.
[0008] To address the aforementioned technical problems, this invention provides the following technical solution: a method for predicting the efficacy of AMD treatment based on feature fusion, comprising, The process involves acquiring optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration (AMD); determining deep learning features and radiomics features based on the OCT image data; selecting input features based on the deep learning features, radiomics features, and demographic characteristics; calculating hidden states using the input features and a temporal analysis network; and analyzing feature change trends based on the hidden states to determine the therapeutic efficacy prediction results for AMD.
[0009] As a preferred embodiment of the AMD efficacy prediction method based on feature fusion described in this invention, the method for determining deep learning features and radiomics features based on the optical coherence tomography (OCT) image data includes: processing the OCT image data and a preset lesion mask to obtain deep learning features; and extracting features from a preset feature domain of the OCT image data to obtain radiomics features.
[0010] As a preferred embodiment of the AMD efficacy prediction method based on feature fusion described in this invention, the following steps are performed: feature selection is performed based on the deep learning features, the radiomics features, and the demographic features to obtain input features. This includes: processing the deep learning features and the radiomics features using a feature correction module to obtain deep learning features and radiomics features processed by the feature correction module; concatenating the deep learning features, radiomics features, and demographic features to obtain total features, and processing the total features using the feature correction module and the Softmax function to obtain feature selection weights; concatenating the deep learning features and radiomics features processed by the feature correction module, and calculating the input features in conjunction with the feature selection weights.
[0011] As a preferred embodiment of the AMD efficacy prediction method based on feature fusion described in this invention, the method for determining the hidden state by combining the input features with a temporal analysis network includes: inputting the input features into the temporal analysis network; and processing the input features using the temporal analysis network to obtain the hidden state and memory unit state at the current time.
[0012] As a preferred embodiment of the AMD efficacy prediction method based on feature fusion described in this invention, the following steps are performed: processing the input features using the temporal analysis network to obtain the hidden state and memory unit state at the current time step includes: performing multiple rounds of interactive operations between the input features at the current time step and the hidden state at the previous time step to obtain interactive features and interactive states; and calculating the hidden state and memory unit state at the current time step based on the interactive features and interactive states, combined with the memory unit state at the previous time step.
[0013] As a preferred embodiment of the AMD efficacy prediction method based on feature fusion described in this invention, the efficacy prediction result for age-related macular degeneration is determined by performing feature change trend analysis based on the hidden states, including: obtaining a first hidden state at the first time point and a second hidden state at the last time point from the hidden states; inputting the first hidden state and the second hidden state into the Transformer subtraction module to obtain the efficacy prediction result for age-related macular degeneration.
[0014] As a preferred embodiment of the AMD efficacy prediction method based on feature fusion described in this invention, the method involves: inputting the first hidden state and the second hidden state into a Transformer subtraction module to obtain the efficacy prediction result for age-related macular degeneration, including: processing the hidden state through a linear layer and a softmax function to obtain multiple attention feature values; concatenating and linearly transforming the multiple attention feature values to obtain the output result of multi-head attention; and determining the efficacy prediction result based on the output result of multi-head attention of the first hidden state and the output result of multi-head attention of the second hidden state.
[0015] This invention provides a predictive system for the treatment efficacy of AMD based on feature fusion.
[0016] To address the aforementioned technical problems, this invention provides the following technical solution: an AMD efficacy prediction system based on feature fusion, comprising: a patient data module for acquiring optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration (AMD); a feature extraction module for determining deep learning features and radiomics features based on the OCT image data; a feature fusion module for performing feature selection based on deep learning features, radiomics features, and demographic characteristics to obtain input features; a temporal correlation feature module for determining hidden states by combining the input features with a temporal analysis network; and an efficacy prediction module for analyzing feature change trends based on the hidden states to determine the efficacy prediction result for AMD.
[0017] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the feature fusion-based AMD efficacy prediction method.
[0018] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the feature fusion-based AMD efficacy prediction method.
[0019] The beneficial effects of this invention are as follows: Addressing the challenges of limited data and significant sample bias in predicting the efficacy of anti-vascular endothelial growth factor (VEGF) therapy for patients with age-related macular degeneration (AMD), this invention proposes a multi-feature fusion long short-term memory network to predict the efficacy of VEGF therapy in AMD patients. By designing a feature extraction network based on 3DResNet and radiomics, high-quality image features are obtained. Furthermore, considering the mutual and individual characteristics of these features, and incorporating demographic features, a feature processing and selection network is proposed to effectively fuse various features. To address the difficulty in obtaining correlations between image features at different time points, this invention integrates image information from various time points using a pre-interactive LSTM network and a Transformer subtraction module, achieving accurate prediction of treatment efficacy. This not only improves the automated processing capabilities of optical coherence tomography (OCT) images but also provides technical support for AMD research and clinical practice. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 The flowchart illustrates the overall steps of a feature fusion-based method for predicting the efficacy of AMD treatment, as provided in one embodiment of the present invention.
[0022] Figure 2 This is a detailed flowchart illustrating a feature fusion-based method for predicting the efficacy of AMD treatment, provided as an embodiment of the present invention.
[0023] Figure 3 This is a diagram of a multi-feature fusion long short-term memory network structure provided in one embodiment of the present invention.
[0024] Figure 4 This is a schematic diagram of an image feature extraction module provided in one embodiment of the present invention.
[0025] Figure 5 This is a structural diagram of a feature processing and selection network provided in one embodiment of the present invention.
[0026] Figure 6 This is a structural diagram of a pre-interactive LSTM network provided in one embodiment of the present invention.
[0027] Figure 7 This is a structural diagram of the Transformer subtraction module provided in one embodiment of the present invention. Detailed Implementation
[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0029] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0030] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0031] Example 1, referring to Figure 1 As one embodiment of the present invention, this embodiment provides a method for predicting the efficacy of AMD based on feature fusion, including: S100: Acquire optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration.
[0032] S200: Determine deep learning features and radiomics features based on optical coherence tomography image data.
[0033] S300: Feature selection is performed based on deep learning features, radiomics features, and demographic features to obtain input features.
[0034] S400: By inputting features and combining them with a temporal analysis network, the hidden state is determined.
[0035] S500: Based on the hidden state, perform characteristic change trend analysis to determine the predictive results of treatment efficacy for age-related macular degeneration.
[0036] It should be noted that age-related macular degeneration (AMD) is a common retinal disease. Although wet AMD accounts for a smaller proportion, the visual impairment caused by wet AMD is more severe than that caused by dry AMD. The clinical features of wet AMD are rapid visual distortion and central vision loss that occur within weeks to months, and in severe cases, it can lead to blindness. Currently, anti-vascular endothelial growth factor (VEGF) therapy has become the main treatment for AMD. However, in predicting the overall efficacy of AMD patients after VEGF therapy, there is currently no outstanding deep learning network that can achieve very good predictive results with a small amount of data. At the same time, existing disease prediction methods usually use radiomics features or deep learning features alone, and few methods can effectively integrate radiomics features, deep learning features, and other clinical information. The changes in the choroidal neovascularization lesion area in optical coherence tomography (OCT) images are particularly important for predicting the efficacy of VEGF therapy in AMD patients. Existing methods face challenges in the extraction and fusion of features from medical image time series data, making it difficult to accurately obtain the correlation between image features at different time points of the patient, resulting in low prediction accuracy.
[0037] Therefore, to address the aforementioned problems of difficulty in feature fusion, insufficient extraction of temporal information, and low accuracy in efficacy prediction, the S100-S500 steps are used to acquire optical coherence tomography (OCT) image data and demographic features of patients at multiple time points. Deep learning features, radiomics features, and demographic features are extracted and fused to construct a temporal analysis network to obtain correlation features between different time points of patients, thereby achieving accurate prediction of the efficacy of anti-vascular endothelial growth factor (VEGF) therapy in AMD patients.
[0038] Example 2, refer to Figures 1-7 As one embodiment of the present invention, based on the previous embodiment, a method for predicting the efficacy of AMD based on feature fusion is provided, including: like Figure 2 The diagram shows a detailed flowchart of a method for predicting the efficacy of AMD based on feature fusion. In this embodiment of the invention, step S100 involves acquiring optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration, including acquiring OCT images and demographic characteristics of wet AMD patients during anti-vascular endothelial growth factor (VEGF) treatment.
[0039] Specifically, the optical coherence tomography (OCT) images were all obtained using a Heidelberg fundus tomography scanner. The resolution of the OCT images was 512×496×19. Before being input into the network, the images were resampled and processed into images of size 512×512×19.
[0040] Furthermore, obtaining demographic characteristics includes: The patient's medical record data is obtained, and features are extracted from the medical record data to obtain the patient's demographic characteristics, including gender, age, and left / right eye labels.
[0041] It should be noted that this invention provides a complete data foundation for subsequent feature extraction and efficacy prediction by acquiring high-quality optical coherence tomography (OCT) image data and standardized demographic characteristics of patients at multiple treatment time points. Through unified image resolution processing, data consistency and comparability are ensured. Compared with incomplete data acquisition methods in existing technologies, this invention solves the problems of unstable data quality and missing temporal information in traditional methods by establishing a standardized multi-time-point data acquisition process. In particular, the high-resolution OCT images obtained through the Heidelberg fundus tomography scanner provide rich details of lesions, and the complete coverage of treatment time points fully reflects the evolution of lesions during treatment. This not only provides high-quality raw data for the extraction of deep learning features and radiomics features but also provides complete time-series information for training the temporal analysis network, effectively ensuring the accuracy and reliability of subsequent efficacy prediction.
[0042] It should be noted that, as Figure 3 The diagram shown illustrates the architecture of the Multi-feature Fusion LSTM Network (MF-LSTM) proposed in this invention. MF-LSTM is an end-to-end deep learning classification and prediction network, comprising an image feature extraction network, a feature processing and selection network, a pre-interactive LSTM network, and a Transformer subtraction module. The images at different time points, i.e., the input images... Where T represents the time node. First, the input is fed into an image feature extraction network to obtain feature information of the optical coherence tomography (OCT) image. Then, it passes through a feature processing and selection network to remove unnecessary features and refine the necessary features. Finally, these features are input into a pre-interactive LSTM network and a Transformer subtraction module to further obtain information on the changes in the patient's images before and after treatment. Finally, the efficacy prediction result is output.
[0043] In this embodiment of the invention, step S200, which determines deep learning features and radiomics features based on optical coherence tomography image data, includes the following steps A1-A2: A1: Deep learning features are obtained by processing optical coherence tomography image data and a preset lesion mask.
[0044] It should be noted that, as Figure 4The diagram shows the image feature extraction module. Using the original input image and the choroidal neovascularization lesion mask and retinal region mask, deep learning features are obtained using 3DResNet, and radiomics features are obtained using radiomics. Then, the deep learning features, radiomics features, and the patient's demographic features (gender, age, and left / right eye labels) are input into the feature processing and selection network module.
[0045] For example, processing optical coherence tomography image data with a preset lesion mask to obtain deep learning features involves using 3DResNet18 as the base network and using the pre-trained weights of the MedicalNet network (trained on 23 medical datasets provided by Tencent YouTu Lab) as the initial weights. The result of multiplying the choroidal neovascularization lesion mask with the original image is used as the network input. The output of the last layer of 3DResNet is passed through a fully connected layer as the extracted deep learning features. The length of the deep learning features is 512.
[0046] A2: Extract features from the preset feature domain of optical coherence tomography image data to obtain radiomics features.
[0047] Specifically, features are extracted from the preset feature domain of optical coherence tomography image data to obtain radiomics features, including the following steps A21-A23: A21: The choroidal neovascularization lesion area and the retinal area are respectively selected as regions of interest, and features are extracted using the original domain, gradient domain, Gaussian-Laplace domain, square domain, and logarithmic domain.
[0048] A22: Remove non-numerical features from the extracted features and perform normalization.
[0049] A23: Extract radiomics features based on different feature domains to obtain radiomics feature vectors.
[0050] For example, in step A23, when extracting radiomics features based on different feature domains, for the original domain, 18 first-order statistical features, 14 shape features, 22 gray-level co-occurrence matrix features, 16 gray-level length matrix features, 16 gray-level interval size matrix features, 5 neighboring gray-level difference matrix features, and 14 gray-level dependence matrix features are extracted, totaling 105 features. For the gradient domain, Gaussian-Laplace domain, square domain, and logarithmic domain, shape features are not calculated, and each domain has 91 features. Therefore, 469 radiomics features can be obtained by calculating the five domains. Among them, 469 radiomics features for the choroidal neovascularization lesion region and 469 radiomics features for the retinal region are obtained respectively from the choroidal neovascularization lesion region and the retinal region according to the above radiomics feature extraction method, resulting in a total of 938 radiomics features.
[0051] It should be noted that this invention achieves comprehensive capture and quantitative expression of key lesion information in optical coherence tomography (OCT) images by employing a pre-trained 3DResNet network combined with a choroidal neovascularization lesion mask for deep learning feature extraction, and a radiomics feature extraction method based on multiple feature domains. The choroidal neovascularization lesion mask guides the deep learning network to focus on key lesion areas, improving the targeting of feature extraction. Compared with single feature extraction methods in existing technologies, this invention solves the problems of insufficient feature information and interpretability in traditional methods by combining deep learning features and radiomics features for dual feature extraction. In particular, through radiomics feature extraction of five different feature domains, it can describe the texture, shape, and statistical characteristics of images from multiple dimensions. Deep learning features can automatically learn high-level abstract representations of images, not only improving the completeness and complementarity of feature expression but also providing high-quality multi-dimensional feature input for subsequent feature fusion, effectively enhancing the feature expression capability for predicting the efficacy of AMD.
[0052] In this embodiment of the invention, step S300 involves feature selection based on deep learning features, radiomics features, and demographic features to obtain input features, including the following steps B1-B3: B1: The deep learning features and radiomics features are processed using the feature correction module to obtain the deep learning features and radiomics features after the feature correction module is applied.
[0053] Specifically, such as Figure 5The diagram shows the structure of the feature processing and selection network. The feature correction module processes the deep learning features and radiomics features to obtain the processed deep learning features and radiomics features. This means that for both radiomics features and deep learning features, a feature correction module is used for the first step of processing. The operational logic of the feature correction module includes the following steps B11-B15: B11: For input feature a, after passing through a fully connected layer, we obtain the first-level feature a1, which can be specifically represented as: ; in, The node weights of the first fully connected layer; This is the bias value for the first fully connected layer.
[0054] B12: The first-level feature a1 is passed through an Exponential Linear Unit Activation (ELU) layer to obtain the second-level feature a2, which can be specifically represented as: ; Furthermore, the exponential linear activation function ELU can be specifically expressed as: ; Wherein, β is a hyperparameter used to control the curve shape of the negative region, and in this invention it is set to 1 based on experience.
[0055] B13: Pass the second-level feature a2 through another fully connected layer to obtain the third-level feature a3, which can be specifically represented as: ; in, These are the node weights of the second fully connected layer; This is the bias amount for the second fully connected layer.
[0056] B14: Gated Linear Units (GLUs) are used to provide additional flexibility and suppress unwanted information in the extracted image features. A GLU can be specifically represented as: ; in, Use the Sigmoid activation function; , These are all weight matrices of GLU. Used to calculate the gating signal. Used for linear transformation paths; , This is the bias matrix of the GLU, used to implement linear transformations; This is an element-wise product.
[0057] B15: Using GLU allows the network to control the contribution of the original input, thus filtering for single feature categories. Simultaneously, to prevent the omission of original information, a residual structure is employed, adding the original input features to the results after passing through the gate unit, followed by layer normalization to obtain the output feature a4, which can be specifically represented as: ; in, For layer normalization.
[0058] B2: The deep learning features, radiomics features, and demographic features are concatenated to obtain the total features. The feature correction module and the Softmax function are then used to process the total features to obtain the feature selection weights.
[0059] Specifically, the total feature is obtained by concatenating deep learning features, radiomics features, and demographic features in order to achieve variable selection. The total feature includes age, gender, and left / right eye information, and the demographic features are restricted to the range of 0-1 through normalization.
[0060] It should be noted that by splicing multimodal features, a preliminary integration of deep learning features, radiomics features, and demographic features was achieved, providing a complete feature foundation for subsequent feature selection. Normalization ensured the consistency of different types of features in terms of numerical range, effectively avoiding the fusion bias problem caused by differences in feature dimensions.
[0061] Specifically, processing the total features using the feature correction module and the Softmax function to obtain the feature selection weights involves inputting the total features into the feature correction module and then passing them through the Softmax function to obtain the feature selection weights v, which can be expressed as follows: ; Where v is the feature selection weight; f is the total features; and FRM is the feature correction module.
[0062] It should be noted that the present invention uses a feature correction module combined with a Softmax function to determine the feature selection weights, thereby achieving the evaluation of the importance of different features. The use of gated linear units can effectively suppress unwanted feature information while retaining key features. The introduction of residual structures prevents the omission of original information and ensures the integrity of information during the feature selection process.
[0063] B3: The deep learning features and radiomics features processed by the feature correction module are concatenated and combined with feature selection weights to obtain the input features.
[0064] Specifically, the deep learning features and radiomics features processed by the feature correction module are concatenated and combined with feature selection weights to obtain the input features, which refer to the radiomics features processed by the feature correction module. and deep learning features After concatenation, each feature is multiplied by its weight to obtain the final features after feature processing and selection network processing. Specifically, it can be expressed as: ; in, and These represent the image omics features and deep learning features processed by the feature correction module, respectively; v represents the feature selection weight; Concat represents the stitching; and t represents the time point number.
[0065] It should be noted that this invention achieves effective integration of deep learning features and radiomics features through weighted fusion based on feature selection weights. Compared with the existing technology that uses a single feature independently, this invention can adjust the weights according to the contribution of different features to the efficacy prediction task, effectively solving the problems of feature fusion difficulties and information redundancy in traditional methods. The preprocessing of the feature correction module ensures feature quality, and the introduction of the weight mechanism enables the model to balance the importance of different modal features. This not only improves the effectiveness and relevance of feature expression, but also provides high-quality fusion feature input for subsequent time series analysis networks, effectively enhancing the feature fusion capability and prediction accuracy of AMD efficacy prediction.
[0066] In this embodiment of the invention, step S400 involves determining the hidden state by inputting features and combining them with a temporal analysis network, including the following steps C1-C2: C1: Input the input features into the time series analysis network.
[0067] In this embodiment of the invention, the time series analysis network refers to a pre-interactive LSTM network.
[0068] Specifically, the fused features at different time points are input into the time series analysis network, including: Given image information at time points T+1, the pre-interactive LSTM network consists of T+1 pre-interactive LSTM modules. The input of the t-th module is the feature data of the image at the current time point after feature processing and selection network processing. The hidden state of the previous moment and the state of the memory unit at the previous moment Where, at t=0, the hidden state of the previous time step. and the state of the memory unit at the previous moment The values are initialized randomly, and the output is the hidden state at the current time. and the current state of the memory unit , where t is the time point number, t=0~T, and t is the time number.
[0069] It should be noted that by constructing a series structure of multiple pre-interactive LSTM modules, the orderly processing of multi-time point fusion features is achieved. Each module is responsible for processing the feature information of one time point. By passing the hidden state and memory unit state of the previous time step, the continuity and integrity of the temporal information are ensured. The random initialization design provides a good starting state for network training.
[0070] C2: The input features are processed using a temporal analysis network to obtain the hidden state and memory unit state at the current time step.
[0071] Specifically, such as Figure 6 The diagram shows the structure of a pre-interactive LSTM network. A temporal analysis network is used to process the input features to obtain the hidden state and memory unit state at the current time step, including the following steps C21-C22: C21: Perform multiple rounds of interactive operations between the current input features and the hidden state from the previous time step to obtain interactive features and interactive states, thereby enhancing the contextual interaction capability.
[0072] Furthermore, in step C21, the input features at the current time step are subjected to multiple rounds of interaction operations with the hidden state at the previous time step to obtain interaction features and interaction states, including: The initialization of interaction features and interaction states can be specifically represented as follows: ; ; in, Initial (Level -1) interaction features; The features of the image at the current time point are obtained after feature processing and selection network processing; This is the initial (level 0) interaction state; This is the hidden state from the previous moment.
[0073] In odd-numbered rounds In this process, the current-level interaction features are generated based on interaction characteristics and interaction states, which can be specifically represented as: ; in, for Level interaction features; for Level interaction features; for Level interaction state; The interaction weight matrix for the i-th level features; Use the Sigmoid activation function; This is element-wise multiplication.
[0074] In even-numbered rounds In this process, the current level interaction state is generated based on interaction features and interaction state, which can be specifically represented as: ; in, for Level interaction features; for Level interaction state; for Level interaction state; for Level-state interaction weight matrix.
[0075] C22: Based on interaction features and interaction states, and combined with the memory unit state of the previous moment, the hidden state and memory unit state of the current moment are calculated.
[0076] Specifically, the hidden state and memory unit state at the current moment are obtained, including: Let the number of iterations be... , It is an odd number, after After a round of interaction, the final interaction features are obtained. and interaction state .
[0077] The final interactive features and interaction state Compared with the previous moment's memory unit state The hidden state at the current time step is obtained by inputting the data into the original LSTM unit for computation. and the current state of the memory unit Specifically, it can be expressed as: ; in, The current hidden state; This represents the current state of the memory cell. This is for the original LSTM module operation.
[0078] It should be noted that by inputting the features and states obtained from multiple rounds of interaction into the original LSTM unit, the pre-interaction mechanism is combined with the traditional LSTM structure. Compared with the existing technology of directly using the original LSTM to process temporal features, this invention effectively enhances the expressive power of features through the pre-interaction process, and solves the problems of insufficient temporal information extraction and weak feature correlation in traditional methods. In particular, the multi-round interaction operation can explore the correlation between features at different time points, and the introduction of the original LSTM unit ensures the preservation of temporal memory. This not only improves the ability to obtain the correlation of image features of patients at different time points, but also provides richer and more accurate temporal correlation features for subsequent efficacy prediction, effectively enhancing the temporal analysis capability of AMD efficacy prediction.
[0079] In this embodiment of the invention, step S500 involves analyzing the characteristic change trend based on the hidden state to determine the therapeutic efficacy prediction result for age-related macular degeneration, including the following steps D1-D2: D1: Obtain the first hidden state at the first time point and the second hidden state at the last time point from the hidden state.
[0080] It should be noted that, considering the main objective of this invention is to predict the improvement in a patient's treatment efficacy, the most important data feature is actually the change in the patient's characteristics at various time points, that is, the trend of change, which can be understood as the difference before and after treatment; therefore, a Transformer subtraction module is added after the output of the pre-interactive LSTM network, such as... Figure 7 As shown, the input to the Transformer subtraction module is the hidden state of the pre-interactive LSTM network at the first time step. and the hidden state at the last point in time. The output is the predictive result of the treatment efficacy for age-related macular degeneration.
[0081] D2: Input the first hidden state and the second hidden state into the Transformer subtraction module to obtain the efficacy prediction results for age-related macular degeneration.
[0082] Specifically, the first and second hidden states are input into the Transformer subtraction module to obtain the efficacy prediction results for age-related macular degeneration, including the following steps D21-D23: D21: The hidden state is processed through a linear layer and a softmax function to obtain multiple attention feature values.
[0083] Specifically, the hidden state is processed through a linear layer and a softmax function to obtain multiple attention feature values, including: The hidden states output by the pre-interactive LSTM network are passed through multiple linear layers to obtain the query matrix. , Key matrix and value matrix ,in For the first One point of attention, =1~ , To determine the number of attention heads, then use the dot product to calculate the first... Attention feature values of each attention head Specifically, it can be expressed as: ; in, Let be the dimension of the vectors in the key matrix; This is the transpose symbol.
[0084] D22: Concatenate and linearly transform multiple attention feature values to obtain the output of multi-head attention.
[0085] Specifically, multiple attention feature values are concatenated and linearly transformed to obtain the output of multi-head attention, including: The results of multiple attention feature values are concatenated and a linear transformation is applied to obtain the output of multi-head attention. Specifically, it can be expressed as: ; in, For the first Attention feature values of each attention head; This is the linear transformation weight matrix.
[0086] D23: Determine the efficacy prediction results based on the output results of multi-head attention in the first hidden state and the output results of multi-head attention in the second hidden state.
[0087] Specifically, the efficacy prediction results are determined based on the output results of multi-head attention in the first hidden state and the output results of multi-head attention in the second hidden state, including: The output prediction probability for each category is obtained by subtracting the output of the multi-head attention at the first time point from the output of the multi-head attention at the last time point, and then processing it through a linear layer and a softmax function.
[0088] In this embodiment of the invention, a focus loss function is employed. Training the network can be specifically represented as: ; in, Indicates the first The probability that a sample is predicted to be the true label; This represents the total number of samples; An adjustable balancing factor is used to control the imbalance between categories, and is empirically set to 0.25 in this invention; This is an adjustment factor used to reduce the loss contribution of easily classified samples, and is set to 2 based on experience in this invention.
[0089] Furthermore, this invention calculates the focus loss for each category separately, and then averages the results across all categories to obtain the final loss function, which can be specifically expressed as: ; in, Number of categories; For the first Focus loss for each category; This represents the average loss.
[0090] It should be noted that by combining linear layers and the softmax function, an effective conversion from hidden states to class prediction probabilities is achieved. Compared with the problems of low accuracy and insufficient interpretability in efficacy prediction in existing technologies, this invention solves the problem of difficulty in accurately obtaining the correlation between image features of patients at different time points in traditional methods by calculating the difference before and after treatment. In particular, the subtraction module can better focus on the changes of sequence features over time. At the same time, the use of linear layers and the softmax function ensures the probabilistic expression and multi-classification ability of the prediction results. This not only improves the accuracy and reliability of AMD efficacy prediction, but also provides clinicians with quantifiable efficacy evaluation basis.
[0091] In summary, this invention addresses the challenges of limited data and significant sample bias in predicting the efficacy of anti-vascular endothelial growth factor (VEGF) therapy in patients with age-related macular degeneration (AMD). It proposes a multi-feature fusion long short-term memory (LSTM) network to predict the therapeutic effect of VEGF therapy in AMD patients. A feature extraction network based on 3DResNet and radiomics is designed to acquire high-quality image features. Furthermore, considering the mutual and individual characteristics of these features, and incorporating demographic features, a feature processing and selection network is proposed to effectively fuse these features. To address the difficulty in obtaining correlations between image features at different time points, a pre-interactive LSTM network and a Transformer subtraction module are used to integrate image information from various time points, enabling accurate prediction of treatment efficacy. This not only improves the automated processing capabilities of optical coherence tomography (OCT) images but also provides technical support for AMD research and clinical practice.
[0092] Example 3, referring to Tables 1 and 2, is an embodiment of the present invention, providing a method for predicting the efficacy of AMD based on feature fusion. To verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0093] This embodiment acquires a dataset of choroidal neovascularization, including optical coherence tomography (OCT) images and demographic characteristics of 88 wet AMD patients from Shanghai First People's Hospital during anti-vascular endothelial growth factor (VEGF) treatment. Specifically, the dataset is divided into four time points: before treatment, during the first VEGF injection, during the second VEGF injection, and during the third VEGF injection, with each treatment interval approximately one month. The OCT images have a resolution of 512×496×19 pixels and were obtained using a Heidelberg fundus tomography scanner. Based on the effect after three treatments and three months, the data are categorized into three groups: effective, moderate, and no response. There are 58 effective cases, 12 moderate cases, and 18 no-response cases. The dataset is divided using a three-fold cross-validation method, with data sizes of 29, 29, and 30 images respectively. Before inputting the data into the network, the images are resampled to a size of 512×512×19 pixels.
[0094] Secondly, the designed multi-feature fusion long short-term memory network was implemented using the open-source deep learning framework PyTorch. The model was trained and validated using an NVIDIA RTX 3090 GPU with 24GB of VRAM. An adaptive moment estimation algorithm was employed for model optimization, and a cosine annealing learning rate was used. The specific formula is as follows: ; in, The initial learning rate was set to 0.01 in the experiment; This refers to the number of iterations during the training process. The total number of iterations for the entire training process is set to 100.
[0095] In addition, the descent exponent power was set to 0.9, the optimizer momentum to 0.9, the weight decay coefficient to 0.0001, the batch size of the training data to 8, the number of interaction rounds r=5 in the pre-interactive LSTM module, and the number of attention heads k in the Transformer subtraction module to 8. During network training, after each iteration, the model was validated on the validation set, and the model parameters with the highest average accuracy in the validation set were retained as the final test network model. In subsequent comparative experiments, the comparison classification prediction algorithm networks all used the focus loss function to achieve network convergence.
[0096] To objectively evaluate the performance of the proposed network in predicting the therapeutic effect of anti-vascular endothelial growth factor therapy, this invention employs five commonly used classification evaluation metrics: precision (Pre), recall (Rec), accuracy (Acc), F1 coefficient (F1), and Kappa coefficient (kappa), with the specific formulas as follows: ; ; ; ; ; in, Observed agreement refers to the proportion of model predictions that match the true labels. Expected agreement is the proportion of predicted results that match the true label under random conditions; TP, TN, FP, and FN are true positives, true negatives, false positives, and false negatives, respectively; T is the number of correct predictions, which is the sum of true positives (TP) and true negatives (TN), i.e., T = TP + TN; F is the number of incorrect predictions, which is the sum of false positives (FP) and false negatives (FN), i.e., F = FP + FN.
[0097] For the five commonly used classification evaluation metrics mentioned above, this embodiment conducted ablation experiments between modules and classification performance tests of different temporal classification networks. Table 1 shows the results of the ablation experiments between modules. The values represent the mean ± standard deviation of the 3-fold classification, and the unit is 1.
[0098] Table 1. Results of ablation experiments between different modules
[0099] In Table 1, the unchecked "Pre-interactive LSTM" option indicates that the original LSTM is used by default to extract temporal features. Table 1 shows that the multi-feature fusion long short-term memory network proposed in this invention has the highest average index and the best overall performance. This indicates that the model designed in this invention effectively combines the feature information extracted by the deep learning network and the feature information from radiomics. Furthermore, the designed feature processing and selection network can better integrate demographic features, achieving an effective combination of the two types of information. Simultaneously, the combination of the proposed pre-interactive LSTM module and the Transformer subtraction module can effectively acquire temporal information and improve prediction accuracy. Comparing the results in the first four rows of Table 1 reveals… Currently, using deep learning features and radiomics features achieves higher accuracy than using deep learning features or radiomics features alone. Adding a feature processing and selection network further improves the overall accuracy. Comparing rows four and five reveals that using a pre-interactive LSTM network effectively enhances the extraction capability of features at various time points compared to a traditional LSTM network, further improving accuracy. Comparing rows five and six shows that using the Transformer subtraction module further enhances the ability to acquire the changing trends of optical coherence tomography (OCT) image features during patient treatment, contributing to further improvements in classification and prediction accuracy.
[0100] Compared with the ablation experiment results shown in Table 1, Table 2 shows the classification performance of different time-series classification networks. The values represent the mean ± standard deviation of the 3-fold classification, and the units are 1.
[0101] Table 2. Classification performance of different time-series classification networks
[0102] As can be seen from Table 2, the multi-feature fusion long short-term memory network designed in this invention achieved the best results on most metrics, reaching 85.7%, 76.9%, 75.4%, 76.2%, and 71.1% in average Acc, Pre, Rec, F1, and kappa, respectively. This indicates that the invention demonstrates very good performance in classification prediction tasks, and achieves the best performance in average Acc, F1, and kappa metrics.
[0103] Therefore, as can be seen from the results in Tables 1 and 2, this invention uses 3DResNet and Long Short-Term Memory (LSTM) networks as the backbone, adds additional radiomics features as auxiliary interpretable features to be fused into the features of the 3DResNet network, and designs a feature processing and selection network to fuse deep learning features, radiomics features, and additional demographic features; proposes a pre-interactive LSTM network to process multi-time point image features, and performs additional fusion on its input to enhance feature extraction capabilities; subsequently, a Transformer-based subtraction module is designed to make the network pay more attention to the changes of sequence features over time, thereby improving classification prediction accuracy; experiments show that the network proposed in this invention can overcome the problem of temporal classification prediction, predict the treatment effect of patients after three treatments based on images of patients before treatment and after three anti-vascular endothelial growth factor (VEGF) treatments, and achieves high classification accuracy, with classification prediction performance superior to other networks.
[0104] Example 4 is an embodiment of the present invention, which provides an AMD efficacy prediction system based on feature fusion, comprising: a patient data module for acquiring optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration; a feature extraction module for determining deep learning features and radiomics features based on the OCT image data; a feature fusion module for performing feature selection based on deep learning features, radiomics features, and demographic characteristics to obtain input features; a temporal correlation feature module for determining hidden states by calculating using the input features in conjunction with a temporal analysis network; and an efficacy prediction module for performing feature change trend analysis based on the hidden states to determine the efficacy prediction result for age-related macular degeneration.
[0105] This embodiment also provides an electronic device applicable to the AMD efficacy prediction method based on feature fusion, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the AMD efficacy prediction method based on feature fusion as proposed in the above embodiment.
[0106] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the AMD efficacy prediction method based on feature fusion as proposed in the above embodiments.
[0107] The storage medium proposed in this embodiment and the method for predicting the efficacy of AMD based on feature fusion proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0108] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0109] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for predicting the efficacy of AMD treatment based on feature fusion, characterized in that: include, To acquire optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration; Deep learning features and radiomics features are determined based on the optical coherence tomography image data; Feature selection is performed based on the deep learning features, the radiomics features, and the demographic features to obtain the input features; The hidden state is determined by combining the input features with a temporal analysis network for computation. Based on the analysis of the characteristic change trends of the hidden state, the therapeutic efficacy prediction results for age-related macular degeneration are determined.
2. The AMD efficacy prediction method based on feature fusion as described in claim 1, characterized in that: Based on the optical coherence tomography image data, deep learning features and radiomics features are determined, including: Deep learning features are obtained by processing the optical coherence tomography image data and the preset lesion mask. Features are extracted from the preset feature domain of the optical coherence tomography image data to obtain radiomics features.
3. The AMD efficacy prediction method based on feature fusion as described in claim 2, characterized in that: Feature selection is performed based on the deep learning features, the radiomics features, and the demographic features to obtain input features, including: The deep learning features and the radiomics features are processed using a feature correction module to obtain the deep learning features and radiomics features after the feature correction module is applied. The deep learning features, the radiomics features, and the demographic features are concatenated to obtain the total features. The total features are then processed using the feature correction module and the Softmax function to obtain the feature selection weights. The deep learning features and radiomics features processed by the feature correction module are concatenated and combined with the feature selection weights to calculate the input features.
4. The AMD efficacy prediction method based on feature fusion as described in claim 3, characterized in that: The hidden state is determined by combining the input features with a temporal analysis network, including: The input features are then fed into the time series analysis network. The input features are processed using the time-series analysis network to obtain the hidden state and memory unit state at the current time.
5. The AMD efficacy prediction method based on feature fusion as described in claim 4, characterized in that: The input features are processed using the time-series analysis network to obtain the hidden state and memory unit state at the current time step, including: The input features at the current time step are combined with the hidden state at the previous time step through multiple rounds of interaction operations to obtain the interaction features and interaction states. Based on the interaction features and interaction state, and combined with the memory unit state of the previous moment, the hidden state and memory unit state of the current moment are calculated.
6. The AMD efficacy prediction method based on feature fusion as described in claim 5, characterized in that: Based on the characteristic change trend analysis of the hidden state, the therapeutic efficacy prediction results for age-related macular degeneration are determined, including: Obtain the first hidden state at the first time point and the second hidden state at the last time point from the hidden states; The first hidden state and the second hidden state are input into the Transformer subtraction module to obtain the therapeutic effect prediction results for age-related macular degeneration.
7. The AMD efficacy prediction method based on feature fusion as described in claim 6, characterized in that: The first hidden state and the second hidden state are input into the Transformer subtraction module to obtain the therapeutic effect prediction results for age-related macular degeneration, including: The hidden state is processed through a linear layer and a softmax function to obtain multiple attention feature values; The multiple attention feature values are concatenated and linearly transformed to obtain the output result of multi-head attention. The therapeutic effect prediction result is determined based on the output results of multi-head attention in the first hidden state and the output results of multi-head attention in the second hidden state.
8. A feature fusion-based AMD efficacy prediction system, employing the feature fusion-based AMD efficacy prediction method as described in any one of claims 1 to 7, characterized in that, include: The patient data module is used to acquire optical coherence tomography (OCT) image data and demographic characteristics of patients with age-related macular degeneration. The feature extraction module is used to determine deep learning features and radiomics features based on optical coherence tomography image data; The feature fusion module is used to select features based on deep learning features, radiomics features, and demographic features to obtain input features; The temporal correlation feature module is used to determine the hidden state by combining input features with a temporal analysis network. The efficacy prediction module is used to analyze the trend of feature changes based on the hidden state to determine the efficacy prediction results for age-related macular degeneration.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the AMD efficacy prediction method based on feature fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the AMD efficacy prediction method based on feature fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Epidermal growth factor receptor mutation state judgment method, medium and electronic equipment
CN112488992A
Method for predicting glucocorticoid treatment TAO by using deep learning in combination with radiomics
CN116597937A
Treatment outcome prediction for neovascular age-related macular degeneration using baseline features
CN117121113A
Anti-vascular endothelial growth factor (VEGF) curative effect prediction apparatus and method
WO2021139446A1