Method for grading and prognosis prediction of senile renal artery stenosis patients

Through multimodal ultrasound imaging deep learning method, grading and prognostic prediction of elderly patients with renal artery stenosis has been solved, and the limitations of evaluation in the existing technology have been achieved, rapid and non-invasive precise grading and prognostic prediction have been achieved, and personalized treatment plans have been assisted.

CN120496843APending Publication Date: 2025-08-15BEIJING HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510598723.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to accurately grading and prognostic prediction of elderly patients with renal artery stenosis. There are limitations in the evaluation based solely on clinical characteristics, and the existing imaging methods are limited or incomplete in the use of elderly patients.

Method used

The multimodal ultrasound imaging deep learning method was adopted to obtain gray-scale ultrasound, Doppler ultrasound and ultrasound contrast images of elderly RAS patients at multiple time points, and the multimodal US image fusion network model was used for grading, and the prognosis prediction was made through multi-time US fusion analysis network model, and accurate evaluation was carried out in combination with clinical characteristics.

Benefits of technology

It has achieved rapid and non-invasive precise grading and prognosis prediction for elderly patients with renal artery stenosis, assisted in the formulation of personalized treatment plans, and reduced the health and economic costs of patients' families.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496843A_ABST
    Figure CN120496843A_ABST
Patent Text Reader

Abstract

The invention provides a method for grading and prognosis prediction of an old patient with renal artery stenosis. The method comprises the following steps: (1) acquiring RAS multi-modal multi-time-point ultrasonic image data of the old; (2) extracting a region of interest of the RAS ultrasonic image; (3) enhancing the image of the region of interest of the RAS ultrasonic image; and (4) grading the RAS by using the trained multi-modal US image fusion network model, and performing prognosis prediction on renal function deterioration by using the trained multi-time-point US fusion analysis network model. The deep learning grading diagnosis and prognosis prediction model of the multi-mode ultrasonic image is used for carrying out feature extraction on the multi-mode multi-time-point ultrasonic image, rapid and noninvasive accurate grading is carried out on RAS, and the prognosis risk of renal function deterioration of different treatment schemes is accurately predicted; therefore, clinical doctors are assisted in formulating accurate and individualized treatment schemes for the old RAS diseases, and the hygienic economic cost of families of patients is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing and analysis, and in particular to a method for grading and predicting the prognosis of elderly patients with renal artery stenosis. Background Art

[0002] Renal artery stenosis (RAS) is a progressive disease, often caused by atherosclerosis, that can lead to hypertension and worsening renal function. Treatment for end-stage patients is very expensive, and medication to control hypertension cannot prevent RAS progression and renal function decline in some patients. Elderly patients experience rapid RAS progression and a poor prognosis, with the rate of progression increasing with age. RAS progression in patients aged 60 to 69 is 1.7 times higher than in those under 60 years of age. Accurately diagnosing RAS grading and predicting RAS prognosis, and guiding timely clinical intervention, are key to improving the prognosis of elderly patients with RAS.

[0003] The diagnosis of RAS involves both anatomic diagnosis (evaluation of the degree of vascular stenosis) and functional diagnosis (evaluation of renal function). Digital subtraction angiography (DSA), CT angiography (CTA), and MR angiography (MRA) only provide anatomical information for RAS but are unable to provide functional diagnosis. Furthermore, limitations such as radiation exposure, invasiveness, and contrast-induced nephropathy restrict their use in elderly patients with RAS. Ultrasound imaging (including grayscale ultrasound, color Doppler ultrasound, and contrast-enhanced ultrasound) is a safe (no renal damage), accurate, dynamic, and one-stop method for simultaneous anatomical diagnosis (assessment of the degree of vascular stenosis) and functional diagnosis (assessment of renal function). Compared with conventional ultrasound, modified CEUS combined with conventional ultrasound hemodynamic monitoring can improve the diagnostic accuracy of RAS from 65% to 90%. However, it requires high operator sensitivity and is highly operator-dependent.

[0004] The prognosis of RAS is associated with clinical characteristics such as age, diabetes, hypertension, cardiac function, and renal function. However, prognostic assessment based solely on clinical characteristics in elderly RAS patients has limitations. Studies have found that color Doppler ultrasound-derived renal artery hemodynamics are also an important factor influencing the prognosis of RAS patients. Ultrasound imaging features are associated with the occurrence of cardiovascular and renal vascular events and the efficacy of interventional treatment in RAS patients during follow-up. Multi-omics feature fusion combining clinical and imaging features can help improve the accuracy of prognostic assessment. Chen et al. analyzed data from 573 RAS patients in the CORAL study and found that the random forest model had the highest accuracy in predicting unfavorable prognosis. Khitan et al. found that the support vector machine algorithm had the highest accuracy in predicting unfavorable prognosis. Professor Li Jianchu and colleagues from Peking Union Medical College Hospital conducted the first study using convolutional neural networks to classify intrarenal artery spectral images, demonstrating reasonable diagnostic accuracy for the classification of severe RAS. Professor Huo Yong from Peking University First Hospital and Professor Jiang Xiongjing from the National Center for Cardiovascular Diseases, Fuwai Hospital, have also developed corresponding prognostic prediction models. The research subjects of these imaging models are mostly non-elderly patients, and they have shortcomings such as incomplete imaging data and inability to predict prognosis, making it difficult to provide accurate and personalized disease monitoring and treatment plans for elderly RAS patients. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for grading and predicting the prognosis of elderly patients with renal artery stenosis in response to the problems existing in the prior art, so as to solve the problem that the prognosis assessment of elderly RAS patients based solely on clinical characteristics is currently limited.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] A method for grading and predicting prognosis of elderly patients with renal artery stenosis, the method comprising:

[0008] (1) Obtain multimodal and multi-time point ultrasound imaging data of elderly RAS;

[0009] (2) extracting the region of interest from the RAS ultrasound image;

[0010] (3) Enhance the image of the region of interest in the RAS ultrasound image;

[0011] (4) The RAS was graded using a trained multimodal US image fusion network model, and the prognosis of renal function deterioration was predicted using a trained multi-time point US fusion analysis network model.

[0012] As a further technical solution, the multimodal and multi-time-point ultrasound imaging data of elderly RAS in step (1) include grayscale ultrasound, Doppler ultrasound and ultrasound contrast-enhanced video images, wherein the grayscale ultrasound image acquisition standards are: a) ≥1 kidney maximum cross-sectional measurement image; b) ≥1 renal artery maximum diameter stenosis cross-sectional measurement image; Doppler ultrasound image acquisition standards are: a) ≥1 renal artery static CDFI image; b) ≥1 renal artery spectrum image; ultrasound contrast-enhanced video acquisition standards: display the entire renal artery in real-time through two-dimensional and color Doppler tracking, measure the two-dimensional lumen diameter and spectrum Doppler parameters, observe the renal artery trunk after entering the contrast-enhancing mode, continuously observe and store the image for 20 to 30 seconds, and extract the ultrasound contrast-enhanced static image from the retained video: ≥1 renal artery maximum diameter stenosis cross-sectional measurement image; multi-time-point ultrasound image acquisition standards: for the same patient, the time interval between the three ultrasound contrast-enhanced tests and image acquisitions cannot be less than 6 months.

[0013] As a further technical solution, in step (2), the ultrasound physician uses Labelme software to roughly outline the ROI along the edge of the lesion with a rectangular frame to complete the extraction of the region of interest.

[0014] As a further technical solution, step (3) specifically includes first adjusting all ultrasound images to a size of 70*70, and then performing random cropping, random scaling, random flipping, random rotation, random translation, and image tensorization.

[0015] As a further technical solution, the network model of the multimodal US image fusion in step (4) is composed of a convolution module and a skip layer link, wherein the convolution module is composed of multiple convolution layers, a layer normalization module and a nonlinear activation function, the convolution layer is used to extract the spatial feature information of the renal artery and the entire kidney area, the layer normalization module is used to complete the internal covariate shift; the nonlinear activation function is used to perform nonlinear transformation.

[0016] As a further technical solution, the specific formula for extracting spatial feature information of the renal artery and the entire kidney area by the convolutional layer is: This formula represents a two-dimensional convolution operation, which is used to extract local features from the input image. X is the modal data of the ultrasound image, with dimensions of height H * width W * number of channels C. W is the convolution kernel, with dimensions of height U * width V * number of input channels C. Y[i,j] is the output of the convolution operation, which refers to the value of the output feature map at position (i,j).

[0017] As a further technical solution, the specific formula used by the layer normalization module to complete the internal covariate shift is: Among them, x i Represents the i-th element of the input feature, μ L represents the mean of the feature, σL represents the standard deviation of the feature, and γ and β represent the learnable scaling and translation parameters.

[0018] As a further technical solution, the specific formula of the nonlinear activation function for nonlinear transformation is: Where x is the input value, GELU enhances the model's ability to capture complex features through nonlinear transformation.

[0019] As a further technical solution, the specific steps of the layer-hopping link are as follows:

[0020] Assume that the input is x and the residual function is F(x), then the output y of the residual block can be expressed as: y = F(x) + x,

[0021] Perform global average pooling on the output feature map of the last residual block and convert the feature map into a vector of fixed length: This formula represents the downsampling operation in a convolutional neural network, which is used to reduce the size of the feature map to H×W×C, where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map. After downsampling, the size becomes (H / 2)×(W / 2)×C', where C' is the new number of channels.

[0022] Finally, the output of the global average pooling layer is connected to the fully connected layer to produce the final score and normalized by the Softmax function: y = Wx + b, The linear transformation formula of the fully connected layer is: W is the weight matrix, b is the bias vector, x is the input feature vector, and z is the output score vector; Softmax normalization function: z i is the original score of the i-th category, K is the total number of categories, and the score is converted into a probability distribution for multi-classification tasks.

[0023] As a further technical solution, the training process of the network model for multimodal US image fusion is as follows:

[0024] (41) According to steps (1) to (3), a multimodal ultrasound imaging dataset of elderly RAS patients who underwent at least three ultrasound contrast examinations during a 12-month follow-up period was constructed;

[0025] (42) The multimodal ultrasound image dataset was randomly divided into a training set and a validation set;

[0026] (43) The parameters of the network model for multimodal US image fusion are initialized based on the pre-trained weights, and then the parameters of the network model for multimodal US image fusion are fine-tuned using the ultrasound image data of the training set.

[0027] As a further technical solution, the multi-time point US fusion analysis network model consists of two parallel convolutional neural network branches. Each branch processes the imaging data from two different time points, compares and fuses the characteristic information of the imaging data from these two time points, and realizes the prognosis prediction of renal function deterioration.

[0028] As a further technical solution, the multi-time point US fusion analysis network model predicts the prognosis of renal function deterioration in the following steps: the image input at each time point enters an independent convolutional neural network, which extracts the feature maps F1 and F2 of each time point to learn the deep feature representation of the images at different time points. The feature maps F1 and F2 extracted by the two branches are then sent to the feature fusion module. The feature fusion module performs a weighted summation of the two feature maps according to the formula F = α*F1+β*F2, calculates the difference between the features, and captures the change information and interaction relationship between the two time points, where α and β are learnable weight parameters; the fused features are sent to the fully connected layer for linear transformation, and finally the Softmax function is used to normalize the model score to make the final prediction of renal function deterioration:

[0029] y=Wx+b,

[0030]

[0031] The linear transformation formula of the fully connected layer is: W is the weight matrix, b is the bias vector, x is the input feature vector, and y is the output score vector;

[0032] Softmax normalization function: z i is the original score of the i-th category, K is the total number of categories, and the score is converted into a probability distribution for multi-classification tasks.

[0033] As a further technical solution, the training process of the multi-time point US fusion analysis network model is as follows:

[0034] (41) According to steps (1) to (3), a multimodal ultrasound imaging dataset of elderly RAS patients who underwent at least three ultrasound contrast examinations during a 12-month follow-up period was constructed;

[0035] (42) The multimodal ultrasound image dataset was randomly divided into a training set and a validation set;

[0036] (43) The parameters of the multi-time point US fusion analysis network model are initialized based on the pre-trained weights, and then the parameters of the multi-time point US fusion analysis network model are fine-tuned using the ultrasound imaging data of the training set.

[0037] Compared with the existing technology, the core of the present invention lies in the deep learning hierarchical diagnosis and prognosis prediction model of multimodal ultrasound images. The model intends to use deep learning methods to extract features from multimodal and multi-time point ultrasound images, quickly and non-invasively accurately grade RAS, and accurately predict the prognostic risk of renal function deterioration under different treatment plans, thereby assisting clinicians in formulating accurate and individualized treatment plans for elderly RAS diseases and reducing the health and economic costs of patients' families. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a system principle block diagram of the present invention;

[0039] Figure 2 This is an image after the region of interest is outlined along the edge of the lesion with a rectangular frame. DETAILED DESCRIPTION

[0040] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Example

[0042] This embodiment provides a method for grading and predicting the prognosis of elderly patients with renal artery stenosis. Figure 1 、 Figure 2 , the method comprises the steps of:

[0043] S1: Constructing a multimodal and multi-time point ultrasound imaging dataset of elderly RAS

[0044] 1) Elderly patients with RAS who underwent CEUS examinations at least three times during the 12-month follow-up period, with multimodal ultrasound imaging data including grayscale ultrasound, Doppler ultrasound, and dynamic CEUS video images;

[0045] 2) Grayscale ultrasound image acquisition criteria: a) ≥1 image of the largest renal cross-section (measuring the long diameter, thick diameter, and parenchymal thickness); b) ≥1 image of the maximum renal artery diameter stenosis ratio cross-section (measuring the maximum inner diameter, minimum inner diameter, and diameter stenosis ratio);

[0046] 3) Doppler ultrasound (CDFI) image acquisition criteria: a) ≥1 renal artery static CDFI image (showing the section with the most abundant blood flow); b) ≥1 renal artery spectral image (measuring peak systolic velocity PSV);

[0047] 4) Contrast-enhanced ultrasound (CEUS) video acquisition standards: Real-time tracking of the entire renal artery using two-dimensional and color Doppler ultrasound, measurement of two-dimensional lumen diameter and spectral Doppler parameters, observation of the main renal artery after entering the angiography mode, continuous observation and storage of images for 20–30 seconds, and instructing the patient to hold their breath; capture of static CEUS images from the remaining video: ≥1 image of the renal artery maximum diameter stenosis section (measurement of the maximum inner diameter, minimum inner diameter, and diameter stenosis ratio);

[0048] 5) Longitudinal cohort multi-time point ultrasound image acquisition standards: For the same enrolled patient, the time interval between the three CEUS examinations and image acquisitions must not be less than 6 months;

[0049] 6) Two senior ultrasound doctors graded RAS according to the diagnostic criteria. If there was disagreement, a third expert would evaluate the case.

[0050] 7) Record the medical history, demographic characteristics, laboratory tests, etc. of elderly patients, and track and record the outcomes of renal function deterioration.

[0051] Based on laboratory examination data such as patient medical history and demographic characteristics, a deep fully connected neural network prediction model was constructed, and the potential mapping relationship between clinical characteristics and the risk of renal function deterioration was explored through multi-layer nonlinear transformation.

[0052] S2: Extraction of Region of Interest in RAS Ultrasound Images

[0053] To prevent redundant pixels around the lesion from interfering with the deep learning model's effective learning of RAS knowledge, this study performed region of interest (ROI) delineation. Ultrasound physicians used Labelme software to roughly outline the ROI along the edge of the lesion with a rectangular frame. Only the outlined portion of the image was retained to complete the ROI extraction. This eliminated redundant information and enabled efficient deep learning model training.

[0054] S3: Image enhancement of regions of interest in RAS ultrasound images

[0055] This paper enhances multimodal US imaging data, including grayscale ultrasound images, CDFI images, and static CEUS images. All ultrasound images are first resized to 70x70 pixels, followed by random cropping, random scaling, random flipping, random rotation, random translation, and image tensorization. Data augmentation effectively prevents network overfitting. All data augmentation steps are performed in Python 3.7 using PyTorch.

[0056] S4: Dataset Partitioning

[0057] For the stenosis grading task, the dataset was randomly split into a training set (n=350) and a validation set (n=88) in a 4:1 ratio. For the renal function assessment task, the dataset was divided into a training set (n=341) and a validation set (n=86). The dataset was randomly partitioned using the train_test_split function in the Sklearn Python library, using a stratified sampling strategy. For the stenosis grading and functional assessment tasks, the stratify parameter of the function was set to the stenosis degree and the radionuclide dynamic phenomenon results, respectively, to maintain the same ratio of target categories in the training and validation sets as in the original dataset. The training set was iteratively adjusted to minimize the training set loss function. The internal validation set was used for hyperparameter selection and tuning.

[0058] S5: Accurate classification of RAS: Network model design for multimodal US image fusion

[0059] A multimodal fusion network framework was developed. First, features from each modality were extracted using ConvNeXt. The model consists of convolutional modules linked by skip layers. Each convolutional module consists of multiple convolutional layers, layer normalization modules, and nonlinear activation functions. The convolutional layers are used to extract spatial features of the renal artery and the entire kidney.

[0060] In order to reduce the size of the feature map and extract the deep semantic features of the ultrasound image, the downsampling operation is used to reduce the dimension of the feature map. The specific formula for extracting the spatial feature information of the renal artery and the entire kidney area by the convolution layer is:

[0061]

[0062] This formula represents a two-dimensional convolution operation, used to extract local features from an input image (or feature map). X is the input feature map (e.g., a modal data of an ultrasound image, in this embodiment, grayscale ultrasound, Doppler ultrasound, and ultrasound contrast imaging video images), with dimensions H*W*C (height*width*number of channels), and W is the convolution kernel (filter), with dimensions U*V*C (height*width*number of input channels). Y[i,j] is the output of the convolution operation, which refers to the value of the output feature map at position (i,j), indicating the degree of match between the local region and the convolution kernel.

[0063] The layer normalization module is used to address internal covariate shift, improve training efficiency, accelerate convergence, and enhance model generalization capabilities:

[0064]

[0065] Among them, x i Represents the i-th element of the input feature, μ L represents the mean of the feature, σL represents the standard deviation of the feature, and γ and β represent the learnable scaling and translation parameters.

[0066] The nonlinear activation function is used to enhance the nonlinear transformation capability of the model and improve the model's ability to capture and represent the complex features of renal images. The model uses the GELU activation function:

[0067]

[0068] Where x is the input value, GELU enhances the model's ability to capture complex features through nonlinear transformation.

[0069] To alleviate the gradient vanishing or gradient exploding problems that occur during deep network training and enable the network to train deep structures more effectively, the model introduces skip-layer links in each convolutional module. Let the input be x and the residual function be F(x) (multiple convolutional layers, layer normalization modules, and nonlinear activation functions), then the output y of the residual block can be expressed as:

[0070] y=F(x)+x,

[0071] Perform global average pooling on the output feature map of the last residual block and convert the feature map into a vector of fixed length:

[0072]

[0073] This formula represents the downsampling operation in a convolutional neural network (such as a pooling layer), which is used to reduce the size of the feature map to (H×W×C), where H is the height of the feature map, W is the width of the feature map, and C is the number of channels in the feature map. After downsampling, the size becomes (H / 2)×(W / 2)×C', where C' is the new number of channels.

[0074] Finally, the output of the global average pooling layer is connected to the fully connected layer to produce the final score and normalized by the Softmax function:

[0075] y=Wx+b,

[0076]

[0077] The linear transformation formula of the fully connected layer is: W is the weight matrix, b is the bias vector, x is the input feature vector, and z is the output score vector. Softmax normalization function: z iis the raw score for class i, and K is the total number of classes. The scores are converted to probability distributions for multi-classification tasks. This model uses a multi-level module to process the region of interest: each module consists of a convolutional layer, layer normalization, and a GELU activation function, connected in parallel with skip layers to form a residual structure. Multiple such modules are cascaded, and features are aggregated through global average pooling. Finally, a fully connected layer outputs the predicted score.

[0078] Finally, the scores of the three modalities are integrated for decision-level fusion to obtain the final RAS diagnosis result.

[0079] S6: Prediction of renal function deterioration: Design of a network model for multi-time point US fusion analysis

[0080] A two-stream convolutional neural network is used to process and analyze imaging data from two different time points of the same patient. This architecture extracts feature information from the image data sets at different time points by independently processing the images at each time point, and then merges these features to achieve the final prediction of renal function deterioration. The two-stream convolutional neural network architecture consists of two parallel convolutional neural network branches, each of which processes image data from two different time points. Comparing and fusing the feature information of the image data at these two time points helps the model predict renal function deterioration more accurately. Specifically, the image input at each time point enters an independent convolutional neural network (such as ResNet). The two convolutional neural networks have the same structure and do not share parameters. They are used to extract features from their respective time points to learn deep feature representations of images at different time points. The feature maps F1 and F2 extracted by the two branches are then sent to the feature fusion module (the two sets of representations F1 and F2 are intermediate results extracted by the neural network):

[0081] F=α*F1+β*F2,

[0082] The two feature maps are weighted and summed to calculate the difference between the features, capturing the change information and interaction between the two time points. (In this method, the generation of the feature "difference" and the "change information and interaction relationship" described below are not achieved through preset explicit mathematical operations, but rather through an implicit mechanism of autonomous learning of the neural network. This mechanism automatically discovers complex spatiotemporal correlation patterns that may be overlooked by clinical explicit methods through a data-driven approach.) α and β are learnable weight parameters. The fused features are fed into a fully connected layer for linear transformation, and finally the Softmax function is used to normalize the model score to predict renal function deterioration:

[0083] y=Wx+b,

[0084]

[0085] The linear transformation formula of the fully connected layer is: W is the weight matrix, b is the bias vector, x is the input feature vector (in this method, the input feature vector is the feature map F1 and F2), and y is the output score vector.

[0086] Softmax normalization function: z i is the original score of the i-th category, K is the total number of categories, and for y, the score is converted into a probability distribution for multi-classification tasks.

[0087] S7: Model construction and validation

[0088] This study employed a transfer learning strategy, initializing model parameters using pre-trained ImageNet weights provided by PyTorch. The model parameters were then fine-tuned using renal ultrasound image data from the training set. The AdamW optimizer was used to optimize network parameters, with an initial learning rate of 0.0005 and a weight decay coefficient of 0.05. A warm start strategy was also employed. During training, the learning rate was adjusted in real time using a cosine descent strategy. The batch size for training was set to 32.

[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be pointed out that any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for grading and predicting prognosis of elderly patients with renal artery stenosis, characterized in that: The method includes: (1) Obtain multimodal and multi-time point ultrasound imaging data of elderly RAS; (2) extracting the region of interest from the RAS ultrasound image; (3) Enhance the image of the region of interest in the RAS ultrasound image; (4) The RAS was graded using a trained multimodal US image fusion network model, and the prognosis of renal function deterioration was predicted using a trained multi-time point US fusion analysis network model.

2. A method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 1, characterized in that: The multimodal and multi-time-point ultrasound imaging data of elderly RAS described in step (1) include grayscale ultrasound, Doppler ultrasound, and ultrasound contrast-enhanced video images, wherein the grayscale ultrasound image acquisition standards are: a) ≥1 kidney maximum cross-sectional measurement image; b) ≥1 renal artery maximum diameter stenosis cross-sectional measurement image; Doppler ultrasound image acquisition standards are: a) ≥1 renal artery static CDFI image; b) ≥1 renal artery spectrum image; ultrasound contrast-enhanced video acquisition standards: display the entire renal artery in real-time through two-dimensional and color Doppler tracking, measure the two-dimensional lumen diameter and spectrum Doppler parameters, observe the renal artery trunk after entering the contrast mode, continuously observe and store the image for 20 to 30 seconds, and extract the ultrasound contrast-enhanced static image from the retained video: ≥1 renal artery maximum diameter stenosis cross-sectional measurement image; multi-time-point ultrasound image acquisition standards: for the same patient, the time interval between the three ultrasound contrast-enhanced tests and image acquisitions cannot be less than 6 months.

3. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 1, characterized in that: In step (2), the ultrasound physician uses Labelme software to roughly outline the ROI along the edge of the lesion with a rectangular frame to complete the extraction of the region of interest.

4. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 1, characterized in that: Step (3) specifically includes first adjusting all ultrasound images to a size of 70*70, and then performing random cropping, random scaling, random flipping, random rotation, random translation, and image tensorization.

5. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 1, characterized in that: The network model of multimodal US image fusion in step (4) is composed of a convolution module and a skip layer link. The convolution module is composed of multiple convolution layers, a layer normalization module and a nonlinear activation function. The convolution layer is used to extract the spatial feature information of the renal artery and the entire kidney area. The layer normalization module is used to complete the internal covariate shift; the nonlinear activation function is used to perform nonlinear transformation.

6. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 5, characterized in that: The specific formula for extracting the spatial feature information of the renal artery and the entire kidney area by the convolution layer is: This formula represents a two-dimensional convolution operation, which is used to extract local features from the input image. X is the modal data of the ultrasound image, with dimensions of height H * width W * number of channels C. W is the convolution kernel, with dimensions of height U * width V * number of input channels C. Y[i,j] is the output of the convolution operation, which refers to the value of the output feature map at position (i,j).

7. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 5, characterized in that: The specific formula used by the layer normalization module to complete the internal covariate shift is: Among them, x i Represents the i-th element of the input feature, μ L represents the mean of the feature, σ L represents the standard deviation of the feature, and γ and β represent the learnable scaling and translation parameters.

8. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 5, characterized in that: The specific formula for nonlinear activation function to perform nonlinear transformation is: Where x is the input value, GELU enhances the model's ability to capture complex features through nonlinear transformation.

9. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 5, characterized in that: The specific steps of skip link are: Assume that the input is x and the residual function is F(x), then the output y of the residual block can be expressed as: y = F(x) + x, Perform global average pooling on the output feature map of the last residual block and convert the feature map into a vector of fixed length: This formula represents the downsampling operation in a convolutional neural network, which is used to reduce the size of the feature map to H×W×C, where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map. After downsampling, the size becomes (H / 2)×(W / 2)×C', where C' is the new number of channels. Finally, the output of the global average pooling layer is connected to the fully connected layer to produce the final score and normalized by the Softmax function: y = Wx + b, The linear transformation formula of the fully connected layer is: W is the weight matrix, b is the bias vector, x is the input feature vector, and z is the output score vector; Softmax normalization function: z i is the original score of the i-th category, K is the total number of categories, and the score is converted into a probability distribution for multi-classification tasks.

10. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 1, characterized in that: The training process of the network model for multimodal US image fusion is as follows: (41) According to steps (1) to (3), a multimodal ultrasound imaging dataset of elderly RAS patients who underwent at least three ultrasound contrast examinations during a 12-month follow-up period was constructed; (42) The multimodal ultrasound image dataset was randomly divided into a training set and a validation set; (43) The parameters of the network model for multimodal US image fusion are initialized based on the pre-trained weights, and then the parameters of the network model for multimodal US image fusion are fine-tuned using the ultrasound image data of the training set.

11. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 1, characterized in that: The multi-time point US fusion analysis network model consists of two parallel convolutional neural network branches. Each branch processes the imaging data from two different time points, compares and fuses the characteristic information of the imaging data from these two time points, and realizes the prognosis prediction of renal function deterioration.

12. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 10, characterized in that: The steps for the multi-time point US fusion analysis network model to predict the prognosis of renal function deterioration are as follows: the image input at each time point enters an independent convolutional neural network, which extracts the feature maps F1 and F2 of each time point to learn the deep feature representation of the images at different time points. The feature maps F1 and F2 extracted by the two branches are then sent to the feature fusion module. The feature fusion module performs a weighted summation of the two feature maps according to the formula F = α*F1+β*F2, calculates the difference between the features, and captures the change information and interaction relationship between the two time points, where α and β are learnable weight parameters; the fused features are sent to the fully connected layer for linear transformation, and finally the Softmax function is used to normalize the model score to make the final prediction of renal function deterioration: y=Wx+b, The linear transformation formula of the fully connected layer is: W is the weight matrix, b is the bias vector, x is the input feature vector, and y is the output score vector; Softmax normalization function: z i is the original score of the i-th category, K is the total number of categories, and the score is converted into a probability distribution for multi-classification tasks.

13. The method for grading and predicting prognosis of elderly patients with renal artery stenosis according to claim 12, characterized in that: The training process of the multi-time point US fusion analysis network model is as follows: (41) According to steps (1) to (3), a multimodal ultrasound imaging dataset of elderly RAS patients who underwent at least three ultrasound contrast examinations during a 12-month follow-up period was constructed; (42) The multimodal ultrasound image dataset was randomly divided into a training set and a validation set; (43) The parameters of the multi-time point US fusion analysis network model are initialized based on the pre-trained weights, and then the parameters of the multi-time point US fusion analysis network model are fine-tuned using the ultrasound imaging data of the training set.