Children bladder ureter reflux prediction method and system

By processing static images of dithiothioate kidneys with a deep learning model based on a multi-head attention mechanism, the problem of insufficient diagnosis of dithiothioate kidney scans was solved, and a non-invasive and accurate diagnosis of vesicoureteral reflux was achieved, which is suitable for long-term health monitoring of pediatric patients.

CN120725976APending Publication Date: 2025-09-30THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510804840.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The existing static renal scan using dithiothreitol succinate is not sensitive and specific enough to diagnose vesicoureteral reflux. It cannot directly display the urine reflux process and relies on voiding cystourethrography, which is invasive and carries radiation risks for children, resulting in inaccurate and unsafe diagnosis.

Method used

A deep learning model based on a multi-head attention mechanism was used to perform background correction, region of interest segmentation and data enhancement on static images of dithiothreitol kidneys. Convolution operations and sliding window transformer modules were used to extract features to achieve automated diagnosis without invasive examinations.

Benefits of technology

It improves the accuracy and consistency of vesicoureteral reflux diagnosis, reduces physical harm and radiation risks to pediatric patients, provides quantifiable diagnostic results, and reduces the need for invasive examinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725976A_ABST
    Figure CN120725976A_ABST
Patent Text Reader

Abstract

A child bladder ureter reflux prediction method and system relate to the field of image processing, and the system comprises an image acquisition unit used for acquiring a dimercaptosuccinic acid kidney static image of a child patient; the image processing unit is used for carrying out region-of-interest division on the image and carrying out preprocessing; the prediction model is trained to obtain the kidney image processed by the image processing unit and output a classification result, and the prediction model comprises an image block feature embedding module; the multi-layer feature extraction module captures local features through a layered window multi-head self-attention mechanism, and an image block merging layer is inserted between every two basic layers to perform down-sampling on a feature map; and the feature fusion layer is used for fusing the features captured by each layer and outputting a classification result. According to the invention, the local and global features in the medical image can be effectively captured, the effective utilization of the medical image data is realized, and the auxiliary diagnosis effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The original application is titled “A Deep Learning Prediction System for Vesicoureteral Reflux in Children.” The application number is 202510008744.6, and the application date is January 3, 2025. Technical Field

[0002] The present invention relates to the field of image processing technology, and in particular to a deep learning prediction system for vesicoureteral reflux in children. Background Art

[0003] Vesicoureteral reflux (VUR) is a common urinary tract disorder in childhood. A static dimercaptosuccinic acid (DMSA) renal scan is a key adjunct to the diagnosis of VUR. As a functional imaging test, DMSA renal scan effectively assesses renal infection or scarring caused by VUR, helping clinicians understand the extent of renal damage and its potential causes. For patients with potential VUR, the results of a static DMSA renal scan can provide important guidance for further voiding cystourethrography (VCUG). When a static DMSA renal scan reveals significant renal damage or scarring, this may indicate more severe VUR, prompting clinicians to recommend a VCUG to confirm the presence and severity of urinary reflux and provide more specific guidance for subsequent treatment decisions.

[0004] However, despite its value in assessing and monitoring renal injury associated with VUR, succinyltransferase (STS) renal scans are not the gold standard for VUR. Succinyltransferase (STS) renal scans cannot directly demonstrate the reflux of urine; their diagnostic value lies primarily in indirectly suggesting the presence of VUR. Therefore, their sensitivity and diagnostic value are limited for early VUR or cases without overt renal injury. Furthermore, STS renal scans have low specificity and may misinterpret renal abnormalities caused by non-VUR factors as VUR-induced injury, potentially leading to unnecessary further testing or treatment.

[0005] Therefore, in the diagnosis and management of VUR, succimer renal static scanning should be considered an assessment tool rather than a sole diagnostic tool. Combining the patient's medical history, clinical manifestations, and especially the results of voiding cystourethrography (VCUG), the gold standard for VUR, allows for a more comprehensive and accurate assessment of VUR and its renal effects, thereby enabling the development of optimal treatment strategies.

[0006] For example, Chinese patent CN115762753A discloses a method for automatic classification of vesicoureteral reflux using deep learning. This method is based on voiding cystourethrography images and uses an integrated deep learning model to automatically classify the severity of vesicoureteral reflux. The method is trained on a dataset annotated by senior doctors, and the prediction results of different deep learning models are integrated using a voting method to form the final classification label. Model preprocessing includes image cropping, standardization, and data augmentation operations. The final model is constructed based on the Residual Network (ResNet) architecture, which can handle the classification of unilateral and bilateral reflux at the same time.

[0007] Invasive examinations that rely on voiding cystourethrography: This method uses voiding cystourethrography images to grade reflux. Although it can provide highly accurate grading results, voiding cystourethrography is radioactive and invasive to patients, making it unsuitable for frequent use, especially for children. In addition, voiding cystourethrography is widely considered the gold standard for diagnosing vesicoureteral reflux, and its diagnostic consistency is high. Even different doctors reach consistent judgments based on the same voiding cystourethrography image. Therefore, this method introduces deep learning technology to diagnose the grade of vesicoureteral reflux based on voiding cystourethrography images. The room for improving diagnostic practicality is very limited, and its auxiliary value for actual clinical work is insufficient. Summary of the Invention

[0008] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a deep learning prediction system for vesicoureteral reflux in children.

[0009] To achieve the above object, the present invention provides the following technical solutions:

[0010] A deep learning prediction system for vesicoureteral reflux in children, comprising:

[0011] An image acquisition unit, used for acquiring a static image of the patient's kidney;

[0012] An image processing unit divides the acquired static image of the kidney with succinimidyl mercaptoacetic acid into regions of interest, performs background correction, and performs quantitative processing to obtain a processed kidney image with risks;

[0013] The prediction model is trained to obtain kidney images processed by the image processing unit and output classification results, and the prediction model includes:

[0014] The image block feature embedding module divides the input image into several non-overlapping image blocks through convolution operations, and maps each image block into an embedding vector of fixed dimension through the convolution layer;

[0015] Several layers of feature extraction modules capture local features through a layered window multi-head self-attention mechanism, and an image block merging layer is inserted between each basic layer to downsample the feature map;

[0016] The feature fusion layer fuses the captured features of each layer and outputs the classification results.

[0017] The image processing unit preprocesses the dithiosuccinate kidney static image based on the following steps:

[0018] S1, background correction of static images of dithiosuccinate kidneys;

[0019] S2, kidney contours were drawn on the background-corrected image, and the image was cropped and scaled;

[0020] S3, select the kidney image on the severely damaged side;

[0021] S4, enhance and normalize the data.

[0022] In S1, the background subtraction method after integration was used for correction. The upper left area of ​​the left kidney and the upper right area of ​​the right kidney were selected as the background area, the region of interest was automatically divided, and the grayscale value of the background area was calculated and subtracted.

[0023] In S2, the automatic threshold selection technology based on the image histogram is used to outline the contours of the two kidneys and calculate the center lines of the two kidneys. The center lines are kept unchanged and the kidney images are symmetrically processed to obtain the original and symmetrical kidney annulus contours:

[0024] Use a rectangular frame to include the outlines of both kidneys and crop the image to an aspect ratio of 2:1;

[0025] Scale the cropped image to a fixed size.

[0026] Each layer of the feature extraction module is provided with a sliding window transformer module.

[0027] The sliding window transformer module uses a 7×7 window size to divide the feature map into several non-overlapping windows, and each window contains several feature points. Multi-head self-attention is performed in each window to calculate the correlation between feature points and capture local features.

[0028] In adjacent sliding window transformer modules, the window position will be moved by half the window size to achieve cross-window information interaction;

[0029] The features processed by the attention mechanism are input into a multi-layer perceptron consisting of two fully connected layers and an activation function for nonlinear transformation.

[0030] In each sliding window transformer module, residual connections and layer normalization are used.

[0031] The feature extraction module consists of 4 stages:

[0032] In stage 1, the input feature dimension is 96, it contains 2 sliding window transformer modules, and the number of attention heads is 3;

[0033] In stage 2, the input feature dimension is 192, it contains two sliding window transformer modules and the number of attention heads is 6;

[0034] In stage 3, the input feature dimension is 384, it contains 6 sliding window transformer modules and the number of attention heads is 12;

[0035] In stage 4, the input feature dimension is 768, it contains 2 sliding window transformer modules, and the number of attention heads is 24.

[0036] The beneficial effects of the present invention are as follows: based on the analysis of static images of kidneys with dithiothioate, there is no need to resort to the traditional invasive examination of voiding cystourethrography, which fundamentally reduces the physical harm to patients, especially children, and reduces the discomfort and medical risks during the examination. At the same time, it does not generate additional radiation risks and is more suitable for long-term health monitoring of children. Traditional methods rely on manual judgment and have certain subjectivity and inconsistency. Our method realizes automated analysis and prediction based on the network model of the multi-head attention mechanism, provides quantifiable and objective diagnostic results, and improves the accuracy and consistency of diagnosis. In our model, through the hierarchical window multi-head self-attention mechanism, key features can be extracted in the local window, and the interaction of cross-window information can be realized. This multi-layer nested structure helps to capture the deep features in the static images of kidneys with dithiothioate, enabling the model to autonomously learn to distinguish the image features of vesicoureteral reflux and effectively deal with the complex structures in the image. The model can automatically learn the importance of features without a large amount of preprocessing operations, achieving accurate classification results. Therefore, our model can quantify doctors' subjective judgments on static images of the kidneys with dithiothreitol, solving problems such as lack of quantitative standards, reliance on subjective judgment, and discomfort and radiation exposure to pediatric patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a design flow chart of the present invention.

[0038] Figure 2 Schematic diagram of the multi-head attention mechanism model of the present invention.

[0039] Figure 3 Schematic diagram of image segmentation of the present invention.

[0040] Figure 4 Images of vesicoureteral reflux, a is an image with vesicoureteral reflux, b is an image without vesicoureteral reflux and abnormal image display.

[0041] Figure 5 Schematic diagram of a sliding window converter module used in an embodiment of the present invention.

[0042] Figure 6 Schematic diagram of the test results of the present invention. DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0044] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0045] like Figure 1 As shown, the present invention provides a deep learning prediction system for vesicoureteral reflux in children, which includes:

[0046] An image acquisition unit, used for acquiring a static image of the patient's kidney of the patient's succimer capsid, wherein the static image of the patient's kidney of the patient's succimer capsid can be input by a doctor;

[0047] The image processing unit divides the acquired kidney static image into regions of interest and performs

[0048] Perform background correction and quantitative processing to obtain a processed kidney image with risks;

[0049] The prediction model is trained to obtain kidney images processed by the image processing unit and output classification results, and the prediction model includes:

[0050] The image block feature embedding module divides the input image into several non-overlapping image blocks through convolution operations, and maps each image block into an embedding vector of fixed dimension through the convolution layer;

[0051] Several layers of feature extraction modules capture local features through a layered window multi-head self-attention mechanism, and an image block merging layer is inserted between each basic layer to downsample the feature map;

[0052] The feature fusion layer fuses the captured features of each layer and outputs the classification results.

[0053] The specific design flow chart is as follows Figure 1 shown.

[0054] The prediction model needs to be trained, and the dataset is constructed:

[0055] A1. Patient Selection Criteria: Children with acute urinary tract infection (UTI) who meet the criteria for the Diagnosis and Treatment of Urinary Tract Infection in Children (2021 Edition) were selected, and their clinical data were complete. These children had static renal examinations with succimer thiazolidsin and had abnormalities confirmed, and had undergone voiding cystourethrography during hospitalization.

[0056] A2. Clinical data collection: Systematically collect the imaging data and report results of patients' succimer renal static images and voiding cystourethrography.

[0057] The image processing unit needs to pre-process the static image of the kidney of dithiothreitol succinate according to the following method:

[0058] B1. Background Correction: To reduce the error caused by the concentration of dithiothiocyanate in normal tissue, background correction was performed using integration followed by background subtraction. The upper left region of the left kidney and the area directly to the right of the right kidney were selected as background regions. The region of interest (ROI) was automatically delineated using programming. The grayscale value of the background region was calculated and subtracted to reduce the influence of background noise and improve image quality.

[0059] B2. Image Cropping and Scaling: Use an automated threshold segmentation algorithm to outline the kidneys, maintaining their centerlines and maintaining symmetry. Use a rectangular box to enclose the kidneys, maintaining appropriate pixel margins to prevent cropping of the region of interest. Crop the image to a fixed aspect ratio and scale it to a uniform size to ensure consistency of the input data.

[0060] B3. Unilateral renal image selection: Quantitative analysis of the regions of interest (ROIs) of both kidneys was performed, and grayscale values ​​were calculated after background correction. In the VUR group, images of the kidney with the most severe damage (high-grade VUR) were selected. Statistical testing was performed to verify the validity of the selection criteria and ensure that the selected images more accurately reflected the actual reflux characteristics.

[0061] B4. Data Augmentation and Normalization: To improve the generalization ability of the model, data augmentation techniques such as random flipping, rotation, cropping, and color jittering are used to increase data diversity. Images are normalized to meet the input requirements of the model.

[0062] The prediction model is a network model based on a multi-head attention mechanism, such as Figure 2 As shown on the left, the method for medical image data processing can efficiently extract and classify medical image data. This method uses a hierarchical structure and a multi-head self-attention mechanism. Figure 2 On the right, multi-scale features are captured and important features are adaptively learned, thereby improving the performance and robustness of the model.

[0063] Model architecture design:

[0064] A1. Feature Embedding Module: The input medical image is divided into image blocks and each block is converted into an embedding vector of fixed dimension through a convolutional layer. This process achieves downsampling, reduces the amount of computation, and produces a feature map with spatial and channel information.

[0065] A2. Attention Mechanism Module: The core layer of this invention is composed of multiple basic units, each of which integrates a feature interaction component and an information integration unit based on the multi-head attention mechanism. These components effectively capture local information and communicate features between different regions by defining feature interactions within specific regions.

[0066] A3. Feature Fusion Layer: An image block merging layer is added between basic layers. Through pooling operations, the feature map is downsampled and the information of adjacent image blocks is fused, increasing the number of channels while maintaining information integrity.

[0067] B. Model features:

[0068] B1. Multi-head Attention Mechanism: The multi-head attention mechanism within each basic unit allows the model to simultaneously focus on different parts of the image, enabling parallel processing of information at multiple scales. This mechanism not only enhances the understanding of local features but also captures different types of dependencies through multiple "heads," increasing the model's ability to learn complex patterns.

[0069] B2. Hierarchical structure: Through hierarchical design, the model can capture multi-scale features and adaptively learn important features, thereby improving the performance and robustness of the model.

[0070] C. Model output:

[0071] C1. Global Feature Fusion: Finally, the abstract features of each layer are fused into the global average pooling layer to extract global information. C2. Classification Result: The classification result is output through the fully connected layer to achieve accurate classification of medical images.

[0072] Through the above method, the present invention can effectively capture local and global features in medical images, realize the effective use of medical imaging data, optimize the training process, improve the accuracy of disease diagnosis, reduce the need for invasive examinations, and has high clinical application value.

[0073] Network model training based on multi-head attention mechanism

[0074] This paper provides a neural network model training process based on a multi-head self-attention mechanism to improve the recognition accuracy of children with vesicoureteral reflux. Through an optimized dataset partitioning strategy and model training process, this method fully utilizes the imaging features of the severely damaged kidney side, improving the model's generalization and predictive performance.

[0075] A. Dataset partitioning strategy:

[0076] The dataset is divided into training set, validation set, and test set. The training set and validation set use the 10-fold cross-validation method to further improve the generalization ability of the model and avoid overfitting.

[0077] B. Model training strategy:

[0078] B1. Training Data Selection: During model training, images of the severely damaged kidney are selected as input. This allows the model to focus on key features related to vesicoureteral reflux and enhance its ability to identify pathological changes.

[0079] B2. Model testing phase: During the model testing phase, imaging data of the patient's bilateral kidneys are input. The model comprehensively analyzes the characteristics of the left and right sides to determine whether vesicoureteral reflux has occurred.

[0080] C. Model optimization and evaluation:

[0081] C1. Optimization algorithm: Use appropriate optimization algorithms and loss functions for model training to improve the convergence speed and accuracy of the model.

[0082] C2. Performance evaluation: By evaluating the performance on the validation set and independent test set, we calculate the model's accuracy and other indicators to ensure that the model has good generalization ability and stability.

[0083] Through the above method, the present invention effectively utilizes the characteristics of medical imaging data, optimizes the training process of the model, improves the recognition accuracy of children with vesicoureteral reflux, reduces the dependence on invasive examinations, and has important clinical application value. Specific embodiment:

[0085] Based on the sliding window transformer (swin_transfomer) neural network.

[0086] In this embodiment,

[0087] Dataset selection criteria

[0088] A1. Patient Screening: 346 children with acute urinary tract infection (UTI) admitted to the Department of Pediatric Nephrology, the Second Affiliated Hospital of Wenzhou Medical University between January 2019 and January 2023 were selected. Patients were required to meet the domestic diagnostic criteria of the "Guidelines for the Diagnosis and Treatment of Urinary Tract Infection in Children (2021 Edition)" and have complete clinical data. All patients underwent static renal examination with dithiothreitol succinate and underwent voiding cystourethrography during hospitalization.

[0089] A2. Clinical data collection: The imaging data and reported results of patients undergoing dithiothreitol succinate renal static scan and voiding cystourethrography were collected.

[0090] Data preprocessing

[0091] B1. Acquisition of static image data of succimer kidney:

[0092] Using centrally purchased instruments, standardized acquisition techniques and protocols, static imaging data of the renal succimerization renal system were collected one hour after radiopharmaceutical injection. All static succimerization renal succimerization images were uploaded to the Picture Archiving and Communication System (PACS) at the Second Affiliated Hospital of Wenzhou Medical University. Images were acquired at the same window width and window level to ensure consistent imaging parameters.

[0093] B2: Background correction method:

[0094] Background correction principle: In order to reduce the error caused by the concentration of dithiothreitol succinate in normal tissue, background subtraction after integration is used for correction.

[0095] Background region selection: Traditional methods typically use the area below the kidney for subtraction. However, related research indicates that the intensity of the upper left region of the left kidney and the area immediately to the right of the right kidney is close to that of the empty kidney. Because dimercaptosuccinic acid and diethylenetriaminepentaacetic acid (DTPA) have similar metabolic patterns in the body, selecting these areas as background regions is more accurate.

[0096] Specific steps: Use the program's automated threshold segmentation algorithm to delineate regions of interest (ROIs), including the kidney region and the corresponding background region. Calculate the grayscale intensity of each ROI. For each kidney ROI, subtract the grayscale value of the corresponding background region to obtain a corrected image.

[0097] B3. Image cropping and scaling (e.g. Figure 3 shown):

[0098] Kidney contouring: The kidneys were delineated using an automatic threshold selection technique based on the image histogram (Otsu's method). The centerlines of the kidneys were calculated and, while maintaining the centerlines, the kidney images were symmetrically processed to obtain the original and symmetrical kidney contours.

[0099] Image cropping: Use a rectangular frame to completely enclose the outline of both kidneys, leaving at least 10 pixels above, below, left, and right to prevent the region of interest from being cropped. Crop the image to an aspect ratio of 2:1.

[0100] Image scaling: Scale the cropped image to a fixed size, 224 × 448 pixels, to unify the image size.

[0101] B4. Unilateral kidney image selection:

[0102] Selection of unilateral kidney images: The python program was used to quantitatively analyze the bilateral regions of interest of the static dimercaptosuccinic acid angiography images, and the grayscale values ​​obtained after background subtraction were calculated. The grayscale values ​​of the bilateral kidneys in the vesicoureteral reflux group (176 cases) after background subtraction were compared using python. It was found that among the 102 cases with different grades of vesicoureteral reflux on the left and right sides, the grayscale values ​​of 95 cases with high-grade VUR were smaller than those on the low-grade side, and the grayscale values ​​of the other 7 cases were not much different on both sides. The McNemar test was performed on these 102 patients to compare the differences in this paired binary data. According to our grayscale value results, the p value of the McNemar test was 3.3*10-12, which was much smaller than the significance level (p<0.05), indicating that the high-grade vesicoureteral reflux side had a lower grayscale value than the low-grade side.

[0103] The presence of a lower grayscale value on the kidney than on the low-grade VUR side was statistically significant. This suggests that, in the VUR group, the kidney on the high-grade VUR side has a lower grayscale value, possibly related to inflammation or other pathological changes. Clinically, we use the high-grade side as the standard for treatment planning. Because the study objective was to identify the presence of VUR using static succimer imaging in children with acute urinary tract infection, we selected renal images from the high-grade VUR side in the VUR group after background subtraction.

[0104] Vesicoureteral reflux group (176 patients): Among 102 patients with different grades of vesicoureteral reflux on the left and right sides, the kidney image on the side with the most severe damage (high-grade vesicoureteral reflux) was selected. The reasons for selecting unilateral kidney images (such as Figure 4 Some cases of vesicoureteral reflux may appear healthy on images, but in fact vesicoureteral reflux has occurred. This is mainly because the other kidney is severely damaged, causing ureteral reflux in the damaged weak kidney. Figure 4

[0105] As shown in (a), there is no reflux in both kidneys, and the images of both kidneys are bright, complete, and without any missing parts. This is the correspondence between normal images and reflux levels. Figure 4 (b) shows high-grade reflux in both kidneys. The left kidney image shows significant image loss, but the right kidney image does not. Therefore, choosing the more severely damaged side better reflects the actual reflux characteristics.

[0106] B5. Data augmentation and normalization:

[0107] Data enhancement: In order to increase data diversity and prevent model overfitting, multiple data enhancement processes are performed on the image:

[0108] Resize images: Resize images to 1.2 times their original size to increase scale variation.

[0109] Random Horizontal Flip: Randomly flip the image horizontally to simulate the position change of the left and right kidneys.

[0110] Random rotation: Randomly rotate the image within the range of ±15 degrees to enhance the model's robustness to angle changes.

[0111] Random cropping and scaling: Randomly crop parts of the image and scale them to the target size to increase image diversity.

[0112] Color dithering: Randomly adjusts the brightness and contrast of an image to simulate changes in imaging conditions.

[0113] Normalization: Convert image data into tensor form. Use predefined mean and standard deviation to standardize the image to meet the input requirements of the model.

[0114] B6: Random seed setting: To ensure the repeatability of experimental results, set the random seed to ensure consistency in data partitioning and model initialization.

[0115] The structure of the prediction model constructed in this embodiment is as follows Figure 5 As shown in the figure, after four rounds of feature extraction and downsampling, each feature extraction module captures local features through a layered windowed multi-head self-attention mechanism. Each layer achieves cross-window feature interaction by shifting windows, enabling the model to effectively integrate global and local information. Furthermore, a patch merging layer is added between each layer to downsample the feature maps. Ultimately, the abstract features from each layer are fused in a global average pooling layer, and the classification results are output through a fully connected layer.

[0116] A. Model architecture design:

[0117] A1. Patch Feature Embedding Module: The model uses preprocessed medical images (224×448 pixels) as input. A 4×4 convolution operation with a stride of 4 is used to divide the input image into several non-overlapping image patches. Each image patch is mapped to an embedding vector of fixed dimension through a convolutional layer, achieving downsampling and reducing computational effort. A normalization layer is added to the embedded feature data to stabilize it and improve model training stability.

[0118] A2. Sliding Window Transformer Module: Using a 7×7 window size, the feature map is divided into several non-overlapping windows, each containing 49 feature points. Window Multi-Head Self-Attention (W-MSA): A multi-head self-attention mechanism is implemented within each window to calculate the correlation between feature points and capture local features. Shifted Window Multi-Head Self-Attention (SW-MSA): Within adjacent Sliding Window Transformer modules, the window position is shifted by half the window size (i.e., leftward and upward), enabling cross-window information exchange and capturing a wider range of features. Multi-Layer Perceptron (MLP): The features processed by the attention mechanism are input into a Multi-Layer Perceptron (MLP) consisting of two fully connected layers and an activation function for nonlinear transformation, enhancing the model's expressiveness. Residual Connections and Layer Normalization: Residual connections and layer normalization are used in each module to prevent gradient vanishing and promote model training stability.

[0119] A3. Patch Merging Layer: This layer is inserted between each base layer and performs downsampling. It merges adjacent 2×2 image patches, reducing the size of the feature map (height and width are halved) while increasing the number of channels (typically doubling them) to maintain feature integrity. The merged features are linearly transformed and normalized to adjust the feature dimension and minimize information loss.

[0120] B. Model structure design:

[0121] B1. Hierarchical Design: The model consists of four stages, each with different input and output feature dimensions. Specifically:

[0122] Stage 1: The input feature dimension is 96, including 2 sliding window transformer modules and 3 attention heads.

[0123] Stage 2: The input feature dimension is 192, including 2 sliding window transformer modules and 6 attention heads.

[0124] Stage 3: The input feature dimension is 384, including 6 sliding window transformer modules and 12 attention heads.

[0125] Stage 4: The input feature dimension is 768, including 2 sliding window transformer modules and 24 attention heads.

[0126] B2. Feature extraction and downsampling: At each stage, feature extraction is performed through the base layer module, and downsampling is performed through the image patch merging layer between stages.

[0127] C. Model output layer:

[0128] C1. Global average pooling layer: The feature map of the last layer is pooled globally to obtain a feature vector of fixed length.

[0129] C2. Fully connected layer: The feature vector is input into the fully connected layer, which outputs the final classification result, which is used to determine whether the child has vesicoureteral reflux.

[0130] D. Model parameter settings:

[0131] D1. Embedding dimension: The initial embedding dimension is set to 96, and the feature dimension is gradually doubled as the stage progresses.

[0132] D2. Window size: The window size is set to 7, which ensures the capture of local features while controlling the computational complexity.

[0133] D3. Activation function: Use Gaussian Error Linear Unit (GELU) activation function in the multilayer perceptron to improve nonlinear expression capabilities.

[0134] D4. Normalization layer: Use the normalization layer (LayerNormalization, LayerNorm) to normalize the features and stabilize the training process.

[0135] The training process of the prediction model is as follows:

[0136] A. Dataset division:

[0137] A1. Initial Partition: A medical imaging dataset of 346 children with acute urinary tract infections was divided into a training set and a validation set in a 9:1 ratio. The training and test sets (90%) were used for model training and cross-validation. The independent validation set (10%) was reserved for testing data used to evaluate final model performance and was not used in training or cross-validation.

[0138] A2. 10-fold cross-validation: 10-fold cross-validation is used on 90% of the training set data: the data is divided into 10 equal parts, one of which is selected as the test set each time (accounting for 9% of the total data), and the remaining nine parts are used as the training set (accounting for 81% of the total data), and this cycle is repeated ten times.

[0139] This method ensures that each data sample is used for validation once, making full use of the data and improving the generalization ability of the model.

[0140] B. Model training:

[0141] B1. Training data preparation:

[0142] During the training process, for each patient, images of the side of the kidney with severe damage are selected as the input data of the model.

[0143] B2. Model training process:

[0144] Parameter setting: Use the pre-built sliding window transformer model, set the learning rate to 0.0001, the batch size to 8, the optimizer to select the optimizer with weight decay (Adam with Weight Decay Regularization, AdamW), and the weight decay to 0.01.

[0145] Training process:

[0146] In each cross-validation fold, the model is trained using the training set, while the test set is used to evaluate the model's performance on unseen data. During training, the cross-entropy loss function is used as the loss metric. An early stopping mechanism is employed, with a patience limit of 20. Training is stopped early when the model's performance on the validation set no longer improves to prevent overfitting. In each training fold, the best-performing model parameters on the validation set are saved for subsequent evaluation.

[0147] C. Model testing:

[0148] C1. Test data preparation: During the model testing phase, for each patient, the image data of both kidneys are input.

[0149] C2. Model Prediction: Unilateral Prediction: The model predicts the left and right kidney images separately, obtaining a unilateral classification result. Bilateral Comprehensive Judgment: If the prediction results for both the left and right kidneys are 0, the comprehensive judgment is 0 (no vesicoureteral reflux); otherwise, it is judged as 1 (vesicoureteral reflux).

[0150] C3. Performance Evaluation: Calculate metrics such as model loss, one-sided accuracy, and two-sided accuracy on the validation set and independent test set. Average the results of ten-fold cross-validation to assess model stability and generalization.

[0151] D. Model optimization and adjustment:

[0152] D1. Learning rate adjustment: Use the stagnation-based learning rate reduction learning rate scheduler (ReduceLearningRateonPlateau, ReduceLROnPlateau) to automatically adjust the learning rate according to the validation set loss to improve model performance.

[0153] D2. Model parameter adjustment: Based on the training and validation results, adjust the model's hyperparameters, such as learning rate, batch size, and training rounds, to optimize model performance.

[0154] Through the above specific implementation steps, the training and testing of the sliding window transformer model were completed, and the effectiveness and reliability of the model in identifying children with vesicoureteral reflux were verified, providing a powerful auxiliary tool for clinical diagnosis.

[0155] Training Results

[0156] After completing the training and validation of the model, we evaluated the performance of the model, such as Figure 6 The specific results are as follows:

[0157] A. 10-fold cross validation training set performance:

[0158] A1. Training set accuracy: During the ten-fold cross-validation process, the model achieved an average accuracy of 0.8947 on the training set. This indicates that the model performs well in learning the characteristics of the training data and is able to effectively fit the training data.

[0159] B. 10-fold cross-validation validation set performance:

[0160] B1. Unilateral Validation Set Accuracy: The model achieved a validation set accuracy of 0.8621 on unilateral kidney images. This indicates that the model has high accuracy in identifying the presence of vesicoureteral reflux in a unilateral kidney.

[0161] B2. Bilateral Validation Set Accuracy: When considering both kidney images, the validation set accuracy was 0.7931. Although this is lower than the unilateral prediction accuracy, the model is still able to accurately determine whether the patient has vesicoureteral reflux.

[0162] C. Test set performance:

[0163] C1. Unilateral test set accuracy: On an independent test set, the model achieved an accuracy of 0.8182 for unilateral kidney images. This demonstrates the model's good predictive power on unseen data and demonstrates its generalization performance.

[0164] C2. Bilateral test set accuracy: When using bilateral kidney images for prediction, the test set accuracy is 0.7879. Although the accuracy is lower, the model still has reference value in practical applications.

[0165] Results Analysis: The sliding window transformer model achieved high accuracy on the training, validation, and test sets. This was particularly true for unilateral kidney image predictions, demonstrating the model's ability to effectively learn and identify features related to vesicoureteral reflux. Unilateral predictions achieved higher accuracy than bilateral predictions. This may be because the model primarily focused on features of the severely damaged side during training, whereas features from the healthy side may interfere with model judgment when predicting bilaterally. The accuracy on the test set was slightly lower than that on the training and validation sets, but remained high, demonstrating the model's good generalization ability and its ability to adapt to unseen case data.

[0166] In addition, we compared several different deep learning models, with experimental results shown in Table 1. While other models also demonstrated success in predicting the likelihood of vesicoureteral reflux in children, they generally underperformed the sliding window transformer model. These results further confirm the unique advantages of the sliding window transformer model.

[0167] We also experimented with different parameter configurations of the sliding window transformer model to optimize its performance and identify the optimal configuration. The experiments covered various configurations, including micro, small, basic, and large. These configurations differed in terms of embedding dimension, number of attention heads, and layer depth. As shown in Table 2, the basic configuration achieved the best performance, surpassing all other configurations.

[0168] The experimental results above demonstrate that the deep learning method based on the sliding window transformer model demonstrates excellent performance in identifying vesicoureteral reflux in children with acute urinary tract infection. The model's stable performance across multiple datasets demonstrates its effectiveness and reliability, providing a valuable auxiliary tool for clinical diagnosis, reducing reliance on invasive testing, and possessing significant clinical application value.

[0169] At the same time, the network architecture based on the multi-head self-attention mechanism of the present invention has been proved through experiments to have a residual network

[0170] 34, residual network 50, Maximum Vision Transformer (MaxVit), Google network, densely connected convolutional network (DesNet) and other deep learning models are used as substitutes, but the prediction accuracy of these deep learning models after training is lower than our network architecture, as shown in the table below.

[0171]

[0172] The network proposed in this paper uses a base parameter size, with options for tiny, small, and large parameter sizes available. The differences between the tiny, small, base, and large versions primarily focus on the embedding dimension, number of attention heads, and layer depth, which affect the model's expressiveness, computational requirements, and applicable scenarios. Experimental comparisons show that the base parameter size yields the best results, as shown in the table below.

[0173]

[0174] The hyperparameters of the model used in the present invention, such as window size, embedding dimension, number of attention heads, layer depth, learning rate, optimizer, etc., can be set appropriately according to the specific situation. In addition, the multi-head self-attention mechanism can also be replaced with other attention mechanisms (such as global self-attention) to adapt to different task requirements. At the same time, the downsampling method and activation function of the image block merging layer can also be replaced or adjusted according to the task requirements to optimize model performance.

[0175] The embodiments should not be regarded as limiting the present invention, but any improvements based on the spirit of the present invention should be within the scope of protection of the present invention.

Claims

1. A method for predicting vesicoureteral reflux in children, characterized in that: include: Obtain a succimer renal image of the child, and correct the background noise of the succimer renal image by integral background subtraction to obtain a corrected renal image. The correction areas are selected at the upper left corner of the left kidney and the upper right corner of the right kidney based on the metabolic characteristics of succimer. The corrected renal image is delineated using an automatic threshold based on the histogram, and symmetry processing is performed around the center line. After symmetry processing is completed, the image is cropped and scaled according to the preset aspect ratio. The renal image on the severely damaged side is selected based on the statistical rule that the grayscale value on the side with high-grade VUR is significantly lower. Image enhancement and normalization are then performed to obtain the preprocessed renal image. The preprocessed kidney image is input into the prediction model for classification prediction to obtain a classification result of the probability of reflux; wherein, the prediction model is constructed based on a sliding window transformer model, and the prediction model includes an image block feature embedding module, several layers of feature extraction modules and a feature fusion layer, and each layer of the feature extraction module is provided with a sliding window transformer module.

2. The method according to claim 1, characterized in that The background noise correction of the dimercaptosuccinate kidney image by integral background subtraction specifically includes: Correction was performed using the post-integration background subtraction method. The upper left area of ​​the left kidney and the upper right area of ​​the right kidney were selected as the background area. The region of interest was automatically divided, and the grayscale value of the background area was calculated and subtracted.

3. The method according to claim 1, characterized in that The process of cropping and scaling the kidney image includes: Use the automatic threshold selection technology based on the image histogram to outline the contours of the two kidneys and calculate the center lines of the two kidneys. Keeping the center lines unchanged, perform symmetry processing on the kidney image to obtain the original and symmetric kidney annulus contours: Use a rectangular frame to include the outlines of both kidneys and crop the image to an aspect ratio of 2:1; Scale the cropped image to a fixed size.

4. The method according to claim 1, wherein The prediction model specifically includes: The image block feature embedding module is used to divide the input image into several non-overlapping image blocks through convolution operations, and map each image block into an embedding vector of fixed dimension through the convolution layer; Several layers of feature extraction modules are used to capture local features through a layered window multi-head self-attention mechanism, and an image block merging layer is inserted between each basic layer to downsample the feature map; The feature fusion layer is used to fuse the features captured by each layer and output the classification results.

5. The method according to claim 4, characterized in that The processing process of the feature extraction module specifically includes: The feature map is divided into several non-overlapping windows, each containing several feature points. Multi-head self-attention is performed within each window to calculate the correlation between feature points and capture local features. In adjacent sliding window transformer modules, the window position is shifted by half the window size to achieve cross-window information exchange. The features processed by the attention mechanism are input into a multi-layer perceptron consisting of two fully connected layers and an activation function for nonlinear transformation. In each sliding window transformer module, residual connections and layer normalization are used.

6. The method according to claim 5, characterized in that: The feature extraction module consists of 4 stages: In stage 1, the input feature dimension is 96, it contains 2 sliding window transformer modules, and the number of attention heads is 3; In stage 2, the input feature dimension is 192, it contains two sliding window transformer modules and the number of attention heads is 6; In stage 3, the input feature dimension is 384, it contains 6 sliding window transformer modules and the number of attention heads is 12; In stage 4, the input feature dimension is 768, it contains 2 sliding window transformer modules, and the number of attention heads is 24.

7. A pediatric vesicoureteral reflux prediction system, characterized in that: include: An image acquisition module is used to obtain a dimercaptosuccinate kidney image of the patient, and correct the background noise of the dimercaptosuccinate kidney image by integral background subtraction to obtain a corrected kidney image; wherein the correction area is selected at the upper left side of the left kidney and the upper right side of the right kidney based on the metabolic characteristics of dimercaptosuccinate; An image preprocessing module is used to outline the bilateral kidney contours of the corrected renal image based on the automatic threshold of the histogram, perform symmetry processing with the center line as the axis, and after the symmetry processing is completed, crop and scale according to a preset aspect ratio. Based on the statistical rule that the grayscale value of the high-grade VUR side is significantly lower, the renal image of the more severely damaged side is selected from the scaled renal image, and image enhancement and normalization are performed to obtain a preprocessed renal image. The prediction module is used to input the preprocessed kidney image into the prediction model for classification prediction to obtain the probability classification result of reflux; wherein, the prediction model is constructed based on the sliding window transformer model, and the prediction model includes an image block feature embedding module, several layers of feature extraction modules and a feature fusion layer, and each layer of the feature extraction module is provided with a sliding window transformer module.

Citation Information

Patent Citations

  • Method for automatically grading bladder ureter backflow through deep learning

    CN115762753A