Method, system and medium for assessing stone friability based on ct images

By constructing a multimodal grouped autoencoder based on CT images, the morphology, texture, and heterogeneity features of stones are extracted, which solves the problem of low accuracy in assessing stone fragility in existing technologies, and achieves more accurate assessment of stone fragility, supporting precise decision-making in clinical treatment strategies.

CN122335778APending Publication Date: 2026-07-03TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
Filing Date
2026-04-08
Publication Date
2026-07-03

Smart Images

  • Figure CN122335778A_ABST
    Figure CN122335778A_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and medium for assessing the fragility of kidney stones based on CT images, belonging to the field of CT image processing technology. The method includes: acquiring CT images with stone border labels and a training set of the maximum fragment volume after ESWL (Extract, Scale, and Lie); extracting stone morphology, texture, and heterogeneity features from the CT images to construct an original multimodal feature vector; constructing a multimodal grouped autoencoder, with the first layer being a group discovery layer; setting a group weight matrix; using mean squared error as the reconstruction loss; and introducing an L2,1 norm sparse regularization term during training to obtain the weight matrix; filtering features based on the absolute value threshold of the weights in each row of the matrix to construct feature groups; calculating group weights based on the variance of each feature group in the training set; for each sample, taking the mean of features within the group, and weighting and fusing them with group weights to obtain a fused feature vector; and finally training a multilayer perceptron regression model using the fused features as input to obtain a fragility assessment model. This method integrates multiple types of features and automatically mines associated feature groups, solving the problem of low accuracy in assessing kidney stone fragility caused by a single reference factor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of CT image processing technology, and more specifically to a method, system, and medium for assessing the fragility of stones based on CT images. Background Technology

[0002] Since its introduction, extracorporeal shock wave lithotripsy (ESWL) has rapidly become a first-line treatment for upper urinary tract stones due to its non-invasive nature, ease of operation, and good tolerability. Especially for proximal ureteral stones, ESWL uses focused shock waves to penetrate human tissue and break the stones into fine particles, allowing them to be naturally excreted with urine, thus avoiding the trauma and complications associated with invasive surgery. Statistics show that ESWL is currently used in over 80% of ureteral stone cases worldwide, and its clinical value is widely recognized. However, despite its widespread application, the uncertainty of its efficacy and potential safety risks remain challenges that urgently need to be addressed in clinical practice. Clinical observations indicate that stones of similar size and location can exhibit significantly different fragmentation effects after ESWL treatment with the same parameters: some stones can be completely fragmented and successfully excreted, while others leave larger fragments, requiring further treatment or conversion to other surgical methods. This inconsistency in treatment efficacy not only increases the medical burden and suffering of patients, but may also delay treatment and even lead to kidney damage. Therefore, accurately determining whether a patient's current stone status is suitable for ESWL treatment before surgery, i.e., how to accurately screen suitable patients for ESWL, has become a key issue in optimizing treatment strategies and improving clinical outcomes.

[0003] In existing technologies, the morphological parameters of kidney stones have been extensively studied and confirmed to be significantly associated with the efficacy of ESWL (Emergency Swing Lift), and are widely used in clinical practice as reference indicators for assessing stone fragility. Morphological parameter-based assessment methods are simple, intuitive, and readily available, and can guide treatment decisions to some extent; for example, larger stones are generally considered to require more impact cycles or higher energy.

[0004] In fact, the physical properties of kidney stones are multi-dimensional. Besides morphology, their internal structural characteristics, especially heterogeneity, are highly correlated with their fragility. Kidney stone heterogeneity reflects the spatial distribution of different density components within the stone. During stone formation, due to fluctuations in urine composition, infection, or metabolic abnormalities, stones often exhibit a layered structure or a multicentric growth pattern, leading to uneven internal density distribution. This heterogeneity causes shock waves to propagate at different density interfaces, resulting in reflection, refraction, and focusing phenomena at the interfaces, thus creating a complex stress field within the stone. The stronger the heterogeneity and the greater the density difference, the more pronounced the local stress concentration at the interfaces, making it easier to initiate and propagate cracks, ultimately promoting stone fragmentation.

[0005] Therefore, considering only morphological parameters while ignoring intrinsic physical properties such as heterogeneity results in low accuracy of existing methods for assessing the fragility of stones, making it difficult to meet the needs of precise clinical decision-making. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides a method for assessing the fragility of kidney stones based on CT images. This method evaluates the morphology, texture, and internal structural heterogeneity of kidney stones by processing existing CT images, and identifies highly correlated features, thus solving the problem of low assessment accuracy caused by a single reference factor in the prior art.

[0007] To achieve the above objectives, the present invention provides the following technical solution.

[0008] A method for assessing the fragility of kidney stones based on CT images includes the following steps: CT images of proximal ureteral stones with border labels and the maximum fragment volume data after one extracorporeal shock wave lithotripsy (ESWL) to reflect the fragility of the stones were acquired to construct a training set. A dataset of original multimodal feature vectors is constructed based on multiple stone morphology features, stone texture features, and stone heterogeneity features extracted from CT images in the training set. A multimodal group autoencoder is constructed, comprising an encoder and a decoder, both of which are symmetric multilayer perceptrons. The first layer of the encoder is a group discovery layer, in which a group weight matrix is ​​set, with each row of the group weight matrix corresponding to a latent neuron. Based on the dataset of the original multimodal feature vectors, the mean squared error is used as the reconstruction error, and a group sparse regularization term related to the group weight matrix is ​​added to optimize the multimodal group autoencoder, obtaining the optimized group weight matrix. The group sparse regularization term is obtained by applying the L2,1 norm to each row of the group weight matrix and summing the L2,1 norms of all rows. Based on the optimized grouping weight matrix, the original features corresponding to the columns whose absolute values ​​in each row weight are greater than a preset threshold are extracted to construct feature groups after each row is grouped; the sum of the variances of each feature group on the dataset of the original multimodal feature vectors is used as the importance score, and the weight of each feature group is determined after normalization to construct the feature group weight vector; the mean value of the features in each feature group is used as the representative value in the dataset of the original multimodal feature vectors, and the feature group weight vector is used to perform weighted fusion to construct the fused feature vector; A multilayer perceptron regression model was constructed and trained based on the fused feature vectors and the corresponding maximum stone volume data in the training set to obtain an evaluation model for stone fragility. The CT image to be evaluated is obtained, and the fusion feature vector in the CT image is extracted based on the optimized grouping weight matrix. The fusion feature vector is then input into the evaluation model to obtain the evaluation result.

[0009] Preferably, constructing the original multimodal feature vector includes the following steps: Multiple stone morphological features, stone texture features, and stone heterogeneity features were extracted from CT images. Among them, the stone morphological features include stone height, stone cross-sectional diameter, maximum cross-sectional area, and stone volume; the stone heterogeneity features include the average HU value; and the stone texture features include skewness, kurtosis, and entropy.

[0010] The data on stone height and cross-sectional diameter are based on a binary mask. The calculation is performed; the maximum cross-sectional area data is obtained by triangulating the mask surface and calculating the sum of the areas of all triangular facets. The stone volume is determined based on the binary mask. The physical volume of a single voxel is determined; mask As a region of interest (ROI), the distribution characteristics of pixel intensity HU values ​​within this region are extracted from the regressed CT image, including the average HU. μ HU skewness γ kurtosis κ Entropy H Among them, skewness describes the asymmetry of the distribution, kurtosis describes the sharpness of the distribution, and entropy characterizes randomness. Multiple stone morphological features, stone texture features, and stone heterogeneity features are concatenated to form an original feature vector, which is then standardized to form a multimodal feature vector. .

[0011] Preferably, the multimodal group autoencoder includes an input layer, an encoder, and a decoder; wherein the encoder is a multilayer perceptron (MLP), the first layer is a group detection layer followed by a ReLU activation function and two fully connected layers, and the weight matrix is... ,in k It is the preset number of potential feature groups. The input feature dimension is defined as follows: the decoder is a multilayer perceptron (MLP) symmetrical to the encoder, which takes the latent representation as input, upsamples it step by step, and finally reconstructs the input feature vector.

[0012] Preferably, the training of the multimodal group autoencoder includes the following steps: Multimodal feature vectors extracted from the training set The dataset is used to train a multimodal grouped autoencoder based on a total loss function combining reconstruction loss and group sparsity regularization: The reconstruction loss is the mean squared error: ; The group sparse regularization term is: ; The total loss function is: .

[0013] Preferably, the construction of the fused feature vector includes the following steps: The weight matrix in the trained group discovery layer Processing is required for the first... j OK Calculate the mean of their absolute values. and standard deviation ; Set a threshold ,in, It's a scaling factor, set to 1.5. The absolute value of the weight is greater than... The original feature indices corresponding to the columns are grouped into a set. ;like If the set is empty, then discard the set.

[0014] For each non-empty feature group Determine the importance score of the group. : ; in, For training set The first in i The sample at the th t Values ​​on each feature; Group importance score for all groups Perform Softmax normalization to obtain the feature group weight vector: , k ′ represents the number of non-empty pairs; Calculate the group to which each sample belongs for each feature group. Mean of the original features : ; Through feature group weight vector The fused feature vector is obtained by weighting and concatenating the means of the original features. .

[0015] Preferably, the construction of the stone fragility assessment model includes the following steps: A multilayer perceptron regression model is constructed, comprising an input layer, three fully connected layers, and an output layer. Each fully connected layer is followed by a ReLU activation function and a Dropout layer. The output layer consists of a single neuron that uses a linear activation function to output the predicted maximum gravel volume. The model is trained based on the training set and the corresponding fused feature vectors. The mean squared error is used as the loss function, and L2 weight decay is added to prevent overfitting, thus obtaining a model for assessing the fragility of kidney stones.

[0016] This invention also proposes a CT image-based system for assessing the fragility of kidney stones, the system comprising: The training set construction module is used to acquire CT images of proximal ureteral stones with border labels and a training set of the maximum stone volume after one extracorporeal shock wave lithotripsy, which is used to reflect the stone fragility assessment index. The original multimodal feature vector construction module is used to construct a dataset of original multimodal feature vectors based on multiple stone morphology features, stone texture features, and stone heterogeneity features extracted from CT images in the training set. A multimodal group autoencoder construction module is used to construct a multimodal group autoencoder, including an encoder and a decoder, both of which are symmetric multilayer perceptrons. The first layer of the encoder is a group discovery layer, and a group weight matrix is ​​set in the group discovery layer, with each row of the group weight matrix corresponding to a latent neuron. Based on the dataset of the original multimodal feature vectors, the mean squared error is used as the reconstruction error, and a group sparse regularization term related to the group weight matrix is ​​added to optimize the multimodal group autoencoder to obtain the optimized group weight matrix. The group sparse regularization term is to apply the L2,1 norm to each row of the group weight matrix and sum the L2,1 norms of all rows. The fusion feature vector generation module is used to extract the original features corresponding to columns whose absolute values ​​in each row of the weights are greater than a preset threshold based on the optimized grouping weight matrix, and construct feature groups after grouping each row; the sum of the variances of each feature group on the dataset of the original multimodal feature vectors is used as the importance score, and the weight of each feature group is determined after normalization, thus constructing the feature group weight vector; the mean value of the features in each feature group is used as the representative value in the dataset of the original multimodal feature vectors, and the feature group weight vector is used to perform weighted fusion to construct the fusion feature vector; The stone fragility assessment model generation module is used to train a model for assessing stone fragility based on the fused feature vectors and the corresponding maximum stone volume data in the training set.

[0017] The present invention also proposes a computer-readable storage medium storing a data processing program, which, when executed by a processor, implements the aforementioned method for assessing the fragility of stones based on CT images.

[0018] The present invention also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned method for assessing the fragility of stones based on CT images.

[0019] The beneficial effects of this invention are: This invention proposes a method for assessing the fragility of kidney stones based on CT images. This method significantly improves the accuracy and interpretability of preoperative assessment of stone fragility during ESWL by constructing an unsupervised feature group discovery and weighted fusion model that integrates multi-dimensional stone features. The method introduces a multimodal grouped autoencoder for unsupervised feature group discovery. By applying L2,1 norm regularization, the model can automatically identify the inherent correlation structure between original features during reconstruction training. This automatic grouping mechanism requires no prior knowledge and is entirely data-driven, effectively uncovering high-order synergistic relationships between features. Subsequently, importance weights are calculated based on the variance of each group on the training set, and the features within each group are weighted and fused to generate new feature vectors with lower dimensionality but richer semantics. This fusion process not only achieves information condensation and noise reduction, but more importantly, each fused feature dimension has clear physical meaning or clinical interpretability, facilitating clinicians' understanding of the model's judgment criteria. Attached Figure Description

[0020] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.

[0021] Figure 2 This is a flowchart illustrating the execution steps of feature group extraction, weighting, and fusion in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0023] Example 1 Relying solely on morphological parameters to assess the fragility of kidney stones has significant limitations. In fact, the physical properties of kidney stones are multi-dimensional. Besides morphology, their internal structural characteristics, especially heterogeneity, are highly correlated with fragility. Stone heterogeneity reflects the spatial distribution of different density components within the stone. During stone formation, due to fluctuations in urine composition, infection, or metabolic abnormalities, stones often exhibit a layered structure or a multicentric growth pattern, leading to uneven internal density distribution. This heterogeneity produces significant mechanical effects under shock wave action: the different propagation speeds of shock waves at different density interfaces cause reflection, refraction, and focusing phenomena at the interfaces, thus creating a complex stress field within the stone. The stronger the heterogeneity and the greater the density difference, the more pronounced the local stress concentration at the interfaces, making it easier to initiate and propagate cracks, ultimately promoting stone fragmentation. Studies have shown that stones with obvious layered structures or porous features are significantly more fragile than stones with a homogeneous and dense structure.

[0024] Therefore, considering only morphological parameters while neglecting intrinsic physical properties such as heterogeneity results in low accuracy of existing methods for assessing the fragility of kidney stones, making it difficult to meet the needs of precise clinical decision-making. Stones with similar morphologies may exhibit completely different fragmentation behaviors due to differences in their internal structures, causing predictive models based on single morphological parameters to frequently fail in practical applications. This technical bottleneck suggests the necessity of incorporating texture features and heterogeneity indicators, reflecting the internal density distribution of kidney stones, into the assessment system. Constructing a comprehensive predictive model that integrates multi-dimensional features is essential for a more comprehensive and accurate characterization of stone fragility, providing a reliable basis for selecting clinical treatment strategies.

[0025] Therefore, this embodiment proposes a method for assessing the fragility of stones based on CT images, the main steps of which are as follows: Figure 1 As shown, the method includes: S1: Obtain CT images of proximal ureteral stones with border labels and the corresponding training set of the maximum stone volume after one extracorporeal shock wave lithotripsy.

[0026] S2: Construct the original multimodal feature vector based on multiple stone morphological features, stone texture features, and stone heterogeneity features extracted from CT images.

[0027] S3: Construct a multimodal group autoencoder, including an encoder and a decoder, which are symmetric multilayer perceptrons. The first layer of the encoder is a group discovery layer, and a group weight matrix is ​​set in the group discovery layer. Each row of the group weight matrix corresponds to a potential neuron. Based on the dataset, the mean squared error is used as the reconstruction error, and a group sparse regularization term related to the group weight matrix is ​​added to train the multimodal group autoencoder to obtain the trained multimodal group autoencoder. Among them, the group sparse regularization term is to apply the L2,1 norm to each row of the group weight matrix and then sum the L2,1 norms of all rows.

[0028] S4: Based on the group weight matrix in the trained multimodal grouped autoencoder, extract the original features corresponding to the columns whose absolute values ​​of the weights in each row are greater than a preset threshold, and construct feature groups for each row; determine the weight of each feature group after normalization based on the sum of the variances of each feature group on the training set, and construct the feature group weight vector; extract the mean value of the features in each feature group as the representative value, and construct the fused feature vector by weighted fusion of the feature group weight vector.

[0029] S5: Construct a multilayer perceptron regression model, including an input layer (for inputting fused feature vectors), three fully connected layers (each followed by a ReLU activation function and a Dropout layer), and an output layer (one neuron, using a linear activation function, outputting the predicted maximum stone volume). Train the model based on the training set and the corresponding fused feature vectors to obtain a model for assessing stone fragility.

[0030] S6: Obtain the CT image to be evaluated, extract the fusion feature vector from the CT image based on the optimized grouping weight matrix, and input it into the evaluation model to obtain the evaluation result.

[0031] Specifically: In S1, non-contrast CT scan data were collected from patients with proximal ureteral stones who met clinical diagnostic criteria. Simultaneously, the maximum diameter (or volume) of the largest stone fragment, assessed by imaging examinations (such as KUB plain film or CT), was accurately recorded from the medical record system after each patient's first extracorporeal shock wave lithotripsy (ESWL) as a true value label for regression prediction. :

[0032] ; In the formula, For the first i CT scan images of a patient with proximal ureteral stones; For image The border markings of the central stone region are usually represented by the coordinates of the bounding box; N This represents the total number of samples in the dataset.

[0033] Furthermore, radiologists use specialized annotation software to meticulously delineate the precise area of ​​the stone layer by layer on the axial images of each CT sequence, based on the border annotations, generating a binary mask. All data The sequences were randomly divided into mutually exclusive training and testing sets based on patient IDs. The raw Henlein unit (HU) values ​​were then truncated to a clinically relevant window to remove irrelevant tissue interference. Finally, each sequence was Z-score normalized to ensure a pixel intensity mean of 0 and a standard deviation of 1.

[0034] S1 is the data cornerstone of the entire project. Its core function is to acquire precisely paired "image-morphological annotation-clinical outcome" data. This includes the precise binary mask of the stones. It forms the basis for all subsequent morphological feature calculations, and the volume of the crushed stone... This is the objective and quantifiable gold standard for measuring the fragility of ESWL. The quality and consistency of the data directly determine the upper limit of the entire modeling task.

[0035] In S2, multiple stone morphological features, stone texture features, and stone heterogeneity features are extracted from CT images. Among them, stone morphological features include stone height, stone cross-sectional diameter, maximum cross-sectional area, and stone volume; stone heterogeneity features include the average HU value (characterizing stone attenuation, obtained from CT images, and belonging to heterogeneity features); stone texture features (characterizing the distribution of different density components in the stone), including skewness, kurtosis, and entropy.

[0036] Specifically, the stone height and cross-sectional diameter data in the stone morphology characteristics are based on binary masks. The calculations are performed using algorithms such as Marching Cubes to triangulate the mask surface based on the maximum cross-sectional area data. The sum of the areas of all triangular faces is calculated, and the stone volume is determined based on the binary mask. The physical volume of a single voxel (determined by the CT pixel pitch and layer thickness) is then determined. Next, the mask is... As a region of interest (ROI), the distribution characteristics of pixel intensity HU values ​​within this region are extracted from the regressed CT image, including the average HU. μ HU skewness γ kurtosis κ Entropy H Among them, skewness describes the asymmetry of the distribution, kurtosis describes the sharpness of the distribution, and entropy characterizes randomness.

[0037] Multiple stone morphological features, stone texture features, and stone heterogeneity features are concatenated to form an original feature vector, which is then standardized to form a multimodal feature vector. .

[0038] Furthermore, S3 constructs a multimodal grouped autoencoder, employing an unsupervised approach to discover the inherent grouping structure among the original features. Specifically:

[0039] A multimodal group autoencoder is constructed, comprising an input layer, an encoder, and a decoder. The encoder is a multilayer perceptron (MLP), with the first layer being a group detection layer followed by a ReLU activation function and two fully connected layers. The weight matrix is ​​as follows. ,in k It is the preset number of potential feature groups. The input feature dimension is denoted as . The decoder is a multilayer perceptron (MLP) symmetric to the encoder, which takes the latent representation as input, progressively upsamples it, and finally reconstructs the input feature vector.

[0040] Multimodal feature vectors extracted from the training set The dataset is used to train a multimodal grouped autoencoder based on a total loss function that combines reconstruction loss and group sparsity regularization.

[0041] The reconstruction loss is the mean squared error: ; The group sparse regularization term is: ; The total loss function is: .

[0042] This invention reconstructs a task-driven network to learn a compact representation of features, while utilizing L2,1 norm regularization as a structured sparsity constraint to force the network to learn a compact representation of features at the group discovery layer. This forms a "sparse row" pattern. After training, The input features corresponding to each row of non-zero elements naturally form a feature set with inherent correlations. This is more interpretable than manually defining or individually evaluating features.

[0043] In step S4, the feature grouping structure is parsed from the trained multimodal grouped autoencoder, and weights are assigned to each group based on the information content of each group. Finally, a new feature vector that integrates the original multimodal information and is more representative is generated.

[0044] Specific steps are as follows Figure 2 As shown: S4.1: Analyze the weight matrix in the trained group discovery layer. Processing is required for the first... j OK Calculate the mean of their absolute values. and standard deviation Set a threshold. ,in, It's a scaling factor, set to 1.5. The absolute value of the weight is greater than... The original feature indices corresponding to the columns are grouped into a set. .like If the set is empty, then discard the set.

[0045] S4.2: For each non-empty feature group Determine the importance score of the group. : ; in, For training set The first in i The sample at the th t The value on each feature.

[0046] S4.3: Group importance score for all groups Perform Softmax normalization to obtain the feature group weight vector: , k ′ represents the number of non-empty pairs.

[0047] S4.4: Calculate the group to which each sample belongs for each feature group. Mean of the original features : ; S4.5: Through the feature set weight vector The fused feature vector is obtained by weighting and concatenating the means of the original features. .

[0048] Step S4 of this invention makes the implicit grouping structure learned in S3 explicit and specific, and assigns importance weights to different groups based on data-driven criteria (within-group variance). The final fused feature vector has a lower dimension, and each dimension represents a combination of features with clinical or physical significance (such as "large volume high heterogeneity group" or "dense homogeneous small stone group"), which significantly improves the representational power of features and the interpretability of the model.

[0049] Finally, S5 proposes a regression model that accurately predicts the maximum fragmentation volume of stones after ESWL, which serves as the final model for assessing stone fragility.

[0050] Specifically, the model is a multilayer perceptron regression model, which includes an input layer (for inputting fused feature vectors), three fully connected layers (each followed by a ReLU activation function and a Dropout layer), and an output layer (one neuron using a linear activation function to output the predicted maximum stone volume). It is trained based on the training set and the corresponding fused feature vectors, using mean squared error (MSE) as the loss function, and adding L2 weight decay to prevent overfitting, thus obtaining a model for assessing the fragility of stones.

[0051] This invention aims to construct an automated and interpretable analysis system for predicting the fragility of proximal ureteral stones after extracorporeal shock wave lithotripsy (ESWL) based on CT images. The entire process begins with the construction of a high-quality, finely annotated dataset. CT images with precise boundary masks of the stones and corresponding labels for the maximum fragmented stone volume after ESWL are collected and rigorously preprocessed and segmented to lay a reliable foundation for supervised learning. Subsequently, the system enters a multimodal feature engineering phase, automatically extracting three types of quantitative features from the CT images and stone masks: morphological features describing the size and appearance of the stone (e.g., volume, sphericity), texture features characterizing the internal density distribution (e.g., skewness, entropy), and heterogeneity features reflecting overall density differences (e.g., mean HU value, standard deviation). These features are standardized and concatenated to form an initial high-dimensional feature vector.

[0052] The core innovation of the process lies in introducing an unsupervised feature structure discovery mechanism. By designing an autoencoder with a special grouping discovery layer and training it using reconstruction loss combined with group sparsity (L2,1 norm) regularization, the model is forced to automatically cluster high-dimensional original features into several feature groups while learning compact encoding. The original features corresponding to the non-zero elements in each row of the weight matrix constitute a group, whose physical meaning may correspond to potential stone type patterns such as "large and loose" or "small and dense." Next, the system analyzes the trained autoencoder, extracts specific feature groups based on weight thresholds, and calculates the importance weights of the features within each group based on the total variance of the features on the training set. Finally, a weighted average is used to fuse the original features of each sample into a low-dimensional fused feature vector representing the information of each group. This step significantly improves the semantic level and discriminative power of the features.

[0053] Finally, a multilayer perceptron regression model with fused features as input was constructed and trained to perform the final fragility prediction (indicating maximum stone volume). This model employs a fully connected network incorporating Dropout and L2 regularization, and utilizes strategies such as early stopping to ensure generalization performance. The entire implementation forms a closed loop: automatically mining clinically meaningful feature combinations from data and then using these combinations for accurate prediction not only provides an accurate assessment tool, but the interpretability of its groupings also helps deepen the understanding of the mechanisms influencing stone fragility, providing potential support for preoperative clinical decision-making.

[0054] The above is one embodiment of the CT image-based method for assessing the fragility of kidney stones. Based on the same idea, this embodiment also provides a corresponding CT image-based system for assessing the fragility of kidney stones. The training set construction module is used to acquire CT images of proximal ureteral stones with border labels and a training set of the maximum stone volume after one extracorporeal shock wave lithotripsy, which is used to reflect the stone fragility assessment index. The original multimodal feature vector construction module is used to construct a dataset of original multimodal feature vectors based on multiple stone morphology features, stone texture features, and stone heterogeneity features extracted from CT images in the training set. A multimodal group autoencoder construction module is used to construct a multimodal group autoencoder, including an encoder and a decoder, both of which are symmetric multilayer perceptrons. The first layer of the encoder is a group discovery layer, and a group weight matrix is ​​set in the group discovery layer, with each row of the group weight matrix corresponding to a latent neuron. Based on the dataset of the original multimodal feature vectors, the mean squared error is used as the reconstruction error, and a group sparse regularization term related to the group weight matrix is ​​added to optimize the multimodal group autoencoder to obtain the optimized group weight matrix. The group sparse regularization term is to apply the L2,1 norm to each row of the group weight matrix and sum the L2,1 norms of all rows. The fusion feature vector generation module is used to extract the original features corresponding to columns whose absolute values ​​in each row of the weights are greater than a preset threshold based on the optimized grouping weight matrix, and construct feature groups after grouping each row; the sum of the variances of each feature group on the dataset of the original multimodal feature vectors is used as the importance score, and the weight of each feature group is determined after normalization, thus constructing the feature group weight vector; the mean value of the features in each feature group is used as the representative value in the dataset of the original multimodal feature vectors, and the feature group weight vector is used to perform weighted fusion to construct the fusion feature vector; The stone fragility assessment model generation module is used to train a model for assessing stone fragility based on the fused feature vectors and the corresponding maximum stone volume data in the training set.

[0055] This embodiment also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 A method for assessing the fragility of stones based on CT images is provided.

[0056] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0057] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for evaluating stone friability based on a CT image, characterized by, Includes the following steps: CT images of proximal ureteral stones with border labels and the maximum fragment volume data after one extracorporeal shock wave lithotripsy (ESWL) to reflect the stone's fragility were acquired to construct a training set. A dataset of original multimodal feature vectors is constructed based on multiple stone morphology features, stone texture features, and stone heterogeneity features extracted from CT images in the training set. A multimodal group autoencoder is constructed, comprising an encoder and a decoder, both of which are symmetric multilayer perceptrons. The first layer of the encoder is a group discovery layer, in which a group weight matrix is ​​set, with each row of the group weight matrix corresponding to a latent neuron. Based on the dataset of the original multimodal feature vectors, the mean squared error is used as the reconstruction error, and a group sparse regularization term related to the group weight matrix is ​​added to optimize the multimodal group autoencoder, obtaining the optimized group weight matrix. The group sparse regularization term is obtained by applying the L2,1 norm to each row of the group weight matrix and summing the L2,1 norms of all rows. Based on the optimized grouping weight matrix, the original features corresponding to the columns whose absolute values ​​in each row weight are greater than a preset threshold are extracted to construct feature groups after each row is grouped; the sum of the variances of each feature group on the dataset of the original multimodal feature vectors is used as the importance score, and the weight of each feature group is determined after normalization to construct the feature group weight vector; the mean value of the features in each feature group is used as the representative value in the dataset of the original multimodal feature vectors, and the feature group weight vector is used to perform weighted fusion to construct the fused feature vector; A multilayer perceptron regression model was constructed and trained based on the fused feature vectors and the corresponding maximum stone volume data in the training set to obtain an evaluation model for stone fragility. The CT image to be evaluated is obtained, and the fusion feature vector in the CT image is extracted based on the optimized grouping weight matrix. The fusion feature vector is then input into the evaluation model to obtain the evaluation result.

2. The CT image-based stone friability assessment method according to claim 1, characterized by, The construction of the original multimodal feature vector includes the following steps: Multiple stone morphological features, stone texture features, and stone heterogeneity features were extracted from CT images; wherein, the stone morphological features include stone height, stone cross-sectional diameter, maximum cross-sectional area, and stone volume; the stone heterogeneity features include the average HU value; and the stone texture features include skewness, kurtosis, and entropy. The data on stone height and cross-sectional diameter are based on a binary mask. The calculation is performed; the maximum cross-sectional area data is obtained by triangulating the mask surface and calculating the sum of the areas of all triangular facets. The stone volume is determined based on the binary mask. The physical volume of a single voxel is determined; mask As the region of interest, the distribution features of pixel intensity HU values ​​within the mapped regression-normalized CT image are extracted, including the average HU value. μ HU skewness γ kurtosis κ Entropy H Among them, skewness describes the asymmetry of the distribution, kurtosis describes the sharpness of the distribution, and entropy characterizes randomness. Multiple stone morphological features, stone texture features, and stone heterogeneity features are concatenated to form an original feature vector, which is then standardized to form a multimodal feature vector. .

3. The method for assessing the fragility of kidney stones based on CT images according to claim 1, characterized in that, The multimodal group autoencoder includes an input layer, an encoder, and a decoder; wherein the encoder is a multilayer perceptron (MLP), the first layer is a group detection layer followed by a ReLU activation function and two fully connected layers, and the weight matrix is ​​as follows. ,in k It is the preset number of potential feature groups. The input feature dimension is defined as follows: the decoder is a multilayer perceptron (MLP) symmetrical to the encoder, which takes the latent representation as input, upsamples it step by step, and finally reconstructs the input feature vector.

4. The method for assessing the fragility of kidney stones based on CT images according to claim 3, characterized in that, The training of the multimodal group autoencoder includes the following steps: Multimodal feature vectors extracted from the training set The dataset is used to train a multimodal grouped autoencoder based on a total loss function combining reconstruction loss and group sparsity regularization: The reconstruction loss is the mean squared error: ; The group sparse regularization term is: ; The total loss function is: .

5. The method for assessing the fragility of kidney stones based on CT images according to claim 1, characterized in that, The construction of the fused feature vector includes the following steps: The weight matrix in the trained group discovery layer Processing is required for the first... j OK Calculate the mean of their absolute values. and standard deviation ; Set a threshold ,in, It is a proportionality coefficient, set to 1.5; the absolute value of the weight is greater than... The original feature indexes corresponding to the columns are grouped into a set. ;like If the set is empty, then discard the set. For each non-empty feature group Determine the importance score of the group. : ; in, For training set The first in i The sample at the th t Values ​​on each feature; Group importance score for all groups Perform Softmax normalization to obtain the feature group weight vector: , k ′ represents the number of non-empty pairs; Calculate the group to which each sample belongs for each feature group. Mean of the original features : ; Through feature group weight vector The fused feature vector is obtained by weighting and concatenating the means of the original features. .

6. The method for assessing the fragility of kidney stones based on CT images according to claim 1, characterized in that, The construction of the stone fragility assessment model includes the following steps: A multilayer perceptron regression model is constructed, comprising an input layer, three fully connected layers, and an output layer. Each fully connected layer is followed by a ReLU activation function and a Dropout layer. The output layer consists of a single neuron that uses a linear activation function to output the predicted maximum gravel volume. The model is trained based on the training set and the corresponding fused feature vectors, using mean squared error as the loss function and adding L2 weight decay to prevent overfitting, thus obtaining a model for assessing the fragility of kidney stones.

7. A system for assessing the fragility of kidney stones based on CT images, characterized in that, The system includes: The training set construction module is used to acquire CT images of proximal ureteral stones with border labels and a training set of the maximum stone volume after one extracorporeal shock wave lithotripsy, which is used to reflect the stone fragility assessment index. The original multimodal feature vector construction module is used to construct a dataset of original multimodal feature vectors based on multiple stone morphology features, stone texture features, and stone heterogeneity features extracted from CT images in the training set. A multimodal group autoencoder construction module is used to construct a multimodal group autoencoder, including an encoder and a decoder, both of which are symmetric multilayer perceptrons. The first layer of the encoder is a group discovery layer, and a group weight matrix is ​​set in the group discovery layer, with each row of the group weight matrix corresponding to a latent neuron. Based on the dataset of the original multimodal feature vectors, the mean squared error is used as the reconstruction error, and a group sparse regularization term related to the group weight matrix is ​​added to optimize the multimodal group autoencoder to obtain the optimized group weight matrix. The group sparse regularization term is to apply the L2,1 norm to each row of the group weight matrix and sum the L2,1 norms of all rows. The fusion feature vector generation module is used to extract the original features corresponding to columns whose absolute values ​​in each row of the weights are greater than a preset threshold based on the optimized grouping weight matrix, and construct feature groups after grouping each row; the sum of the variances of each feature group on the dataset of the original multimodal feature vectors is used as the importance score, and the weight of each feature group is determined after normalization, thus constructing the feature group weight vector; the mean value of the features in each feature group is used as the representative value in the dataset of the original multimodal feature vectors, and the feature group weight vector is used to perform weighted fusion to construct the fusion feature vector; The stone fragility assessment model generation module is used to train a model for assessing stone fragility based on the fused feature vectors and the corresponding maximum stone volume data in the training set.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a data processing program, which, when executed by a processor, implements the method for assessing the fragility of stones based on CT images as described in any one of claims 1 to 6.

9. A computer device, characterized in that, The invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method for assessing the fragility of stones based on CT images as described in any one of claims 1 to 6.