A deep learning-based bladder volume prediction method

By combining deep learning with image and text information, a bladder volume prediction method is developed. This method automatically selects features and integrates multimodal data, solving the adaptability problem of traditional methods in complex scenarios and achieving efficient and accurate prediction of bladder volume and risk level.

CN119722781BActive Publication Date: 2025-10-21CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411788533.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-10-21
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Traditional methods for measuring bladder volume are poorly adaptable in complex scenarios, rely heavily on operational experience, and are difficult to meet the high-efficiency and accurate diagnostic needs of modern medicine.

Method used

A deep learning-based bladder volume prediction method is adopted. Through an image information extraction module, an LCIA attention mechanism module, a text information embedding module, and an RFEA multimodal feature fusion module, combined with a dynamic multi-task loss function, features are automatically selected and image and text information is fused to improve prediction accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of bladder volume prediction and risk level classification, can handle image blur and multimodal data in complex scenarios, reduces training time, and provides more accurate bladder function assessment and disease risk prediction support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722781B_ABST
    Figure CN119722781B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of medical image processing, and particularly relates to a bladder volume prediction method based on deep learning, which comprises the following steps: establishing a bladder volume and risk level multi-task prediction model, training input images and text information, and predicting the bladder volume and risk level through the trained model. By fusing image and text multi-modal information, the application designs an efficient deep learning network architecture, significantly improves the accuracy and robustness of bladder volume prediction and risk level classification, the image information extraction module combines with the LCIA attention mechanism, effectively captures the details of local and global features, the text information embedding module strengthens the expression ability of non-image information through low-dimensional representation, and the RFEA multi-modal feature fusion module dynamically adjusts the importance of multi-modal data, so that the image and text features can synergistically act, thereby improving the overall prediction performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and specifically relates to a bladder volume prediction method based on deep learning. Background Art

[0002] Bladder volume measurement is a crucial foundation for the diagnosis and treatment of urinary tract diseases and is widely used in clinical settings such as urinary retention, bladder dysfunction, prostate disease, and postoperative rehabilitation assessment. Traditional measurement methods primarily include catheterization and ultrasound scanning. Catheterization involves inserting a catheter to directly empty the bladder and measure urine volume. While highly accurate, it is highly invasive and can lead to patient discomfort and infection. Ultrasound scanning, however, has become the mainstream method due to its non-invasive, rapid, and real-time nature. However, its accuracy is significantly affected by operator experience and image quality, particularly in the presence of anatomical variations, body shape differences, or image blur, which can lead to significant errors. With the advancement of medical technology, the demand for automation and intelligent technology is becoming increasingly prominent. Traditional measurement methods are unable to meet the demands of rapid, accurate, and large-scale diagnosis and treatment. In this context, deep learning technology, with its powerful data-driven capabilities and feature extraction advantages, offers an innovative approach to bladder volume measurement. Deep learning algorithms can automatically extract the bladder margin and calculate bladder capacity from ultrasound images, eliminating reliance on manual experience. This not only improves measurement efficiency and accuracy, but also can handle blurry images and multimodal data in complex scenarios. Furthermore, multimodal learning, incorporating information such as the patient's age, gender, and medical history, further enhances the model's diagnostic capabilities and generalizability. Therefore, the application of deep learning in bladder volume measurement not only addresses the shortcomings of traditional methods but also provides important support for the development and promotion of intelligent medical devices, demonstrating significant clinical and application value.

[0003] Therefore, the application of deep learning in bladder volume measurement not only addresses the shortcomings of traditional methods but also provides important support for the development and promotion of intelligent medical devices, demonstrating significant clinical and application value. However, traditional methods have poor adaptability in complex scenarios and are highly dependent on operator experience, making them unable to meet the needs of modern medicine for high efficiency and accurate diagnosis. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a bladder volume prediction method based on deep learning, comprising:

[0005] Establish a multi-task prediction model for bladder volume and risk level, train the input image and text information, and use the trained model to predict bladder volume and risk level;

[0006] The bladder volume prediction model includes: an image information extraction module, an LCIA attention mechanism module, a text information embedding module, and an RFEA multimodal feature fusion module;

[0007] The training of the bladder volume prediction model comprises:

[0008] S1: Apply feature selection algorithms to automatically select the most relevant feature subsets from multimodal datasets to improve model prediction performance and training efficiency of multimodal data;

[0009] S2: The image in the most relevant feature subset is input into the image information extraction module and the LCIA attention mechanism module for processing in sequence to obtain image feature information;

[0010] The image information extraction module obtains basic image features and provides initial information input for subsequent modules;

[0011] The LCIA attention mechanism module combines local area features and channel features in the basic image features to more accurately extract effective information;

[0012] S3: The text information in the most relevant feature subset is transformed into low-dimensional text through the text information embedding module to obtain text feature information. This allows the text information to be processed in the neural network and work together with image features, thereby improving the model's accuracy in bladder volume prediction and risk level classification.

[0013] S4: Through the RFEA multimodal feature fusion module, image feature information and text feature information are weighted and fused to improve the performance of the model;

[0014] S5: We designed a dynamic multi-task loss function to automatically balance the model’s attention between the loss values ​​of different tasks, enabling the model to better cope with volume scale differences and category imbalance, thereby improving overall performance and robustness.

[0015] Beneficial effects of the present invention:

[0016] By fusing image and text multimodal information, the present invention designs an efficient deep learning network architecture, significantly improving the accuracy and robustness of bladder volume prediction and risk level classification; the image information extraction module combines the LCIA attention mechanism to effectively capture the details of local regions and global features; the text information embedding module enhances the expressive power of non-image information through low-dimensional representation; the RFEA multimodal feature fusion module dynamically adjusts the importance of multimodal data, enabling image and text features to work synergistically, thereby improving overall prediction performance; the dynamic multi-task loss function adaptively balances the losses of bladder volume prediction and risk level classification, significantly solving the problems of volume scale difference and category imbalance, and enhancing the model's adaptability to practical applications.

[0017] In addition, the present invention introduces a multimodal feature selection method (MMFS), which improves data analysis efficiency and prediction capabilities by automatically screening the most relevant features. Compared with traditional methods, the present invention can more comprehensively capture bladder-related lesion information, while optimizing computational efficiency and reducing training time. In actual clinical applications, the present invention provides more accurate support for bladder function assessment, disease risk prediction, and treatment plan formulation, opening up a new path to improving the intelligence level and diagnostic quality of medical image analysis, and has important scientific research value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0020] A bladder volume prediction method based on deep learning, comprising:

[0021] Establish a multi-task prediction model for bladder volume and risk level, train the input image and text information, and use the trained model to predict bladder volume and risk level;

[0022] The bladder volume prediction model includes: an image information extraction module, an LCIA attention mechanism module, a text information embedding module, and an RFEA multimodal feature fusion module;

[0023] The training of the bladder volume prediction model comprises:

[0024] S1: Applying a feature selection algorithm to automatically select the most relevant feature subset from a multimodal dataset, thereby improving model prediction performance and training efficiency of the multimodal data; the multimodal dataset consists of bladder data;

[0025] S2: The image in the most relevant feature subset is input into the image information extraction module and the LCIA attention mechanism module for processing in sequence to obtain image feature information;

[0026] The image information extraction module obtains basic image features and provides initial information input for subsequent modules;

[0027] The LCIA attention mechanism module combines local area features and channel features in the basic image features to more accurately extract effective information;

[0028] S3: The text information in the most relevant feature subset is transformed into low-dimensional text through the text information embedding module to obtain text feature information. This allows the text information to be processed in the neural network and work together with image features, thereby improving the model's accuracy in bladder volume prediction and risk level classification.

[0029] S4: Through the RFEA multimodal feature fusion module, image feature information and text feature information are weighted and fused to improve the performance of the model;

[0030] S5: We designed a dynamic multi-task loss function to automatically balance the model’s attention between the loss values ​​of different tasks, enabling the model to better cope with volume scale differences and category imbalance, thereby improving overall performance and robustness.

[0031] Step 1: Before model training, the multimodal feature selection method (MMFS) is first applied to automatically screen the input multimodal data. By analyzing the correlation between images, text, and other patient information, the most predictive features for bladder volume prediction and risk classification are selected. This reduces the interference of redundant features on model training, improves computational efficiency, and enhances final prediction performance.

[0032] Step 2: The image information is fed into the image feature extraction module. The core task of this module is to extract deep-level feature information from the bladder ultrasound image, supporting subsequent bladder volume prediction and risk classification. Using a multi-layer convolutional neural network (CNN), the image feature extraction module gradually extracts key information from the input image and forms a feature vector that effectively represents the image content.

[0033] Step 2.1: Input the output of the image information extraction module into the LCIA attention mechanism module. The core task of the LCIA attention mechanism module is to improve the expressiveness of image features. This module combines local region features and channel features of the image through weighted operations, allowing the model to focus more on key parts relevant to the task.

[0034] Step 3: The corresponding text information is simultaneously input into the text information embedding module. The main function of the text information embedding module is to convert the patient's unstructured text information (such as gender, age, medical history, region, etc.) into a low-dimensional vector representation, so that it can be integrated and processed together with the image features in the deep learning model. In this way, the model can simultaneously utilize the characteristics of image and text data, comprehensively considering the patient's multi-dimensional background information, thereby improving the accuracy of bladder volume prediction and the effectiveness of risk classification.

[0035] The text information is converted into a low-dimensional vector representation through the text information embedding module, including:

[0036]

[0037] Among them, V(t i ) represents the low-dimensional vector representation obtained by transformation, V(w j ) represents the word embedding vector, w j Represents words in the symptom text description, n i Represents text t i The number of words in V(t i ) represents the text t i The vector representation of .

[0038] Step 4: The output from the text information embedding module and the output from step 2.1 are simultaneously input into the RFEA multimodal feature fusion module. The RFEA multimodal feature fusion module is responsible for weighted fusion of features from the image extraction module and the text information embedding module. The fused features are processed by two task branches: the bladder volume prediction branch uses a fully connected layer to regress the output volume value; and the risk level prediction branch uses a fully connected layer and a softmax activation function for classification. Through this module, the model can comprehensively utilize information from multiple data sources to improve the prediction performance of bladder volume and risk level. The weighted fusion of different modalities helps the model better handle the complementarity of image and text information, providing more comprehensive diagnostic support.

[0039] Step 5: A dynamic multi-task loss function is designed to balance the losses of different tasks. For the bladder volume prediction and risk level classification tasks, the loss function can automatically adjust the focus according to the nature and objectives of the task, ensuring that the model can achieve the best performance when dealing with volume scale differences and category imbalance. Through multi-task learning, the model can not only complete the prediction of multiple tasks, but also share useful knowledge between tasks, thereby improving overall performance and robustness. The dynamic multi-task loss function is expressed as:

[0040] L all =λ1L v+λ2L r

[0041] L v =αNAME+(1-α)MSE

[0042]

[0043] Among them, L all represents the dynamic multi-task loss function, L v represents the bladder volume prediction loss function, L r represents the risk level classification loss function, λ1 and λ2 represent the dynamic weights of bladder volume prediction and risk level classification prediction, α represents the weight for controlling the mean absolute error and mean square error loss, NAME represents the standardized mean absolute error, and MSE represents the mean square error; y i represents the true volume, y' i represents the prediction volume, ∈ represents a small constant to prevent the denominator from being zero, N represents the number of samples when calculating the loss, and C represents the number of categories; p ic represents the predicted probability that the i-th sample belongs to the c-th class, α c represents the class weight, and γ represents the focusing factor that reduces the loss of easy-to-classify samples.

[0044] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A bladder volume prediction method based on deep learning, characterized in that: include: Establish a multi-task prediction model for bladder volume and risk level, train the input image and text information, and use the trained model to predict bladder volume and risk level; The multi-task prediction model for bladder volume and risk level includes: an image information extraction module, an LCIA attention mechanism module, a text information embedding module, and an RFEA multimodal feature fusion module; Training of the bladder volume prediction model includes: S1: Apply multimodal feature optimization selection methods to automatically select the most relevant feature subsets from multimodal datasets, improving model prediction performance and training efficiency of multimodal data; Apply multimodal feature optimization selection methods to automatically select the most relevant feature subsets from multimodal datasets, improving model prediction performance and training efficiency of multimodal data, including: Initializing feature subsets: Randomly selecting multiple feature subsets from a multimodal dataset based on the data dimensions and feature types to ensure that different feature combinations are covered; the multimodal dataset includes: images, text, and clinical data; Evaluate and select the optimal subset: Use the mean square error in the regression task and the accuracy in the classification task as evaluation criteria to evaluate the performance of each feature subset and select the feature subset with the best fitness value as the basis for subsequent optimization; Optimize feature sets: By combining features and removing redundancy, we can optimize feature subsets. The optimized feature subsets will focus more on the core information of the task, improving the model's predictive capabilities and computational efficiency. Iterative optimization: gradually optimize the feature subset through multiple iterations. After each iteration, adjust the feature subset according to the new evaluation results and fitness until the predetermined performance target or the maximum number of iterations is achieved; S2: The image in the most relevant feature subset is input into the image information extraction module and the LCIA attention mechanism module for processing in sequence to obtain image feature information; The image information extraction module obtains basic image features and provides initial information input for subsequent modules; The LCIA attention mechanism module combines local area features and channel features in the basic image features to more accurately extract effective information; S3: The text information in the most relevant feature subset is transformed into low-dimensional text through the text information embedding module to obtain text feature information. This allows the text information to be processed in the neural network and work together with image features, thereby improving the model's accuracy in bladder volume prediction and risk level classification. The text information in the most relevant feature subset is transformed into a low-dimensional form through the text information embedding module to obtain text feature information, including: Among them, V(t i ) represents the low-dimensional vector representation obtained by transformation, V(w j ) represents the word embedding vector, w j Represents words in the symptom text description, n i Represents text t i The number of words in V(t i ) represents the text t i Vector representation of ; S4: Through the RFEA multimodal feature fusion module, the image feature information and text feature information are weighted and fused to complete the bladder volume and risk level prediction; The RFEA multimodal feature fusion module performs weighted fusion of image feature information and text feature information to complete bladder volume and risk level prediction, including: Image and text information stitching: stitching image feature information and text feature information together to obtain a joint feature vector. The stitched feature contains global information from both the image and the text, providing a basis for subsequent correlation analysis. Correlation analysis: The joint feature vector is input into the fully connected layer for processing, and a correlation weight matrix is ​​calculated to represent the correlation between image and text information, and provide a weighted basis for subsequent feature fusion; Multimodal information fusion: Image feature information and text feature information are weighted and fused through a correlation weight matrix to ensure that they are appropriately combined according to their importance in specific tasks; The fused features are processed by two task branches: the bladder volume prediction branch uses a fully connected layer to regress and output the volume value; the risk level prediction branch uses a fully connected layer and a softmax activation function to achieve classification; S5: We designed a dynamic multi-task loss function to automatically balance the model's attention between the loss values ​​of different tasks. This allows the model to better cope with volume scale differences and class imbalance, improving overall performance and robustness. The dynamic multi-task loss function includes: L all =λ1L v +λ2L r L v =αNAME+(1-α)MSE Among them, L all represents the dynamic multi-task loss function, L v represents the bladder volume prediction loss function, L r represents the risk level classification loss function, λ1 and λ2 represent the dynamic weights of bladder volume prediction and risk level classification prediction, α represents the weight for controlling the mean absolute error and mean square error loss, NAME represents the standardized mean absolute error, and MSE represents the mean square error; y i represents the true volume, y' i represents the prediction volume, ∈ represents a small constant to prevent the denominator from being zero, N represents the number of samples when calculating the loss, and C represents the number of categories; p ic represents the predicted probability that the i-th sample belongs to the c-th class, α c represents the class weight, and γ represents the focusing factor that reduces the loss of easy-to-classify samples.

2. The method for predicting bladder volume based on deep learning according to claim 1, characterized in that: The image information extraction module consists of five cascaded convolutional layers, which extracts and refines features layer by layer, gradually transforming low-level edge and texture information into high-level global semantic features to obtain basic image features.

3. The method for predicting bladder volume based on deep learning according to claim 1, characterized in that: The LCIA attention mechanism module combines local region features and channel features in the basic image features to more accurately extract effective information, including: Local area attention calculation: The basic features of the image are divided into multiple regions to obtain local area features. The local area features are weighted by dynamically generating convolution kernels to enhance the feature expression ability of key areas. Channel attention calculation: Use global pooling and a fully connected network to weight each channel of the basic features of the image, generate the weight of each channel, and highlight the channel features that contribute to the task; Region and channel interaction: Combine local region attention and channel attention to generate joint weights to enhance multi-dimensional feature expression and obtain the final image feature information.

Citation Information

Patent Citations

  • Named entity recognition method based on comparative learning and multi-modal semantic interaction

    CN117574904A

  • Intelligent lung cancer metastasis prediction system based on GCAVE-GAN and multi-mode fusion

    CN118628462A