RNN-based pulmonary artery segmentation method, system and device, and medium
By combining the RNN method of U-Net, Mamba network and nmODE module, a small amount of reliable labeled data is used to perform pulmonary artery segmentation, the problem of low segmentation accuracy caused by incomplete label data is solved, the segmentation accuracy of small peripheral branches and the robustness of the model are improved, and it is suitable for real-time diagnosis.
Patent Information
- Application Number
- CN202510821989.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In the prior art, difficulty in labeling peripheral branches of pulmonary artery leads to incomplete label data, affecting the accuracy of pulmonary artery segmentation, especially in poor segmentation effect of tiny peripheral branches.
The pulmonary artery segmentation method based on RNN is adopted, and the U-Net network and Mamba network are combined with the nmODE module to train through a small amount of reliable annotation data, combined with the EM algorithm to infer potential labels, and to construct scoring and annotation propensity functions to improve segmentation accuracy.
It significantly improves the segmentation accuracy of the small peripheral branches, reduces the annotation workload, enhances the robustness of the model to complex structures, and is suitable for real-time or near-real-time clinical applications, supporting the accurate diagnosis of pulmonary artery-related diseases.
Smart Images

Figure CN120339629A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, relates to pulmonary artery segmentation, and particularly relates to a pulmonary artery segmentation method, system, device and medium based on RNN. Background Art
[0002] The pulmonary artery is the main blood vessel that transports blood from the right ventricle to the lungs, and its tree-like structure is of great significance anatomically. The accurate segmentation of the pulmonary artery is not only crucial for the diagnosis and treatment of diseases such as pulmonary arterial hypertension (PAH), but also plays a key role in evaluating the function, pathological changes of pulmonary blood circulation, and the formulation of surgical plans.
[0003] In recent years, with the rapid development of artificial intelligence technology, a large number of medical image segmentation methods based on deep learning have been proposed. The most classic one is the U-Net structure, which consists of an encoder, a decoder, and skip connections. The encoder part captures the features and context of the input image through a series of convolutional and downsampling operations, gradually reducing the size of the feature map; the decoder, on the contrary, gradually enlarges the feature map to the original size through a series of convolutional and upsampling operations to restore spatial information; the skip connections directly connect the corresponding layers of the decoder and the encoder, allowing the network to retain high-resolution information from the encoder feature map, which helps to retain details during the segmentation process. The U-Net structure is still the mainstream structure of medical image segmentation models to date, including its application in pulmonary artery segmentation.
[0004] The invention patent application with the application number 202311359989.0 discloses a pulmonary artery segmentation method and system based on context attention and feature fusion, which includes the following steps: In the training stage, obtain the original CT image and perform image preprocessing, extract image blocks, construct and train a 3D U-Net network based on context attention and feature fusion. In the inference stage, obtain the CT image to be inferred and perform image preprocessing, extract image blocks, record the CT image to which the block belongs and its position in the CT image, obtain the pulmonary artery segmentation prediction value based on the trained 3D U-Net network, combine the pulmonary artery segmentation prediction values of all image blocks belonging to the same CT to obtain the pulmonary artery segmentation prediction value of the whole CT image, and use hard segmentation processing to extract the largest connected component of the pulmonary artery to obtain the segmentation result; Among them, the 3D U-Net network includes: an encoder, a decoder, and a skip connection structure, and the encoder is connected to the decoder; The encoder includes a context attention convolution block and a downsampling operation module, and extracts abstract features from the input image block through convolution and downsampling operations, reduces the size of the feature map, increases the number of channels of the feature map, and obtains the feature map; The decoder includes a context attention convolution block and an upsampling operation module, and restores the feature map to the same size as the input image block through convolution and upsampling; The skip connection structure connects the low-level feature map in the encoder to the high-level feature map in the decoder.
[0005] Similar to the above-mentioned invention patent application, although the prior art can achieve the segmentation of the pulmonary artery through the U-Net network. However, it still faces some problems in the segmentation of the pulmonary artery tree structure, especially in the segmentation of small peripheral branches. The peripheral branches of the pulmonary artery are usually sparsely distributed and have subtle changes, making them show a low contrast in the image, and it is often difficult to ensure the accuracy of the annotation of small branches. Due to the difficulty in annotating the peripheral branches of the pulmonary artery in the prior art, most of the annotation data of the pulmonary artery tree structure in medical images is incomplete. Therefore, how to solve the problem of incomplete label data in the segmentation of the pulmonary artery tree structure and solve the problem of low segmentation accuracy of the pulmonary artery (especially small peripheral branches) caused by incomplete label data has become a major challenge in the current analysis of pulmonary artery medical images and urgently needs to be solved. Summary of the Invention
[0006] The purpose of the present invention is to: In order to solve the technical problem that the incomplete label data affects the accuracy of pulmonary artery segmentation due to the difficulty in annotating the peripheral branches of the pulmonary artery in the prior art, provide a pulmonary artery segmentation method, system, device and medium based on RNN, which uses a small amount of reliable annotations to segment the pulmonary artery tree structure, effectively improves the segmentation accuracy of small peripheral branches, and reduces the work burden of annotators.
[0007] To achieve the above object, the present invention specifically adopts the following technical solutions: A pulmonary artery segmentation method based on RNN, comprising the following steps: Step S1, obtaining sample data; Obtain lung CT sample images covering different slices and branch structures, and label the pulmonary artery main trunk and some branches in some lung CT sample images to obtain labeled lung CT sample images and unlabeled lung CT sample images; Step S2, constructing a pulmonary artery segmentation model; The pulmonary artery segmentation model includes a U-Net network and two Mamba networks with selective state spaces. The nmODE module is introduced in the intermediate high-dimensional representation stage of the U-Net network. One Mamba network is used to construct a scoring function for predicting the probability that each voxel is a foreground, and the other Mamba network is used to construct a labeling tendency function for estimating the possibility that the voxel is manually labeled as a foreground; The lung CT image is input into the U-Net network, and the original features are extracted after the downsampling stage. The original features are dynamically mapped by the nmODE module and then converted into a feature representation with long-term memory. After the upsampling stage, they are respectively input into the two Mamba networks, and the Mamba networks output the predicted true structure and labeling status; Step S3, training the pulmonary artery segmentation model; Use the labeled lung CT sample images and unlabeled lung CT sample images in step S1 to train the pulmonary artery segmentation model constructed in step S2; Step S4, real-time segmentation; Obtain the lung CT image to be segmented and input it into the pulmonary artery segmentation model trained in step S3, and the pulmonary artery segmentation model outputs the pulmonary artery segmentation result.
[0008] Furthermore, data augmentation is performed on the lung CT sample images obtained in step S1. The specific method is as follows: For the labeled lung CT sample images, 50% of the lung CT sample images are horizontally flipped, and all or part of the remaining lung CT sample images are randomly rotated, and the rotation angle range is from -15° to 15°; For the unlabeled lung CT sample images, different regions of the image are randomly selected for cropping, the local structure in the image is changed through non-linear transformation, and the image is partially randomly occluded.
[0009] Further, in step S2, the U-Net network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a first nmODE module, and a second nmODE module. The image is input into the first convolutional layer, and the output of the first convolutional layer is input into the second convolutional layer after downsampling. The output of the second convolutional layer is input into the third convolutional layer after downsampling. The output of the third convolutional layer is input into the fourth convolutional layer after downsampling. The output of the fourth convolutional layer is input into the fifth convolutional layer after downsampling. The fifth convolutional layer is combined with the first nmODE module, and the original features output by the fifth convolutional layer are dynamically mapped by the first nmODE module and then converted into feature representations with long-term memory ; Feature representation After being upsampled, it is input into the sixth convolutional layer. The sixth convolutional layer is combined with the second nmODE module, and the features output by the sixth convolutional layer are dynamically mapped by the second nmODE module and then converted into feature representations with long-term memory ; Feature representation After being upsampled, it is input into the seventh convolutional layer. The output of the seventh convolutional layer is input into the eighth convolutional layer after upsampling. The output of the eighth convolutional layer is input into the ninth convolutional layer after upsampling. The output of the ninth convolutional layer is used as the input of two Mamba networks
[0010] Furthermore, the ordinary differential equations of the first nmODE module and the second nmODE module are as follows: ; ; where represents a function related to time t, represents the output function of the derivative function; represents a constant used to control the influence of external input on the mapping stability of the model; represents the external input; represents learnable weight parameters used to adjust the influence of input features on the final output; represents the input features; represents the bias term used to adjust the input features when calculating the external input influence
[0011] Further, in step S2, the Mamba network includes a first linear projection layer, a tenth convolutional layer, a second linear projection layer, a convolutional layer, an activation function, a selective SSM layer, and a third linear projection layer The output of the U-Net network is used as the input of the Mamba network and is respectively input into the first linear projection layer and the second linear projection layer. The output of the first linear projection layer is sequentially passed through the tenth convolutional layer, the activation function, and the selective SSM layer, and then concatenated with the output of the second linear projection layer as the input of the third linear projection layer. The third linear projection layer outputs the predicted true structure and the annotation status .
[0012] Furthermore, in step S3, when training the pulmonary artery segmentation model, the total loss function is expressed as: ; ; ; wherein, , both represent weights, represents the scoring loss, represents the propensity loss, represents the predicted label of the th sample, represents the scoring function, represents the annotation propensity function, represents the th sample's input feature, represents the th sample's annotation status, , represent the optimized parameters of the two loss functions, represents the number of samples.
[0013] Even further, in step S3, when training the pulmonary artery segmentation model, the two Mamba networks are calculated iteratively through the EM algorithm. The E-step is responsible for inferring the posterior probability of the labels, and the M-step updates the parameters by maximizing the likelihood function. Specifically: In the E-step, the calculation formula for the posterior probability is: ; wherein, represents the conditional probability that the label is given the feature , represents the annotation status probability of the feature ; represents the conditional probability of the annotation status, indicating whether the feature is labeled as the positive class; In the M-step, according to the posterior probability , the maximum log-likelihood function is used to update the model parameters, and the goal is to maximize the joint probability of the labels and the probability of the annotation status; the formula for the maximization process is expressed as:
[0014] ; Among them, represents the log-likelihood function of the parameter model, represents the number of samples, represents the th sample's true label estimated probability, represents the th sample's true label expected value, represents the conditional probability of the annotation status, represents the given input feature and parameter for the th sample's true label predicted probability, represents the given input feature, true label and parameter for the th sample's annotation status
[0015] An RNN-based pulmonary artery segmentation system, comprising: A sample data acquisition module, configured to acquire pulmonary CT sample images covering different slices and branch structures, and label the pulmonary artery main trunk and some branches in some of the pulmonary CT sample images to obtain labeled pulmonary CT sample images and unlabeled pulmonary CT sample images; A pulmonary artery segmentation model construction module, where the pulmonary artery segmentation model includes a U-Net network and two Mamba networks with selective state spaces. The nmODE module is introduced in the intermediate high-dimensional representation stage of the U-Net network. One Mamba network is used to construct a scoring function for predicting the probability of each voxel being a foreground, and the other Mamba network is used to construct an annotation tendency function for estimating the possibility that the voxel is manually labeled as a foreground; The pulmonary CT image is input into the U-Net network, and the original features are extracted after the downsampling stage. The original features are dynamically mapped by the nmODE module and then converted into a feature representation with long-term memory. After the upsampling stage, they are respectively input into the two Mamba networks, and the Mamba networks output the predicted true structure and annotation status; A pulmonary artery segmentation model training module, which is used to train the pulmonary artery segmentation model constructed by the pulmonary artery segmentation model construction module by using the labeled pulmonary CT sample images and unlabeled pulmonary CT sample images in the sample data acquisition module; A real-time segmentation module, which is used to obtain the pulmonary CT image to be segmented and input it into the pulmonary artery segmentation model trained by the pulmonary artery segmentation model training module, and the pulmonary artery segmentation model outputs the pulmonary artery segmentation result.
[0016] A computer device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0017] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the steps of the above method.
[0018] The beneficial effects of the present invention are as follows: 1. In the present invention, the pulmonary artery segmentation method is based on the memory-enhanced multi-head attention mechanism (Mamba) and neural ordinary differential equations (nmODE), which can effectively handle the complex topological relationships of the pulmonary artery tree structure. Especially in the segmentation tasks of small peripheral branches and distal regions, the segmentation accuracy is significantly improved, and the false negative rate is reduced; in addition, the sample images include labeled samples and unlabeled samples, and both are used as sample data during the training of the pulmonary artery segmentation model. The dual-branch structure adopted by the model can combine the EM algorithm to infer potential labels during the training process, enabling the model to segment the pulmonary artery tree structure with a small amount of reliable annotations, effectively improving the segmentation accuracy of small peripheral branches, reducing the workload of annotators, and effectively solving the problem of low segmentation accuracy of pulmonary arteries (especially small peripheral branches) caused by incomplete label data.
[0019] 2. In the present invention, by introducing the weakly supervised learning method, the dependence on complete annotated data is reduced, allowing efficient training through simple annotations (such as main branch annotations), and the model self-infers the small branch regions, thus greatly reducing the annotation workload and improving the annotation efficiency.
[0020] 3. In the present invention, by combining the advantages of Mamba and nmODE, it can capture both global information and local details, use the memory-enhanced mechanism to maintain stable feature representations, further enhance the robustness of the model to complex structures, improve the resistance to image noise and interference, and ensure the accurate segmentation of pulmonary artery branches.
[0021] 4. In the present invention, the pulmonary artery segmentation method has a low computational cost and higher efficiency, is applicable to real-time or near-real-time clinical applications, can help doctors accurately diagnose pulmonary artery-related diseases, especially in the early detection, diagnosis, and treatment of pulmonary arterial hypertension (PAH), and provides strong support for formulating personalized treatment plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic flow chart of the present invention; Figure 2 is a schematic structural diagram of the U-Net network in the present invention; Figure 3 is a schematic structural diagram of the Mamba network in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0024] Therefore, based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0025] Embodiment 1 This embodiment provides an RNN-based pulmonary artery segmentation method for segmenting pulmonary arteries in lung CT images. As Figure 1 shown, it includes the following steps: Step S1, obtaining sample data; Obtain lung CT sample images covering different slices and branch structures, and label the main pulmonary artery and some branches in some lung CT sample images to obtain labeled lung CT sample images and unlabeled lung CT sample images.
[0026] For labeled sample data, the inclusion criteria are as follows: 1) The patient undergoes pulmonary artery-related imaging examinations before receiving thin-section chest CT scans (0.75–1.5 mm); 2) There are clear pulmonary artery annotation data, especially in the main pulmonary artery and its branches; 3) The annotation data are provided by experienced medical experts, and each case contains complete pulmonary artery region annotation information. The exclusion criteria are as follows: 1) Lack of key clinical data (such as age, gender, disease history, etc.); 2) No CT scan is performed or the pulmonary artery region in the CT image is difficult to distinguish and annotate, such as pulmonary artery adhesion at the hilum or severe atelectasis; 3) The data quality is poor, resulting in the annotation being unable to accurately reflect the actual situation of the pulmonary artery; 4) The image quality of the CT scan is low and unable to clearly show the structure of the pulmonary artery, especially the small branches.
[0027] For unlabeled sample data, the inclusion criteria are as follows: 1) The patient undergoes pulmonary artery-related imaging examinations before receiving thin-section chest CT scans (0.75–1.5 mm); 2) It contains unlabeled pulmonary artery image data, which come from the same clinical group and the image quality meets the standards. The exclusion criteria are as follows: 1) Lack of key clinical data (such as age, gender, disease history, etc.); 2) No CT scan is performed or the pulmonary artery region in the CT image is difficult to distinguish and annotate, such as pulmonary artery adhesion at the hilum or severe atelectasis; 3) The data quality is poor, resulting in the pulmonary artery image information being unable to be clearly extracted and analyzed.
[0028] Regarding the inclusion criteria for the above sample data, lung CT images of 203 patients diagnosed with pulmonary nodule diseases were collected from four medical centers and two equipment manufacturers. The resolution of these images was standardized to 512*512 pixels, the pixel pitch was 0.5 to 0.95 mm, and the slice thickness was 1 mm. All data were anonymized to ensure patient privacy. The dataset was divided into a training set, a validation set, and a test set, including 100, 30, and 73 respectively; the specific parameters of the dataset such as age, number of slices, and pixel pitch were known, and detailed information such as the minimum value, maximum value, and mean value was provided. During the annotation process, each CT image in the dataset was manually marked by experienced radiologists and semi-automatically annotated using MIMICS software through the region-growing method. First, the radiologists adjusted the window width and window level to ensure that the pulmonary artery structure was clearly visible; then, the radiologists iteratively and manually selected seed points to obtain a rough pulmonary artery structure mask; subsequently, all radiologists fine-tuned the mask results of the pulmonary artery structure and checked the annotations with each other, and finally used the voting method to merge the annotation results of 5 radiologists. All annotation processes were strictly reviewed to ensure the accuracy and consistency of the annotations.
[0029] To increase the diversity of sample data and improve the generalization ability of the model, different data augmentation strategies are applied to both labeled images and unlabeled images. The specific data augmentation methods are as follows: For the labeled pulmonary CT sample images, 50% of the pulmonary CT sample images are horizontally flipped to simulate the pulmonary artery images from different perspectives; all or part of the remaining pulmonary CT sample images are randomly rotated within the range of -15° to 15°, so that the model can learn the features of the pulmonary artery at different angles and avoid relying on a specific direction.
[0030] For the unlabeled pulmonary CT sample images, different regions of the images are randomly selected for cropping to simulate the local detail changes in clinical images; the local structures in the images are changed through non-linear transformation, and the images are partially randomly occluded to simulate the common missing information problems in clinical data, improving the model's adaptability to data incompleteness.
[0031] By comprehensively using the above enhancement techniques, the model is exposed to more diverse pulmonary artery structure variations, noise interferences, and incomplete data during the training process, significantly enhancing the model's robustness; it also enables the model to better adapt to different qualities, perspectives, missing information, and noise situations in clinical images, thereby improving the accuracy and generalization ability of pulmonary artery segmentation.
[0032] Step S2, construct a pulmonary artery segmentation model; The pulmonary artery segmentation model includes a U-Net network and two Mamba networks with selective state spaces. The nmODE module is introduced in the intermediate high-dimensional representation stage of the U-Net network. One Mamba network is used to construct a scoring function for predicting the probability of each voxel being the foreground , and the other Mamba network is used to construct a labeling tendency function for estimating the likelihood that the voxel is manually labeled as the foreground , and its structure is as Figure 2 shown.
[0033] The pulmonary CT image is input into the U-Net network, and the original features are extracted after the downsampling stage. The original features are dynamically mapped by the nmODE module and converted into a feature representation with long-term memory , and then after the upsampling stage, they are respectively input into the two Mamba networks, and the Mamba networks output the predicted true structure and the labeling status .
[0034] As Figure 2As shown, the U-Net network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a first nmODE module, and a second nmODE module. An image is input into the first convolutional layer, and the output of the first convolutional layer is input into the second convolutional layer after downsampling. The output of the second convolutional layer is input into the third convolutional layer after downsampling. The output of the third convolutional layer is input into the fourth convolutional layer after downsampling. The output of the fourth convolutional layer is input into the fifth convolutional layer after downsampling. The fifth convolutional layer is combined with the first nmODE module, and the original features output by the fifth convolutional layer are dynamically mapped by the first nmODE module and then converted into feature representations with long-term memory. ; Feature representation After upsampling, it is input into the sixth convolutional layer. The sixth convolutional layer is combined with the second nmODE module, and the features output by the sixth convolutional layer are dynamically mapped by the second nmODE module and then converted into feature representations with long-term memory. ; Feature representation After upsampling, it is input into the seventh convolutional layer. The output of the seventh convolutional layer is input into the eighth convolutional layer after upsampling. The output of the eighth convolutional layer is input into the ninth convolutional layer after upsampling. The output of the ninth convolutional layer is used as the input to two Mamba networks.
[0035] As Figure 3 shown, the Mamba network includes a first linear projection layer, a tenth convolutional layer, a second linear projection layer, a convolutional layer, an activation function, a selective SSM layer, and a third linear projection layer; The output of the U-Net network is used as the input to the Mamba network and is respectively input into the first linear projection layer and the second linear projection layer. The output of the first linear projection layer passes through the tenth convolutional layer, the activation function, and the selective SSM layer in sequence and is then concatenated with the output of the second linear projection layer as the input to the third linear projection layer. The third linear projection layer outputs the predicted true structure and the annotation status .
[0036] The U-Net network is used as a feature extraction module. It is constructed based on the U-Net network and utilizes the differentiable mapping of nmODE. The U-Net network consists of four downsampling stages, and each stage includes a 3*3*3 convolutional layer, followed by instance normalization and the Leaky-ReLU activation function. The nmODE module is a continuous dynamic system, and a stable mapping can be established between its data space and representation space through a global attractor. Since the image is sparse in the low-dimensional space, the nmODE module is applied to the high-dimensional space to fully utilize its powerful mapping ability. The nmODE module is applied to the high-dimensional data after the feature extraction of the U-Net network. Among them, the ordinary differential equations of the first nmODE module and the second nmODE module are: ; ; Among them, represents a function related to time t, represents the output function is the derivative function of represents a constant, which is used to control the influence of external input on the stability of the model mapping; represents the external input; represents the learnable weight parameter, which is used to adjust the influence of the input feature on the final output; represents the input feature; represents the bias term, which is used to adjust the input feature when calculating the influence of the external input .
[0037] is a non-linear term. By adding the external input , it can dynamically adjust the features according to different input situations; given the input feature and the initial value , the nmODE module will give an output that changes with time t. As , the output will converge to the attractor, and the attractor represents the stable state in the feature space.
[0038] Step S3, training the pulmonary artery segmentation model; Use the labeled pulmonary CT sample images and unlabeled pulmonary CT sample images in Step S1 to train the pulmonary artery segmentation model constructed in Step S2.
[0039] When training the pulmonary artery segmentation model, use the weak supervision learning method for training. Specifically: Assume that each sample is a voxel instance in the feature space, and the annotation is unsegmented. Therefore, there is a label The true label corresponding to this voxel. Additionally, define as a random variable indicating whether this voxel is labeled as the positive class (i.e., indicating being labeled as the positive class).
[0040] During the annotation process, annotation tendency is introduced, where the annotation status has a dependency relationship with the features of the instance and the label. Specifically, the annotation process can be modeled by the following formula: ; where, represents the conditional probability of the label given the feature , represents the probability of the annotation status given the label and the feature , represents the conditional probability of the annotation status.
[0041] During the annotation process, the probability can be modeled by a tendency function that describes the relationship between the annotation status and the features and the label, and is expressed by the formula: ; where, represents the parameter of this function, which is used for feature learning in the annotation process.
[0042] According to the above definition, the goal of weak supervision learning is to find the relationship between the annotation status and the true label through inference and optimization. During the training process, the EM algorithm is used to infer the latent label , and the annotation data is used to optimize the parameters of the model. Specifically: Given the feature and the observable annotation status , the true label is latent. Using the conditional joint distribution given by the following formula, the goal of the model is to construct a segmentation model by inferring the label . The specific conditional joint distribution can be expressed by the following formula: ; This formula shows how to infer the joint distribution of the latent label and the annotation status based on the sample features and annotation information given the training set.
[0043] Based on the above weak supervised learning, the EM algorithm is adopted to enhance the label inference ability of the model and improve the segmentation accuracy. In the E-step, the EM algorithm infers the label by calculating the posterior probability of each instance. , and infers the label ; Then in the M-step, the parameters of the model are optimized; By alternately executing these two steps, the model is gradually optimized and infers the most likely label, thereby improving the pulmonary artery segmentation accuracy.
[0044] For the conditional probability , which is the posterior probability given the input feature , it is determined by the following scoring function , and there is the following fact: ; Among them, is the probability that the input feature is labeled as the positive class. To estimate the parameter , the following maximization problem needs to be solved: ; By taking the logarithm of the right side of the above formula, maximizing the likelihood function is equivalent to: ; Then, through the alternating iteration of the E-step and the M-step, the EM algorithm can, in the case of incomplete annotation, realize the effective training of the segmentation model by inferring the label y and optimizing the model parameters. In this process, the E-step is responsible for inferring the posterior probability of the label, and the M-step updates the model parameters by maximizing the likelihood function, as follows: In the E-step, the calculation formula for the posterior probability is: ; Among them, represents the conditional probability that the label is given the feature , represents the annotation status probability of the feature ; represents the conditional probability of the annotation status, indicating whether the feature is labeled as the positive class; In the M-step, according to the posterior probability obtained in the E-step, the model parameters are updated by maximizing the log-likelihood function, and the goal is to maximize the joint probability of the label and the probability of the annotation status; The formula for the maximization process is expressed as:
[0045] ; Among them, represents the log-likelihood function of the parametric model, represents the number of samples, represents the estimated probability of the true label of the th sample, represents the expected value of the true label of the th sample, represents the conditional probability of the annotation status, represents the predicted probability of the true label of the th sample given the input feature and the parameter th sample, represents the predicted probability of the annotation status of the th sample given the input feature , the true label and the parameter th sample. represents the predicted probability of the annotation status of the th sample.
[0046] This step updates the model by optimizing the parameters, corresponding to the model for assigning labels and the model for annotation status respectively.
[0047] When training the pulmonary artery segmentation model, the features extracted by the U-Net network are passed as input to the Mamba network to construct the scoring function and the annotation propensity function . And the Mamba network helps the model to pay more attention to small branches and complex structures through memory enhancement and multi-head attention mechanisms. According to the results of the EM algorithm, two parameters need to be updated, corresponding to two loss functions. Therefore, the total loss function is expressed as: ; ; ; where , both represent weights, represents the scoring loss, represents the propensity loss, represents the predicted label of the th sample, represents the scoring function, represents the annotation propensity function, represents the input feature of the th sample, represents the annotation status of the th sample. , represents the optimized parameters of two loss functions, and represents the number of samples.
[0048] Step S4, real-time segmentation; Obtain the lung CT image to be segmented and input it into the pulmonary artery segmentation model trained in step S3, and the pulmonary artery segmentation model outputs the pulmonary artery segmentation result.
[0049] Embodiment 2 This embodiment provides a pulmonary artery segmentation system based on RNN, which includes: A sample data acquisition module, which is used to acquire lung CT sample images covering different slices and branch structures, and label the main pulmonary artery and some branches in some lung CT sample images to obtain labeled lung CT sample images and unlabeled lung CT sample images; A pulmonary artery segmentation model construction module, where the pulmonary artery segmentation model includes a U-Net network and two Mamba networks with selective state spaces. The nmODE module is introduced in the intermediate high-dimensional representation stage of the U-Net network. One Mamba network is used to construct a scoring function for predicting the probability of each voxel being the foreground, and the other Mamba network is used to construct a labeling tendency function for estimating the possibility that the voxel is manually labeled as the foreground; The lung CT image is input into the U-Net network, and the original features are extracted after the downsampling stage. The original features are dynamically mapped by the nmODE module and then converted into a feature representation with long-term memory. After the upsampling stage, they are respectively input into the two Mamba networks, and the Mamba networks output the predicted true structure and labeling status; A pulmonary artery segmentation model training module, which is used to train the pulmonary artery segmentation model constructed by the pulmonary artery segmentation model construction module by using the labeled lung CT sample images and unlabeled lung CT sample images in the sample data acquisition module; A real-time segmentation module, which is used to obtain the lung CT image to be segmented and input it into the pulmonary artery segmentation model trained by the pulmonary artery segmentation model training module, and the pulmonary artery segmentation model outputs the pulmonary artery segmentation result.
[0050] For more specific content of each module, reference can be made to the records in Embodiment 1.
[0051] Embodiment 3 A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the pulmonary artery segmentation method based on RNN.
[0052] Among them, the computer device can be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.
[0053] The memory at least includes one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or D interface display memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system installed on the computer device and various application software, such as the program code of the RNN-based pulmonary artery segmentation method. In addition, the memory can also be used to temporarily store various types of data that have been output or will be output.
[0054] In some embodiments, the processor can be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is usually used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the RNN-based pulmonary artery segmentation method.
[0055] Embodiment 4 A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the RNN-based pulmonary artery segmentation method.
[0056] Among them, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor to cause the at least one processor to execute the steps of the RNN-based pulmonary artery segmentation method as described above.
[0057] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server or network device, etc.) to execute the RNN-based pulmonary artery segmentation method described in the embodiments of the present application.
Claims
1. A pulmonary artery segmentation method based on RNN, characterized in that Including the following steps: Step S1, obtaining sample data; Obtaining lung CT sample images covering different slices and branch structures, and annotating the main pulmonary artery and some branches in some of the lung CT sample images to obtain labeled lung CT sample images and unlabeled lung CT sample images; Step S2, constructing a pulmonary artery segmentation model; The pulmonary artery segmentation model includes a U-Net network and two Mamba networks with selective state spaces. The nmODE module is introduced in the intermediate high-dimensional representation stage of the U-Net network. One Mamba network is used to construct a scoring function for predicting the probability of each voxel being the foreground, and the other Mamba network is used to construct a labeling tendency function for estimating the possibility that the voxel is manually labeled as the foreground; The lung CT image is input into the U-Net network. After the downsampling stage, the original features are extracted. The original features are dynamically mapped by the nmODE module and converted into feature representations with long-term memory. After the upsampling stage, they are respectively input into the two Mamba networks, and the Mamba networks output the predicted true structure and labeling status; Step S3, training the pulmonary artery segmentation model; Using the labeled lung CT sample images and unlabeled lung CT sample images in Step S1 to train the pulmonary artery segmentation model constructed in Step S2; Step S4, real-time segmentation; Obtaining the lung CT image to be segmented and inputting it into the pulmonary artery segmentation model trained in Step S3, and the pulmonary artery segmentation model outputs the pulmonary artery segmentation result.
2. The pulmonary artery segmentation method based on RNN according to claim 1, wherein Performing data augmentation on the lung CT sample images obtained in Step S1. The specific method is: For the labeled lung CT sample images, 50% of the lung CT sample images are horizontally flipped, and all or part of the remaining lung CT sample images are randomly rotated, and the rotation angle range is from -15° to 15°; For the unlabeled lung CT sample images, different regions of the image are randomly selected for cropping, the local structure in the image is changed through non-linear transformation, and the image is partially randomly occluded.
3. The pulmonary artery segmentation method based on RNN according to claim 1, wherein In step S2, the U-Net network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a first nmODE module, and a second nmODE module. An image is input into the first convolutional layer, and the output of the first convolutional layer is input into the second convolutional layer after downsampling. The output of the second convolutional layer is input into the third convolutional layer after downsampling. The output of the third convolutional layer is input into the fourth convolutional layer after downsampling. The output of the fourth convolutional layer is input into the fifth convolutional layer after downsampling. The fifth convolutional layer is combined with the first nmODE module, and the original features output by the fifth convolutional layer are dynamically mapped by the first nmODE module and then converted into feature representations with long-term memory ; Feature representation After upsampling, it is input into the sixth convolutional layer. The sixth convolutional layer is combined with the second nmODE module. The features output by the sixth convolutional layer are dynamically mapped by the second nmODE module and then converted into feature representations with long-term memory ; Feature representation After upsampling, it is input into the seventh convolutional layer. The output of the seventh convolutional layer is input into the eighth convolutional layer after upsampling. The output of the eighth convolutional layer is input into the ninth convolutional layer after upsampling. The output of the ninth convolutional layer is used as the input of two Mamba networks 4. The pulmonary artery segmentation method based on RNN according to claim 3, wherein The ordinary differential equations of the first nmODE module and the second nmODE module are: ; ; Among them, represents a function related to time t, represents the output function 's derivative function; represents a constant used to control the influence of external input on the stability of the model mapping; represents the external input; represents learnable weight parameters used to adjust the influence of input features on the final output; represents the input features; represents the bias term used to adjust the input features when calculating the external input 's influence.
5. The pulmonary artery segmentation method based on RNN according to claim 3, characterized in that In Step S2, the Mamba network includes a first linear projection layer, a tenth convolutional layer, a second linear projection layer, a convolutional layer, an activation function, a selective SSM layer, and a third linear projection layer; The output of the U-Net network is used as the input of the Mamba network and is respectively input into the first linear projection layer and the second linear projection layer. The output of the first linear projection layer passes through the tenth convolutional layer, the activation function, and the selective SSM layer in sequence, and then is concatenated with the output of the second linear projection layer and used as the input of the third linear projection layer. The third linear projection layer outputs the predicted true structure and the annotation status .
6. The pulmonary artery segmentation method based on RNN according to claim 1, wherein, In step S3, when training the pulmonary artery segmentation model, the total loss function is expressed as: ; ; ; Among them, and both represent weights, represents the scoring loss, represents the tendency loss, represents the th predicted label of the sample, represents the scoring function, represents the annotation tendency function, represents the th input feature of the sample, represents the th annotation status of the sample, and represent the optimized parameters of the two loss functions, represents the number of samples.
7. The pulmonary artery segmentation method based on RNN according to claim 6, characterized in that, In Step S3, when training the pulmonary artery segmentation model, the two Mamba networks are calculated iteratively through the EM algorithm. The E step is responsible for inferring the posterior probability of the label, and the M step updates the parameters by maximizing the likelihood function. Specifically: In the E step, the calculation formula for the posterior probability is: ; Among them, represents the conditional probability when the given feature has the label represents the annotation status probability of the feature ; represents the conditional probability of the annotation status, indicating whether the feature is labeled as the positive class; In the M step, according to the posterior probability obtained in the E step , the model parameters are updated by maximizing the log-likelihood function. The goal is to maximize the joint probability of the labels and the probability of the annotation states. The formula for the maximization process is as follows: ; Among them, represents the log-likelihood function of the parametric model, represents the number of samples, represents the true label of the th sample, represents the true label of the th sample, represents the conditional probability of the annotation state, represents the given input feature and parameter true label of the th sample, represents the given input feature true label and parameter predicted probability of the annotation state of the th sample.
8. A pulmonary artery segmentation system based on RNN, characterized in that, Including: A sample data acquisition module for obtaining lung CT sample images covering different slices and branch structures, and annotating the main pulmonary artery and some branches in some of the lung CT sample images to obtain labeled lung CT sample images and unlabeled lung CT sample images; The pulmonary artery segmentation model construction module, where the pulmonary artery segmentation model includes a U-Net network and two Mamba networks with selective state spaces. The nmODE module is introduced in the intermediate high-dimensional representation stage of the U-Net network. One Mamba network is used to construct a scoring function for predicting the probability of each voxel being the foreground, and the other Mamba network is used to construct a labeling tendency function for estimating the likelihood that the voxel is manually labeled as the foreground; The lung CT image is input into the U-Net network. After the downsampling stage, the original features are extracted. The original features are dynamically mapped by the nmODE module and then converted into a feature representation with long-term memory. After the upsampling stage, they are respectively input into the two Mamba networks, and the Mamba networks output the predicted true structure and labeling status; The pulmonary artery segmentation model training module is used to train the pulmonary artery segmentation model constructed by the pulmonary artery segmentation model construction module using the labeled lung CT sample images and unlabeled lung CT sample images in the sample data acquisition module; The real-time segmentation module is used to obtain the lung CT image to be segmented and input it into the pulmonary artery segmentation model trained by the pulmonary artery segmentation model training module, and the pulmonary artery segmentation model outputs the pulmonary artery segmentation result.
9. A computer device, characterized in that: It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of a pulmonary artery segmentation method based on RNN as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Stores a computer program. When the computer program is executed by the processor, the processor executes the steps of a pulmonary artery segmentation method based on RNN as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Pulmonary artery segmentation method and system for CT images based on contextual attention and feature fusion
CN117593518B
CT image pulmonary artery segmentation method and system based on context attention and feature fusion
CN117593518A
Pulmonary nodule analysis method, system and device based on large model and medium
CN119941731A
Image segmentation method and system based on parallel network framework and dynamic fusion, terminal and storage medium
CN119942130A
Cited By
Lung CT image nodule detection method based on selection state space and convolution model
CN121095227A
Lung ct image nodule detection method based on selection state space and convolution model
CN121095227B