Lung ventilation estimation method based on semi-supervised learning and 4D CT spatial-temporal characteristics
By introducing semi-supervised learning and spatiotemporal features in 4D CT, using the 3D U-net network and Mean-Teacher framework, and combining training with and without labeled data, the problems of low 4D CT imaging accuracy and high labeling cost are solved, and high-precision ventilation images are generated.
Patent Information
- Application Number
- CN202510979171.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-12
AI Technical Summary
Existing deep learning-based lung ventilation estimation methods have low accuracy in 4D CT imaging and rely on expensive radioactive tracers and limited labeled data, resulting in model overfitting and inability to effectively generate high-precision ventilation images.
A method based on semi-supervised learning and 4D CT spatiotemporal features was adopted. By constructing a 3D U-net network, adding a temporal displacement mechanism and a cross-sectional attention mechanism, and using the Mean-Teacher semi-supervised framework, combined with labeled and unlabeled data for training, high-precision ventilation images were generated.
It achieves higher imaging accuracy and lower labeling cost, solves the problem of underutilization of temporal information in 4D CT, prevents network overfitting, and generates high-precision ventilation images.
Smart Images

Figure CN120635675A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a lung ventilation estimation method based on semi-supervised learning and 4D CT spatiotemporal features. Background Art
[0002] Pulmonary ventilation estimates derived from four-dimensional computed tomography (4D CT) can accurately quantify local lung ventilation performance, providing critical support for personalized radiotherapy and minimizing the adverse effects of radiation on areas with good lung function. Currently, most clinical ventilation estimation methods involve patients inhaling radioactive gas and using imaging equipment to observe the distribution of the gas in the patient's lungs to obtain corresponding lung ventilation images. However, these methods have slow imaging speeds and require expensive specialized equipment and tracer gases during the imaging process, which severely limits their clinical application.
[0003] To address this clinical challenge, a feasible approach is to calculate ventilation images using the spatiotemporal information related to respiration contained in 4D CT. Traditional ventilation estimation methods primarily track the Hu value or lung volume changes in registered 4D CT images to calculate ventilation images. However, the effectiveness of these methods is highly dependent on the accuracy of image registration, and this uncertainty cannot meet clinical needs. To eliminate the impact of image registration errors, deep learning-based imaging methods have become a more promising option.
[0004] With the continuous development of deep learning technology, deep learning methods have been widely applied in various medical image processing fields. Existing deep learning-based ventilation estimation methods typically use encoder-decoder networks (such as U-net or V-net) to generate lung ventilation images from 3D or 4D CT images. Compared with traditional ventilation estimation methods, neural networks can more effectively extract complex spatial information during the respiratory process, resulting in more accurate ventilation images. Although this method achieves excellent performance, it still has limitations. First, existing encoder-decoder networks simply concatenate 3D CT images of different phases and feed them into the network. Temporal information in 4D CT images remains to be exploited, resulting in low imaging accuracy. Second, in clinical practice, 4D CT images of the lungs are relatively easy to obtain, while the corresponding ventilation images are difficult to obtain due to high cost and special requirements for radioactive tracers. Therefore, existing deep learning-based ventilation estimation methods can only rely on fully supervised learning with limited labeled data, which leads to model overfitting. Due to these limitations, it is impossible to generate high-accuracy ventilation images based on 4D CT images of the lungs, and thus cannot effectively estimate ventilation. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a lung ventilation estimation method based on semi-supervised learning and 4D CT spatiotemporal features to solve the problems existing in the above-mentioned prior art.
[0006] To achieve the above objectives, the present invention provides a method for estimating pulmonary ventilation based on semi-supervised learning and 4D CT spatiotemporal features, comprising:
[0007] Acquiring labeled data and unlabeled data, wherein the labeled data set includes a lung 4D CT differential image and its corresponding ventilation image, and the unlabeled data set includes only the lung 4D CT differential image;
[0008] Constructing a ventilation estimation network and training it using a Mean-Teacher semi-supervised framework with labeled and unlabeled data; the ventilation estimation network is a 3D U-net network architecture, and a time shift mechanism and a cross-sectional attention mechanism are added to the 3D U-net network architecture;
[0009] The lung 4D CT image is processed by the trained ventilation estimation network to obtain a corresponding ventilation image.
[0010] Optionally, the process of acquiring the lung 4D CT differential image includes:
[0011] Acquire several phase images of the lung 4D CT, perform difference calculation on adjacent phase images in chronological order to obtain several differential phase images, retain the first phase image as the anchor image of the lung 4D CT differential image, integrate the differential phase image and the anchor phase image to obtain the lung 4D CT differential image, wherein the phase images of the lung 4D CT are 3D CT images at different times.
[0012] Optionally, in the ventilation estimation network, the encoder and the decoder are jump-connected via a cross-sectional attention mechanism.
[0013] Optionally, the corresponding layer structures of the encoder and decoder are jump-connected through the cross-sectional attention mechanism. In the jump connection, the cross-sectional attention mechanism, the first ReLU function, the first normalization layer, the second ReLU function and the second normalization layer are processed in sequence to obtain the processing results, and the processing results are passed to the corresponding layer structure of the decoder, wherein there is a residual connection between the cross-sectional attention mechanism and the first ReLU function, the first normalization processing and the second ReLU function.
[0014] Optional, cross-sectional attention data processing includes:
[0015]
[0016] Where Q and K represent the query and key in the attention, ⊙ is the matrix multiplication, QK T (m,n) is the similarity between the mth slice in Q and the nth slice in K, and softmax(·) indicates the normalization of the weights.
[0017] Optionally, in the ventilation estimation network, each convolution module in the encoder is residually connected through a residual block, and a time displacement mechanism is added to the residual block, wherein the input features of the current convolution module are fused with the first 1 / 4 features in the channel dimension of the previous and next adjacent phases, wherein the features of the previous and next adjacent phases are the convolution input features of the same layer adjacent to the current phase.
[0018] Optionally, the relevant fusion process of the time shift mechanism is:
[0019]
[0020] Where T is the total number of phases, Represents the characteristics of a single phase, F i P Indicates the original characteristics of the current phase, f i P-1 With f i P+1 They represent the features obtained by the front and back phase exchange, respectively, O represents the zero matrix, P represents the phase index, P = 1 represents the anchor point phase, and P > 1 represents the corresponding differential phase.
[0021] Optionally, the process of training the ventilation estimation network includes:
[0022] The Mean-Teacher semi-supervised framework includes a teacher network and a student network, wherein both the teacher network and the student network adopt a ventilation estimation network; the network parameters of the teacher network are updated by exponential moving average according to the network parameters of the student network;
[0023] The unlabeled data is processed by the teacher network to obtain the teacher network output results. Based on the total loss of the semi-supervised framework, the student network is trained by the teacher network output results and labeled data to obtain a trained ventilation estimation network.
[0024] Optionally, the parameter update process of the teacher network is:
[0025]
[0026] Among them, η represents the smoothing coefficient, τ is the current training number, represents the network parameters of the student network, Represents the network parameters of the teacher network.
[0027] Alternatively, the total loss L of the semi-supervised framework is:
[0028]
[0029] Among them, L S represents the supervision loss, L U1 and L U2 represents the consistency loss, λ(τ) is the weight of the consistency loss, τ is the current number of training times, is the output of the student network, is the teacher network output, L mse is the mean square error loss, y i Indicates a label.
[0030] Compared with the prior art, the present invention has the following advantages and technical effects:
[0031] The pulmonary ventilation estimation method described in this paper, based on semi-supervised learning and 4D CT spatiotemporal features, addresses the problem that current ventilation estimation networks cannot fully utilize 4D CT spatiotemporal features, achieving higher imaging accuracy. Furthermore, the semi-supervised training approach addresses the issue of scarce labeled data and prevents network overfitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0033] Figure 1 This is a structural diagram of a pulmonary ventilation estimation network based on semi-supervised learning and 4D CT spatiotemporal features according to an embodiment of the present invention;
[0034] Figure 2 This is a schematic diagram of the cross-sectional attention mechanism of an embodiment of the present invention. DETAILED DESCRIPTION
[0035] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0036] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0037] The present invention discloses a lung ventilation estimation method based on semi-supervised learning and 4D CT spatiotemporal features, belonging to the field of image processing technology. The method comprises: 4D CT differential image design, which converts implicit temporal information between phases into explicit voxel-level motion signals, thereby facilitating the network to fully learn temporal information; a temporal feature enhancement strategy of differential time shifting, which enables the network to effectively extract multi-phase features with zero computational effort; a cross-sectional attention mechanism, which is used to enhance the network's learning ability in the Superior-Inferior (SI) direction; and the introduction of a Mean-Teacher semi-supervised framework, which effectively utilizes limited labeled data and a large amount of unlabeled data to train the network, thereby reducing the manpower and time costs of labeling images.
[0038] The present invention provides a method for estimating pulmonary ventilation based on semi-supervised learning and 4D CT spatiotemporal features, which includes:
[0039] S1) acquiring a 4D CT phase image and ventilation image of the patient during breathing;
[0040] S2) preprocessing the 4D CT phase image and ventilation image;
[0041] S3) performing differential design on the pre-processed 4D CT phase, and producing the result into a labeled 4D CT differential image dataset;
[0042] S4) collecting a large number of existing public 4D CT images, wherein the number of public 4D CT images is much larger than the number of acquired patients and does not include corresponding ventilation images;
[0043] S5) performing S2)-S3) operations on the existing public 4D CT phase image in S4) to obtain an unlabeled 4D CT difference image dataset;
[0044] S6) Using the labeled and unlabeled 4D CT difference image datasets described in S3 and S5, a ventilation estimation network based on the spatiotemporal features of lung 4D CT was trained under the Mean-Teacher semi-supervised framework.
[0045] S7) Processing the 4D CT image using the trained ventilation estimation network to generate a ventilation image corresponding to high imaging accuracy.
[0046] Specifically, the specific process of S3) is:
[0047] The difference between the n phases of a single 4D CT is made to obtain n-1 differential phase images, and the first phase image of the 4DCT is used as the anchor phase image to provide the initial spatial and density information of the lungs.
[0048] During a single 4D CT acquisition process, 3D phase images at different time points will be scanned and recorded according to different times.
[0049] Specifically, the ventilation estimation network based on the spatiotemporal characteristics of lung 4D CT in S6) includes:
[0050] Ventilation estimation network with 3D U-net as backbone;
[0051] Incorporating multi-phase information into a single 3D volume unit by temporal shifting in an encoder of the network;
[0052] A jump connection is made between the encoder and decoder of the network through cross-sectional attention to highlight the depth information of the image.
[0053] Specifically, the specific process of the time shift is:
[0054] The first 1 / 4 of the features within each phase channel dimension are selected and bidirectionally swapped in the time dimension, integrating the current feature with the features of the previous and next phases. The temporal shift is performed within the residual block, and residual connections are used to fuse the pre- and post-shift features, avoiding feature loss caused by the shift. The first 1 / 4 of the features within the channel dimension are: assuming a feature has c channels, then the first c / 4 of the features within the channel are selected.
[0055] Specifically, the cross-sectional attention mechanism includes:
[0056] Each individual 3D volume unit is split into multiple 2D cross-sectional slices along the SI direction, and the cross-relationships between the cross-sections are calculated through a multi-head attention mechanism, making the network's fitting ability in the SI direction stronger, which is consistent with the topological structure of the lungs.
[0057] The 3D volume unit is the 3D feature output by the phase image at each layer in the encoder and decoder.
[0058] To address the above issues, the present invention proposes a semi-supervised learning method for estimating ventilation volume based on the spatiotemporal features of lung 4D CT. The method includes the following steps:
[0059] like Figure 1As shown in the figure, a ventilation estimation network based on the spatiotemporal features of 4D lung CT is constructed, using the 3DU-net network as its backbone. Furthermore, temporal shifting is used to integrate multi-phase information into a single 3D volume unit in the encoder. Cross-sectional attention is used to establish skip connections between the encoder and decoder, highlighting the depth information of the image. Furthermore, to address the limitation of insufficient labeled data, the network is trained using the Mean-Teacher semi-supervised framework. This framework effectively utilizes limited labeled data and a large amount of unlabeled data, reducing labeling costs.
[0060] Optionally, a 3D U-net backbone network is constructed by introducing a residual block into the convolution blocks of the encoder and decoder, where the residual block includes 3D convolution, batch normalization, ReLU, and time displacement.
[0061] Optionally, a differential design is performed on the 4D CT. The method includes: performing differentials between n phases of a single 4D CT to obtain n-1 differential phase images, and using the first phase image of the 4D CT as the anchor phase image to provide initial spatial and density information of the lungs. The process can be expressed as:
[0062]
[0063] Where i is the data index, P is the phase index, and x i is the difference image, is a single phase of a 4D CT image.
[0064] An optional time shift mechanism is constructed by selecting 1 / 4 of the features of each phase and performing a bidirectional exchange in the time dimension to integrate the current features into the features of the previous and next phases. The calculation expression is:
[0065]
[0066] Where T is the total number of phases, Represents the characteristics of a single phase, F i P Indicates the original characteristics of the current phase, f i P-1 With f i P+1 They represent the features obtained before and after phase exchange, and O represents the zero matrix.
[0067] Each phase is in the time dimension, which is the time axis. Each phase of 4D CT is arranged on the time axis in sequential time.
[0068] Under the time displacement mechanism, each differential image (differential phase) corresponds to an encoder, and the encoders of different differential images are parallel structures. For different layers in the encoder, the features of each layer are transmitted to the adjacent previous and next encoders through the parallel encoder structure. 1 / 4 features under the channel dimension in a certain layer of a certain encoder are selected and fused with 1 / 4 features of the corresponding layers of the two adjacent differential images as the feature input of this layer of the encoder.
[0069] Optional, such as Figure 2 As shown in Figure 2, within the cross-sectional attention mechanism, the attention weight calculation expression is:
[0070]
[0071] Where Q and K represent the query and key in the attention, ⊙ is the matrix multiplication, QK T (m,n) is the similarity between the mth slice in Q and the nth slice in K, and softmax(·) indicates the normalization of the weights.
[0072] exist Figure 2 In [1], PE means the following: Position Encoding (PE): Since the network itself cannot identify the position of the slice, position encoding is essential. As with the Transformer, a learnable position encoding PE is first added to it. For slice p and all elements with an even number of channels 2j:
[0073]
[0074] For slice p and all elements with odd number of channels 2j+1:
[0075]
[0076] Optionally, the total loss L of the semi-supervised framework includes the supervision loss L S With consistency loss L U , the calculation expression is:
[0077]
[0078] Among them, L U1 and L U2 represents the consistency loss, λ(τ) is the weight of the consistency loss, τ is the current number of training times, is the output of the student network, is the teacher network output, L mse is the mean square error loss.
[0079] Optional, network parameters of the student model The update is performed by gradient descent, and the parameters of the teacher model Then through The exponential moving average is updated. The update process is as follows:
[0080]
[0081] Among them, η represents the smoothing coefficient and τ is the current number of training times. Updates are more stable, reflecting long-term average trend.
[0082] 4D CT images and their corresponding ventilation images are collected during the patient's breathing period and preprocessed. A differential design is performed on the preprocessed 4D CT, and the results are used to create a labeled 4D CT difference image dataset. A large number of existing public 4D CT images are collected, of which the number of public 4D CT images is much greater than the number of patients collected and does not contain corresponding ventilation images. These 4D CT images without corresponding ventilation images are preprocessed to obtain an unlabeled 4DCT difference image dataset. Using these labeled and unlabeled 4D CT difference image datasets, a ventilation estimation network based on the spatiotemporal features of lung 4D CT is trained under the Mean-Teacher semi-supervised framework.
[0083] During the testing phase, the trained student network is used to infer 4D CT data to obtain accurate ventilation images.
[0084] The ventilation volume is estimated by using the ventilation image to obtain an estimated ventilation volume result.
[0085] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A pulmonary ventilation estimation method based on semi-supervised learning and 4D CT spatiotemporal features, characterized by: include: Acquire a labeled dataset and an unlabeled dataset, wherein the labeled dataset includes a lung 4D CT differential image and its corresponding ventilation image, and the unlabeled dataset includes only the lung 4D CT differential image; A ventilation estimation network was constructed and trained using a Mean-Teacher semi-supervised framework using both labeled and unlabeled datasets. The ventilation estimation network employed a 3D U-net architecture, in which a temporal shift mechanism and a cross-sectional attention mechanism were added. The lung 4D CT image is processed by the trained ventilation estimation network to obtain a corresponding ventilation image.
2. The method according to claim 1, characterized in that The process of acquiring the lung 4D CT differential image includes: Acquire several phase images of the lung 4D CT, perform difference calculation on adjacent phase images in chronological order to obtain several differential phase images, retain the first phase image as the anchor image of the lung 4D CT differential image, integrate the differential phase image and the anchor phase image to obtain the lung 4D CT differential image, wherein the phase images of the lung 4D CT are 3D CT images at different times.
3. The method according to claim 1, characterized in that In the ventilation estimation network, the encoder and decoder are jump-connected through a cross-sectional attention mechanism.
4. The method according to claim 1, wherein The corresponding layer structures of the encoder and the decoder are jump-connected through the cross-sectional attention mechanism. In the jump connection, the cross-sectional attention mechanism, the first ReLU function, the first normalization layer, the second ReLU function and the second normalization layer are processed in sequence to obtain the processing results, and the processing results are passed to the corresponding layer structure of the decoder, wherein there is a residual connection between the cross-sectional attention mechanism and the first ReLU function, the first normalization processing and the second ReLU function.
5. The method according to claim 1, wherein The data processing process of the cross-sectional attention mechanism includes: Where Q and K represent the query and key in the attention, ⊙ is the matrix multiplication, QK T (m,n) is the similarity between the mth slice in Q and the nth slice in K, and softmax(·) indicates the normalization of the weights.
6. The method according to claim 1, characterized in that In the ventilation estimation network, each convolution module in the encoder is residually connected through a residual block, and a time displacement mechanism is added to the residual block, wherein the input features of the current convolution module are fused with the first 1 / 4 features in the channel dimension of the previous and next adjacent phases, wherein the features of the previous and next adjacent phases are the convolution input features of the same layer adjacent to the current phase.
7. The method according to claim 6, characterized in that The relevant fusion process of the time displacement mechanism is: Where T is the total number of phases, Represents the characteristics of a single phase, F i P Indicates the original characteristics of the current phase, f i P-1 With f i P+1 They represent the features obtained by the front and back phase exchange, respectively. O represents the zero matrix, P represents the phase index, P = 1 represents the anchor point phase, and P > 1 represents the corresponding differential phase.
8. The method according to claim 1, characterized in that The process of training the ventilation estimation network includes: The Mean-Teacher semi-supervised framework includes a teacher network and a student network, wherein both the teacher network and the student network adopt a ventilation estimation network; the network parameters of the teacher network are updated by exponential moving average according to the network parameters of the student network; The unlabeled dataset is processed by the teacher network to obtain the teacher network output results. Based on the total loss of the semi-supervised framework, the student network is trained by the teacher network output results and the labeled dataset to obtain a trained ventilation estimation network.
9. The method according to claim 1, characterized in that The parameter update process of the teacher network is: Among them, η represents the smoothing coefficient, τ is the current training number, represents the network parameters of the student network, Represents the network parameters of the teacher network.
10. The method according to claim 1, characterized in that The total loss L of the semi-supervised framework is: Among them, L S represents the supervision loss, L U1 and L U2 represents the consistency loss, λ(τ) is the weight of the consistency loss, τ is the current number of training times, is the output of the student network, is the teacher network output, L mse is the mean square error loss, y i Indicates a label.