Method for displaying focus activity in lung of lung puncture under guidance of CT (Computed Tomography)
By training the respiratory posture-image generation model, dynamic intrapulmonary lesions images are generated in real time, which solves the problem of puncture accuracy with high mobility in the lung lesions, improves the puncture efficiency and success rate, and reduces the risk of complications.
Patent Information
- Application Number
- CN202510060587.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When lung puncture is performed under CT guidance, the up and down movement of the lesions in the lung is high, making it difficult for the puncture needle to accurately align the lesions, increasing the risks of pneumothorax and pleural reactions, and reducing the success rate of puncture.
By training the respiratory posture-image generation model, the condition generation model is used to generate dynamic intrapulmonary lesions in real time, guiding the y-axis direction of the puncture needle to achieve accurate puncture of the lesions.
The puncture efficiency and success rate were significantly improved, the puncture-related pneumothoracic and pleural response risks were reduced, and the radiation dose was reduced in patients.
Smart Images

Figure CN120078454A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to medical technology, and particularly to a method for displaying the mobility of pulmonary lesions under CT guidance during lung puncture. Background Art
[0002] Under normal circumstances, the vital capacity of adult men is about 3500 - 4000 ml, and that of adult women is about 2500 - 3000 ml. The mobility of the lower lung border is 6 - 8 cm. Although ultrasound has the advantage of real-time observation, CT has better display resolution and is the preferred mediating method for percutaneous biopsy of pulmonary lesions. Due to the uncontrollable up-and-down movement of lung lesions near the diaphragm or close to the pleura in the lower lungs, even with pre-operative breathing training and breath-holding scanning under voice control, each needle insertion operation is performed at a "fixed" time point (such as manually judging the end of exhalation or the beginning of inhalation), and the puncture needle cannot be aligned with the lesion that moves with breathing. For lesions with a large up-and-down mobility in the lungs, even through methods such as breathing training, due to reasons such as poor patient cooperation and the time difference between the needle insertion operation and the CT scan, it is very difficult for the puncture needle to align with the lesion in the y-axis plane, resulting in adjustments to the puncture route and an increase in the number of scans, increasing the risks of pneumothorax, pleural reaction, etc. At the same time, there is a certain time difference (at least 1 minute, inevitable) between the CT scan and the next operation, making this needle insertion can be defined as having a difference from the original designed route and being a somewhat aimless needle insertion (blind puncture) to a certain extent. It is very likely that the puncture route needs to be re-regulated, inevitably increasing the scanning time and the radiation dose received by the patient, increasing the probability of the puncture needle passing through the pleura, further increasing adverse complications such as pneumothorax and pleural reaction, and reducing the puncture success rate. Currently, no effective methods and instruments have been published to achieve such operations.
[0003] Therefore, how to effectively display the amplitude and rhythm of the patient's breathing and the lesion, and guide the operator's needle insertion route, is of great significance for improving the puncture efficiency, success rate and reducing complications. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a method for displaying the mobility of pulmonary lesions under CT guidance during lung puncture. The present invention uses the coding features of the breathing posture as the condition of the generation model, and uses the conditional generation model to generate dynamic corresponding images of the pulmonary lesions in real time, guiding the y-axis needle insertion direction of the puncture needle outside the body and guiding the doctor to perform external puncture.
[0005] It can effectively improve the puncture efficiency and success rate, and significantly reduce the risks such as pneumothorax and pleural reaction during puncture.
[0006] Technical Solution: The method for displaying the mobility of pulmonary lesions under CT guidance during lung puncture according to the present invention includes the following steps: As Figure 2 shown:
[0007] Step S1: Collection of the training dataset.
[0008] Furthermore, collect the case CT images and the corresponding breathing postures, and regard them as a pair of breathing posture-image pairs. Specifically, use a strain gauge sensor to collect the moving distance d of the thoracic surface when the subject breathes. This moving distance d is the breathing posture corresponding to the CT image. Crop the collected images into H×W. Use the collected data to train the model.
[0009] Step S2: Training of the breathing posture-image generation model.
[0010] Furthermore, the network framework for generating real-time images according to the breathing posture is as Figure 3 shown. During the training process, we input the image x into the encoder E(·) in the autoencoder. Using this encoder, we compress the image and obtain the feature fi in its latent space. Subsequently, we input the breathing posture corresponding to this image, that is, the thoracic displacement distance d, into the breathing posture encoder τ θ (·), and obtain the representation f b of the breathing posture, as Figure 4 shown. Then, we perform the noise addition process in the diffusion model on the feature f i of this image in the latent space.
[0011] To enable the network to generate specific images according to the breathing posture, during the forward noise addition process, we use the posture feature f b as the guiding condition for image generation. Specifically, use the cross-attention mechanism to fuse the posture feature f b and the image feature fi into the Unet in the diffusion model. We use the noise-added feature x t at the previous moment as the input, and use Unet to predict the noise ∈ θ (x t , t, τ θ (d)). We use the noise ∈~N(0,1) added to the image feature f i as the supervision, train this Unet and supervise it through the L LDM loss function:
[0012]
[0013] Step S3: Inference of the breathing posture-image generation model.
[0014] Furthermore, during the reverse denoising process of inference, we use the trained Unet ∈ θ (·), and use the noise-added feature x tTaking ∼N(0, I) and the respiratory displacement d of the input target object as inputs, the noise ∈ at this moment is predicted. t Then, according to the principle of the diffusion model, the feature x at time t - 1 is calculated. t-1 The calculation of this feature is as follows:
[0015]
[0016] x is a sampling point in the standard normal distribution, x ∼ N(0, I). From the above formula, the time step x is continuously iterated to realize the feature (noise is 0) x at time 0. 0 At this time, x 0 is the image feature generated by the model in the latent space according to the input respiratory posture condition. We input this feature into the decoder D(·) of the autoencoder to realize the fine-grained decoding of the latent feature, so as to realize the CT image of the pulmonary lesion under this respiratory posture. The real-time generation of the image.
[0017] Step S4: Display of the generated image.
[0018] Furthermore, the image generated by the model is displayed in real time according to the respiratory frequency.
[0019] Furthermore, the system network architecture of the present invention includes the following modules, Figure 1 as shown:
[0020] Image autoencoder. Among them, encoder E(·) is responsible for compressing the target image into the latent space to perform high-dimensional feature modeling and information compression on the image, reducing the computational complexity of the model; decoder D(·) is responsible for finely decoding the features extracted from the latent space, so as to realize the feature reconstruction into an image.
[0021] Respiratory posture encoder. This module encodes according to the collected respiratory posture parameters to obtain the feature f of the posture information. b It sends the posture encoding into the diffusion model as the generation condition of the target image to realize the generation of the pulmonary lesion image under this specific posture. Specifically, a strain gauge sensor is used to collect the moving distance d on the surface of the chest during the object's breathing. This moving distance d is the respiratory posture under the corresponding CT image. The maximum displacement distance is evenly divided into n intervals, and according to the interval where the real-time displacement distance is located. The posture is encoded using one-hot encoding to obtain y = [0, 0, 1,..., 0], where y has n dimensions corresponding to different intervals, 0 indicates not in the corresponding interval, and 1 indicates the interval where it is located. Then, the posture is sent into the designed posture encoder τ. θAmong them, as shown in Figure 4 . First, a learnable embedding matrix is randomly initialized . This matrix defines learnable high-dimensional features of dimension D for n different intervals of breathing postures. According to the one-hot encoded index of the breathing posture, the corresponding high-dimensional feature F n = yF b is used to represent the corresponding posture in the latent space, and its dimension is 1×D. Then, F n is further fed into a fully connected layer (MLP) to further enhance the feature expression ability and obtain the final posture feature f b .
[0022] Conditional diffusion model. The operating principle of the diffusion model includes a forward noise addition process and a reverse denoising process, as shown in Figure 3 . Its main body contains a Unet network ∈ θ (·) used to predict noise or intermediate feature states. It generates a specific image according to Bayes' theorem given a normally distributed noise. The operating principle of the diffusion model can be summarized as:
[0023] a) Forward noise addition process Q(x t |x t-1 ): Given the original input x 0 satisfying the data distribution Q(x 0 ), the forward noise addition process continuously superimposes Gaussian noise ∈ t ~ N(0, I) at time step t, gradually transforming it into completely random Gaussian noise. The noise addition process can be expressed as:
[0024]
[0025] where α t is the noise addition coefficient obeying a specific functional relationship, and x t is the image representation at time step t. As the time step t increases, the α t coefficient continuously decreases until it approaches 0, while the original data distribution continuously approaches Gaussian noise. Using this noise addition process, we can train a Unet model ∈ θ (·) based on the noise-added input and image state at time step t. Subsequently, given time step t, Unet can predict the noise added at that time.
[0026] b) Reverse denoising process P(x t-1 |x t ): The purpose of this process is to continuously infer the denoised input x t at time step t - 1 based on the trained noise predictor Unet and the randomly input normal distribution vector x t-1 , until the input z at time step 0 is obtained0 At this time, the noise is completely removed, which is the finally generated image. This process can be expressed as:
[0027]
[0028] where x is a sampling point in the standard normal distribution, used to make the process differentiable when constructing the distribution, x ∼ N(0, I); σ t is the variance of the normal distribution satisfied by z t at time t,
[0029] On this basis, the breathing posture is added as a condition to the diffusion model, so as to guide the generation of the intrapulmonary lesion images in this posture.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] For patients who cannot cooperate with breathing training and intrapulmonary lesions with a large range of up and down movement near the diaphragm or close to the pleura in the lower lungs, the present invention can effectively display the up and down movement range and rhythm of the intrapulmonary lesions, so as to guide the operator to accurately puncture the lesions, and has high-efficiency and practical clinical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a system module diagram;
[0033] Figure 2 is a method flow schematic diagram;
[0034] Figure 3 is a schematic diagram of the trained model structure;
[0035] Figure 4 is a schematic diagram of the breathing posture encoder structure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be further described below.
[0037] The method for displaying the mobility of intrapulmonary lesions during CT-guided lung puncture in this embodiment includes the following steps:
[0038] Step S1: Collection of the training dataset. 10,000 pairs of breathing postures - images are collected manually. Among them, 8,000 pairs are used as the training set, and the remaining 2,000 pairs are used as the validation set. We use a strain gauge sensor to collect the moving distance d of the thoracic surface during the subject's breathing, and regard it as the breathing posture under this CT image. We take the diameter of the thoracic surface at the end of the exhalation action of the human body as the starting point of the displacement, and the diameter of the thoracic surface at the end of the inhalation as the end point, so as to quantify the breathing posture during the human breathing process. To unify the input size of the model, the collected images are cropped into 512×512.
[0039] Step S2: Training of the breathing posture - image generation model. The structure of the image generation model for breathing posture regulation is as Figure 3 shown. During the training process, we input the image x into the encoder in the auto - encoder. This encoder is based on the variational auto - encoder (VAE) to achieve image compression and obtain the features fi in the latent space. Subsequently, we input the breathing posture d corresponding to this image into the designed breathing posture encoder τ θ (·), as Figure 4 shown. This encoder extracts the posture representation f b · Then, in the latent space, the features f i of the input image are continuously subjected to the forward noise - adding process.
[0040] As Figure 3 shown, during the forward noise - adding process, we send the breathing posture feature f b and the image feature f i into the Unet ∈ θ (·) in the diffusion model. This Unet network contains multiple cross - attention modules, and uses the cross - attention mechanism to continuously integrate the posture features into the image features. The cross - attention mechanism can be expressed as follows:
[0041]
[0042] Among them, Q is f b after linear mapping, is the mapping matrix; K is f i after linear mapping, is the mapping matrix; V is f i after linear mapping, is the mapping matrix; m is the number of layers of the cross - attention module in the Unet.
[0043] Use the noise - added feature x t at the previous moment as the input, and use the Unet to predict the noise ∈ θ (x t, t, τ θ (d)). We use the noise ∈~N(0, 1) added to the image feature f i as supervision to train the Unet and supervise it through the L LDM loss function:
[0044]
[0045] Step S3: Inference of the breathing posture-image generation model. In the reverse process of inference, according to the trained Unet, the noisy feature x t ~N(0, I) at the sampled t moment and the breathing posture d at this moment are used as inputs to predict the noise ∈ t at this moment. According to the diffusion model principle, the feature x t-1 at the t - 1 moment is obtained:
[0046]
[0047] x is a random sampling point in the standard normal distribution, x~N(0, I). From the above formula, continuous iteration is carried out to realize the noisy feature (noise is 0) x 0 at the 0 moment. At this time, x 0 is the latent image feature generated by the model according to the breathing posture. Then the model inputs this feature into the decoder D(·) of the autoencoder to realize the decoding of the latent feature, and thus the image under this breathing posture is generated.
[0048] The above is only the preferred embodiment of the present invention and does not impose any limitation on the present invention. Any person skilled in the art, within the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification and other changes to the technical solution and technical content disclosed by the present invention, all of which belong to the content of the technical solution of the present invention and still fall within the protection scope of the present invention.
Claims
1. A method for displaying the activity of pulmonary lesions during lung puncture under CT guidance, characterized in that: The steps include: Step S1: Collect case CT images and corresponding breathing postures as training data sets; Step S2: training a breathing posture-image generation model; Step S3: Reasoning the breathing posture-image generation model; Step S4: Display the generated image.
2. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 1, characterized in that: The step S1: collecting the moving distance d of the chest surface when the subject breathes, where the moving distance d corresponds to the breathing posture under the corresponding CT image, and cutting the collected image into H×W.
3. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 2, characterized in that: Step S2: input the image x into encoderE(·) in the image autoencoder, compress the image and obtain the image feature f in its latent space i Then, the respiratory posture corresponding to the image, that is, the chest displacement distance d, is input into the respiratory posture encoder τ θ (·), we get the representation of breathing posture f b , and then, the feature f of the image in the latent space i Performs the noise addition process in the diffusion model.
4. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 3, characterized in that: In the forward process of adding noise, the posture feature f b As a guiding condition for image generation, the cross-attention mechanism is used to transform the pose feature f b And the image feature f i Fusion into the Unet in the diffusion model, using the noise feature x of the previous moment t As input, use Unet to predict the noise ∈ θ (x t , t, τ θ (d)) to image feature f i The added noise ∈~N(0,1) is used as supervision to train the Unet and pass L LDM The loss function is supervised:
5. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 3, characterized in that: The image autoencoder includes an encoder E(·) and a decoder D(·). The encoder E(·) is responsible for compressing the target image into a latent space to perform high-dimensional feature modeling and information compression on the image and reduce the computational complexity of the model; the decoder D(·) is responsible for fine-grained decoding of the features extracted from the latent space to realize feature reconstruction into an image.
6. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 3, characterized in that: The respiratory posture encoder τ θ (·) Encode the collected breathing posture parameters to obtain the characteristics of posture information f b , which transmits the posture code to the diffusion model as the generation condition of the target image, and realizes the generation of the lung lesion image under the specific posture.
7. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 6, characterized in that: The strain gauge sensor is used to collect the moving distance d of the chest surface when the subject breathes. The moving distance d corresponds to the breathing posture under the corresponding CT image. The maximum displacement distance is evenly divided into n intervals. According to the interval where the real-time displacement distance is located, the posture is encoded using one-hot coding to obtain y=[0,0,1,...,0], where y has n dimensions corresponding to different intervals, 0 indicates that it is not in the corresponding interval, and 1 indicates the interval. Then, the posture is sent to the designed posture encoder τ θ (·) First, a learnable embedding matrix is randomly initialized This matrix defines learnable high-dimensional features of dimension D for each breathing posture in n different intervals. The high-dimensional feature F corresponding to the unique hot encoding index of the breathing posture is n =yF b , with a dimension of 1×D, is used to represent the corresponding posture in the latent space, and then, F n It is sent to a fully connected layer MLP to further enhance the feature expression ability and obtain the final posture feature f b .
8. The method for displaying the activity of pulmonary lesions during CT-guided lung puncture according to claim 6, characterized in that: The operation principle of the diffusion model includes the forward denoising process and the reverse denoising process. The main body includes a Unet network ∈ θ (·) is used to predict noise or characteristic intermediate states.
9. The method for displaying the activity of lesions in lungs during CT-guided lung puncture according to claim 8, characterized in that: The forward noise adding process Q(x t |x t-1 ): Given the original input x0 that satisfies the data distribution Q(x0), the forward noise addition process continuously superimposes the Gaussian noise ∈ at time t t ~N(0,I), so that it gradually transforms into completely random Gaussian noise. The noise adding process can be expressed as: Among them, α t is the noise factor that obeys a specific functional relationship, x t is the image representation at time t. As the time step t increases, α t The coefficient decreases continuously until it approaches 0, and the original data distribution approaches Gaussian noise. By using this noise adding process, a Unet model ∈ can be trained based on the noise adding input and image state at time t. θ (·), subsequently given a time t, Unet can predict the noise added at that time; The reverse denoising process P(x t-1 |x t ): The purpose of this process is to train the noise predictor Unet and the normal distribution vector x of the random input t , and continue to reason forward to obtain the denoised input x at time t-1 t-1 , until the input z0 at time 0 is obtained, at which time the noise is completely removed, which is the final generated image. The process can be expressed as: Where x is a sampling point in the standard normal distribution, which is used to make the process differentiable when constructing the distribution, x~N(0,I); σ t is time t t The variance of the normal distribution satisfied by 10. The method for displaying the activity of pulmonary lesions during CT-guided lung puncture according to claim 9, characterized in that: Step S3: In the reverse denoising process of inference, Unet∈ is obtained according to the training θ (·), the noise feature x at time t obtained by sampling t ~N(0, I) and the input target object's respiratory displacement d are used as input to predict the noise ∈ at that moment t Then, according to the diffusion model principle, the characteristic x at time t-1 is calculated t-1 , the feature is calculated as follows: x is a sampling point in the standard normal distribution, x~N(0,I). According to the above formula, the time step x is continuously iterated to achieve the feature x0 at time 0 (noise is 0). At this time, x0 is the image feature generated by the model in the latent space according to the input respiratory posture condition. The feature is input into the decoderD(·) of the autoencoder to achieve fine-grained decoding of the latent feature, thereby achieving the CT image of the lung lesion under the respiratory posture. Real-time generation.