Three-dimensional bronchial image generation device
Through the three-dimensional bronchial image generation device, the global and local feature extraction units combined with wavelet transform are used to solve the problem of low accuracy of bronchial tree structure reconstruction, and the accuracy and safety of bronchoscopic operation are improved.
Patent Information
- Application Number
- CN202511028709.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing technologies have problems with low edge reconstruction accuracy and poor generalization when generating bronchial tree structures. Especially when there are large individual differences, it is difficult to accurately identify and locate the target bronchus, resulting in insufficient accuracy and safety of bronchoscopic navigation.
A 3D bronchial image generation device uses a global feature extraction unit and a local feature extraction unit combined with a wavelet transform and self-attention transform module to generate an accurate bronchial tree structure. The device includes a deep self-attention transform network and a wavelet-based U-shaped network. It uses multi-level self-attention calculation and wavelet transform algorithms to extract global and local features step by step, reducing image aliasing and improving edge reconstruction accuracy.
It achieves accurate reconstruction of the bronchial tree structure, provides important reference information, and improves the accuracy, efficiency, and safety of bronchoscopic operations.
Smart Images

Figure CN120526063B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and artificial intelligence technology, and more specifically, to a three-dimensional bronchial image generating device. Background Art
[0002] Bronchoscopy is a vital tool for diagnosing and treating respiratory diseases. It involves inserting a tiny camera through the mouth and into the trachea, allowing for a comprehensive examination of the patient's airway while under anesthesia. This standard procedure has been in place for many years, but it still has many limitations. During examinations or treatments, doctors use the bronchoscope to visualize tissues and organs within the body and perform procedures accordingly. Therefore, ensuring that the bronchoscope can reach its intended target location along its intended route is crucial for these procedures.
[0003] However, bronchoscopic navigation faces many difficulties in practical applications, which limit its wider and more efficient application. These difficulties are mainly reflected in precise positioning, individual differences, technical requirements for operation, equipment and cost. The bronchial tree is complex in structure, with branches from the main bronchi to the lobar bronchi and segmental bronchi. The branches become thinner and more numerous as they go to the periphery. The bronchial trees of different people vary in morphology, branching angles and lengths, which makes it challenging for the navigation system to accurately identify and locate the target bronchus. It is easy to make misjudgments or find it difficult to find the best path. Even experienced doctors may have positioning errors. Therefore, before performing bronchoscopic navigation treatment, how to obtain the accurate bronchial tree structure and provide doctors with important reference information based on the accurate bronchial tree structure to improve the accuracy, efficiency and safety of bronchoscopic operation has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] In view of this, the present invention provides a three-dimensional bronchial image generating device.
[0005] The present invention provides a three-dimensional bronchial image generation device, comprising: a global feature extraction unit, for inputting three-dimensional medical image data corresponding to the bronchi into I self-attention transformation modules connected in series to perform multi-level self-attention calculations, and outputting global features corresponding to each of the above self-attention transformation modules; a local feature extraction unit, for using I wavelet downsampling modules to perform convolution calculations and wavelet transforms on the respective input data, respectively, to obtain local features corresponding to each of the above wavelet downsampling modules, wherein the input data of the first wavelet downsampling module is the above-mentioned three-dimensional medical image data, and the input data of the i-th wavelet downsampling module is calculated based on the i-1th local feature output by the i-1th wavelet downsampling module and the i-1th global feature output by the i-1th self-attention transformation module. ; A feature fusion unit is used to use I wavelet upsampling modules to perform deconvolution calculations and wavelet inverse transforms on the respective input data, and output the fusion features corresponding to each of the above wavelet upsampling modules, so as to generate a three-dimensional bronchial image according to the I fusion features, wherein the input data of the first wavelet upsampling module is calculated based on the I local features output by the I wavelet downsampling module and the I global features output by the I self-attention transform module, and the input data of the i-th wavelet upsampling module is calculated based on the I-i+1 local features output by the I-i+1 wavelet downsampling module, the I-i+1 global features output by the I-i+1 self-attention transform module, and the i-1 fusion features output by the i-1 wavelet upsampling module, i=2,.....,I.
[0006] According to an embodiment of the present invention, a global feature extraction unit inputs 3D medical image data corresponding to the bronchi into one self-attention transform module connected in series, performing multi-level self-attention calculations. This module extracts global long-range dependencies from the 3D medical image data, generating multi-scale global features. A local feature extraction unit utilizes one wavelet downsampling module to perform convolution and wavelet transform on each input data. This downsamples the bronchial-related input data using wavelet transform algorithms within the wavelet downsampling modules at different levels, simultaneously removing noise and unimportant details from the corresponding input data and extracting multiple local features containing highly accurate local information. The feature fusion unit uses I wavelet upsampling modules to perform deconvolution calculation and wavelet inverse transform on each input data. Each wavelet upsampling module is used to upsample the local features, global features and fusion features of the same scale related to the bronchi, while further refining the feature information corresponding to the bronchi. The detailed features of the three-dimensional image corresponding to the bronchi are accurately retained step by step, so that the final segmentation result can accurately locate the specific position and boundary of the bronchi, thereby reducing the image aliasing effect and improving the edge reconstruction accuracy, obtaining an accurate bronchial tree structure with strong applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The above and other objects, features and advantages of the present invention will become more apparent from the following description of the embodiments of the present invention with reference to the accompanying drawings.
[0008] Figure 1 A schematic structural diagram of a neural network model used in a three-dimensional bronchial image generation device according to an embodiment of the present invention is shown.
[0009] Figure 2 A flow chart of a method for generating a three-dimensional bronchial image according to an embodiment of the present invention is shown.
[0010] Figure 3 A schematic structural diagram of a wavelet downsampling module according to an embodiment of the present invention is shown.
[0011] Figure 4 A structural diagram of a self-attention transformation module according to an embodiment of the present invention is shown.
[0012] Figure 5 A block diagram of a three-dimensional bronchial image generating apparatus according to an embodiment of the present invention is shown.
[0013] Figure 6 A block diagram of an electronic device suitable for implementing the above-described three-dimensional bronchial image generation method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0015] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0016] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0017] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0018] In the embodiments of the present invention, the collection, updating, analysis, processing, use, transmission, provision, creation, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken to prevent unauthorized access to user personal information data and maintain the security of user personal information and network security.
[0019] In the embodiment of the present invention, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0020] Three-dimensional bronchial images directly extracted from corresponding CT (computed tomography) images of the bronchi suffer from issues such as ghosting, aliasing, and unclear image quality. To accurately determine the bronchial tree structure, related technologies typically generate two-dimensional medical images from these corresponding CT (computed tomography) images and perform correlation processing on these two-dimensional images to construct the bronchial tree structure.
[0021] In related technologies, there are mainly the following methods for bronchial tree reconstruction.
[0022] 1. Threshold-based algorithms, including global thresholding and adaptive thresholding. Global thresholding is computationally simple and fast, but it is less adaptable to images with uneven grayscale, and can easily lead to incomplete bronchial edges or missegmentation. Adaptive thresholding can better adapt to local image variations and achieve better segmentation results for images with uneven grayscale, but it is computationally intensive, and parameter selection can affect the results.
[0023] 2. Region-based algorithms, including region growing and region merging. The region growing method requires the selection of seed points within the bronchus. Its advantages are good reconstruction effects on continuous bronchial regions with relatively uniform grayscale and simple calculations. Its disadvantages are that it is sensitive to seed point selection, easily affected by noise, and may suffer from overgrowth or undergrowth problems. The region merging method initially divides the image into multiple small regions, and then merges similar regions based on similarity metrics between regions, such as grayscale, texture, and other features, to ultimately form a bronchial tree structure. It can handle bronchi with complex shapes and has a certain degree of robustness to noise, but its computational complexity is high, and the choice of similarity metric affects the merging results.
[0024] 3. Model-based algorithms. Algorithms based on anatomical models rely on anatomical knowledge of the bronchial tree to establish a geometric model of the bronchi, including prior models of parameters such as branch angles, lengths, and diameters. The bronchial tree is reconstructed by matching and fitting image data to the model. However, this reliance on prior knowledge limits the model's universality and makes it difficult to adapt to individual differences.
[0025] Fourth, deep learning-based algorithms. Algorithms based on generative adversarial networks (GANs) consist of a generator and a discriminator. The generator attempts to produce realistic bronchial tree images, while the discriminator determines whether the generated images are realistic. Adversarial training is used to optimize the generator and achieve bronchial tree reconstruction. While GAN-based algorithms can produce more realistic and detailed bronchial tree images, the training process is complex and prone to problems such as pattern collapse.
[0026] 5. Algorithms based on mathematical morphology. Dilation and erosion algorithm: The dilation operation expands the bronchial region, while the erosion operation contracts it. By repeatedly performing dilation and erosion operations, noise and small interference areas are removed, and broken parts of the bronchi are connected to reconstruct the bronchial tree. It can effectively handle noise and small voids and enhance the continuity of the bronchi, but it may change the shape and size of the bronchi, so the structural elements and the number of operations need to be carefully selected. Opening and closing operation algorithm: The opening operation first erodes and then dilates to remove small bumps and noise in the image; the closing operation first dilates and then erodes to fill small voids in the image and connect broken bronchi. It can be used to smooth bronchial edges, remove noise, and repair minor defects in the bronchi, but may have limited effect on complex bronchial structures.
[0027] Based on the above analysis, it can be seen that when related technologies generate two-dimensional medical images based on CT images corresponding to the bronchi and perform relevant processing on the two-dimensional medical images to construct the bronchial tree structure, there are problems such as low edge reconstruction accuracy and poor generalization. For some objects, the obtained bronchial tree structure may also have image ghosting, aliasing, and unclear images.
[0028] In view of this, an embodiment of the present invention provides a three-dimensional bronchial image generation device, which can be applied to the fields of image processing and artificial intelligence technology.
[0029] According to an embodiment of the present invention, a three-dimensional bronchial image generation device includes: a global feature extraction unit for inputting three-dimensional medical image data corresponding to the bronchi into I self-attention transformation modules connected in series to perform multi-level self-attention calculations and output global features corresponding to each self-attention transformation module; a local feature extraction unit for using I wavelet downsampling modules to perform convolution calculations and wavelet transforms on the respective input data to obtain local features corresponding to each wavelet downsampling module, wherein the input data of the first wavelet downsampling module is the three-dimensional medical image data, and the input data of the i-th wavelet downsampling module is calculated based on the i-1th local feature output by the i-1th wavelet downsampling module and the i-1th global feature output by the i-1th self-attention transformation module; A feature fusion unit is used to use I wavelet upsampling modules to perform deconvolution calculations and wavelet inverse transforms on their respective input data, and output the fusion features corresponding to each wavelet upsampling module, so as to generate a three-dimensional bronchial image based on the I-th fusion features, wherein the input data of the first wavelet upsampling module is calculated based on the I-th local features output by the I-th wavelet downsampling module and the I-th global features output by the I-th self-attention transform module, and the input data of the i-th wavelet upsampling module is calculated based on the I-i+1th local features output by the I-i+1th wavelet downsampling module, the I-i+1th global features output by the I-i+1th self-attention transform module, and the i-1th fusion features output by the i-1th wavelet upsampling module, i=2,.....,I.
[0030] Figure 1 A schematic structural diagram of a neural network model used in a three-dimensional bronchial image generation device according to an embodiment of the present invention is shown.
[0031] like Figure 1 As shown, the neural network model may include a deep self-attention transformation network 110 and a wavelet transform-based U-shaped network 120. The input data of the neural network model is three-dimensional medical image data corresponding to the bronchi.
[0032] The deep self-attention transformation network 110 may include cascaded I self-attention transformation modules. The wavelet transform-based U-type network 120 may include cascaded I wavelet downsampling modules and cascaded I wavelet upsampling modules.
[0033] According to the embodiment of the present invention, I can be selected according to actual conditions and is not limited here. For example, I can be 4, 6, 8, etc.
[0034] For example, I can be 4, and the four cascaded self-attention transformation modules included in the self-attention transformation network 110 can be a first self-attention transformation module 111, a second self-attention transformation module 112, a third self-attention transformation module 113 and a fourth self-attention transformation module 114.
[0035] The four cascaded wavelet down-sampling modules included in the wavelet transform-based U-shaped network 120 may be a first wavelet down-sampling module 121, a second wavelet down-sampling module 122, a third wavelet down-sampling module 123, and a fourth wavelet down-sampling module 124. The four cascaded wavelet up-sampling modules included in the wavelet transform-based U-shaped network 120 may be a first wavelet up-sampling module 125, a second wavelet up-sampling module 126, a third wavelet up-sampling module 127, and a fourth wavelet up-sampling module 128.
[0036] According to an embodiment of the present invention, 3D medical image data corresponding to the bronchi can first be input into two branch networks: a deep self-attention transform network 110 and a wavelet-based U-shaped network 120. The deep self-attention transform network 110 is then used as a feature extractor to extract global long-range dependencies. The wavelet-based U-shaped network 120 is used as the backbone network, comprising one cascaded wavelet downsampling module and one cascaded wavelet upsampling module, to model and extract local context information.
[0037] According to an embodiment of the present invention, a wavelet-based U-shaped network 120 serves as a backbone network, and its structure includes an encoder and a decoder. One wavelet downsampling module serves as the encoder in the wavelet-based U-shaped network 120, and one wavelet upsampling module serves as the decoder in the wavelet-based U-shaped network 120.
[0038] According to an embodiment of the present invention, the three-dimensional bronchial image generation device provided by the present invention uses a coding mode in the backbone network, and the input data is fused with the feature data output by the deep self-attention transform network. At the same time, the up and down sampling modules in the backbone network use wavelet transform to replace traditional convolution, deconvolution, pooling, interpolation and other modules to refine the segmentation and remove the aliasing effect, and finally output the three-dimensional bronchial tree structure image through the volume rendering algorithm.
[0039] Figure 2 A flow chart of a method for generating a three-dimensional bronchial image according to an embodiment of the present invention is shown.
[0040] According to an embodiment of the present invention, Figure 2 The three-dimensional bronchial image generation method shown is applied to the three-dimensional bronchial image generation device of an embodiment of the present invention.
[0041] like Figure 2As shown, the method for generating a three-dimensional bronchial image may include operations S201 to S203.
[0042] In operation S201, the global feature extraction unit inputs the three-dimensional medical image data corresponding to the bronchus into I self-attention transformation modules connected in series to perform multi-level self-attention calculations, and outputs the global features corresponding to each self-attention transformation module.
[0043] For example, the compressed CT image data can be processed using relevant software to obtain three-dimensional medical image data corresponding to the bronchi, wherein the three-dimensional medical image data includes three-dimensional data corresponding to the bronchi and three-dimensional data corresponding to structures adjacent to the bronchi.
[0044] For example, when I is 4, the three-dimensional medical image data corresponding to the bronchus can be input Figure 1 The four self-attention transformation modules connected in series perform multi-level self-attention calculations. The first self-attention transformation module 111 can perform self-attention calculations on the three-dimensional medical image data and output a first global feature. The second self-attention transformation module 112 can perform self-attention calculations on the first global feature and output a second global feature. The third self-attention transformation module 113 can perform self-attention calculations on the second global feature to obtain a third global feature. The fourth self-attention transformation module 114 can perform self-attention calculations on the third global feature and output a fourth global feature.
[0045] According to an embodiment of the present invention, a global feature extraction unit is used to input the three-dimensional medical imaging data corresponding to the bronchi into I self-attention transformation modules connected in series to perform multi-level self-attention calculations, and a technical means of outputting the global features corresponding to each self-attention transformation module is used. The three-dimensional medical imaging data can be hierarchically extracted using I self-attention transformation modules to obtain global features with multiple scales, thereby achieving the ability of global modeling and extracting global long-distance dependencies.
[0046] In operation S202, the local feature extraction unit uses I wavelet downsampling modules to perform convolution calculation and wavelet transform on the respective input data to obtain local features corresponding to each wavelet downsampling module.
[0047] Among them, the input data of the first wavelet downsampling module is three-dimensional medical image data, and the input data of the i-th wavelet downsampling module is calculated based on the i-1th local feature output by the i-1th wavelet downsampling module and the i-1th global feature output by the i-1th self-attention transformation module.
[0048] For example, the i-1th local feature output by the i-1th wavelet downsampling module and the i-1th global feature output by the i-1th self-attention transformation module can be added to obtain the input data of the i-th wavelet downsampling module.
[0049] For example, when I is 4, four wavelet downsampling modules can be used to perform convolution calculations and wavelet transforms on their respective input data. The first wavelet downsampling module 121 can perform convolution calculations and wavelet transforms on the three-dimensional medical image data corresponding to the bronchi, outputting a first local feature. The second wavelet downsampling module 122 can perform convolution calculations and wavelet transforms on the sum of the first local feature output by the first wavelet downsampling module 121 and the first global feature output by the first self-attention transform module 111, outputting a second local feature. The third wavelet downsampling module 123 can perform convolution calculations and wavelet transforms on the sum of the second local feature output by the second wavelet downsampling module 122 and the second global feature output by the second self-attention transform module 112, outputting a third local feature. The fourth wavelet downsampling module 124 can perform convolution calculations and wavelet transforms on the sum of the third local feature output by the third wavelet downsampling module 123 and the second global feature output by the third self-attention transform module 113, outputting a fourth local feature.
[0050] According to an embodiment of the present invention, a local feature extraction unit utilizes I wavelet downsampling modules to perform convolution calculations and wavelet transforms on respective input data, so that a wavelet transform algorithm can be embedded in a model network structure. While downsampling the three-dimensional bronchial data to reduce the data dimension, the noise in the three-dimensional bronchial data can be removed step by step according to the wavelet transform algorithms in the wavelet downsampling modules at different levels. The characteristics of the wavelet transform algorithm are utilized to suppress noise and unimportant details, and local information with higher precision is extracted step by step to obtain a plurality of local features including local information with higher precision.
[0051] According to an embodiment of the present invention, the local feature extraction unit utilizes I wavelet downsampling modules to perform convolution calculation and wavelet transform on the respective input data, so that the network model can focus on more abstract and advanced feature expressions in each downsampling operation, providing a rich local feature foundation for subsequent feature fusion and segmentation tasks.
[0052] In operation S203, the feature fusion unit uses the wavelet upsampling modules to perform deconvolution and inverse wavelet transform on the input data, outputting fused features corresponding to each wavelet upsampling module. A 3D bronchial image can be generated based on the first fused feature among the fused features corresponding to each wavelet upsampling module.
[0053] The input data of the first wavelet upsampling module is calculated based on the first local feature output by the first wavelet downsampling module and the first global feature output by the first self-attention transformation module. The input data of the i-th wavelet upsampling module is calculated based on the i-i+1th local feature output by the i-i+1th wavelet downsampling module, the i-i+1th global feature output by the i-i+1th self-attention transformation module, and the i-1th fused feature output by the i-1th wavelet upsampling module, where i = 2, ..., 1. The three-dimensional bronchial image includes a three-dimensional bronchial tree structure.
[0054] For example, the first local feature output by the first wavelet downsampling module and the first global feature output by the first self-attention transformation module can be added to obtain the input data of the first wavelet upsampling module.
[0055] For example, the I-i+1th local feature output by the I-i+1th wavelet downsampling module and the I-i+1th global feature output by the I-i+1th self-attention transformation module can be added, and the added feature and the i-1th fusion feature output by the i-1th wavelet upsampling module can be channel-spliced to obtain the input data of the i-th wavelet upsampling module.
[0056] According to an embodiment of the present invention, one wavelet upsampling module can combine local and global features at different levels to obtain a richer feature representation. Each wavelet upsampling module can combine the feature data output by the corresponding feature layer in the encoder to perform feature fusion and refinement. Through feature fusion, the 3D image details corresponding to the bronchi are gradually restored.
[0057] For example, when I is 4, four wavelet upsampling modules can be used to perform deconvolution calculations and inverse wavelet transforms on their respective input data. The first wavelet upsampling module 125 can perform deconvolution calculations and inverse wavelet transforms on the sum of the fourth local feature output by the fourth wavelet downsampling module 124 and the fourth global feature output by the fourth self-attention transform module 114 to obtain a first fused feature.
[0058] The third local feature output by the third wavelet downsampling module 123 and the third global feature output by the third self-attention transform module 113 are added together, and the added feature is channel-joined with the first fused feature output by the first wavelet upsampling module 125 to obtain input data for the second wavelet upsampling module 126. The second wavelet upsampling module 126 performs deconvolution calculation and inverse wavelet transform on the input data of the second wavelet upsampling module 126, and outputs a second fused feature.
[0059] The second local feature output by the second wavelet downsampling module 122 and the second global feature output by the second self-attention transform module 112 are added together, and the added feature is channel-joined with the second fused feature output by the second wavelet upsampling module 126 to obtain input data for the third wavelet upsampling module 127. The third wavelet upsampling module 127 performs deconvolution calculation and inverse wavelet transform on the input data of the third wavelet upsampling module 127, and outputs a third fused feature.
[0060] The first local feature output by the first wavelet downsampling module 121 and the first global feature output by the first self-attention transform module 111 are added together, and the added feature is then channel-joined with the third fused feature output by the third wavelet upsampling module 127 to obtain input data for the fourth wavelet upsampling module 128. The fourth wavelet upsampling module 128 performs deconvolution and inverse wavelet transform on the input data of the fourth wavelet upsampling module 128, and outputs a fourth fused feature, so that a three-dimensional bronchial image can be generated based on the fourth fused feature.
[0061] According to an embodiment of the present invention, each of the wavelet upsampling modules can upsample local features, global features, and fused features at the same scale while further refining the feature information. The wavelet upsampling modules accurately preserve the detailed features of the image corresponding to the bronchi at each level, enabling the final segmentation result to accurately locate the specific location and boundaries of the bronchi.
[0062] According to an embodiment of the present invention, a global feature extraction unit inputs 3D medical image data corresponding to the bronchi into one self-attention transform module connected in series, performing multi-level self-attention calculations. This module extracts global long-range dependencies from the 3D medical image data, generating multi-scale global features. A local feature extraction unit utilizes one wavelet downsampling module to perform convolution and wavelet transform on each input data. This downsamples the bronchial-related input data using wavelet transform algorithms within the wavelet downsampling modules at different levels, simultaneously removing noise and unimportant details from the corresponding input data and extracting multiple local features containing highly accurate local information. The feature fusion unit uses I wavelet upsampling modules to perform deconvolution calculation and wavelet inverse transform on each input data. Each wavelet upsampling module is used to upsample the local features, global features and fusion features of the same scale related to the bronchi, while further refining the feature information corresponding to the bronchi. The detailed features of the three-dimensional image corresponding to the bronchi are accurately retained step by step, so that the final segmentation result can accurately locate the specific position and boundary of the bronchi, thereby reducing the image aliasing effect and improving the edge reconstruction accuracy, obtaining an accurate bronchial tree structure with strong applicability.
[0063] The three-dimensional bronchial image generation method provided in an embodiment of the present invention processes the three-dimensional medical imaging data corresponding to the bronchi to obtain an accurate bronchial tree structure. Based on the accurate bronchial tree structure, it can provide doctors with important reference information to improve the accuracy, efficiency and safety of bronchoscopic operations.
[0064] According to an embodiment of the present invention, Figure 2 In operation S202 shown, the local feature extraction unit uses I wavelet downsampling modules to perform convolution calculation and wavelet transform on the respective input data to obtain the local features corresponding to each wavelet downsampling module. The operation may include the following operations: for each wavelet downsampling module, using the first convolution submodule included in the wavelet downsampling module to perform convolution calculation on the input data of the wavelet downsampling module to obtain a first convolution feature; using the second convolution submodule included in the wavelet downsampling module to perform convolution calculation on the first convolution feature to obtain a second convolution feature; performing wavelet transform and threshold denoising on the second convolution feature to obtain a local feature.
[0065] According to an embodiment of the present invention, a local feature extraction unit performs convolution calculation on the input data of each wavelet downsampling module using the first convolution submodule included in the wavelet downsampling module to obtain a first convolution feature, and then performs convolution calculation on the first convolution feature using the second convolution submodule included in the wavelet downsampling module to obtain a second convolution feature. The second convolution feature is subjected to wavelet transform and threshold denoising to obtain a local feature. This can downsample the three-dimensional medical imaging data corresponding to the bronchi, gradually reducing the spatial size of the feature map while increasing the number of channels to extract local feature information at different scales. At the same time, each downsampling operation using the wavelet downsampling module causes the network model to focus on more abstract and advanced feature expressions, providing a rich local feature foundation for subsequent feature fusion and segmentation tasks.
[0066] According to an embodiment of the present invention, the I wavelet downsampling modules have the same structure and function. Figure 1 Taking the first wavelet down-sampling module 121 in FIG. 1 as an example, the specific structure and function of each wavelet down-sampling module are further explained.
[0067] For example, the local feature extraction unit may input the 3D medical image data corresponding to the bronchi into the two convolution submodules included in the first wavelet downsampling module 121, perform two consecutive convolution calculations, and obtain a second convolution feature. Subsequently, the second convolution feature is input into the wavelet transform sampling layer in the second convolution submodule for wavelet transform and threshold denoising to obtain a local feature.
[0068] Figure 3A schematic structural diagram of a wavelet downsampling module according to an embodiment of the present invention is shown.
[0069] like Figure 3 As shown, the first wavelet downsampling module 121 includes a first convolution submodule 1211 and a second convolution submodule 1212 .
[0070] The first convolution submodule 1211 can perform a convolution calculation on the input data of the first wavelet downsampling module 121 to obtain a first convolution feature. The second convolution submodule 1212 can perform a convolution calculation on the first convolution feature to obtain a second convolution feature, and perform a wavelet transform and threshold denoising on the second convolution feature to obtain a local feature, which is now a first local feature.
[0071] The following will be Figure 3 Taking the first convolution submodule 1211 included in the first wavelet downsampling module 121 as an example, the local feature extraction unit performs convolution calculation on the input data of the wavelet downsampling module for each wavelet downsampling module using the first convolution submodule included in the wavelet downsampling module to obtain the first convolution feature for further interpretation.
[0072] like Figure 3 As shown, the first convolution submodule 1211 may include a first convolution layer 1211_1, a first normalization layer 1211_2, and a first linear rectification activation function layer 1211_3.
[0073] The local feature extraction unit uses the first convolution submodule 1211 included in the first wavelet downsampling module 121 to perform convolution calculation on the input data of the first wavelet downsampling module 121 to obtain the first convolution feature, which may include: using the first convolution layer 1211_1 to perform convolution calculation on the input data of the wavelet downsampling module 121 to obtain a first initial convolution feature; using the first normalization layer 1211_2 to normalize the first initial convolution feature to obtain a first normalized feature; using the first linear rectification activation function layer 1211_3 to perform nonlinear mapping on the first normalized feature to obtain a first convolution feature.
[0074] The following will be Figure 3 Taking the second convolution submodule 1212 included in the first wavelet downsampling module 121 as an example, the local feature extraction unit uses the second convolution submodule included in the wavelet downsampling module to perform convolution calculation on the first convolution feature to obtain the second convolution feature, and performs wavelet transform and threshold denoising on the second convolution feature to obtain the local feature for further interpretation.
[0075] like Figure 3As shown, the second convolution submodule 1212 includes a second convolution layer 1212_1, a second normalization layer 1212_2, a second linear rectification activation function layer 1212_3 and a wavelet transform sampling layer 1212_4.
[0076] The local feature extraction unit uses the second convolution submodule 1212 included in the first wavelet downsampling module 121 to perform convolution calculation on the first convolution feature to obtain a second convolution feature, and performs wavelet transform and threshold denoising on the second convolution feature. Obtaining the local feature may include: using the second convolution layer 1212_1 to perform convolution calculation on the first convolution feature to obtain a second initial convolution feature; using the second normalization layer 1212_2 to normalize the second initial convolution feature to obtain a second normalized feature; using the second linear rectification activation function layer 1212_3 to perform nonlinear mapping on the second normalized feature to obtain a second convolution feature; using the wavelet transform sampling layer 1212_4 to perform wavelet transform on the second convolution feature to obtain multiple frequency band data, and according to a preset threshold, performing feature extraction on the multiple frequency band data to obtain local features.
[0077] According to an embodiment of the present invention, both the first convolutional layer 1211_1 and the second convolutional layer 1212_1 are 3D convolutional layers. The convolution kernels of the first convolutional layer 1211_1 and the second convolutional layer 1212_1 are both 3*3*3 in size, where the first 3 represents the height of the convolution kernel, the second 3 represents the width of the convolution kernel, and the third 3 represents the depth of the convolution kernel. The first convolutional layer 1211_1 can be used to perform a 3D convolution calculation on the input data of the first wavelet downsampling module 121 to obtain a first initial convolution feature. The second convolutional layer 1212_1 can be used to perform a 3D convolution calculation on the first convolution feature to obtain a second initial convolution feature.
[0078] According to an embodiment of the present invention, the local feature extraction unit uses the first convolution layer 1211_1 to perform 3D convolution calculation on the input data of the first wavelet downsampling module 121 to obtain a first initial convolution feature, thereby downsampling the input data of the first wavelet downsampling module 121 using the first convolution layer 1211_1, reducing the dimension of the input data, and obtaining a first initial convolution feature with a lower data dimension.
[0079] According to an embodiment of the present invention, the multiple frequency band data include frequency band data included in a first frequency band component obtained after the second convolution feature is processed by the first filter coefficient, and frequency band data respectively included in seven second frequency band components obtained after the second convolution feature is processed by the second filter coefficient.
[0080] According to an embodiment of the present invention, the first filter coefficient may be a low-pass filter coefficient. The second filter coefficient may be a high-pass filter coefficient. The first filter coefficient and the second filter coefficient are both filter coefficients used by the wavelet transform sampling layer 1212_4 when performing a wavelet transform on the second convolution feature. The first frequency band component is a low-frequency component. The second frequency band component is a high-frequency component.
[0081] For example, based on a preset threshold, threshold denoising can be performed on the frequency band data of each of the seven second frequency band components. Frequency band data greater than the preset threshold can be removed from the frequency band data of each of the seven second frequency band components. Local features can then be obtained based on the frequency band data of the first frequency band component and the frequency band data of each of the seven second frequency band components that have undergone threshold denoising. The preset threshold can be selected based on actual circumstances and is not limited herein.
[0082] According to an embodiment of the present invention, the local feature extraction unit performs a wavelet transform on the second convolution feature using wavelet transform sampling layer 1212_4 to obtain multiple frequency band data. Feature extraction is then performed on the multiple frequency band data based on a preset threshold to obtain local features. This reduces the data dimension of the second convolution feature, downsamples the second convolution feature, and utilizes the characteristics of the wavelet transform to suppress noise. Furthermore, high-frequency components in the multiple frequency band data are filtered after threshold denoising to remove some noise and unimportant details, resulting in highly accurate local information corresponding to the bronchi, i.e., local features.
[0083] According to an embodiment of the present invention, corresponding to the encoder, the decoder performs an upsampling operation on the input data, and restores the feature map reduced in the downsampling process to the original image, that is, the size of the three-dimensional image formed by the three-dimensional medical imaging data corresponding to the bronchus.
[0084] According to an embodiment of the present invention, the decoder restores the compressed feature map to its original size through operations such as deconvolution and upsampling. It then combines the feature data transmitted by the encoder for further feature fusion and refinement. Its deconvolution module is similar to the convolution module in the encoder. The decoder's input data passes through the deconvolution layer in the decoder's deconvolution module and then enters the inverse wavelet transform sampling layer for image size restoration. The final output is a feature map of a U-shaped network combined with a wavelet transform, namely a three-dimensional bronchial image.
[0085] According to an embodiment of the present invention, Figure 2In operation S203 shown, the feature fusion unit uses I wavelet upsampling modules to perform deconvolution calculation and wavelet inverse transform on the respective input data, and outputs the fusion features corresponding to each wavelet upsampling module. The operation may include the following operations: for each wavelet upsampling module, use the first deconvolution submodule included in the wavelet upsampling module to perform deconvolution calculation on the input data of the wavelet upsampling module to obtain a first deconvolution feature; use the second deconvolution submodule included in the wavelet upsampling module to perform deconvolution calculation on the first deconvolution feature to obtain a second deconvolution feature; and perform wavelet inverse transform on the second deconvolution feature to obtain a fusion feature.
[0086] According to an embodiment of the present invention, the first deconvolution submodule may include a deconvolution layer, a normalization layer, and a linear rectification activation function layer. The feature fusion unit performs deconvolution calculation on the input data of the wavelet upsampling module using the first deconvolution submodule included in the wavelet upsampling module for each wavelet upsampling module, and obtaining the first deconvolution feature may include: using the deconvolution layer, normalization layer, and linear rectification activation function layer included in the first deconvolution submodule to sequentially perform deconvolution, normalization, and linear rectification operations on the input data of the wavelet upsampling module to obtain the first deconvolution feature. Among them, the input data of the deconvolution layer is the input data of the wavelet upsampling module, the input data of the normalization layer is the output data of the deconvolution layer, and the input data of the linear rectification activation function layer is the output data of the normalization layer.
[0087] According to an embodiment of the present invention, the second deconvolution submodule differs from the first deconvolution submodule in that: in addition to including a deconvolution layer, a normalization layer, and a linear rectification activation function layer, the second deconvolution submodule also includes a wavelet inverse transform sampling layer. The feature fusion unit uses the second deconvolution submodule included in the wavelet upsampling module to perform deconvolution calculation on the first deconvolution feature, and obtaining the second deconvolution feature may include: using the deconvolution layer, normalization layer, and linear rectification activation function layer included in the second deconvolution submodule to perform deconvolution, normalization, and linear rectification operations on the first deconvolution feature in sequence to obtain the second deconvolution feature. The feature fusion unit performs an inverse wavelet transform on the second deconvolution feature to obtain a fused feature, which includes: using the inverse wavelet transform sampling layer included in the second deconvolution submodule to perform an inverse wavelet transform on the second deconvolution feature to obtain a fused feature.
[0088] According to an embodiment of the present invention, the feature fusion unit uses the deconvolution layer, normalization layer, and linear rectification activation function layer included in the second deconvolution submodule to sequentially perform deconvolution, normalization, and linear rectification operations on the first deconvolution feature to obtain a second deconvolution feature, thereby achieving the use of the deconvolution operation to sample the first deconvolution feature back to a high-resolution feature. The feature fusion unit uses the wavelet inverse transform sampling layer included in the second deconvolution submodule to perform wavelet inverse transform on the second deconvolution feature to obtain a fused feature, thereby achieving the upsampling of the second deconvolution feature back to a high-resolution feature.
[0089] According to an embodiment of the present invention, the deconvolution layer included in the first deconvolution submodule and the deconvolution layer included in the second deconvolution submodule are both 3D deconvolution layers.
[0090] According to an embodiment of the present invention, a deep self-attention transformer network focuses on learning long-range dependencies between three-dimensional medical imaging data corresponding to bronchi.
[0091] According to an embodiment of the present invention, the attention mechanism used by the deep self-attention transformation network originates from the human visual system and is a key technology in deep learning, which allows the corresponding neural network to selectively focus on important information. The weights obtained by the attention mechanism represent the importance of the information at each position. The attention mechanism allows the corresponding neural network model to automatically focus on key information when processing data, so that in the 3D bronchial segmentation task, the deep self-attention transformation network can focus on the data in the bronchial area, reduce the interference of irrelevant information such as background, and improve the accuracy of segmentation. When processing complex 3D medical imaging data, it can effectively reduce the amount of calculation, improve computing efficiency, and enable the corresponding neural network model to run better under limited resources. It also enhances the interpretability of the corresponding neural network model. Through the attention weight, the focus of the corresponding neural network model on different areas can be intuitively understood, which facilitates the analysis of the decision-making process of the corresponding neural network model.
[0092] According to an embodiment of the present invention, Figure 2The operation shown, the global feature extraction unit inputs the three-dimensional medical imaging data corresponding to the bronchus into I self-attention transformation modules connected in series to perform multi-level self-attention calculations, and outputs the global features corresponding to each self-attention transformation module. It can include the following operations: for each self-attention transformation module, the window-based multi-head self-attention sub-module included in the self-attention transformation module is used to perform non-overlapping division on the input data of the self-attention transformation module to obtain multiple first window data, and multi-head self-attention calculations are performed on the multiple first window data respectively and the calculation results are spliced to obtain self-attention features, wherein the input data of the first self-attention transformation module is three-dimensional medical imaging data, and the input data of the i-th self-attention transformation module is the i-1-th global feature output by the i-1-th self-attention transformation module; the self-attention features are overlapped and divided using the moving window-based multi-head self-attention sub-module included in the self-attention transformation module to obtain multiple second window data, and multi-head self-attention calculations are performed on the multiple second window data respectively and the calculation results are spliced to obtain global features.
[0093] According to an embodiment of the present invention, the I self-attention transformation modules have the same structure and function. Figure 1 Taking the first self-attention transformation module 111 in the example, the specific structure and function of each self-attention transformation module are further explained.
[0094] Figure 4 A structural diagram of a self-attention transformation module according to an embodiment of the present invention is shown.
[0095] like Figure 4 As shown, the first self-attention transformation module 111 may include a window-based multi-head self-attention submodule 1111 and a moving window-based multi-head self-attention submodule 1112.
[0096] The global feature extraction unit inputs the three-dimensional medical imaging data corresponding to the bronchi into the first self-attention transformation module 111 for self-attention calculation, and outputs the global features corresponding to the first self-attention transformation module 111, which may include the following operations: using the window-based multi-head self-attention sub-module 1111 included in the first self-attention transformation module 111 to perform non-overlapping division on the input data of the first self-attention transformation module 111 to obtain multiple first window data, performing multi-head self-attention calculation on the multiple first window data respectively and splicing the calculation results to obtain self-attention features; using the moving window-based multi-head self-attention sub-module 1112 included in the first self-attention transformation module 111 to perform overlapping division on the self-attention features to obtain multiple second window data, performing multi-head self-attention calculation on the multiple second window data respectively and splicing the calculation results to obtain global features.
[0097] The following will be Figure 4 Taking the window-based multi-head self-attention sub-module 1111 included in the first self-attention transformation module 111 as an example, for each self-attention transformation module in the global feature extraction unit, the window-based multi-head self-attention sub-module included in the self-attention transformation module is used to perform non-overlapping division on the input data of the self-attention transformation module to obtain multiple first window data, and multi-head self-attention calculations are performed on the multiple first window data respectively and the calculation results are spliced to obtain self-attention features for further interpretation.
[0098] According to an embodiment of the present invention, the number of feature extraction layers included in the window-based multi-head self-attention submodule 1111 and the moving window-based multi-head self-attention submodule 1112 can be selected according to actual conditions and is not limited here. For example, the number of feature extraction layers included in the window-based multi-head self-attention submodule 1111 and the moving window-based multi-head self-attention submodule 1112 can be 12 layers.
[0099] like Figure 4 As shown, the window-based multi-head self-attention submodule 1111 includes a third normalization layer 1111_1, a window-based multi-head self-attention layer 1111_2, a fourth normalization layer 1111_3 and a first fully connected layer 1111_4.
[0100] The global feature extraction unit uses the window-based multi-head self-attention submodule 1111 included in the first self-attention transformation module 111 to perform non-overlapping division on the input data of the first self-attention transformation module 111 to obtain multiple first window data, and performs multi-head self-attention calculation on the multiple first window data respectively and splices the calculation results to obtain the self-attention feature, which may include: using the third normalization layer 1111_1 to normalize the input data of the first self-attention transformation module 111 to obtain a third normalized feature; performing non-overlapping division on the third normalized feature to obtain multiple first window data, and using the window-based multi-head self-attention submodule 1111 to obtain the first window data. The attention layer 1111_2 performs multi-head self-attention calculation on multiple first window data and splices the calculation results to obtain a first initial self-attention feature; performs residual connection on the input data of the first self-attention transformation module 111 and the first initial self-attention feature to obtain a first residual connection feature; uses the fourth normalization layer 1111_3 to normalize the first residual connection feature to obtain a fourth normalized feature; uses the first fully connected layer 1111_4 to fully connect the fourth normalized feature to obtain a first fully connected feature; performs residual connection on the first fully connected feature and the first residual connection feature to obtain a self-attention feature.
[0101] For example, when performing non-overlapping partitioning on the third normalized feature, the window size can be 4×4×4, obtaining a plurality of first window data with non-overlapping data, where the first 4 is the height of the window, the second 4 is the width of the window, and the third 4 is the depth of the window.
[0102] For example, the input data of the self-attention transformation module can be written as ,The third normalization layer of the self-attention transformation module first performs layer normalization operation on it. Layer normalization is to normalize each input data in the feature dimension, as shown in formula (1).
[0103] (1);
[0104] in, yes The mean of yes The variance of This is to prevent the denominator from reaching a minimum value of zero. In the process of training the neural network model used in the three-dimensional bronchial image generation method, this step can stabilize the data distribution and improve the training effect.
[0105] According to an embodiment of the present invention, by aggregating the input data of the same layer and calculating the mean and variance, the input data of each layer is normalized. This can make the distribution of the input data of each layer in the neural network model relatively stable. At the same time, when training the neural network model, this normalization process can accelerate the learning speed of the neural network model.
[0106] According to an embodiment of the present invention, the multidimensional third normalized features that have undergone layer normalization are divided into multiple non-overlapping multidimensional first window data, so as to split the large-size multidimensional data into multiple small window data, narrow the calculation scope, and reduce the amount of calculation.
[0107] For example, the process of using the window-based multi-head self-attention layer 1111_2 to perform multi-head self-attention calculations on each first window data is as follows: the number of heads is set to , for the first window data , respectively through The linear transformation of the group is obtained Group query (Query, denoted as ), key (Key, denoted as ) and value (Value, denoted as )matrix, ; Calculate the attention score of each group ,in, yes (or ) dimension, used to scale the dot product result to avoid the dot product result being too large and causing the normalized function gradient to disappear; the attention score of each group Use a normalization function, such as the softmax function, to normalize and obtain the attention weight ; The attention weight Multiply with the corresponding value matrix and sum to get the output of each head .Bundle Output of the head Splice by channel dimension to get splicing features , and then through a linear transformation , get the calculation results of the window-based multi-head self-attention within the window ,Right now .
[0108] For example, the process of splicing the calculation results corresponding to multiple first window data using the window-based multi-head self-attention layer 1111_2 is as follows: Transform the input data of the self-attention module The structure of the splicing is restored to restore the shape of the 3D feature map and obtain the first initial self-attention feature .
[0109] For example, a residual connection can be used to With input data Add together to get the first residual connection feature ,in, Used to mark the data before processing by the self-attention transformation module, Used to mark the data processed by the self-attention transformation module. Then, Input Feedforward Neural Network (FFN), FFN includes the fourth normalization layer 1111_3 and the first fully connected layer 1111_4. Perform layer normalization to obtain the fourth normalized feature. Then use the two fully connected layers in the first fully connected layer 1111_4 (with ReLU (Rectified Linear Unit, rectified linear unit) activation function in the middle) to fully connect the fourth normalized feature to obtain the first fully connected feature. Finally, the first fully connected feature is connected with the first residual feature. Perform residual connection to obtain the final output self-attention feature.
[0110] According to an embodiment of the present invention, the fully connected layer may include a fully connected feedforward network, which is mainly composed of two layers of fully connected linear units, with a Gaussian error linear unit used as an activation function in the middle.
[0111] According to an embodiment of the present invention, the number of channels of the hidden sublayer in the window-based multi-head self-attention layer included in each of the I self-attention transformation modules is {48, 96, 192, 384}, and the number of sublayers is {2, 2, 6, 2}.
[0112] According to an embodiment of the present invention, when the global feature extraction unit uses a window-based multi-head self-attention sub-module to perform window-based multi-head self-attention calculation on the input data, although the amount of calculation is reduced after the window is divided, the information between each window cannot be shared. Therefore, a moving window-based multi-head self-attention sub-module can be used to perform moving window-based multi-head self-attention calculation on the input data, and by moving the window, content interaction between windows can be achieved.
[0113] Specifically, the moving window shifts voxels in the input data in a cyclic manner, using half the window size as the shift length to change the position of the voxels in the input data. The shifted input data is then divided into different windows through window partitioning to ensure that the shifted window and the original window each contain feature information from different locations of the input data, thereby achieving information exchange between different windows. Simultaneously, the feature extraction advantages of the multi-head self-attention mechanism are leveraged on the moving window to model the global dependency between input and output. The moving window-based multi-head self-attention submodule can effectively extract feature information from the input data, thereby capturing long-term dependencies in the input data and achieving global modeling capabilities.
[0114] The following will be Figure 4 Taking the multi-head self-attention sub-module 1112 based on moving window included in the first self-attention transformation module 111 as an example, the self-attention features are overlapped and divided by the multi-head self-attention sub-module based on moving window included in the self-attention transformation module to obtain multiple second window data, and multi-head self-attention calculations are performed on the multiple second window data respectively and the calculation results are spliced to obtain global features for further interpretation.
[0115] like Figure 4 As shown, the moving window-based multi-head self-attention submodule 1112 may include a fifth normalization layer 1112_1, a moving window-based multi-head self-attention layer 1112_2, a sixth normalization layer 1112_3 and a second fully connected layer 1112_4.
[0116] The global feature extraction unit uses the multi-head self-attention submodule 1111 based on the moving window included in the first self-attention transformation module 111 to perform overlapping division on the self-attention feature to obtain multiple second window data, and performs multi-head self-attention calculation on the multiple second window data respectively and splices the calculation results. The global feature can include: using the fifth normalization layer 1112_1 to normalize the self-attention feature to obtain the fifth normalized feature; overlapping division on the fifth normalized feature to obtain multiple second window data, and using the multi-head self-attention layer 1112_1 based on the moving window to obtain the global feature. 112_2 performs multi-head self-attention calculation on multiple second window data and splices the calculation results to obtain a second initial self-attention feature; performs residual connection on the self-attention feature and the second initial self-attention feature to obtain a second residual connection feature; uses the sixth normalization layer 1112_3 to normalize the second residual connection feature to obtain a sixth normalized feature; uses the second fully connected layer 1112_4 to fully connect the sixth normalized feature to obtain a second fully connected feature; performs residual connection on the second fully connected feature and the second residual connection feature to obtain a global feature.
[0117] For example, when overlapping partitioning is performed on the fifth normalized feature, the window size may be 4×4×4, and a second window of data with multiple overlapping data is obtained, where the first 4 is the height of the window, the second 4 is the width of the window, and the third 4 is the depth of the window.
[0118] According to an embodiment of the present invention, the fifth normalization layer 1112_1, the moving window-based multi-head self-attention layer 1112_2, the sixth normalization layer 1112_3 and the second fully connected layer 1112_4 have similar structures and functions to the third normalization layer 1111_1, the window-based multi-head self-attention layer 1111_2, the fourth normalization layer 1111_3 and the first fully connected layer 1111_4, respectively, and for simplicity, they are not repeated here.
[0119] According to an embodiment of the present invention, before processing three-dimensional medical image data corresponding to bronchi using a three-dimensional bronchial image generation method, it is necessary to pre-train a neural network model used in the three-dimensional bronchial image generation method.
[0120] For example, 3D medical imaging data corresponding to the bronchi of multiple different subjects can be collected, and necessary preprocessing can be performed on the data to remove unclear, blurred, or other non-compliant data. Then, true 3D bronchial images can be constructed from the multiple 3D medical imaging data that meet the requirements, thereby obtaining a 3D medical imaging data dataset. The subjects can be humans. The 3D medical imaging data dataset includes multiple samples, each of which includes 3D medical imaging data corresponding to the bronchi of the subject and a true 3D bronchial image corresponding to the subject.
[0121] For each of a plurality of samples included in a 3D medical image data dataset, 3D medical image data corresponding to a bronchi of a subject included in each sample may be input into a neural network model, and a predicted 3D bronchial image may be output using the neural network model. The predicted 3D bronchial image and a true 3D bronchial image corresponding to the subject may be input into a loss function formula based on the Dice coefficient to calculate a loss value, and model parameters of the neural network model may be updated based on the loss value.
[0122] When the number of training times of the neural network model using multiple samples reaches the preset number or the prediction accuracy of the neural network model reaches the preset accuracy, the trained neural network model can be used for Figure 2 Neural network model for the three-dimensional bronchial image generation method.
[0123] According to the three-dimensional bronchial image generation method provided by an embodiment of the present invention, the main network architecture is used to extract local features of the three-dimensional medical imaging data corresponding to the bronchi. The global features extracted by the self-attention transformation module are then integrated, and the segmentation is refined using the wavelet transformation module to remove the aliasing effect. Finally, the bronchial tree structure image is output through the volume rendering algorithm.
[0124] According to the three-dimensional bronchial image generation method provided by an embodiment of the present invention, a deep self-attention transform cross algorithm is used to extract the three-dimensional bronchial tree structure from the three-dimensional medical imaging data corresponding to the bronchi, and a wavelet transform is embedded to remove the image aliasing effect in bronchial reconstruction, which can ultimately be used to assist doctors in bronchial approach navigation.
[0125] The three-dimensional bronchial image generation method provided according to an embodiment of the present invention has higher accuracy than the existing technology. It can generate three-dimensional bronchial tree images end-to-end based on three-dimensional medical imaging data corresponding to the bronchi, solving the problem that CT images cannot intuitively reflect the bronchial tree structure, thereby assisting doctors in guiding treatment.
[0126] Based on the above-mentioned 3D bronchial image generation method, an embodiment of the present invention provides a 3D bronchial image generation device. Figure 5 The device is described in detail.
[0127] Figure 5 A block diagram of a three-dimensional bronchial image generating apparatus according to an embodiment of the present invention is shown.
[0128] like Figure 5 As shown, the three-dimensional bronchial image generating device 500 of this embodiment includes a global feature extraction unit 510 , a local feature extraction unit 520 and a feature fusion unit 530 .
[0129] The global feature extraction unit 510 is used to input the three-dimensional medical image data corresponding to the bronchi into I self-attention transformation modules connected in series to perform multi-level self-attention calculations and output the global features corresponding to each self-attention transformation module.
[0130] The local feature extraction unit 520 is configured to perform convolution calculations and wavelet transforms on the input data of each of the wavelet downsampling modules, respectively, to obtain local features corresponding to each wavelet downsampling module. The input data of the first wavelet downsampling module is 3D medical image data, and the input data of the i-th wavelet downsampling module is calculated based on the i-1th local feature output by the i-1th wavelet downsampling module and the i-1th global feature output by the i-1th self-attention transform module.
[0131] Feature fusion unit 530 is configured to perform deconvolution and inverse wavelet transform on the input data of each of the I wavelet upsampling modules, and output fused features corresponding to each wavelet upsampling module, so as to generate a three-dimensional bronchial image based on the first fused features. The input data of the first wavelet upsampling module is calculated based on the first local features output by the first wavelet downsampling module and the first global features output by the first self-attention transform module. The input data of the i-th wavelet upsampling module is calculated based on the i-i+1th local features output by the i-i+1th wavelet downsampling module, the i-i+1th global features output by the i-i+1th self-attention transform module, and the i-1th fused features output by the i-1th wavelet upsampling module, where i = 2, ..., 1.
[0132] It should be noted that the three-dimensional bronchial image generation device part in the embodiment of the present invention corresponds to the three-dimensional bronchial image generation method part in the embodiment of the present invention. The description of the three-dimensional bronchial image generation device part specifically refers to the three-dimensional bronchial image generation method part, which will not be repeated here.
[0133] Figure 6 A block diagram of an electronic device suitable for implementing the above-described three-dimensional bronchial image generation method according to an embodiment of the present invention is shown. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0134] like Figure 6As shown, an electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0135] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the programs in the ROM 602 and / or RAM 603 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.
[0136] According to an embodiment of the present invention, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.
[0137] According to an embodiment of the present invention, the method flow according to an embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present invention are executed. According to an embodiment of the present invention, the system, device, apparatus, module, unit, etc. described above can be implemented by a computer program module.
[0138] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0139] According to embodiments of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0140] For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603 .
[0141] An embodiment of the present invention also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present invention. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the method provided by the embodiment of the present invention.
[0142] When the computer program is executed by the processor 601, the above functions defined in the system / device of the embodiment of the present invention are performed. According to the embodiment of the present invention, the above-described systems, devices, modules, units, etc. can be implemented by computer program modules.
[0143] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0144] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0146] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present invention is defined by the accompanying embodiments and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present invention.
Claims
1. A three-dimensional bronchial image generation device, characterized in that: The device comprises: a global feature extraction unit, configured to input the three-dimensional medical image data corresponding to the bronchi into I self-attention transformation modules connected in series, perform multi-level self-attention calculations, and output the global features corresponding to each of the self-attention transformation modules; a local feature extraction unit, configured to perform convolution calculation and wavelet transform on respective input data using I wavelet downsampling modules, respectively, to obtain local features corresponding to respective wavelet downsampling modules, wherein the input data of the first wavelet downsampling module is the three-dimensional medical image data, and the input data of the i-th wavelet downsampling module is calculated based on the i-1th local feature output by the i-1th wavelet downsampling module and the i-1th global feature output by the i-1th self-attention transform module; A feature fusion unit is used to use I wavelet upsampling modules to perform deconvolution calculations and wavelet inverse transforms on their respective input data, and output the fusion features corresponding to each of the wavelet upsampling modules, so as to generate a three-dimensional bronchial image based on the I fusion features, wherein the input data of the first wavelet upsampling module is calculated based on the I local features output by the I wavelet downsampling module and the I global features output by the I self-attention transform module, and the input data of the i-th wavelet upsampling module is calculated based on the i-i+1 local features output by the i-i+1 wavelet downsampling module, the i-i+1 global features output by the i-i+1 self-attention transform module, and the i-1 fusion features output by the i-1 wavelet upsampling module, i=2,.....,I.
2. The device according to claim 1, characterized in that The local feature extraction unit uses I wavelet downsampling modules to perform convolution calculation and wavelet transform on the respective input data, and obtains the local features corresponding to each of the wavelet downsampling modules, including: For each wavelet downsampling module, using a first convolution submodule included in the wavelet downsampling module to perform convolution calculation on input data of the wavelet downsampling module to obtain a first convolution feature; The second convolution submodule included in the wavelet downsampling module is used to perform convolution calculation on the first convolution feature to obtain a second convolution feature, and the second convolution feature is subjected to wavelet transform and threshold denoising to obtain the local feature.
3. The device according to claim 1, characterized in that The feature fusion unit uses I wavelet upsampling modules to perform deconvolution calculation and wavelet inverse transform on the respective input data, and outputs the fusion features corresponding to each of the wavelet upsampling modules, including: For each wavelet upsampling module, using a first deconvolution submodule included in the wavelet upsampling module to perform deconvolution calculation on input data of the wavelet upsampling module to obtain a first deconvolution feature; The first deconvolution feature is deconvolved by using a second deconvolution submodule included in the wavelet upsampling module to obtain a second deconvolution feature; and the second deconvolution feature is inversely transformed by wavelet to obtain the fusion feature.
4. The device according to claim 1, characterized in that The global feature extraction unit inputs the three-dimensional medical image data corresponding to the bronchus into the I self-attention transformation modules connected in series to perform multi-level self-attention calculations, and outputs the global features corresponding to each of the self-attention transformation modules, including: For each self-attention transformation module, the input data of the self-attention transformation module is non-overlappedly divided using the window-based multi-head self-attention submodule included in the self-attention transformation module to obtain multiple first window data, and the multiple first window data are respectively subjected to multi-head self-attention calculation and the calculation results are spliced to obtain self-attention features, wherein the input data of the first self-attention transformation module is the three-dimensional medical image data, and the input data of the i-th self-attention transformation module is the i-1th global feature output by the i-1th self-attention transformation module; The self-attention features are overlapped and divided using the multi-head self-attention sub-module based on moving windows included in the self-attention transformation module to obtain multiple second window data. Multi-head self-attention calculations are performed on the multiple second window data respectively and the calculation results are spliced to obtain the global features.
5. The device according to claim 2, characterized in that The first convolution submodule includes a first convolution layer, a first normalization layer and a first linear rectification activation function layer; The local feature extraction unit performs convolution calculation on the input data of each wavelet downsampling module using the first convolution submodule included in the wavelet downsampling module to obtain the first convolution feature including: Using the first convolution layer to perform convolution calculation on the input data of the wavelet downsampling module to obtain a first initial convolution feature; Normalizing the first initial convolutional features using the first normalization layer to obtain first normalized features; The first normalized features are nonlinearly mapped using the first linear rectification activation function layer to obtain the first convolution features.
6. The device according to claim 5, characterized in that The second convolution submodule includes a second convolution layer, a second normalization layer, a second linear rectification activation function layer and a wavelet transform sampling layer; The local feature extraction unit uses the second convolution submodule included in the wavelet downsampling module to perform convolution calculation on the first convolution feature to obtain a second convolution feature, and performs wavelet transform and threshold denoising on the second convolution feature to obtain the local feature including: Performing a convolution calculation on the first convolution feature using the second convolution layer to obtain a second initial convolution feature; Normalizing the second initial convolutional features using the second normalization layer to obtain second normalized features; Performing nonlinear mapping on the second normalized features using the second linear rectification activation function layer to obtain the second convolution features; The second convolution feature is subjected to wavelet transform by using the wavelet transform sampling layer to obtain a plurality of frequency band data, and feature extraction is performed on the plurality of frequency band data according to a preset threshold to obtain the local feature.
7. The device according to claim 4, characterized in that The window-based multi-head self-attention submodule includes a third normalization layer, a window-based multi-head self-attention layer, a fourth normalization layer and a first fully connected layer; The global feature extraction unit performs non-overlapping partitioning of the input data of each self-attention transformation module using the window-based multi-head self-attention submodule included in the self-attention transformation module to obtain a plurality of first window data, performs multi-head self-attention calculations on the plurality of first window data respectively, and splices the calculation results to obtain self-attention features including: Normalizing the input data of the self-attention transformation module using the third normalization layer to obtain a third normalized feature; Performing non-overlapping division on the third normalized feature to obtain the plurality of first window data, performing multi-head self-attention calculation on the plurality of first window data respectively using the window-based multi-head self-attention layer and concatenating the calculation results to obtain a first initial self-attention feature; Performing a residual connection on the input data of the self-attention transformation module and the first initial self-attention feature to obtain a first residual connection feature; Normalizing the first residual connection feature using the fourth normalization layer to obtain a fourth normalized feature; Performing a full connection on the fourth normalized feature using the first fully connected layer to obtain a first fully connected feature; Perform a residual connection on the first fully connected feature and the first residual connection feature to obtain the self-attention feature.
8. The device according to claim 4, characterized in that The moving window-based multi-head self-attention submodule includes a fifth normalization layer, a moving window-based multi-head self-attention layer, a sixth normalization layer and a second fully connected layer; The global feature extraction unit uses the multi-head self-attention submodule based on the moving window included in the self-attention transformation module to perform overlapping division on the self-attention feature to obtain multiple second window data, performs multi-head self-attention calculation on the multiple second window data respectively and splices the calculation results to obtain the global feature including: Normalizing the self-attention feature using the fifth normalization layer to obtain a fifth normalized feature; Perform overlapping division on the fifth normalized feature to obtain the plurality of second window data, perform multi-head self-attention calculation on the plurality of second window data respectively using the multi-head self-attention layer based on the moving window and splice the calculation results to obtain a second initial self-attention feature; Performing a residual connection on the self-attention feature and the second initial self-attention feature to obtain a second residual connection feature; Normalizing the second residual connection feature using the sixth normalization layer to obtain a sixth normalized feature; Performing a full connection on the sixth normalized feature using the second fully connected layer to obtain a second fully connected feature; Performing a residual connection on the second fully connected feature and the second residual connection feature to obtain the global feature.
9. The device according to claim 6, characterized in that The multiple frequency band data include frequency band data included in a first frequency band component obtained after the second convolution feature is processed by the first filter coefficient, and frequency band data included in each of seven second frequency band components obtained after the second convolution feature is processed by the second filter coefficient.
10. The device according to claim 6, characterized in that The sizes of the convolution kernels of the first convolution layer and the second convolution layer are both 3*3*3.
Citation Information
Patent Citations
U-shaped image segmentation network based on convolution enhanced cross self-attention deformer
CN115908805A
Remote sensing image building extraction method and system based on local-global features
CN119785218A
Cited By
Respiratory tract endoscope image definition enhancement processing method
CN122023185A
A method for enhancing the definition of images of endoscopic examinations of the respiratory tract
CN122023185B