Method, device and medium for reconstructing biological tubular structures in electron microscopy micrographs
Through the biotubular structure segmentation network with self-attention mechanism and the tracking network of skeleton point prediction, the segmentation and tracking errors of biological tubular structure reconstruction in electron microscope images are solved, and a higher accuracy of three-dimensional reconstruction is achieved.
Patent Information
- Application Number
- CN202411102017.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-08-12
AI Technical Summary
The prior art has problems of inaccurate segmentation and mistracking in the reconstruction of biological tubular structures in electron microscopy images, especially in poor image quality or complex structures, which affects the accuracy of reconstruction.
A biotubular structure segmentation network with self-attention mechanism is used for segmentation, and a long-distance dependency is captured using the self-attention mechanism, and a skeleton point set prediction and connection are carried out through a biotubular structure tracking network with skeleton point prediction to build a tree structure.
It improves the segmentation accuracy and tracking accuracy of the biological tubular structure, reduces missing points, obtains more accurate three-dimensional reconstruction results, stabilizes network training, and improves the overall performance of reconstruction.
Smart Images

Figure CN119048673B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electron microscopy image processing, and particularly to a method for reconstructing biological tubular structures in large-scale electron microscopy images. Background Art
[0002] Tracheas and blood vessels play crucial roles in the respiratory system. The research on their structures and functions is essential for deeply understanding respiratory diseases and developing new treatment methods. The development of three-dimensional reconstruction technology has provided a deeper understanding of typical branched tubular structures such as tracheas and blood vessels. This technology aims to reconstruct the tubular circuits of tracheas or blood vessels in biological electron microscopy images, providing an important tool for biomedical analysis. It not only helps scientists understand the natural changes of these structures more deeply but also provides strong assistance to doctors in pathological diagnosis. By three-dimensionally reconstructing tracheas and blood vessels, researchers can more clearly observe structural abnormalities, which is crucial for studying the pathogenesis of vascular-related diseases such as asthma, tracheitis, and atherosclerosis. The research on tracheas and blood vessels not only helps people deeply understand the complexity of the respiratory system but also provides important scientific bases for the prevention and treatment of related diseases.
[0003] With the development of high-speed and high-resolution microscopy imaging technology, whole-brain ultra-high-resolution scanning has become possible. In 2018, the Howard Hughes Medical Institute published the first complete nano-scale electron microscopy scan image of a Drosophila brain (Full Adult Fly Brain, abbreviated as the FAFB dataset) in the journal Cell. Its physical resolution reached 4nm×4nm×40nm. The whole-brain image consisted of 7,062 slice images in total, and the total data volume reached 106TB (ZHENG Z, et al., “A complete electron microscopy volume of the brain of adult Drosophila melanogaster”, in Cell, 2018). At the nano-scale resolution, it is possible to more clearly observe the details of microscopic structures such as tracheas, neuronal bundles, and cell bodies. This high-resolution imaging ability provides rich data resources for performing fine three-dimensional reconstruction of tracheas on large-scale datasets and in-depth biological research.
[0004] In recent years, researchers have invested a lot of research work on the reconstruction of biological tubular tree structures (such as trachea, blood vessels, neurons, etc.), which can be divided into segmentation and tracking steps. In segmentation, most existing methods use 3D CNN series methods to segment three-dimensional medical image data. This type of method uses a three-dimensional convolution kernel to process three-dimensional images, effectively utilizes the z-axis spatial information of biomedical images, and deeply mines the characteristics of three-dimensional space. Limited by the size of the three-dimensional perception field, this type of method has weak processing capabilities for global context information. In recent years, some methods using the Transformer model have been applied to medical organ segmentation tasks. With its powerful self-attention mechanism, the Transformer model not only focuses on local details when processing sequence data, but also takes into account global context information, and can more effectively capture long-distance dependencies, which is crucial for continuity and accuracy in tubular structure segmentation tasks.
[0005] Since image segmentation only generates pixel-by-pixel or voxel-by-voxel category labels to form segmentation masks, it cannot directly express the three-dimensional structural characteristics of the target, such as connectivity, tubular bifurcation, axial length, etc. Therefore, it is necessary to track biological tubular structures, connect tissue fragments in the segmentation results, and reconstruct tubular structure pathways for subsequent biological analysis. Existing traditional tracking methods often rely on manually designed features, selected seed points, track voxels belonging to specific biological tubular structures, and reconstruct the morphology of biological tubular structures. Such methods are highly dependent on the quality of image segmentation, and often have poor effects in areas with poor image quality, which easily leads to a large number of breaks and incorrect connections in the tracking results, and subsequent artificial introduction of priors is required for correction. In recent years, some methods have used the power of neural networks to improve tracking accuracy. Most of these methods divide tracking into two steps: point set prediction and point set connection. These methods use neural networks to extract features from biological electron microscope microscopic images, learn and identify multi-scale and multi-level feature representations in images, improve tracking performance in areas with poor image quality and complex structures, and show better robustness. However, for areas with weak signals, the performance of many methods is affected. For example, when PointNeuron performs point set prediction, it samples point clouds from the image and loses image information in areas with weaker signals. In contrast, NRTR uses the DETR method based on target detection, see the literature "Carion N, et al., End-to-end object detection with Transformers, in ECCV, 2020", to predict point sets from optical microscope images, thereby retaining more image details. However, the basic version of DETR used by NRTR is unstable during training and converges slowly. In addition, when processing some dense electron microscope data, NRTR frequently generates connection errors in the point set connection stage, affecting the accuracy of the final tubular structure reconstruction.
[0006] In view of this, the present invention is hereby provided. Summary of the Invention
[0007] The object of the present invention is to provide a method, device and medium for reconstructing biological tubular structures in electron microscopy images, which can obtain more accurate three-dimensional tubular structure reconstruction results from biological electron microscopy images, thereby solving the above technical problems existing in the prior art.
[0008] The object of the present invention is achieved by the following technical solutions:
[0009] A method for reconstructing biological tubular structures in electron microscopy images, characterized by comprising:
[0010] Step 1: Divide the original electron microscopy image into blocks to obtain block images. Use a biological tubular structure segmentation network with a self-attention mechanism to capture the long-range dependencies of biological tubular structures in the original electron microscopy image, and perform biological tubular structure segmentation on each block image one by one to obtain a biological tubular structure probability segmentation image based on the long-range dependencies.
[0011] Step 2: Perform preprocessing on the biological tubular structure probability segmentation image obtained in Step 1 to obtain image blocks for tracking. Use a biological tubular structure tracking network with skeleton point prediction to predict the skeleton point set of the biological tubular structure for the obtained image blocks for tracking, and construct a tree-like biological tubular structure by connecting each skeleton point of the skeleton point set of the biological tubular structure according to the connection direction to obtain the reconstruction result of the biological tubular structure.
[0012] A processing device, comprising:
[0013] At least one memory for storing one or more programs;
[0014] At least one processor capable of executing the one or more programs stored in the memory, and when the one or more programs are executed by the processor, enabling the processor to implement the method of the present invention.
[0015] A readable storage medium storing a computer program, which can implement the method of the present invention when executed by a processor.
[0016] Compared with the prior art, the method, device and medium for reconstructing biological tubular structures in electron microscopy images provided by the present invention have the following beneficial effects:
[0017] By using a biological tubular structure segmentation network with a self-attention mechanism, the biological tubular structures in the original electron microscopy micrographs are accurately segmented to obtain a probability segmentation image of the biological tubular structures. The image patches for tracking obtained after preprocessing the probability segmentation image of the biological tubular structures are predicted by a biological tubular structure tracking network with skeleton point prediction to obtain a set of skeleton points of the biological tubular structures. The skeleton points in the set of skeleton points of the biological tubular structures are connected according to the connection direction to construct a tree-like biological tubular structure, and the reconstruction result of the biological tubular structure is obtained. Since the biological tubular structure segmentation network with a self-attention mechanism uses the self-attention mechanism to effectively process the long-range dependence information of the biological tubular structures in the original electron microscopy micrographs, the accuracy of segmentation is improved and the missed segmentation is reduced. Since the biological tubular structure tracking network with skeleton point prediction is adopted, the features of interest can be selected from the features output by the Transformer encoder to generate skeleton proposal points, which are used as the initial queries of the Transformer decoder, optimizing the network structure, stabilizing the network training, and improving the network performance. In addition, the connection direction of each skeleton point is predicted during the prediction of the set of skeleton points, and the process of connecting the skeleton points in the set of skeleton points is guided by the connection direction to obtain a more accurate reconstruction result. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0019] Figure 1 It is a flowchart of the method for reconstructing biological tubular structures in electron microscopy micrographs provided by the embodiments of the present invention.
[0020] Figure 2 It is a flowchart of the biological tubular structure tracking in the reconstruction method provided by the embodiments of the present invention.
[0021] Figure 3 It is a schematic diagram for comparing the reconstruction results of the biological tubular structures of the reconstruction method provided by the embodiments of the present invention; among them, (a) is a schematic diagram of the biological tubular structures of the segmented image patches; (b) is a schematic diagram of the reconstruction result of the biological tubular structures by the manual annotation method; (c) is a schematic diagram of the reconstruction result of the biological tubular structures by the NGPST method; (d) is a schematic diagram of the reconstruction result of the biological tubular structures by the NRTR method; (e) is a schematic diagram of the reconstruction result of the biological tubular structures by the method of the present invention.
[0022] Figure 4 It is a schematic diagram of the whole-brain reconstruction result of the reconstruction method provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Combined with the specific content of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below; obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments, which does not constitute a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present invention.
[0024] First, the terms that may be used in this article are described as follows:
[0025] The term "and / or" means that either or both of the two can be realized. For example, X and / or Y means that it includes both the case of "X" or "Y" and the three cases of "X and Y".
[0026] The description of terms such as "comprising", "including", "containing", "having" or other similar semantics should be interpreted as non-exclusive inclusion. For example: including a certain technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction condition, processing condition, parameter, algorithm, signal, data, product or article, etc.) should be interpreted as not only including the explicitly listed certain technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.
[0027] The term "consisting of" means excluding any unexplicitly listed technical feature elements. If this term is used in a claim, this term will make the claim a closed type, making it not contain technical feature elements other than the explicitly listed technical feature elements, except for related conventional impurities. If this term only appears in a certain clause of the claim, then it only limits the elements explicitly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.
[0028] Unless otherwise clearly specified or limited, terms such as "installed", "connected", "joined", "fixed", etc. should be understood in a broad sense. For example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this article can be understood according to specific situations.
[0029] The orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of description and to simplify the description, rather than explicitly or implicitly indicating that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this article.
[0030] The solution provided by the present invention will be described in detail below. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. In the embodiments of the present invention, those not specified in specific conditions are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. For the reagents or instruments not specified in the production manufacturer in the embodiments of the present invention, they are all conventional products that can be obtained by purchasing in the market.
[0031] As Figure 1 shown, an embodiment of the present invention provides a method for reconstructing a biological tubular structure in an electron microscopy microscopic image, including:
[0032] Step 1: Divide the original electron microscopy microscopic image into blocks to obtain block images. Use a biological tubular structure segmentation network with a self-attention mechanism to capture the long-range dependence relationship of the biological tubular structure in the original electron microscopy microscopic image, and perform biological tubular structure segmentation on each block image one by one to obtain a biological tubular structure probability segmentation image based on the long-range dependence relationship;
[0033] Step 2: Perform preprocessing on the biological tubular structure probability segmentation image obtained in Step 1 to obtain image blocks for tracking. Use a biological tubular structure tracking network with skeleton point prediction to predict the skeleton point set of the biological tubular structure for the obtained image blocks for tracking, and connect each skeleton point of the skeleton point set of the biological tubular structure according to the connection direction to construct a tree-like biological tubular structure, and obtain the reconstruction result of the biological tubular structure.
[0034] It can be known that the above reconstruction result of the biological tubular structure is a reconstruction data of the biological tubular structure, which can be saved in a standard text SWC file. In the modern biological field, this biological tubular structure is often defined as a tree-like skeleton. The tree-like skeleton is composed of numerous skeleton points, and each skeleton point contains its three-dimensional spatial position, radius size, and connection information with other nodes. A tubular structure path is constructed through the connection information between the skeleton points.
[0035] Preferably, in step 2 of the above reconstruction method, the probability segmentation image of the biological tubular structure obtained in step 1 is pre-processed in the following manner to obtain an image patch for tracking. The image patch for tracking is predicted by a biological tubular structure tracking network with skeleton point prediction to obtain a set of skeleton points of the biological tubular structure. According to the connection direction, each skeleton point of the set of skeleton points of the biological tubular structure is connected to construct a tree-like biological tubular structure, and the reconstruction result of the biological tubular structure is obtained, including:
[0036] Step 21: Using the segmentation probability of the biological tubular structure region in the probability segmentation image of the biological tubular structure segmented in step 1, a segmented three-dimensional segmentation probability image is obtained. Multiple three-dimensional segmentation probability images are stitched into an overall three-dimensional segmentation probability image, and the obtained overall three-dimensional segmentation probability image is downsampled and cut into at least one fixed-size image patch for tracking.
[0037] Step 22: Input each image patch for tracking obtained in step 21 into a biological tubular structure tracking network with skeleton point prediction for skeleton point prediction, and obtain the three-dimensional spatial position and connection direction of the skeleton points of the biological tubular structure in the image patch for tracking. The three-dimensional spatial position is composed of coordinates and radius.
[0038] Step 23: Using the three-dimensional spatial position and connection direction of the obtained skeleton points of the biological tubular structure to connect each skeleton point, complete the three-dimensional construction of the tree-like structure of the biological tubular structure in each image patch for tracking. Finally, the reconstruction results of the tree-like structures of the biological tubular structures in multiple image patches for tracking are stitched together to obtain the reconstruction result of the biological tubular structure corresponding to the overall three-dimensional segmentation probability image.
[0039] Preferably, in the above reconstruction method, the biological tubular structure segmentation network with self-attention mechanism adopts a Swin UNETR network with self-attention mechanism.
[0040] The physical resolution of the input electron microscopy micrograph of the Swin UNETR network with self-attention mechanism is 16nm×16nm×40nm, and the voxel size is 128×128×16.
[0041] The output of the Swin UNETR network with self-attention mechanism is a probability image of whether each voxel belongs to the biological tubular structure, that is, a probability segmentation image of the biological tubular structure.
[0042] The biological tubular structure tracking network with skeleton point prediction includes: an image feature extraction network, a skeleton point set prediction network based on Transformer, and a skeleton point set connection module. Among them,
[0043] The image feature extraction network can take as input the tracking image patches obtained by preprocessing the probability segmentation image of the biological tubular structure obtained in step 1, and extract low-resolution image features as output;
[0044] The Transformer-based skeleton point set prediction network is communicatively connected to the output end of the image feature extraction network. It can take the low-resolution image features output by the image feature extraction network as the input token sequence, calculate the associations between the image patch tokens in the input token sequence, generate the skeleton proposal point feature sequence of the biological tubular structure, and predict the classification probabilities, three-dimensional spatial positions, and connection directions of each skeleton point in the skeleton point sequence of the biological tubular structure. The skeleton point set is composed of the classification probabilities, three-dimensional spatial positions, and connection directions of each skeleton point, and the three-dimensional spatial position is composed of three-dimensional spatial coordinates and a radius;
[0045] The skeleton point set connection module is communicatively connected to the output end of the Transformer-based skeleton point set prediction network. It can take the skeleton point set predicted by the Transformer-based skeleton point set prediction network as the input, and connect each skeleton point in the skeleton point set based on the predicted connection directions of each skeleton point to generate the tree structure of the final biological tubular structure.
[0046] Preferably, in the above reconstruction method, the image feature extraction network is composed of a three-dimensional residual network and a feature pyramid network connected in sequence. Among them, the three-dimensional residual network takes as input the tracking image patches obtained by preprocessing the probability segmentation image of the biological tubular structure obtained in step 1, can perform three-dimensional residual processing and then output to the feature pyramid network for low-resolution image feature extraction to obtain low-resolution image features;
[0047] The Transformer-based skeleton point set prediction network is composed of a Transformer encoder, a skeleton proposal point set prediction sub-module, a Transformer decoder, and a skeleton point set prediction sub-module; among them,
[0048] The Transformer encoder is communicatively connected to the image feature extraction network. It can take as input the position encoding of each pixel position of the low-resolution image features output by the image feature extraction network and the pixel feature sequence composed of all pixel features of the low-resolution image features. Through the self-attention mechanism, it maps the pixel feature sequence and the position encoding of each pixel position to a feature sequence containing the position information and content information of all image features, and outputs the feature sequence;
[0049] The proposed skeleton point prediction sub-module is communicatively connected to the output end of the Transformer encoder. Taking the feature sequence output by the Transformer encoder as input, it can respectively predict the classification probability, three-dimensional spatial position, and connection direction of the proposed skeleton points of the biological tubular structure. The three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius;
[0050] The Transformer decoder is communicatively connected to the Transformer encoder and the proposed skeleton point set prediction sub-module respectively. Taking the top N proposed skeleton points of the biological tubular structure with the largest probabilities among the classification probabilities output by the proposed skeleton point set prediction sub-module as the proposed skeleton points, it selects N features corresponding to these N proposed skeleton points from the feature sequence and recombines them into a proposed feature sequence for biological tubular structure point prediction as the initial object query matrix. Using the position encodings of these N proposed skeleton points and the feature sequence output by the Transformer encoder as keys and values for each layer of the Transformer decoder, it utilizes the cross-attention mechanism to capture the dependencies between the target biological tubular structure points and all positions in the image, and outputs the target biological tubular structure point features with the same shape as the proposed feature sequence for biological tubular structure point prediction;
[0051] The proposed skeleton point set prediction sub-module is communicatively connected to the Transformer decoder. Taking the target biological tubular structure point features output by the Transformer decoder as input, it obtains the classification probability, three-dimensional spatial position, and angle of each proposed skeleton point. The three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius, and the proposed skeleton points obtained form a proposed skeleton point set.
[0052] Preferably, in the above reconstruction method, in the proposed skeleton point set prediction network based on Transformer, the Transformer encoder is composed of multiple encoder layers connected in sequence, and each encoder layer is composed of a multi-head self-attention module and a multi-layer perceptron module;
[0053] The proposed skeleton point set prediction sub-module is composed of six multi-layer perceptrons. Among them, the first two multi-layer perceptrons are respectively used to predict the classification probability and three-dimensional spatial position of the proposed skeleton points of the biological tubular structure, and the last four multi-layer perceptrons are used to predict the horizontal angle and vertical angle corresponding to the connection direction;
[0054] The Transformer decoder is composed of multiple decoder layers connected in sequence, and each decoder layer is composed of a multi-head self-attention sub-module, a multi-head cross-attention sub-module, and a multi-layer perceptron sub-module.
[0055] Preferably, in the above reconstruction method, the skeleton point set prediction sub-module consists of six multi-layer perceptrons. Among them, the first two multi-layer perceptrons are respectively used to predict the classification probability and three-dimensional spatial position of the biological tubular structure points, and the last four multi-layer perceptrons are used to predict the horizontal angle and vertical angle corresponding to the connection direction.
[0056] Preferably, in the above reconstruction method, the total loss function of the biological tubular structure tracking network with skeleton point prediction is:
[0057]
[0058] where the meanings of the parameters are as follows: denotes the sum of the matching losses of all predicted skeleton proposal points and their matching points in the ground truth; ; denotes the sum of the matching losses of all predicted skeleton points and their matching points in the ground truth; ; denotes the skeleton proposal point set composed of N biological tubular structure points with the largest probabilities selected by the skeleton proposal point set prediction sub-module of the Transformer-based skeleton point set prediction network of the biological tubular structure tracking network with skeleton point prediction using the Top-K strategy; denotes the skeleton proposal point set and the optimal matching with the ground truth skeleton point set Y; denotes the skeleton point set composed of N skeleton points output by the Transformer-based skeleton point set prediction network; denotes the skeleton point set and the optimal matching with the ground truth skeleton point set Y; ω1 and ω2 are hyperparameters respectively;
[0059] In the above total loss, are all calculated as follows:
[0060]
[0061] where c i , L i , and Ai respectively represent the classification probability, three-dimensional spatial position, and angle of the ground truth Y i . The three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius; respectively represent the classification probability, three-dimensional spatial position, and angle of the prediction result . The three-dimensional spatial position consists of coordinates and a radius; ω cls , ω box , ω iou , and ω dir are hyperparameters respectively; Represents the classification loss, set as the cross-entropy loss function; and both represent the position loss, is the L1 loss function, is the IoU loss function for spheres; is the direction loss, The calculation method is as follows:
[0062]
[0063] Among them, respectively represent the true direction d of the direction of each skeleton point i corresponding horizontal angle classification, horizontal angle within-class offset value, vertical angle classification, and vertical angle within-class offset; respectively represent the predicted direction of the horizontal angle classification, horizontal angle within-class offset value, vertical angle classification, and vertical angle within-class offset; Represents the angle classification loss, set as the cross-entropy loss function; Represents the angle within-class offset regression loss, using the L1 loss function, and only calculates the difference between the predicted value and the true value in the offset values within the angle target interval.
[0064] Preferably, in step 2 of the above reconstruction method, the skeleton point set of the biological tubular structure is connected according to the connection direction in the following manner to construct a tree-like biological tubular structure, and a reconstructed image of the biological tubular structure is obtained, including:
[0065] Delete the skeleton points with a classification probability less than a predetermined threshold from the skeleton point set of the biological tubular structure, use the remaining skeleton points as the selected skeleton points, calculate the predicted connection direction of each skeleton point according to the angles predicted by the selected skeleton points, establish a point-edge graph for all the remaining skeleton points, and establish an edge relationship between two skeleton points with a distance less than the distance threshold and an angle difference between the two points and the predicted direction less than the angle threshold. Then, use the angle difference and the distance between points as the edge weight e i,j , e i,j is:
[0066] e i,j = ω3dist(P i , P j ) +
[0067] ω4min{[1 - cos(d i , dir i,j )], [1 - cos(d j , dir j,i )]} (9);
[0068] Among them, dist(P i , Pj ) represents the Euclidean space distance between skeleton point i and skeleton point j; (P i , P j ) represents the positions of skeleton point i and skeleton point,; cos(d i , dir i,j ) represents the cosine distance between the predicted direction of skeleton point i and the orientation from skeleton point i to skeleton point,; (d i , dir i,j ) represents the predicted direction of skeleton point i and the orientation from skeleton point i to skeleton point,; cos(d j , dir j,i ) represents the cosine distance between the predicted direction of skeleton point j and the orientation from skeleton point j to skeleton point i; (d j , dir j,i ) represents the predicted direction of skeleton point j and the orientation from skeleton point j to skeleton point i; ω3 and ω4 are both hyperparameters; min{} represents taking the minimum value;
[0069] The constructed point-edge graph is loop-removed by using a minimum spanning tree to obtain the tree structure of the final biological tubular structure.
[0070] The embodiment of the present invention further provides a processing device, including:
[0071] At least one memory for storing one or more programs;
[0072] At least one processor capable of executing the one or more programs stored in the memory, and when the one or more programs are executed by the processor, enabling the processor to implement the above method.
[0073] The embodiment of the present invention further provides a readable storage medium storing a computer program, which can implement the above method when executed by a processor.
[0074] In summary, the method of the embodiment of the present invention uses a biological tubular structure segmentation network with a self-attention mechanism to accurately segment the biological tubular structure from the original electron microscopy image to obtain a probability segmentation image of the biological tubular structure. Through a biological tubular structure tracking network with skeleton point prediction, the skeleton point set of the biological tubular structure is predicted from the image patches for tracking obtained by preprocessing the probability segmentation image of the biological tubular structure. The skeleton point set of the biological tubular structure is connected according to the connection direction to construct a tree-like biological tubular structure, and a reconstructed image of the biological tubular structure is obtained. The Swin UNETR network is used as the biological tubular structure segmentation network with a self-attention mechanism, and the self-attention mechanism is used to effectively process the long-range dependence information of the biological tubular structure in the original electron microscopy image, improving the accuracy of segmentation and reducing missed segmentation. Since a biological tubular structure tracking network with skeleton point prediction is adopted, interesting features can be selected from the output feature sequence output by the Transformer encoder to generate skeleton proposal points, which are used as the initial queries of the Transformer decoder, optimizing the network structure, stabilizing network training, and improving network performance. In addition, the connection direction of the skeleton point set is predicted in the prediction of the skeleton point set, and the connection direction is used to guide the connection process of each skeleton point in the skeleton point set to obtain a more accurate reconstructed structure. Experiments were carried out on the FAFB electron microscopy data for the method of the present invention. By using a small amount of data to train the biological tubular structure segmentation network with a self-attention mechanism and the biological tubular structure tracking network with skeleton point prediction, on the test data set, the method of the present invention has the best performance in all indicators of the reconstructed image compared with the mainstream traditional NGPST method and the deep learning NRTR method, and there is a significant improvement in all indicators.
[0075] In order to more clearly show the technical solutions provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the method for reconstructing biological tubular structures in electron microscopy images provided by the embodiments of the present invention.
[0076] Embodiment 1
[0077] As Figure 1As shown in the figure, an embodiment of the present invention provides a method for reconstructing biological tubular structures in electron microscopy images. Taking the reconstruction of Drosophila trachea based on large-scale electron microscopy images of Transformer as an example, it specifically includes: trachea segmentation and trachea tracking. Due to problems such as the complex diversity of trachea morphology and uneven quality of electron microscopy images, major challenges are posed to trachea reconstruction. To address these problems, the present invention adopts a trachea segmentation method based on the Swin UNETR network (see the literature: Hatamizadeh A, et al., “Swin UNETR: Swin Transformers for semantic segmentation of brain tumors in MRI images”, in BrainLes, 2021) and a trachea tracking method based on DETR and connection directions. The overall method steps of the present invention are as follows:
[0078] First, considering that the trachea has a large radius change and a long span, the commonly used 3D CNN method is limited by the size of the three-dimensional perception field and has a weak ability to process global context information. It is prone to missing segmentation in the segmentation of long and thin tracheas, resulting in trachea breaks in the subsequent reconstruction process and affecting the accuracy of the reconstruction results. The Swin UNETR network is used as a biological tubular structure segmentation network with a self-attention mechanism. The self-attention mechanism is used to effectively process the long-range dependence information of the trachea, improving the accuracy of trachea segmentation and reducing missing segmentation. In addition, the existing traditional biological tubular structure tracking methods are greatly affected by image quality and it is difficult to reconstruct accurate whole-brain tracheas. Often, it is necessary to artificially introduce biological prior knowledge for correction. The latest deep learning tracking method NRTR directly predicts a set of skeleton points and then uses a simple connection method to construct a tree structure. However, its network training is unstable, and directly applying it to electron microscopy trachea data will produce incorrect connections at dense trachea locations. Based on the NRTR method (see the literature: Wang Y, et al., “NRTR: Neuron reconstruction with Transformer from 3D optical microscopy images”, in IEEE TMI, 2024), the present invention adds a newly designed skeleton proposal point prediction sub-module to the biological tubular structure tracking network with skeleton point prediction. From the output feature sequence output by the Transformer encoder, select the features of interest to generate skeleton proposal points, which serve as the initial queries for the Transformer decoder, optimize the network structure, stabilize network training, and improve network performance; and additionally predict the connection directions of each skeleton point in the skeleton point set during the prediction of the skeleton point set. Utilize the connection directions to guide the connection process of each skeleton point in the skeleton point set to obtain more accurate reconstruction results.
[0079] Specifically, in the embodiments of the present invention, the Swin UNETR network with a self-attention mechanism is used to segment the trachea from the original electron microscopy micrograph to obtain a probabilistic trachea segmentation image. After preprocessing the probabilistic trachea segmentation image to obtain image patches for tracking, a biological tubular structure tracking network with skeleton point prediction is then used to predict the image patches to obtain a set of skeleton points of the trachea structure in the image patches. The set of skeleton points specifically includes the three-dimensional spatial positions of the skeleton points (composed of three-dimensional spatial coordinates and radii) and the connection directions, and the skeleton points in the set of skeleton points are connected according to the connection directions to construct a tree-like structure reconstruction result of the trachea structure.
[0080] Next, each part will be described in detail.
[0081] (1) Swin UNETR network with a self-attention mechanism (i.e., a biological tubular structure segmentation network with a self-attention mechanism):
[0082] Trachea segmentation is the first step in the trachea reconstruction process, and the accuracy of trachea segmentation often determines the accuracy of subsequent reconstruction results. Considering the large morphological differences in the trachea of the Drosophila whole brain, such as large differences in diameter range and long trachea spans, etc., the trachea segmentation network needs to maintain spatial continuity in slender tracheas. For this reason, the Swin UNETR network with a self-attention mechanism is adopted in the present invention. It not only retains the U-shaped structure and skip connection design of U-Net, effectively fusing multi-scale image features, but also uses a powerful self-attention mechanism to effectively capture long-range dependencies in the image sequence, thereby grasping global information and generating a fine and spatially continuous segmentation result. Considering that the complete image is too large, it needs to be divided into blocks and input into the Swin UNETR network with a self-attention mechanism, and the segmentation results are obtained and stitched together to form a whole brain segmentation result. Specifically, the physical resolution of the input image of the Swin UNETR network with a self-attention mechanism in this embodiment is 16 nm × 16 nm × 40 nm, the voxel size is 128 × 128 × 16, and the output is the probability of whether each voxel belongs to the trachea.
[0083] (2) Biological tubular structure tracking network with skeleton point prediction:
[0084] Due to problems such as the dense distribution in some regions of Drosophila trachea and the presence of noise in the images, the trachea tracking task of electron microscopy images at high resolution faces the challenge of tracking errors. Traditional tracking methods for modern biological tissues start from the perspective of sequential tracking or from the perspective of global graph optimization. The mainstream tracking method for biological tissues, NGPST, automatically samples seed points based on the segmented image, tracks to neighboring positions according to the image intensity, and determines termination and bifurcation based on artificially set rules to reconstruct the morphology of biological tissues. These methods rely heavily on image quality and often produce multiple errors in regions with poor image quality. Subsequent work often requires artificially introducing biological prior knowledge to correct the tracking results, which requires manually adjusting the correction rules according to the correction results and is often time-consuming and laborious.
[0085] To address the above problems, the present invention proposes a method for reconstructing biological tubular structures based on a DETR (Carion N, et al., End-to-end object Detection with Transformers, in ECCV, 2020) detection network and connection directions, which can learn and identify multi-scale and multi-level feature representations in the segmented image and improve the tracking performance in regions with poor image quality and regions with complex structures. Given an electron microscopy slice image, first, use the Swin UNETR network to predict the trachea segmentation probability to obtain a three-dimensional segmentation probability image, and splice multiple segmentation results into an overall three-dimensional segmentation probability image; secondly, downsample the overall three-dimensional segmentation probability and cut it into multiple image patches of a fixed size, and input each image patch into a biological tubular structure tracking network with skeleton point prediction to detect the three-dimensional spatial positions and connection directions of the skeleton points in the image patch; then, use the three-dimensional spatial position and connection direction information of the skeleton point set to complete the construction of the tubular tree structure, and finally splice the reconstruction results of multiple image patches to obtain the trachea reconstruction result corresponding to the overall three-dimensional segmentation probability image.
[0086] The present invention proposes a method for tracking biological tubular structures in electron microscopy images guided by connection directions, specifically including: an image feature extraction network, a Transformer-based skeleton point set prediction network, and a skeleton point set connection module; the method process is as Figure 2 shown,
[0087] Among them, the Transformer-based skeleton point set prediction network includes: a Transformer encoder, a Transformer decoder, a skeleton proposal point set prediction sub-module, and a skeleton point set prediction sub-module. Among them, the image feature extraction network takes the tracking image patches obtained by dividing the overall three-dimensional segmentation probability image of the trachea into blocks as input, extracts image features, and then takes the image features of each image patch as words (tokens) and inputs them into the Transformer encoder. The Transformer encoder calculates the dependencies between the various tokens in the input image features and outputs a series of skeleton proposal points. Among them, the Transformer decoder takes the output of the skeleton proposal point set prediction sub-module as the query target, and through the attention mechanism with the output features of the Transformer encoder, generates a skeleton point feature sequence, thereby generating a set of skeleton points containing probability, three-dimensional spatial position, and connection direction. Then, based on the predicted connection direction of each skeleton point, the skeleton points in the skeleton point set are connected to generate the reconstruction result of the tree structure of the final trachea structure.
[0088] 1) Each network structure:
[0089] Compared with the voxel-by-voxel high-resolution trachea segmentation result, the tree structure of the trachea structure realizes a more simplified point set representation by abstracting multiple voxels in the segmented image into a spherical point. In order to input the whole-brain segmentation image into the biological tubular structure tracking network with skeleton point prediction, it is necessary to first downsample the segmentation image and crop and divide it into smaller image patches. The biological tubular structure tracking network with skeleton point prediction of the present invention reads the cropped probability segmentation image patches and predicts the skeleton point set of the trachea structure from them. The point set can be expressed as where W, H, and D are the width, height, and depth of the image patch, N is a hyperparameter that needs to be greater than the number of skeleton points of the true trachea structure corresponding to any image patch, (x j , y j , z j ), r j and d j are the three-dimensional spatial position (composed of three-dimensional spatial coordinates and radius) and connection direction of the normalized skeleton point respectively, c j ∈[0, 1] represents the classification probability that the predicted point belongs to the skeleton point, and d j is the direction vector of the three-dimensional prediction direction. Specifically, after multiple experiments, in this embodiment, the physical resolution of the input image of the biological tubular structure tracking network with skeleton point prediction is selected to be 128nm×128nm×128nm (this image can be obtained by interpolating and sampling the three-dimensional probability segmentation image), and the voxel size is 64×64×64.
[0090] For three-dimensional connection direction prediction, in actual training, first, the three-dimensional direction d j is split into the horizontal angle ψ j and the vertical angle φ j . The prediction is performed separately for these two angles, and the calculation method is as follows.
[0091] d j =(d j,x , d j,y ,, d j,z ) (1);
[0092]
[0093] The present invention follows the method of direction target detection and adopts the method of "segmented prediction" (binning) to discretize the continuous angle into λ intervals, and each interval represents an angle range of size . Taking ψ j as an example, first predict the probability that the angle belongs to each interval Then predict the specific offset value of the angle within each interval (which needs to be normalized according to ), and φ j is similar. In this way, the azimuth angle (ψ j , φ j ) is converted into and predict A j . In the inference stage, after obtaining , calculate the specific angle according to the interval where the maximum probability of these two angles is located and the offset within the interval, and then obtain the predicted direction d j according to the inverse process of the above formulas (1) to (3).
[0094] The image feature extraction network includes: a three-dimensional residual network and a feature pyramid network. The input of this image feature extraction network is a single-channel segmentation probability image After passing through the image feature extraction network, low-resolution image features are obtained Input the low-resolution image feature I f into the skeleton point set
[0095] prediction network based on Transformer. Take each pixel of the low-resolution image feature as a token, and all pixels are input into the Transformer encoder as a sequence. The sequence is where t represents the length of the sequence, that is, the image size of the low-resolution image feature I f , and t = D0×H0×W0. At the same time, considering the spatial position in the image, calculate the position encoding of each feature position, and the position encoding is The positional encoding is implemented using cosine positional encoding. The sequence S and the positional encoding P are jointly used as the input to the Transformer encoder.
[0096] The Transformer encoder, as the feature extractor of the Transformer-based skeleton point set prediction network, consists of K encoder layers. Each encoder layer is composed of a multi-head self-attention module and a multi-layer perceptron (MLP). Since the Transformer architecture is permutation-invariant, the output of the upper encoder layer and the positional encoding are input into the next encoder layer each time. The Transformer encoder maps the sequence S and P to the feature sequence Z through the self-attention mechanism, and inputs the feature sequence Z into the skeleton proposal point set prediction sub-module and the Transformer decoder.
[0097] The Transformer decoder relies on target queries to model the relationship between the skeleton point set and the segmented image. To obtain better initial query points, the present invention proposes a query point dynamic generation process. In the NRTR method, the query points are generated through a learnable linear layer. However, within a single iteration, for different images, these query point inputs are usually static, that is, they are the same. In the initial learning stage of the network, this linear layer often uses all-zero or random initialization to generate, resulting in unclear meanings of the query points, which in turn affects the training stability and convergence speed. To further utilize the Transformer encoder, the present invention selects some tokens features from the feature sequence Z output by the last encoder layer of the Transformer encoder as priors to enhance the Transformer decoder queries. The present invention calculates the scores of all features by connecting a skeleton proposal point set prediction sub-module after the Transformer encoder, and selects the top N (the same number as the query points) encoder features from them. Specifically, the skeleton proposal point set prediction sub-module predicts the classification probability that each feature position point belongs to the tracheal structure skeleton point, and uses the features of the N tokens with the highest probability as the proposal points. The features of these proposal points and the corresponding positional encoding are used as the target queries of the Transformer decoder, so that the target queries are associated with the possible target regions in the segmented image at the beginning, enabling the Transformer decoder to more effectively focus on these specific regions. The present invention constrains the skeleton proposal point set prediction sub-module and the Transformer encoder by calculating the loss between the prediction result of the skeleton proposal point set prediction sub-module and the ground truth. By such a method, the training of the Transformer-based skeleton point set prediction network can be made more stable, converge faster, and obtain better prediction results.
[0098] The skeleton proposal point set prediction sub-module includes six multi-layer perceptrons. The first two multi-layer perceptrons are used to predict the classification probability and three-dimensional spatial position of the skeleton proposal points respectively, and the last four multi-layer perceptrons are used to predict two angles corresponding to the connection direction of each skeleton proposal point, including the classification probabilities of different angular intervals of the horizontal angle ψ and the vertical angle φ and the offset values within the intervals. In the skeleton proposal point set prediction sub-module, the feature sequence Z output by the Transformer encoder passes through the skeleton proposal point set prediction sub-module to output the probabilities of t skeleton proposal points Three-dimensional spatial position Angle The dynamic query of the present invention will select some information from these information according to the probabilities C of t skeleton proposal points E as the target initial query of the Transformer decoder. Specifically, the present invention adopts the Top-K strategy to select the first N prediction points with the largest probabilities in the probabilities C of t skeleton proposal points E as the query proposal points, and selects the corresponding N features from the feature sequence Z output by the Transformer encoder to recombine them into a sequence Q 0 , which is input into the Transformer decoder as the initial object query matrix of the Transformer decoder. In addition, the position encoding P of these N proposal points 0 is input into each layer of the Transformer decoder.
[0099] The Transformer decoder models the relationship between the skeleton point target and the image through the object query (Object Query). It consists of K decoder layers, and each decoder layer is composed of a multi-head self-attention module, a multi-head cross-attention module (Multi-head Cross-attention) and a multi-layer perceptron. The predicted feature Q of the skeleton proposal point set prediction sub-module 0 and P 0 , based on the feature sequence Z output by the encoder, are input into the Transformer decoder. Since the Transformer decoder is also permutation-invariant, each layer of the decoder also needs to input the position encoding. The Transformer decoder uses the cross-attention mechanism to capture the dependency relationship between the target skeleton points and all positions in the image, and outputs the target skeleton point feature Q with the same dimension as the sequence Q 0 .
[0100] Input the target skeleton point feature Q into the skeleton point set prediction sub-module to obtain the detailed information of the final tracheal structure skeleton points. The network architectures of the skeleton point set prediction sub-module and the skeleton proposal point set prediction sub-module are the same, but the specific network parameters are different and need to be learned during network training. The output feature sequence of the Transformer decoder passes through the skeleton point set prediction sub-module to output the classification probabilities of N skeleton points Three-dimensional spatial position Angle As the final skeleton point set prediction result, it is used to construct the tree structure of the tracheal structure subsequently.
[0101] 2) Loss function:
[0102] The loss function for training the biological tubular structure tracking network with skeleton point prediction is as follows:
[0103] The skeleton point set prediction sub-module outputs a predicted skeleton point set of a fixed size, which contains N predicted skeleton points; in addition, in order to make full use of the Transformer encoder information, the Transformer encoder outputs the features of t voxel points, and N skeleton proposal points are selected in the skeleton proposal point set prediction sub-module and input into the Transformer decoder.
[0104] The loss function of the present invention then calculates using the N skeleton proposal points of the skeleton proposal point set prediction sub-module and the N predicted skeleton points of the skeleton point set prediction sub-module. First, these skeleton point sets need to be matched with the real point set, and then the loss is calculated. For the convenience of introduction, it is first assumed that the predicted skeleton point set has been matched with the real point set, and the matching loss is introduced, and then the matching method is given to calculate the final total loss.
[0105] (21) Each loss function:
[0106] The skeleton proposal point set prediction sub-module will output N skeleton proposal points, and the skeleton point set prediction sub-module will output N skeleton points. Assuming there are M skeleton points in the ground truth, it is required that N > M. First, supplement N - M points with empty categories in the ground truth so that there are finally N points in the ground truth, and one-to-one matching is performed with the N predicted skeleton points. For a predicted skeleton point and its matching point Y in the ground truth i , the loss calculation is as follows:
[0107]
[0108] Among them, c i , L i , A i respectively represent the ground truth Y iThe classification probability, three-dimensional spatial position, and angle, where the three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius; respectively represent the prediction results The classification probability, three-dimensional spatial position, and angle, where the three-dimensional spatial position consists of coordinates and a radius; ω cls ω box ω iou ω dir are hyperparameters respectively; represents the classification loss, which is set as the cross-entropy loss function; and both represent the position loss, is the L1 loss function, is the sphere intersection over union loss function; is the direction loss, The calculation method is as follows:
[0109]
[0110] where, respectively represent the true direction values d of each skeleton point i corresponding horizontal angle classification, horizontal angle within-class offset value, vertical angle classification, vertical angle within-class offset; respectively represent the predicted direction of the horizontal angle classification, horizontal angle within-class offset value, vertical angle classification, vertical angle within-class offset; represents the angle classification loss, which is set as the cross-entropy loss function; is the angle within-class offset regression loss, which is set as the L1 loss, and only calculates the gap between the predicted value and the true value in the offset values where the angle target interval is located.
[0111] Next, the loss of the predicted skeleton point set is described. For the true skeleton point set Y and the predicted skeleton point set It is assumed that their optimal matching has been obtained The loss calculation is as follows:
[0112]
[0113] When the true point is a skeleton point, the classification, position, and direction losses will be calculated. When the true point is empty, only the classification loss is calculated. The loss of the two point sets is the sum of the losses of each point. The next section will introduce how to determine the optimal matching between point sets.
[0114] 3) Point set matching:
[0115] The above described how to calculate the loss in the case of determining the point set matching. Next, how to find the optimal matching between point sets will be described. Here, the definition of point set matching is that for the ground truth skeleton point set Y with M points, and the predicted skeleton point set with N points For each skeleton point in the predicted skeleton point set, how to find the most accurate and unique matching point from the ground truth point set. The present invention uses the Hungarian algorithm for matching. First, N - M empty set points are added to the ground truth point set. After the number of points in the two sets is the same, then consider all one-to-one corresponding point set matching schemes to minimize the optimal matching loss. The calculation formula is as follows:
[0116]
[0117] Using the above method, the optimal matching scheme between the predicted skeleton point set and the ground truth skeleton point set is found. On the one hand, in order to make the final output as close as possible to the ground truth skeleton point set, it is necessary to constrain the output of the skeleton point set prediction sub-module; on the other hand, in order to make the initial target query of the Transformer decoder as close as possible to the ground truth skeleton point set, it is also necessary to constrain the output of the skeleton proposal point set prediction sub-module. Therefore, the final loss of this skeleton proposal point set prediction sub-module is:
[0118]
[0119] Among them, represents the set of N skeleton proposals with the largest probabilities selected by the skeleton proposal point set prediction sub-module using the Top-K strategy; represents the skeleton proposal point set and the optimal matching with the ground truth skeleton point set Y; represents the set of N skeleton points output by the skeleton point set prediction sub-module, represents the predicted skeleton point set and the optimal matching with the ground truth skeleton point set Y. ω1 and ω2 are hyperparameters.
[0120] 4) Point set connection:
[0121] The skeleton point set connection module connects the predicted skeleton point set into a tree structure according to the predicted direction of each skeleton point. After obtaining the skeleton point set through the Transformer-based skeleton point set prediction network, the skeleton point set needs to be processed and connected into the final reconstructed result of the tree structure. First, a probability threshold of 0.5 is used to filter the skeleton points, and the skeleton points with a classification probability less than 0.5 are deleted. Then, the predicted connection direction of each skeleton point is calculated based on the angle information predicted by the skeleton points. A point-edge graph is established for all the remaining skeleton points. An edge relationship is established between two skeleton points whose distance is less than the distance threshold and the angular difference between the two points' directions and the predicted direction is less than the angle threshold. Then, the angular difference and the distance between points are used as edge weights, that is:
[0122] e i,j =ω3dist(P i ,P j )+
[0123] ω4min{[1 - cos(d i ,dir i,j )],[1 - cos(d j ,dir j,i )]} (9);
[0124] Among them, (P i ,P j ) represents the positions of skeleton point i and skeleton point j; (d i ,dir i,j ) represents the predicted direction of skeleton point i and the orientation from skeleton point i to skeleton point j; dist(P i ,P j ) represents the Euclidean space distance between skeleton point i and skeleton point j; cos(d i ,dir i,j ) represents the cosine distance between the predicted direction of skeleton point i and the orientation from skeleton point i to skeleton point j; ω3 and ω4 are both hyperparameters; since there cannot be loops in the trachea structure skeleton SWC file (that is, there is only one path between two skeleton points), the minimum spanning tree is used to remove the loops from the constructed point-edge graph to obtain the tree structure of the final trachea structure.
[0125] 5) Overall framework:
[0126] The method of the present invention is mainly divided into two steps. In the first step, the Swin UNETR network with self-attention mechanism is used to extract the tracheal structure region from the original electron microscopy micrographs. Through the self-attention mechanism of the Transformer network, long-range dependencies in the image are captured, and a more accurate and spatially continuous tracheal structure segmentation result is obtained. In the second step, a biological tubular structure tracking network with skeleton point prediction is used to perform tracheal tracking by a connection direction-guided biological tubular structure tracking method. This method can extract multi-level and multi-scale image features, can adaptively process regions with poor image quality and regions with complex biological structures, and generate better and more accurate tracking results. By using the connection direction of the skeleton point set predicted by the biological tubular structure tracking network with skeleton point prediction to guide the connection process, a more accurate tracheal structure reconstruction result is obtained.
[0127] Table 1 shows the comparison of the accuracy of tracheal reconstruction results of each method for Drosophila electron microscopy micrographs
[0128]
[0129] Table 2 shows the ablation experiment results of the method of the present invention in the Drosophila tracheal tracking task
[0130]
[0131] As can be seen from Table 1 above, compared with the traditional tracking method NGPST and the deep learning method NRTR, the method of the present invention obtains the best results in all indicators and has a large improvement. As can be seen from Table 2 above, the method of the present invention uses the skeleton proposal point prediction sub-module to increase the Transformer decoder query and utilize the direction prediction information to guide the tracheal structure connection, realizing a more accurate reconstruction result.
[0132] In the embodiment of the present invention, a biological tubular structure segmentation network with self-attention mechanism is used to segment the Drosophila trachea in the Drosophila whole-brain electron microscopy micrographs, obtaining a tracheal tracking result with more accurate three-dimensional spatial positions of skeleton points and better spatial continuity, and solving the problem of missing segmentation of slender tracheas. The present invention realizes the accurate three-dimensional tracking of the tracheal structure in the electron microscopy micrographs through a biological tubular structure tracking network with skeleton point prediction based on a detection network and connection direction prediction. The present invention proposes a method for dynamically generating query points in the skeleton point prediction sub-module, making the network training more stable and obtaining better prediction results. And a tree-like connection structure of the tracheal structure is constructed according to the connection direction of the predicted skeleton points, obtaining an accurate three-dimensional reconstruction result. According to the method of the present invention, an accurate three-dimensional reconstruction result of the trachea can be obtained under the annotation of a small amount of tracheal data in the Drosophila whole-brain electron microscopy micrographs, see Figure 4. Under experimental verification, the method of the present invention has achieved state-of-the-art three-dimensional reconstruction performance of biological tubular structures.
[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0134] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art section of this article is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.
Claims
1. A method for reconstructing biological tubular structures in electron microscopy microscopic images, characterized in that, Including: Step 1: Divide the original electron microscopy micrograph into blocks to obtain block images. Use a biological tubular structure segmentation network with a self-attention mechanism to capture the long-range dependencies of biological tubular structures in the original electron microscopy micrograph through the self-attention mechanism. Segment the biological tubular structures in each block image one by one, and obtain a biological tubular structure probability segmentation image based on the long-range dependencies. Step 2: Preprocess the biological tubular structure probability segmentation image obtained in Step 1 to obtain image blocks for tracking. Use a biological tubular structure tracking network with skeleton point prediction to predict the skeleton point set of the biological tubular structure for the obtained image blocks for tracking. Connect each skeleton point in the skeleton point set of the biological tubular structure according to the connection direction to construct a tree-like biological tubular structure, and obtain the reconstruction result of the biological tubular structure.
2. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 1, wherein In Step 2, the biological tubular structure probability segmentation image obtained in Step 1 is preprocessed in the following manner to obtain image blocks for tracking. Use a biological tubular structure tracking network with skeleton point prediction to predict the skeleton point set of the biological tubular structure for the obtained image blocks for tracking. Connect each skeleton point in the skeleton point set of the biological tubular structure according to the connection direction to construct a tree-like biological tubular structure, and obtain the reconstruction result of the biological tubular structure, including: Step 21: Use the segmentation probability of the biological tubular structure region in the biological tubular structure probability segmentation image segmented in Step 1 to obtain a block three-dimensional segmentation probability image. Stitch multiple block three-dimensional segmentation probability images into an overall three-dimensional segmentation probability image. Downsample the obtained overall three-dimensional segmentation probability image and cut it into at least one tracking image block of a fixed size. Step 22: Input each tracking image block obtained in Step 21 into a biological tubular structure tracking network with skeleton point prediction for skeleton point prediction to obtain the three-dimensional spatial position and connection direction of the skeleton points of the biological tubular structure in the tracking image block. The three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius. Step 23: Use the three-dimensional spatial position and connection direction of the obtained skeleton points of the biological tubular structure to connect each skeleton point, complete the three-dimensional construction of the tree-like structure of the biological tubular structure for each tracking image block. Finally, stitch the reconstruction results of the tree-like structures of the biological tubular structures of multiple tracking image blocks to obtain the reconstruction result of the biological tubular structure corresponding to the overall three-dimensional segmentation probability image.
3. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 1 or 2, characterized in that, The biological tubular structure segmentation network with a self-attention mechanism uses a Swin UNETR network with a self-attention mechanism. The physical resolution of the input electron microscopy micrograph of the Swin UNETR network with a self-attention mechanism is 16nm×16nm×40nm, and the voxel size is 128×128×16. The output of the Swin UNETR network with a self-attention mechanism is a probability image indicating whether each voxel belongs to a biological tubular structure, that is, a biological tubular structure probability segmentation image. The biological tubular structure tracking network with skeleton point prediction includes: an image feature extraction network, a skeleton point set prediction network based on Transformer, and a skeleton point set connection module. Among them, The described image feature extraction network can take as input the tracking image patches obtained by preprocessing the probability segmentation image of the biological tubular structure obtained in step 1, and extract low-resolution image features as output; The described Transformer-based skeleton point set prediction network is communicatively connected to the output end of the image feature extraction network. It can use the low-resolution image features output by the image feature extraction network as the input token sequence, calculate the associations between the image patch tokens in the input token sequence, generate a skeleton proposal point feature sequence of the biological tubular structure, and predict the classification probability, three-dimensional spatial position, and connection direction of each skeleton point in the skeleton point sequence of the biological tubular structure. The skeleton point set is composed of the classification probability, three-dimensional spatial position, and connection direction of each skeleton point, and the position consists of coordinates and a radius; The described skeleton point set connection module is communicatively connected to the output end of the Transformer-based skeleton point set prediction network. It can take the skeleton point set predicted by the Transformer-based skeleton point set prediction network as input, and connect each skeleton point in the skeleton point set based on the predicted connection direction of each skeleton point to generate the tree structure of the final biological tubular structure.
4. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 3, wherein The described image feature extraction network consists of a three-dimensional residual network and a feature pyramid network connected in sequence. Among them, the three-dimensional residual network takes as input the tracking image patches obtained by preprocessing the probability segmentation image of the biological tubular structure obtained in step 1, can perform three-dimensional residual processing and then output to the feature pyramid network for low-resolution image feature extraction to obtain low-resolution image features; The described Transformer-based skeleton point set prediction network consists of a Transformer encoder, a skeleton proposal point set prediction sub-module, a Transformer decoder, and a skeleton point set prediction sub-module; among them, The described Transformer encoder is communicatively connected to the image feature extraction network. It can use the position encoding at each pixel position of the low-resolution image features output by the image feature extraction network and the pixel feature sequence composed of all pixel features of the low-resolution image features as the common input. Through the self-attention mechanism, it maps the pixel feature sequence and the position encoding at each pixel position to a feature sequence containing the position information and content information of all image features, and outputs the feature sequence; The described skeleton proposal point prediction sub-module is communicatively connected to the output end of the Transformer encoder. It can use the feature sequence output by the Transformer encoder as input, and respectively predict the classification probability, three-dimensional spatial position, and connection direction of the skeleton proposal points of the biological tubular structure. The three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius; The Transformer decoder is communicatively connected to the Transformer encoder and the skeleton proposal point set prediction sub-module respectively. It can use the top N biological tubular structure skeleton points with the largest probabilities in the classification probabilities output by the skeleton proposal point set prediction sub-module as skeleton proposal points, select N features corresponding to these N skeleton proposal points from the feature sequence, recombine them into a biological tubular structure point prediction feature proposal sequence as the initial object query matrix, and use the position encodings of these N skeleton proposal points and the feature sequence output by the Transformer encoder as keys and values for each layer of the Transformer decoder, and utilize the cross-attention mechanism to capture the dependencies between the target biological tubular structure points and all positions in the image, and output target biological tubular structure point features with the same shape as the biological tubular structure point prediction feature proposal sequence. The skeleton point set prediction sub-module is communicatively connected to the Transformer decoder. It can use the target biological tubular structure point features output by the Transformer decoder as input to obtain the classification probabilities, three-dimensional spatial positions and angles of each skeleton point. The three-dimensional spatial position consists of three-dimensional spatial coordinates and a radius, and the obtained skeleton points form a skeleton point set.
5. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 4, wherein In the Transformer-based skeleton point set prediction network, the Transformer encoder is composed of multiple encoder layers connected in sequence, and each encoder layer is composed of a multi-head self-attention module and a multi-layer perceptron module. The skeleton proposal point set prediction sub-module is composed of six multi-layer perceptrons. Among them, the first two multi-layer perceptrons are respectively used to predict the classification probabilities and three-dimensional spatial positions of the biological tubular structure skeleton points, and the last four multi-layer perceptrons are used to predict the horizontal angle and vertical angle corresponding to the connection direction. The Transformer decoder is composed of multiple decoder layers connected in sequence, and each decoder layer is composed of a multi-head self-attention sub-module, a multi-head cross-attention sub-module and a multi-layer perceptron sub-module.
6. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 4, wherein, The skeleton point set prediction sub-module is composed of six multi-layer perceptrons. Among them, the first two multi-layer perceptrons are respectively used to predict the classification probabilities and three-dimensional spatial positions of the biological tubular structure points, and the last four multi-layer perceptrons are used to predict the horizontal angle and vertical angle corresponding to the connection direction.
7. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 3, wherein The total loss function of the biological tubular structure tracking network with skeleton point prediction is: Among them, the meanings of the parameters are as follows: represents all predicted skeleton proposal points and its matching points in the ground truth the sum of the matching losses; represents all predicted skeleton points and its matching points in the ground truth the sum of the matching losses; represents the set of skeleton proposal points composed of N skeleton points with the largest probabilities selected by the Top-K strategy by the skeleton proposal point set prediction sub-module of the Transformer-based skeleton point set prediction network of the biological tubular structure tracking network with skeleton point prediction; represents the set of skeleton proposal points the optimal match with the ground truth skeleton point set Y; represents the set of skeleton points composed of N skeleton points output by the Transformer-based skeleton point set prediction network; represents the set of skeleton points the optimal match with the ground truth skeleton point set Y; ω1 and ω2 are hyperparameters respectively; Among the above total losses, all are calculated as follows and are: Among them, c i , L i , A i respectively represent the classification probability, three-dimensional spatial position and angle of the true value Y i . The three-dimensional spatial position is composed of three-dimensional spatial coordinates and a radius; respectively represent the classification probability, three-dimensional spatial position and angle of the prediction result . The three-dimensional spatial position is composed of three-dimensional spatial coordinates and a radius; ω cls , ω box , ω iou , ω dir are hyperparameters respectively; represents the classification loss, which is set as the cross-entropy loss function; and both represent the position loss, is the L1 loss function, is the sphere intersection over union loss function; is the direction loss, The calculation method is as follows: Among them, respectively represent the true direction d of each skeleton point i corresponding horizontal angle classification, horizontal angle within-class offset value, vertical angle classification, and vertical angle within-class offset; respectively represent the predicted direction of the horizontal angle classification, horizontal angle within-class offset value, vertical angle classification, and vertical angle within-class offset; represents the angle classification loss, set as the cross-entropy loss function; represents the angle within-class offset regression loss, using the L1 loss function, and only calculates the gap between the predicted value and the true value in the offset values where the angle target interval is located.
8. The method for reconstructing a biological tubular structure in an electron microscope microscopic image according to claim 1 or 2, characterized in that, In step 2, the skeleton point set of the biological tubular structure is connected according to the connection direction in the following manner to construct a tree-like biological tubular structure, and a biological tubular structure reconstruction image is obtained, including: Delete the skeleton points with classification probabilities less than a predetermined threshold from the set of skeleton points of the biological tubular structure, and use the remaining skeleton points as the selected skeleton points. Calculate the predicted connection direction of each skeleton point based on the angles predicted by the selected skeleton points. Establish a point-edge graph for all the remaining skeleton points, establish an edge relationship between two skeleton points whose distance is less than the distance threshold and the angle difference between the two points' directions and the predicted direction is less than the angle threshold, and then use the angle difference and the distance between points as the edge weight e i,j , e i,j is: e i,j = ω3dist(P i , P j ) + ω4min{[1 - cos(d i , dir i,j )], [1 - cos(d j , dir j,i )]} (9); Among them, dist(P i , P j ) represents the Euclidean space distance between skeleton point i and skeleton point j; (P i , P j ) represents the positions of skeleton point i and skeleton point j; cos(d i , dir i,j ) represents the cosine distance between the predicted direction of skeleton point i and the orientation from skeleton point i to skeleton point j; (d i , dir i,j ) represents the orientation from the predicted direction of skeleton point i to skeleton point j; cos(d j , dir j,i ) represents the cosine distance between the predicted direction of skeleton point j and the orientation from skeleton point j to skeleton point i; (d j , dir j,i ) represents the orientation from the predicted direction of skeleton point j to skeleton point i; ω3 and ω4 are both hyperparameters; min{} represents taking the minimum value; The minimum spanning tree is used to remove the loops from the constructed point-edge graph to obtain the tree-like structure of the final biological tubular structure.
9. A processing device, characterized in that, Including: At least one memory for storing one or more programs; At least one processor capable of executing the one or more programs stored in the memory. When the one or more programs are executed by the processor, the processor can implement the method according to any one of claims 1-8.
10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Tubular structure image segmentation method
CN108510506A
Method and device for segmenting lung blood vessels in CT image, terminal equipment and medium
CN116740081A