Image segmentation method, electronic equipment and storage medium
Through the anisotropic convolution blocks and long convolution kernels of the image segmentation network, the problem of insufficient detection of linear structure tissues is solved, the accurate segmentation of small tissues in medical images is achieved, and the surgical risk is reduced.
Patent Information
- Application Number
- CN202510545759.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-12
Smart Images

Figure CN120635133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and in particular to an image segmentation method, electronic equipment, storage medium and computer program product. Background Art
[0002] During medical surgery, it is often necessary to avoid some small tissues when operating on the human body, including some tissues with small diameters and linear structures, such as nerves and blood vessels. In the related art, there is a lack of detection methods for detecting such small diameter linear tissues. Taking thyroid surgery as an example, it is often necessary to avoid the recurrent laryngeal nerve during thyroid surgery. In the existing technology, doctors often monitor the functional status of the recurrent laryngeal nerve by monitoring the electrical signals of the recurrent laryngeal nerve. However, the electrical signals lack intuitive spatial position information, making it difficult to prevent mechanical damage to the recurrent laryngeal nerve during surgery. Summary of the Invention
[0003] The present invention is proposed in view of the above problems. This solution can more accurately segment small tissues with linear structures in medical images, which helps doctors determine the specific locations of these small tissues based on the segmentation results of the medical images.
[0004] According to one aspect of the present invention, an image segmentation method is provided for segmenting objects with linear structures in medical images. The method includes: obtaining an image to be segmented collected for a target object, the target object having a linear structure; inputting the image to be segmented into an image segmentation network to obtain an image segmentation result of the image to be segmented, the image segmentation result being used to indicate a predicted position of the target object in the image to be segmented; wherein the image segmentation network includes a plurality of network stages connected in sequence and a segmentation head connected to the last network stage of the plurality of network stages, the plurality of network stages being used to extract feature maps of the image to be segmented, the input feature map of the latter network stage in every two adjacent network stages of the plurality of network stages including a feature map obtained by upsampling and / or downsampling at least a portion of the output feature map of the previous network stage, the segmentation head being used to determine the image segmentation result of the image to be segmented based on the feature maps output by the plurality of network stages, each of the plurality of network stages including at least one anisotropic convolution block, the anisotropic convolution block including a plurality of anisotropic convolution kernels with different directions, the plurality of anisotropic convolution kernels being elongated strip convolution kernels.
[0005] Optionally, the image to be segmented is input into an image segmentation network to obtain an image segmentation result of the image to be segmented, including: performing the following operations in each anisotropic convolution block of the image segmentation network: using multiple anisotropic convolution kernels in the anisotropic convolution block to perform convolution operations on the input feature map of the input anisotropic convolution block respectively to obtain convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels; and performing feature fusion on the convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels to obtain the output feature map of the anisotropic convolution block.
[0006] Optionally, feature fusion is performed on the convolution feature maps corresponding one-to-one to multiple anisotropic convolution kernels to obtain the output feature map of the anisotropic convolution block, including: weighted summation of the convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels according to the target weights corresponding one-to-one to the multiple anisotropic convolution kernels, and inputting the summation result into a preset activation function to obtain the output feature map of the anisotropic convolution block output by the preset activation function.
[0007] Optionally, multiple anisotropic convolution kernels in the anisotropic convolution block are used to perform convolution operations on the input feature map of the input anisotropic convolution block, including: for each anisotropic convolution kernel in the multiple anisotropic convolution kernels, determining the initial sampling point position of the anisotropic convolution kernel for the sampling point of the input feature map; adding a target offset to the initial sampling point position to obtain a new sampling point position; sampling on the input feature map according to the new sampling point position, and convolving the feature value obtained by the sampling with the anisotropic convolution kernel to obtain a convolution feature map corresponding to the anisotropic convolution kernel.
[0008] Optionally, the target object is the recurrent laryngeal nerve, and the aspect ratio of the anisotropic convolution kernel is greater than a preset value; and / or, the directions of the multiple anisotropic convolution kernels include directions corresponding to at least two angles from 0° to 180°, wherein 0° and 180° represent the horizontal direction of the input feature map, and 90° represents the vertical direction of the input feature map.
[0009] Optionally, the image segmentation network is trained in the following manner: obtaining training data, the training data including training images and annotated segmentation results, the training images including target objects, the annotated segmentation results being used to indicate the annotated position of the target objects in the training images; inputting the training images into the image segmentation network to obtain predicted segmentation results of the training images, the predicted segmentation results being used to indicate the predicted position of the target objects in the training images; adjusting the parameters of the image segmentation network based on the annotated segmentation results and the predicted segmentation results until the image segmentation network meets preset requirements.
[0010] Optionally, the parameters of the image segmentation network are adjusted based on the labeled segmentation results and the predicted segmentation results, including: calculating a loss value based on the labeled segmentation results and the predicted segmentation results, the loss value including a topology break perception loss value and / or a gradient direction consistency loss value; adjusting the parameters of the image segmentation network based on the loss value; wherein the labeled segmentation result includes the position of the labeled connected domain in the training image, and the position of the labeled connected domain in the training image represents the labeled position of the target object in the training image; the predicted segmentation result includes the position of the predicted connected domain in the training image, and the position of the predicted connected domain in the training image represents the predicted position of the target object in the training image; the topology break perception loss value is determined by the ratio between the intersection area and the total area, and the topology break perception loss value is negatively correlated with the intersection area, the intersection area is the area of the intersection area between the predicted connected domain and the labeled connected domain matching the predicted connected domain, and the total area is the sum of the area of the predicted connected domain and the area of the corresponding matching labeled connected domain; the gradient direction consistency loss value is determined by the cosine value between the gradient map corresponding to the predicted segmentation result and the gradient map corresponding to the labeled segmentation result, and the gradient direction consistency loss value is negatively correlated with the cosine value.
[0011] Optionally, when there are multiple predicted connected domains, before calculating the loss value based on the labeled segmentation result and the predicted segmentation result, the method also includes: for each predicted connected domain of the multiple predicted connected domains, matching the predicted connected domain with the labeled connected domain indicated by the labeled segmentation result; wherein the labeled connected domain with the largest overlapping area with the predicted connected domain is the labeled connected domain that matches the predicted connected domain; or, the labeled connected domain with the largest intersection-over-union ratio with the predicted connected domain is the labeled connected domain that matches the predicted connected domain.
[0012] Optionally, after inputting the image to be segmented into the image segmentation network, the method further includes: outputting the image segmentation result of the image to be segmented to a display interface of a display device for display.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising: a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used to execute the above-mentioned image segmentation method when the processor is executed.
[0014] According to yet another aspect of the present invention, a storage medium is provided, on which program instructions are stored. The program instructions are used to execute the above-mentioned image segmentation method when running.
[0015] According to yet another aspect of the present invention, a computer program product is provided, comprising computer program instructions, which are used to execute the above-mentioned image segmentation method when run.
[0016] The above technical solution segments an image to be segmented containing a target object with a linear structure by adopting an image segmentation network in which each network stage includes anisotropic convolution blocks. This can utilize the directional selectivity of the anisotropic convolution kernels in multiple directions contained in the anisotropic convolution block to effectively suppress the interference of the isotropic background in the image to be segmented, and can have a strong extraction capability for linear features in a specific direction, which is conducive to accurately segmenting the target object; in addition, by adopting a long anisotropic convolution kernel, the image segmentation network can have a strong response capability to the linear features of the target object with a linear structure; in addition, by adopting anisotropic convolution blocks in each network stage, it is conducive to ensuring that the linear texture features of the target object at all levels are more completely captured, thereby ensuring accurate segmentation of the target object in the image to be segmented; on the other hand, the input feature map of each network stage can include a feature map obtained by upsampling and / or downsampling the output feature map of the previous network stage, which is conducive to the image segmentation network gradually extracting more low-resolution feature maps, thereby simultaneously capturing the global semantic information and local detail information of the image to be segmented.
[0017] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other objects, features, and advantages of the present invention will become more apparent through a more detailed description of the embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings are provided to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and are not intended to limit the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0019] Figure 1 A schematic flow chart of an image segmentation method according to an embodiment of the present invention is shown;
[0020] FIG2( a ) shows a schematic diagram of an image to be segmented according to an embodiment of the present invention;
[0021] FIG2( b ) shows the image segmentation result obtained by inputting the image to be segmented in FIG2( a ) into the image segmentation network;
[0022] Figure 3 A schematic diagram of the structure of an image segmentation network according to an embodiment of the present invention is shown;
[0023] Figure 4A schematic diagram of an anisotropic convolution block according to one embodiment of the present invention is shown;
[0024] FIG5( a ) shows a schematic diagram of a labeled segmentation result according to an embodiment of the present invention;
[0025] FIG5( b ) shows a schematic diagram of a predicted segmentation result according to an embodiment of the present invention;
[0026] Figure 6 A schematic block diagram of an image segmentation apparatus according to an embodiment of the present invention is shown;
[0027] Figure 7 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present invention more apparent, exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present invention.
[0029] To at least partially address the above technical issues, embodiments of the present invention provide an image segmentation method, apparatus, electronic device, storage medium, and computer program product. This solution can relatively accurately segment fine linear structures in medical images, helping doctors determine the specific locations of these fine structures based on the segmentation results.
[0030] See also Figure 1 FIG2 is a schematic flow chart of an image segmentation method according to an embodiment of the present invention. According to one aspect of the present invention, an image segmentation method is provided for segmenting linear objects in medical images. The method includes steps S110 and S120.
[0031] For example, the medical image may be an ultrasound image acquired using an ultrasound system, an endoscopic image acquired using an endoscope system, etc. The medical image may include linear objects such as the sciatic nerve, median nerve, recurrent laryngeal nerve, and capillaries. Linear structures generally have a large aspect ratio. For example, the recurrent laryngeal nerve has an aspect ratio generally greater than 10:1. The image segmentation method of the embodiment of the present invention may be applicable to image segmentation of medical images containing linear objects.
[0032] In step S110 , an image to be segmented is acquired for a target object, where the target object is a linear structure.
[0033] Exemplarily, the target object may be a biological tissue with a linear structure, such as the recurrent laryngeal nerve, capillaries, etc. in the aforementioned embodiment. The image to be segmented may be a medical image containing a partial area or the entire area of the target object, and the shape of the target object in the image to be segmented may be a linear shape. Taking the target object being the recurrent laryngeal nerve as an example, the medical image that may contain the recurrent laryngeal nerve may be, for example, a thyroid image. The thyroid image may include a thyroid longitudinal section (i.e., sagittal plane) image and a thyroid transverse section (i.e., cross-sectional) image, wherein the shape of the recurrent laryngeal nerve in the thyroid longitudinal section image is usually linear, and the shape in the thyroid transverse section image is usually point-shaped. It can be understood that the former can be used as the image to be segmented in the embodiment of the present invention.
[0034] In step S120, the image to be segmented is input into an image segmentation network to obtain an image segmentation result for the image to be segmented, the image segmentation result being used to indicate the predicted location of the target object in the image to be segmented. The image segmentation network includes a plurality of network stages connected sequentially and a segmentation head connected to the last network stage of the plurality of network stages. The plurality of network stages are used to extract feature maps of the image to be segmented. The input feature map of the subsequent network stage in each of two adjacent network stages of the plurality of network stages includes a feature map obtained by upsampling and / or downsampling at least a portion of the output feature map of the previous network stage. The segmentation head is used to determine the image segmentation result of the image to be segmented based on the feature maps output by the plurality of network stages. Each of the plurality of network stages includes at least one anisotropic convolution block, the anisotropic convolution block including a plurality of anisotropic convolution kernels having different directions, the plurality of anisotropic convolution kernels being elongated convolution kernels. An anisotropic convolution kernel may refer to a single convolution kernel having direction-dependence, whose scale or response characteristics vary in different directions. That is, an anisotropic convolution kernel has different feature extraction capabilities in different directions and is not limited to convolution in only one direction.
[0035] Exemplarily, an image to be segmented is input into an image segmentation network, which can output an image segmentation result for the image to be segmented. The image segmentation result can be, for example, a binary image, where the pixel value of the target object region can be 1, and the pixel values of other regions outside the target object region can be 0. The image segmentation result can also be a marking of the target object region in the image to be segmented, for example, by highlighting the edge of the target object region, or by highlighting the target object region. The image segmentation result can indicate the predicted position of the target object in the image to be segmented. Please refer to Figures 2(a) and 2(b). Figure 2(a) is a schematic diagram of an image to be segmented according to one embodiment of the present invention, and Figure 2(b) is the image segmentation result obtained by inputting the image to be segmented in Figure 2(a) into the image segmentation network. In this embodiment, the image to be segmented is an ultrasound image of a longitudinal section of the thyroid gland, and the target object contained therein is the recurrent laryngeal nerve. The image segmentation result is a binary image, and in the image segmentation result, the position of the white curve in the binary image can serve as the predicted position of the recurrent laryngeal nerve in the longitudinal section of the thyroid gland.
[0036] Exemplarily, the network architecture adopted by the image segmentation network can be, for example, a deep experimental network (DeepLab), a U-net (U-Net), a segmentation network (SegNet), a high-resolution network (HRNet), etc., which can include a network architecture adopted by a neural network of multiple network stages connected in sequence. In other words, the image segmentation network of an embodiment of the present invention may include multiple network stages connected in sequence. Each network stage may include one or more anisotropic convolution blocks, and each anisotropic convolution block may include multiple anisotropic convolution kernels with different directions, and the sizes of the multiple anisotropic convolution kernels may be the same. For example, an anisotropic convolution block may include: an anisotropic convolution kernel of 0°, an anisotropic convolution kernel of 45°, an anisotropic convolution kernel of 90°, and an anisotropic convolution kernel of 135°. For another example, an anisotropic convolution block may include: an anisotropic convolution kernel of 0°, an anisotropic convolution kernel of 25°, and an anisotropic convolution block of 65°. It should be noted that 0° and 180° can both represent the horizontal direction of the input feature map, and 90° can represent the vertical direction of the input feature map. Each anisotropic convolution block can specifically include two or more anisotropic convolution kernels, and the directions of each anisotropic convolution kernel can be different. Generally speaking, the rotation center of the convolution kernel is the center point of the convolution kernel (i.e., the geometric center). For example, the rotation center of a convolution kernel with a size of a*b can be the position corresponding to (a / 2, b / 2). An anisotropic convolution kernel can be obtained by rotating a convolution kernel with a size of a*b around the rotation center of the convolution kernel in the direction of feature extraction. Those skilled in the art will understand that the basic principle of convolution kernel rotation is to rotate the two-dimensional matrix to reposition each element in the convolution kernel, which will not be elaborated here.
[0037] Exemplarily, the shape of the anisotropic convolution kernel can be a long strip, such as a 1*6, 1*7, or 1*8 long strip convolution kernel. It can be understood that each network stage can also include one or more isotropic convolution blocks, and each isotropic convolution block can include at least one isotropic convolution kernel, such as a 3*3 square convolution kernel. In a network stage, the isotropic convolution block and the anisotropic convolution block can be located in parallel convolution paths, that is, the convolution paths of the isotropic convolution block and the anisotropic convolution block can be different. Alternatively, the isotropic convolution block and the anisotropic convolution block can be connected in a preset order on the same convolution path. For example, the anisotropic convolution block and the isotropic convolution block can be connected in sequence, first the anisotropic convolution block extracts the features of the input feature map, and then the isotropic convolution block extracts features from the feature map output by the anisotropic convolution block. It should be noted that the receptive field of an isotropic convolution kernel (e.g., a 3*3 square convolution kernel) is symmetrically distributed, allowing for relatively uniform extraction of feature information from all directions of the image. However, isotropic convolution kernels struggle to effectively capture the linear texture features of the target object. Furthermore, in some medical images, isotropic background noise may exist around the target object, such as thyroid tissue and vascular artifacts near the recurrent laryngeal nerve (i.e., the target object) in thyroid ultrasound images. Neural networks containing only isotropic convolution kernels can easily confuse the target object region with the background region. In particular, when the signal-to-noise ratio (SNR) of the target object region in the image to be segmented is low, it is difficult to accurately segment the target object in the image to be segmented. Therefore, by embedding anisotropic convolution kernels in multiple directions into the image segmentation network, the directional sensitivity of the image segmentation network can be improved and the interference of isotropic background in the image to be segmented can be suppressed. Furthermore, because the anisotropic convolution kernels used are long strips, the image segmentation network is more responsive to the linear features of the target object, which is a specific shape. On the other hand, since each network stage is distributed at different network depths, if anisotropic convolution blocks are only embedded in some network stages with shallower depths, the feature levels extracted by these network stages will be shallower, and it may also be difficult to capture the linear texture features of the target object. Therefore, by embedding anisotropic convolution kernels in each of multiple network stages, the segmentation performance of the image segmentation network for objects with linear structures can be ensured.
[0038] Exemplarily, in an image segmentation network, for every two adjacent network stages, the feature map input to the subsequent network stage may include a feature map obtained by upsampling at least part of the feature map output by the previous network stage, or may include a feature map obtained by downsampling at least part of the feature map output by the previous network stage, or may include a feature map obtained by upsampling and downsampling at least part of the feature map output by the previous network stage. The resolution of the feature map obtained by downsampling is smaller than the original feature map, and the resolution of the feature map obtained by upsampling is larger than the original feature map. It can be understood that the feature map input to the subsequent network stage may also include a feature map directly output by the previous network stage. The image segmentation network may also include a segmentation head connected to the last network stage, and the segmentation head may determine and output the corresponding image segmentation result based on the feature map output by the last network stage.
[0039] See also Figure 3 As shown in FIG, it is a schematic diagram of the structure of an image segmentation network according to an embodiment of the present invention. Figure 3 In the embodiment shown, the network architecture used by the image segmentation network is the HRNet network architecture, which includes three network stages. In HRNet, each network stage may include four convolution blocks (the four convolution blocks in each network stage are in Figure 3(Simplified as a bold arrow in the figure). For each network stage, the network stage may include one anisotropic convolution block and three isotropic convolution blocks, or the network stage may include two anisotropic convolution blocks and two isotropic convolution blocks, or the network stage may include three anisotropic convolution blocks and one isotropic convolution block. The embodiment of the present invention does not limit the number and arrangement order of anisotropic convolution blocks and isotropic convolution blocks in the network stage. The isotropic convolution block and the anisotropic convolution block can be located in the same convolution path, and the connection order thereof can be, for example, the order of connecting the anisotropic convolution block first and the isotropic convolution block second, or the order of connecting the isotropic convolution block first and the anisotropic convolution block second. The first network stage (i.e., stage 1) includes one anisotropic convolution block, the second network stage (i.e., stage 2) connected to the first network stage includes two anisotropic convolution blocks, and the third network stage (i.e., stage 3) connected to the second network stage includes three anisotropic convolution blocks. The image segmentation network may further include a segmentation head connected to the third network stage. After the image to be segmented is input into the segmentation network, feature extraction may be performed on the image to be segmented to obtain an initial feature map ①, which may be input into stage 1. In stage 1, feature map ① may be convolved with an isotropic convolution kernel to obtain feature map ②, which may be input into the anisotropic convolution block in stage 1 to obtain feature map ③, and feature map ③ and the feature map obtained by downsampling feature map ③ may be input into stage 2. In stage 2, feature map ③ and the feature map obtained by downsampling feature map ③ may be convolved with an isotropic convolution kernel to obtain feature maps ④ and ⑤, respectively. By using anisotropic convolution blocks to convolve feature maps ④ and ⑤, feature maps ⑥ and ⑦ can be obtained. Figure 7 It can be input into stage 3, and feature map ⑥ and feature map ⑦ can be downsampled and upsampled respectively to obtain their respective corresponding feature maps. In stage 3, the isotropic convolution block can be used to convolve and fuse the feature maps obtained by upsampling feature map ⑥ and feature map ⑦ to obtain feature map ⑧; the isotropic convolution block can also be used to convolve and fuse the feature maps obtained by the first downsampling of feature map ⑦ and feature map ⑥ to obtain feature map ⑨; the isotropic convolution block can also be used to convolve and fuse the feature maps obtained by downsampling feature map ⑦ and feature map ⑥ after the second downsampling to obtain feature map ⑩. Then, feature map ⑧, feature map ⑨ and feature map ⑩ are convolved using anisotropic convolution blocks to obtain feature maps Feature Map And feature maps Feature Map And feature maps Can be upsampled separately, feature map Feature Map The feature map and feature map obtained by upsampling The feature map obtained by upsampling can be input to the segmentation head, which outputs the corresponding image segmentation result based on the input feature map.
[0040] The above technical solution segments an image to be segmented containing a target object with a linear structure by adopting an image segmentation network in which each network stage includes anisotropic convolution blocks. This can utilize the directional selectivity of the anisotropic convolution kernels in multiple directions contained in the anisotropic convolution block to effectively suppress the interference of the isotropic background in the image to be segmented, and can have a strong extraction capability for linear features in a specific direction, which is conducive to accurately segmenting the target object; in addition, by adopting a long anisotropic convolution kernel, the image segmentation network can have a strong response capability to the linear features of the target object with a linear structure; in addition, by adopting anisotropic convolution blocks in each network stage, it is conducive to ensuring that the linear texture features of the target object at all levels are more completely captured, thereby ensuring accurate segmentation of the target object in the image to be segmented; on the other hand, the input feature map of each network stage can include a feature map obtained by upsampling and / or downsampling the output feature map of the previous network stage, which is conducive to the image segmentation network gradually extracting more low-resolution feature maps, thereby simultaneously capturing the global semantic information and local detail information of the image to be segmented.
[0041] Optionally, the image to be segmented is input into an image segmentation network to obtain an image segmentation result of the image to be segmented, including: performing the following operations in each anisotropic convolution block of the image segmentation network: using multiple anisotropic convolution kernels in the anisotropic convolution block to perform convolution operations on the input feature map of the input anisotropic convolution block respectively to obtain convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels; and performing feature fusion on the convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels to obtain the output feature map of the anisotropic convolution block.
[0042] Exemplarily, each anisotropic convolution block in the image segmentation network can perform feature extraction on the feature map (i.e., input feature map) input to the anisotropic convolution block. Specifically, each anisotropic convolution kernel in the anisotropic convolution block can perform a convolution operation on the input feature map input to the anisotropic convolution block to obtain a corresponding convolution feature map. The obtained convolution feature maps can be feature fused, and the anisotropic convolution block can output an output feature map obtained by feature fusion. For example, an anisotropic convolution block may include four anisotropic convolution kernels with different directions. The four anisotropic convolution kernels can perform convolution operations on the input feature map input to the anisotropic convolution block respectively to obtain four convolution feature maps with the same resolution. Feature fusion of the four convolution feature maps can obtain the output feature map output by the anisotropic convolution block. Please refer to Figure 4 As shown in FIG, it is a schematic diagram of an anisotropic convolution block according to an embodiment of the present invention. Figure 4 The anisotropic convolution block shown can include four long anisotropic convolution kernels, each of which corresponds to a different direction. The four anisotropic convolution kernels can each perform a convolution operation on the input feature map, and the four convolution feature maps obtained by each convolution are fused to obtain the output feature map of the anisotropic convolution block.
[0043] The above technical solution fuses the convolution feature maps obtained by each anisotropic convolution kernel contained in each anisotropic convolution block, which can fuse the information of the input feature map in various directions, which is beneficial to remove artifacts or noise caused by the anisotropic convolution kernel in a single direction to a certain extent, and can reduce the computational burden in subsequent network stages.
[0044] Optionally, feature fusion is performed on the convolution feature maps corresponding one-to-one to multiple anisotropic convolution kernels to obtain the output feature map of the anisotropic convolution block, including: weighted summation of the convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels according to the target weights corresponding one-to-one to the multiple anisotropic convolution kernels, and inputting the summation result into a preset activation function to obtain the output feature map of the anisotropic convolution block output by the preset activation function.
[0045] Exemplarily, for each anisotropic convolution block, the multiple anisotropic convolution kernels contained in the anisotropic convolution block each have a corresponding target weight, and the target weight is usually a learnable weight. According to the target weight corresponding to each anisotropic convolution kernel, the corresponding output convolution feature map can be weighted summed, and the obtained weighted summation result can be input into a preset activation function. The output result of the activation function is the output feature map of the anisotropic convolution block. Specifically, taking the anisotropic convolution kernels contained in the anisotropic convolution block as an example, the directions of 0°, 45°, 90°, and 135° include, the output feature map can be expressed by the following formula (1):
[0046]
[0047] Where, F dir represents the output feature map, F θ Represents the convolution feature map corresponding to the anisotropic convolution kernel with direction θ, W θ represents the target weight of the anisotropic convolution kernel with direction θ, and σ represents the preset activation function.
[0048] In the above technical solution, anisotropic convolution kernels in different directions can have corresponding target weights, which is conducive to dynamically adjusting the emphasis on features in each direction according to the specific situation of the image to be segmented, and is conducive to improving the segmentation accuracy of the image segmentation network.
[0049] Optionally, multiple anisotropic convolution kernels in the anisotropic convolution block are used to perform convolution operations on the input feature map of the input anisotropic convolution block, including: for each anisotropic convolution kernel in the multiple anisotropic convolution kernels, determining the initial sampling point position of the anisotropic convolution kernel for the sampling point of the input feature map; adding a target offset to the initial sampling point position to obtain a new sampling point position; sampling on the input feature map according to the new sampling point position, and convolving the feature value obtained by the sampling with the anisotropic convolution kernel to obtain a convolution feature map corresponding to the anisotropic convolution kernel.
[0050] For example, the target object usually has a morphological feature of local curvature, such as the local curvature of the nerve direction. In order to adapt to the local curvature of the target object, the anisotropic convolution kernel can perform a convolution operation on the input feature map in a deformable convolution manner. Specifically, for each sampling point on the input feature map, an additional offset (i.e., a target offset) can be added to the sampling point position (i.e., the initial sampling point position) where the sampling point is located to adjust the sampling point position of the anisotropic convolution kernel. The target offset is a learnable offset. More specifically, the image segmentation network can also include a convolution layer for generating a target offset. These convolution layers can be embedded in one or more network stages of the image segmentation network and can be optimized during the training process. The convolution layer for generating the target offset can generate two target offsets for each sampling point, namely, target offsets in the horizontal and vertical directions of the input feature map. By superimposing the target offset on the original sampling point position, a new sampling point position can be obtained. For example, the original sampling point position is (x, y), and the target offset is equal to (△x, △y), where △x is the target offset in the horizontal direction and △y is the target offset in the vertical direction. Then the new sampling point position is (x+△x, y+△y). Assuming that the size of an anisotropic convolution kernel is h*w, it can be understood that the anisotropic convolution kernel can output 2*h*w target offsets. In the training stage, the initialization value of the target offset can be 0. During the training process, the target offset can be gradually optimized. After the image segmentation network completes the training, the target offset may be negative, 0 or positive. The anisotropic convolution kernel can be convolved with the feature value of the new sampling point position on the input feature map to obtain the corresponding convolution feature map. It can be understood that the isotropic convolution kernel in the image segmentation network can also optionally use a deformable convolution method to perform feature operations with the corresponding input feature map. The embodiment of the present invention does not make specific limitations on the convolution method of the isotropic convolution kernel.
[0051] The above technical solution increases the target offset to the initial sampling point position of the input feature map of the anisotropic convolution kernel, which can better adapt to the local curvature characteristics of the target object, thereby obtaining a more accurate image segmentation result.
[0052] Optionally, the target object is the recurrent laryngeal nerve, and the aspect ratio of the anisotropic convolution kernel is greater than a preset value; and / or the directions of the multiple anisotropic convolution kernels include directions corresponding to at least two angles between 0° and 180°, wherein 0° and 180° represent the horizontal direction of the input feature map, and 90° represents the vertical direction of the input feature map. The preset value corresponding to the aspect ratio can be a value such as 2 or 3. For different target objects, the preset value can be adaptively adjusted according to the length of the target object.
[0053] For example, the recurrent laryngeal nerve is usually a slender linear structure in the longitudinal section image of the thyroid gland, with an aspect ratio greater than 10:1 and a diameter of approximately 2 mm, which is much smaller than other types of nerves, such as the sciatic nerve and the median nerve (about 10 mm in diameter). The image segmentation method of the embodiment of the present invention can be used to segment the longitudinal section image of the thyroid gland containing the recurrent laryngeal nerve. In other words, the target object can be the recurrent laryngeal nerve, and the image to be segmented can be the longitudinal section image of the thyroid gland. In this case, an anisotropic convolution kernel with an aspect ratio greater than a preset value can be used. The size of the anisotropic convolution kernel can be, for example, 1*6 and 6*1, or 1*7 and 7*1, or 1*8 and 8*1. It is more preferable to use anisotropic convolution kernels with a size of 1*7 and 7*1. Moreover, the direction of each anisotropic convolution kernel can be a direction corresponding to any angle between 0° and 180°. When the target object is the recurrent laryngeal nerve, anisotropic convolution kernels of 0°, 45°, 90° and 135° can be preferred. 0° and 180° both represent the horizontal direction of the input feature map, and 90° represents the vertical direction of the input feature map. Correspondingly, 45° and 135° can represent the two diagonal directions of the input feature map. It can be understood that the horizontal direction of the input feature map is also the horizontal direction of the image to be segmented, and the vertical direction of the input feature map is also the vertical direction of the image to be segmented.
[0054] The image segmentation network used in the embodiment of the present invention has a high sensitivity to small nerves with linear structures, and is therefore particularly suitable for image segmentation of the recurrent laryngeal nerve. In addition, the use of the above-mentioned anisotropic convolution kernel of a specific size and a specific direction in the image segmentation network can better adapt to the morphological characteristics of the recurrent laryngeal nerve, so that when the recurrent laryngeal nerve is included in the image to be segmented, a more accurate image segmentation result can be obtained.
[0055] Optionally, the image segmentation network is trained in the following manner: obtaining training data, the training data including training images and annotated segmentation results, the training images including target objects, the annotated segmentation results being used to indicate the annotated position of the target objects in the training images; inputting the training images into the image segmentation network to obtain predicted segmentation results of the training images, the predicted segmentation results being used to indicate the predicted position of the target objects in the training images; adjusting the parameters of the image segmentation network based on the annotated segmentation results and the predicted segmentation results until the image segmentation network meets preset requirements.
[0056] For example, an image segmentation network can be trained using training images captured for a target object and the corresponding annotated segmentation results. Similarly, the training image includes the target object, and the annotated segmentation results can indicate the annotated location of the target object in the training image. Inputting the training image into the image segmentation network can generate a corresponding predicted segmentation result. This predicted segmentation result can indicate the predicted location of the target object in the training image as predicted by the image segmentation network, which may deviate significantly from the annotated location. Please refer to Figures 5(a) and 5(b). Figure 5(a) is a schematic diagram of an annotated segmentation result according to one embodiment of the present invention, and Figure 5(b) is a schematic diagram of a predicted segmentation result according to one embodiment of the present invention. Specifically, the training image corresponding to the annotated segmentation result in Figure 5(a) and the training image corresponding to the predicted segmentation result in Figure 5(b) are the same training image. More specifically, Figure 5(a) is the annotated segmentation result of a training image containing the recurrent laryngeal nerve, and Figure 5(b) is the predicted segmentation result obtained by inputting the training image into the image segmentation network. In this embodiment, the curve smoothness of the predicted segmentation result is less than the curve smoothness of the annotated segmentation result, and the directions of the two curves differ significantly. The image segmentation network can be adjusted by annotating the segmentation results and predicting the segmentation results. When the deviation between the predicted segmentation result and the corresponding annotated segmentation result meets the allowable prediction deviation, the image segmentation network can be considered to meet the preset requirements.
[0057] The above technical solution can train the image segmentation network quickly and efficiently, thereby ensuring that the performance of the trained image segmentation network can meet the required segmentation accuracy.
[0058] Optionally, the parameters of the image segmentation network are adjusted based on the labeled segmentation results and the predicted segmentation results, including: calculating a loss value based on the labeled segmentation results and the predicted segmentation results, the loss value including a topology break perception loss value and / or a gradient direction consistency loss value; adjusting the parameters of the image segmentation network based on the loss value; wherein the labeled segmentation result includes the position of the labeled connected domain in the training image, and the position of the labeled connected domain in the training image represents the labeled position of the target object in the training image; the predicted segmentation result includes the position of the predicted connected domain in the training image, and the position of the predicted connected domain in the training image represents the predicted position of the target object in the training image; the topology break perception loss value is determined by the ratio between the intersection area and the total area, and the topology break perception loss value is negatively correlated with the intersection area, the intersection area is the area of the intersection area between the predicted connected domain and the labeled connected domain matching the predicted connected domain, and the total area is the sum of the area of the predicted connected domain and the area of the corresponding matching labeled connected domain; the gradient direction consistency loss value is determined by the cosine value between the gradient map corresponding to the predicted segmentation result and the gradient map corresponding to the labeled segmentation result, and the gradient direction consistency loss value is negatively correlated with the cosine value.
[0059] For example, the labeled segmentation result may indicate the location of one or more labeled connected domains in the training image. Each labeled connected domain may represent a target object. Accordingly, the location of the labeled connected domain in the training image may represent the labeled location of the target object in the training image. Similarly, the predicted segmentation result may indicate the location of one or more predicted connected domains in the training image. Each predicted connected domain may represent a target object predicted by the image segmentation network. Accordingly, the location of the predicted connected domain in the training image may represent the predicted location of the target object in the training image. For example, a topology break perception loss value may be calculated based on the labeled segmentation result and the predicted segmentation result. Specifically, for each predicted connected domain, the area of the intersection between the predicted connected domain and the matched labeled connected domain may be recorded as the intersection area. The topology break perception loss value may be negatively correlated with the intersection area, with the larger the intersection area, the smaller the topology break perception loss value. The sum of the area of the predicted connected domain and the area of the matched labeled connected domain may be recorded as the total area. The topology break perception loss value may be specifically determined based on the ratio between the intersection area and the total area. More specifically, the topology breakage perception loss value can be expressed by the following formula (2):
[0060]
[0061] Where, L break represents the topology break perception loss value, c represents the cth predicted connected domain, C represents the total number of predicted connected domains, P c represents the area of the cth predicted connected domain, G c represents the area P of the labeled connected domain that matches the cth predicted connected domain c ·G c Represents the area of the intersection between the c-th predicted connected component and the corresponding matched labeled connected component.
[0062] For example, a gradient directional consistency loss value can also be calculated based on the labeled segmentation result and the predicted segmentation result. The gradient directional consistency loss value is usually negatively correlated with the cosine value between the gradient map corresponding to the predicted segmentation result and the gradient map corresponding to the labeled segmentation result. More specifically, the gradient directional consistency loss value can be expressed by the following formula (3):
[0063]
[0064] Where, L grad Represents the gradient direction consistency loss value, Represents the gradient map corresponding to the predicted segmentation result, Represents the gradient map corresponding to the labeled segmentation result. The gradient map can usually be obtained by using the Sobel operator to perform gradient operations on the predicted segmentation result and the labeled segmentation result.
[0065] For example, the loss value also includes the topology fracture perception loss value L break And the gradient direction consistency loss value L grad When , the loss value can be expressed by the following formula (4):
[0066] L=λ1L break +λ2L gard (4)
[0067] Where λ1 represents the weight of the topology break perception loss, and λ2 represents the weight of the gradient direction consistency loss. The specific values of λ1 and λ2 can be defined by the user according to actual needs.
[0068] By introducing a topology-fracture-aware loss, the above-mentioned technical solution can effectively optimize the problem of segmentation results that may contain fractures when segmenting the target object. In particular, taking the recurrent laryngeal nerve as an example, when a thyroid lesion invades the gland, it may cause compression and adhesion between the recurrent laryngeal nerve and the lesion, which may appear as a fracture in the image. In this case, when the target object appears fractured in the image to be segmented, adjusting the image segmentation network parameters based on the topology-fracture-aware loss can effectively improve the performance of the image segmentation network and largely avoid the appearance of fractured target objects in the image segmentation results. Since the topology-fracture-aware loss is negatively correlated with the area of the intersection between the predicted connected domain and the matched annotated connected domain, the smaller the intersection area, the greater the degree of fracture and the higher the topology-fracture-aware loss. During the training process, in order to reduce the topology-fracture-aware loss, the area of the intersection between the predicted connected domain and the matched annotated connected domain is gradually increased through parameter adjustment, thereby enabling the image segmentation network to better recognize linear structures. On the other hand, by introducing the gradient direction consistency loss, it can ensure that the shape of the target object in the image segmentation result is highly consistent with the actual shape. Since the gradient direction consistency loss value is negatively correlated with the cosine value between the gradient map corresponding to the predicted segmentation result and the gradient map corresponding to the labeled segmentation result, the smaller the cosine value, the greater the difference in the gradient direction between the predicted segmentation result and the labeled segmentation result. During the training process, in order to reduce the gradient direction consistency loss value, the cosine value between the gradient map corresponding to the predicted segmentation result and the gradient map corresponding to the labeled segmentation result can be gradually increased by adjusting the parameters, so that the image segmentation result obtained by the image segmentation network can be better aligned with the actual target object in the gradient direction.
[0069] For example, the following Table 1 shows the Dice coefficients of the predicted connected domain and the labeled connected domain output by HRNet after adjusting the parameters of HRNet containing different types of convolution kernels using different loss values.
[0070] Table 1
[0071]
[0072]
[0073] As can be seen from Table 1 above, after only Dice Loss is used to adjust the parameters of HRNet without anisotropic convolution blocks, the corresponding Dice coefficient is 0.78, and the performance of the trained HRNet is poor. Those skilled in the art will understand that the larger the Dice coefficient, the larger the overlap area between the predicted connected domain and the labeled connected domain, and the better the segmentation performance of the image segmentation network. After only Dice Loss is used to adjust the parameters of HRNet containing anisotropic convolution blocks, the corresponding Dice coefficient is 0.82, which shows that the anisotropic convolution block can effectively capture the characteristics of linear structures. After only topological break perception loss is used to adjust the parameters of HRNet containing anisotropic convolution blocks, the corresponding Dice coefficient is 0.85, which shows that the introduction of topological break perception loss can effectively enhance the image segmentation network's maintenance of the continuity of linear structures. After tuning the HRNet parameters using the topology-fracture-aware loss and the gradient-direction-consistency loss for anisotropic convolutional blocks, the corresponding Dice coefficient was 0.87. This demonstrates that the gradient-direction-consistency loss can also effectively optimize image segmentation networks. Therefore, in this experiment, the image segmentation network obtained by tuning the HRNet parameters using the topology-fracture-aware loss and the gradient-direction-consistency loss for anisotropic convolutional blocks achieved the best performance.
[0074] Optionally, when there are multiple predicted connected domains, before calculating the loss value based on the labeled segmentation result and the predicted segmentation result, the method also includes: for each predicted connected domain of the multiple predicted connected domains, matching the predicted connected domain with the labeled connected domain indicated by the labeled segmentation result; wherein the labeled connected domain with the largest overlapping area with the predicted connected domain is the labeled connected domain that matches the predicted connected domain; or, the labeled connected domain with the largest intersection-over-union ratio with the predicted connected domain is the labeled connected domain that matches the predicted connected domain.
[0075] Exemplarily, the number of predicted connected domains can be multiple. For each predicted connected domain, the predicted connected domain can be matched with each annotated connected domain. In some embodiments, the annotated connected domain with the largest overlapping area with the predicted connected domain can be used as the annotated connected domain that matches the predicted connected domain. In other embodiments, the annotated connected domain with the largest intersection-over-union (IOU) with the predicted connected domain can be used as the annotated connected domain that matches the predicted connected domain. It should be noted that if the overlapping area of each annotated connected domain with the predicted connected domain is 0 (or the intersection-over-union ratio is 0), the aforementioned intersection area can be regarded as 0.
[0076] The above technical solution can quickly determine the labeled connected domains that match each predicted connected domain, so that the loss value can be efficiently calculated to train the image segmentation network.
[0077] Optionally, after inputting the image to be segmented into the image segmentation network, the method further includes: outputting the image segmentation result of the image to be segmented to a display interface of a display device for display.
[0078] Exemplarily, the image segmentation results of the image to be segmented can be displayed on a display interface of a display device. The image segmentation results can include a binary image output by the image segmentation network, and pixel positions with a pixel value of 1 can be used as the predicted position of the target object in the image to be segmented. The image segmentation results can also include a position marker superimposed on the image to be segmented based on the predicted position indicated by the binary image, and the image position indicated by the position marker can be used as the predicted position of the target object in the image to be segmented. For example, the position marker can be a highlighted rectangular frame, and the area within the rectangular frame can be used as the position range of the predicted position of the target object in the image to be segmented. In other words, the predicted position of the target object is within the rectangular frame. For another example, the position marker can be a highlighted bold line, the shape of the bold line is consistent with the shape of the target object, and the image position of the bold line can be used as the predicted position of the target object in the image to be segmented. In other embodiments, the image segmentation results can be predicted coordinates in the image coordinate system used by the image to be segmented (this image coordinate system is also the image coordinate system used by the binary image output by the image segmentation network), and the predicted position of the target object can be represented by the predicted coordinates. In this case, the display interface may include an image display area and a non-image display area. The image display area may display a binary image output by the image segmentation network and / or the original image to be segmented, and the non-image display area may display predicted coordinates. It should be noted that the above-mentioned display operation may be real-time. During surgery, the collected medical images may be input into the image segmentation network in real time, and the image segmentation results may be displayed in real time on the display interface of the display device. The doctor may avoid the target object based on the real-time displayed image segmentation results.
[0079] The above technical solution outputs the image segmentation results to the display interface of the display device for display, which makes it convenient for users to locate the position of the target object according to the displayed image segmentation results. In particular, during surgery, doctors can refer to the image segmentation results to determine the actual physical position of the target object, thereby avoiding the target object as much as possible.
[0080] See also Figure 6 , which is a schematic block diagram of an image segmentation apparatus according to one embodiment of the present invention. In another aspect of the present invention, an image segmentation apparatus is provided for segmenting objects with linear structures in medical images. The apparatus 600 includes:
[0081] An acquisition module 610 is used to acquire an image to be segmented that is collected for a target object, where the target object is a linear structure;
[0082] An input module 620 is used to input the image to be segmented into the image segmentation network to obtain an image segmentation result of the image to be segmented, where the image segmentation result is used to indicate the predicted position of the target object in the image to be segmented;
[0083] In which, the image segmentation network includes multiple network stages connected in sequence and a segmentation head connected to the last network stage of the multiple network stages, the multiple network stages are used to extract feature maps of the image to be segmented, the input feature map of the latter network stage in every two adjacent network stages of the multiple network stages includes a feature map obtained by upsampling and / or downsampling at least part of the output feature map of the previous network stage, the segmentation head is used to determine the image segmentation result of the image to be segmented based on the feature maps output by the multiple network stages, each network stage of the multiple network stages includes at least one anisotropic convolution block, the anisotropic convolution block includes multiple anisotropic convolution kernels with different directions, and the multiple anisotropic convolution kernels are long strip convolution kernels.
[0084] See also Figure 7 As shown, it is a schematic block diagram of an electronic device according to an embodiment of the present invention. On the other hand, an electronic device is also provided, and the electronic device 700 includes: a processor 710 and a memory 720, wherein the memory 720 stores computer program instructions, and the computer program instructions are used by the processor 710 to execute the above-mentioned image segmentation method when running.
[0085] Exemplarily, the electronic device may be, for example, an ultrasound system or an endoscope system.
[0086] According to another aspect of the present invention, a storage medium is also provided, on which program instructions are stored. When the program instructions are executed by a computer or processor, the computer or processor executes the corresponding steps of the above-mentioned image segmentation method according to the embodiment of the present invention, and is used to implement the corresponding modules in the above-mentioned image segmentation device according to the embodiment of the present invention or the corresponding modules in the above-mentioned image segmentation device. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above-mentioned storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0087] According to yet another aspect of the present invention, a computer program product is provided, comprising computer program instructions, which are used to execute the above-mentioned image segmentation method when run.
[0088] A person skilled in the art can understand the specific implementation and beneficial effects of the above-mentioned image segmentation device, electronic device, storage medium and computer program product by reading the above-mentioned detailed description of the image segmentation method. For the sake of brevity, they will not be repeated here.
[0089] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present invention. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as claimed in the appended claims.
[0090] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0091] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not performed.
[0092] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0093] Similarly, it should be understood that in order to streamline the present invention and aid in understanding one or more of the various inventive aspects, in the description of exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach to the present invention should not be interpreted as reflecting the intention that the claimed invention requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present invention.
[0094] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0095] Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of the present invention and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0096] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules in the image segmentation apparatus according to an embodiment of the present invention. The present invention can also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0097] It should be noted that the above embodiments illustrate rather than limit the invention, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0098] The above is merely a description of specific embodiments of the present invention, and the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention are intended to be covered by the scope of protection of the present invention. The scope of protection of the present invention shall be based on the scope of protection of the claims.
Claims
1. An image segmentation method, characterized in that: For segmenting an object with a linear structure in a medical image, the method comprises: Acquire an image to be segmented that is collected for a target object, wherein the target object has a linear structure; Inputting the image to be segmented into an image segmentation network to obtain an image segmentation result of the image to be segmented, wherein the image segmentation result is used to indicate a predicted position of the target object in the image to be segmented; In which, the image segmentation network includes a plurality of network stages connected in sequence and a segmentation head connected to the last network stage of the plurality of network stages, the plurality of network stages are used to extract the feature map of the image to be segmented, the input feature map of the latter network stage in every two adjacent network stages of the plurality of network stages includes a feature map obtained by upsampling and / or downsampling at least part of the output feature map of the previous network stage, the segmentation head is used to determine the image segmentation result of the image to be segmented based on the feature maps output by the plurality of network stages, each network stage of the plurality of network stages includes at least one anisotropic convolution block, the anisotropic convolution block includes a plurality of anisotropic convolution kernels with different directions, and the plurality of anisotropic convolution kernels are long strip convolution kernels.
2. The method according to claim 1, characterized in that Inputting the image to be segmented into an image segmentation network to obtain an image segmentation result of the image to be segmented includes: The following operations are performed in each of the anisotropic convolution blocks of the image segmentation network: Using multiple anisotropic convolution kernels in the anisotropic convolution block to perform convolution operations on the input feature map of the anisotropic convolution block, so as to obtain convolution feature maps corresponding to the multiple anisotropic convolution kernels one by one; The convolution feature maps corresponding to the multiple anisotropic convolution kernels are feature fused to obtain the output feature map of the anisotropic convolution block.
3. The method according to claim 2, characterized in that The step of performing feature fusion on the convolution feature maps corresponding one-to-one to the multiple anisotropic convolution kernels to obtain the output feature map of the anisotropic convolution block includes: According to the target weights corresponding to the multiple anisotropic convolution kernels, the convolution feature maps corresponding to the multiple anisotropic convolution kernels are weighted summed, and the summation result is input into a preset activation function to obtain the output feature map of the anisotropic convolution block output by the preset activation function.
4. The method according to claim 2, characterized in that The step of using the multiple anisotropic convolution kernels in the anisotropic convolution block to perform convolution operations on the input feature maps input into the anisotropic convolution block includes: For each anisotropic convolution kernel in the plurality of anisotropic convolution kernels, Determining initial sampling point positions of the anisotropic convolution kernel for sampling points of the input feature map; Adding a target offset to the initial sampling point position to obtain a new sampling point position; Sampling is performed on the input feature map according to the new sampling point position, and a convolution operation is performed on the eigenvalues obtained by sampling and the anisotropic convolution kernel to obtain a convolution feature map corresponding to the anisotropic convolution kernel.
5. The method according to any one of claims 1 to 4, characterized in that The target object is the recurrent laryngeal nerve, and the aspect ratio of the anisotropic convolution kernel is greater than a preset value; and / or, The directions of the multiple anisotropic convolution kernels include directions corresponding to at least two angles from 0° to 180°, wherein 0° and 180° represent the horizontal direction of the input feature map, and 90° represents the vertical direction of the input feature map.
6. The method according to any one of claims 1 to 4, characterized in that The image segmentation network is trained as follows: Acquire training data, the training data including a training image and a labeled segmentation result, the training image including the target object, the labeled segmentation result being used to indicate a labeled position of the target object in the training image; Inputting the training image into the image segmentation network to obtain a predicted segmentation result of the training image, wherein the predicted segmentation result is used to indicate a predicted position of the target object in the training image; The parameters of the image segmentation network are adjusted based on the labeled segmentation result and the predicted segmentation result until the image segmentation network meets the preset requirements.
7. The method according to claim 6, characterized in that The adjusting the parameters of the image segmentation network based on the labeled segmentation result and the predicted segmentation result includes: Calculating a loss value based on the labeled segmentation result and the predicted segmentation result, the loss value including a topology break perception loss value and / or a gradient direction consistency loss value; Adjusting parameters of the image segmentation network based on the loss value; In which, the labeled segmentation result includes the position of the labeled connected domain in the training image, and the position of the labeled connected domain in the training image represents the labeled position of the target object in the training image; the predicted segmentation result includes the position of the predicted connected domain in the training image, and the position of the predicted connected domain in the training image represents the predicted position of the target object in the training image; the topology fracture perception loss value is determined by the ratio between the intersection area and the total area, and the topology fracture perception loss value is negatively correlated with the intersection area, the intersection area is the area of the intersection area between the predicted connected domain and the labeled connected domain matching the predicted connected domain, and the total area is the sum of the area of the predicted connected domain and the area of the corresponding matching labeled connected domain; the gradient direction consistency loss value is determined by the cosine value between the gradient map corresponding to the predicted segmentation result and the gradient map corresponding to the labeled segmentation result, and the gradient direction consistency loss value is negatively correlated with the cosine value.
8. The method according to claim 7, characterized in that When the number of the predicted connected components is multiple, before calculating the loss value based on the labeled segmentation result and the predicted segmentation result, the method further includes: For each predicted connected domain of the plurality of predicted connected domains, matching the predicted connected domain with the labeled connected domain indicated by the labeled segmentation result; The labeled connected domain with the largest overlapping area with the predicted connected domain is the labeled connected domain that matches the predicted connected domain; or the labeled connected domain with the largest intersection-and-union ratio with the predicted connected domain is the labeled connected domain that matches the predicted connected domain.
9. The method according to any one of claims 1 to 4, characterized in that After inputting the image to be segmented into the image segmentation network, the method further includes: The image segmentation result of the image to be segmented is output to a display interface of a display device for display.
10. An electronic device comprising a processor and a memory, characterized in that: The memory stores computer program instructions, which are used by the processor to execute the image segmentation method according to any one of claims 1 to 9 when the processor is running the computer program instructions.
11. A storage medium having program instructions stored thereon, characterized in that: The program instructions are used to execute the image segmentation method according to any one of claims 1 to 9 when running.
12. A computer program product comprising computer program instructions, characterized in that The computer program instructions are used to execute the image segmentation method according to any one of claims 1 to 9 when run.
Citation Information
Patent Citations
Anisotropic convolution-based image classification method and system
CN111126494A
Image segmentation method based on deep learning and anisotropic active contour
CN114581392A
Model training method and device based on recurrent laryngeal nerves and recurrent laryngeal nerve recognition method and device
CN117392475A
Paravascular anomaly segmentation method based on double decoders and local feature enhancement network
CN118172369A
3D Anisotropic Hybrid Network: Transferring Convolutional Features from 2D Images to 3D Anisotropic Volumes
US20190130562A1