Image Processing Method and System Based on Convolutional Neural Network
By using multiple convolutional layers and spatial attention maps to modify the coordinate information in the image processing method of CNN, the deviation and variance problems of prediction results in ultrasonic image segmentation are solved, and higher segmentation accuracy and stability are achieved, while reducing the calculation cost.
Patent Information
- Application Number
- CN202180102421.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-14
AI Technical Summary
The existing CNN-based image processing method has deviations and variance of prediction results in ultrasonic image segmentation, and the calculation cost is high, making it difficult to obtain stable prediction results.
Multiple convolutional layers are used to perform feature extraction operations, and coordinate information is modified through spatial attention maps to generate weighted coordinate maps to improve the output feature map quality of the convolutional layer.
It improves the accuracy and stability of image segmentation, reduces the deviation and variance of segmentation results, and reduces the calculation cost.
Smart Images

Figure CN118043858B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to methods and systems for image processing based on a convolutional neural network (CNN). Background Art
[0002] Convolutional neural networks (CNNs) are a well-known class of artificial neural networks in the art and have been applied to various fields for prediction purposes, particularly image processing for various prediction applications such as image segmentation and image classification. Although CNNs can generally be understood to be applicable to various fields of various prediction applications, using CNNs in various prediction applications may not always provide satisfactory prediction results (e.g., not accurate enough in image segmentation or image classification), and obtaining satisfactory prediction results can be difficult or challenging.
[0003] For example, medical ultrasound imaging is a safe and non-invasive real-time imaging modality that uses high-frequency sound waves to provide images of human structures. Compared with other imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI), ultrasound imaging is relatively inexpensive, portable, and more widespread, and is thus widely regarded as the stethoscope of the 21st century. However, ultrasound images can be obtained from a handheld probe, so they are operator-dependent and susceptible to a large number of artifacts such as severe speckle noise, shadows, and blurred boundaries. This increases the difficulty of segmenting the tissue structure of interest (e.g., anatomical structure) from adjacent tissues. Many traditional methods (e.g., active contours, graph cuts, superpixels, and deep models (e.g., fully convolutional network (FCN), U-Net, etc.)) have been proposed and applied to ultrasound image segmentation. However, due to the noise characteristics of ultrasound images, such traditional methods usually produce poor results. Although deep models have made great progress compared with traditional methods, accurately segmenting soft tissue structures from ultrasound images remains a challenging task.
[0004] Another problem associated with the segmentation of ultrasound images using a single deep model is that they usually produce results with high bias due to blurred boundaries and textures, and high variance due to noise and inhomogeneity. To reduce bias and variance, multi-model integration methods such as bagging algorithms, boosting methods, etc. have been proposed. However, the computational cost of training multiple models for integration is expensive. To solve this problem, it has previously been proposed to train the model once while saving multiple sets of model weights on the optimization path through learning rate annealing. However, this method still requires multiple runs of the inference process. To solve this problem, many multi-stage prediction-refinement deep models (e.g., Hourglass Net, CU-Net, R 3-Net, BASNet), to predict and progressively refine the segmentation results through their cascaded modules. Although such a strategy may be able to reduce segmentation bias, its impact on variance is limited, which means that their average performance across the entire dataset may seem good, but they are unlikely to produce stable predictions for different input images.
[0005] Accordingly, there is a need to provide a method and system for CNN-based image processing that seeks to overcome or at least improve one or more problems associated with conventional methods and systems for CNN-based image processing, particularly to enhance or improve the prediction ability (e.g., the accuracy of prediction results) associated with CNN-based image processing (such as but not limited to image segmentation). The present invention has arisen in this context. Summary of the Invention
[0006] According to a first aspect of the present invention, there is provided a method for CNN-based image processing using at least one processor, the method comprising:
[0007] Receiving an input image;
[0008] Performing a plurality of feature extraction operations using a plurality of convolutional layers of a CNN based on the input image to respectively generate a plurality of output feature maps; and
[0009] Generating an output image of the input image based on the plurality of output feature maps of the plurality of convolutional layers,
[0010] wherein, for each of the plurality of feature extraction operations, performing the feature extraction operation using a convolutional layer comprises:
[0011] Generating an output feature map of the convolutional layer based on the input feature map received by the convolutional layer and a plurality of weighted coordinate maps;
[0012] Generating a plurality of weighted coordinate maps based on the plurality of coordinate maps and a spatial attention map; and generating a spatial attention map based on the input feature map received by the convolutional layer for modifying the coordinate information of each of the plurality of coordinate maps to generate a plurality of weighted coordinate maps.
[0013] According to a second aspect of the present invention, there is provided a system for CNN-based image processing, the system comprising: a memory; and at least one processor communicatively coupled to the memory, the processor being configured to execute the CNN-based image processing method according to the first aspect of the present invention.
[0014] According to a third aspect of the present invention, there is provided a computer program product embodied in one or more non-transitory computer-readable storage media, the computer program product comprising executable instructions that can be executed by at least one processor to perform the CNN-based image processing method according to the first aspect of the present invention.
[0015] According to a fourth aspect of the present invention, there is provided a method for using a CNN to segment tissue structures in an ultrasound image, the method using at least one processor, the method comprising:
[0016] performing the CNN-based image processing method according to the above first aspect of the present invention, wherein
[0017] the input image is an ultrasound image including tissue structures; and
[0018] the output image has the segmented tissue structures and is the result of inferring the input image using a CNN.
[0019] According to a fifth aspect of the present invention, there is provided a CNN-based image processing system, the system comprising: a memory; and at least one processor communicatively coupled to the memory, the processor being configured to perform the method for using a CNN to segment tissue structures in an ultrasound image according to the above fourth aspect of the present invention.
[0020] According to a sixth aspect of the present invention, there is provided a computer program product embodied in one or more non-transitory computer-readable storage media, the computer program product including executable instructions that can be executed by at least one processor to perform the method for using a CNN to segment tissue structures in an ultrasound image according to the above fourth aspect of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Embodiments of the present invention will be better understood and become apparent to those of ordinary skill in the art from the following written description, which is given by way of example only and in conjunction with the accompanying drawings, in which:
[0022] Figure 1 A schematic flow chart of a CNN-based image processing method according to various embodiments of the present invention is depicted;
[0023] Figure 2 A schematic block diagram of a CNN-based image processing system according to various embodiments of the present invention is depicted;
[0024] Figure 3 A schematic block diagram of an exemplary computer system according to various embodiments of the present invention, which exemplary computer system can be used to implement or implement a CNN-based image processing system is depicted;
[0025] Figure 4A and 4B An exemplary network architecture of an example CNN according to various example embodiments of the present invention is depicted;
[0026] Figure 5A table (Table 1) showing various exemplary embodiments of the present invention is presented, and Table 1 shows an exemplary detailed configuration of the prediction module and the refinement module of an exemplary CNN;
[0027] Figure 6 A schematic block diagram of a residual U-shaped block (RSU) according to various exemplary embodiments of the present invention is depicted;
[0028] Figure 7A and 7B A schematic block diagram of a residual block ( Figure 7A ) and an RSU ( Figure 7B ) according to various exemplary embodiments is depicted;
[0029] Figure 8A and 8B A schematic block diagram of a conventional coordinate convolution (Coord Conv) ( Figure 8A ) and an attention coordinate convolution (AC-Conv) ( Figure 8B ) according to various exemplary embodiments of the present invention is depicted;
[0030] Figure 9A and 9B A schematic block diagram of a conventional cascaded refinement module and a parallel refinement module according to various exemplary embodiments of the present invention is depicted;
[0031] Figure 10 A schematic diagram of a thyroid and ultrasound scanning protocol according to various exemplary embodiments of the present invention, and a corresponding ultrasound image with manually marked thyroid lobe coverage are depicted;
[0032] Figure 11 A table (Table 2) showing the volume and the number of corresponding slices (images) in each subset of ultrasound images according to various exemplary embodiments of the present invention is depicted;
[0033] Figure 12 A table (Table 3) showing a quantitative evaluation or comparison of an exemplary CNN with other existing technology segmentation models on transverse (TRX) and sagittal (SAG) test sets according to various exemplary embodiments of the present invention is depicted;
[0034] Figures 13A to 13L Sample segmentation results of TRX thyroid images using an exemplary CNN according to various exemplary embodiments of the present invention are shown;
[0035] Figures 14A to 14L Sample segmentation results of SAG thyroid images using an exemplary CNN according to various exemplary embodiments of the present invention are shown;
[0036] Figure 15A and Figure 15BSuccess rate curves of an example CNN and other prior art models according to various example embodiments of the present invention on TRX images and SAG images are respectively shown; and
[0037] Figure 16 A table (Table 4) depicting an ablation study conducted on different convolutional blocks and refinement architectures is depicted. Detailed implementation manners
[0038] Various embodiments of the present invention provide a method and system for image processing based on a convolutional neural network (CNN), particularly a method and system for image processing based on a deep CNN. A CNN is a type or kind of artificial neural network, which can also be referred to as a CNN model, or simply as a model. For example, as described in the background art, although a CNN can generally be understood to be applicable to various fields of various prediction applications, using a CNN in various prediction applications may not always provide satisfactory prediction results (for example, not accurate enough in image segmentation or image classification), and it may be difficult or challenging to obtain satisfactory prediction results. For example, ultrasound images including tissue structures (such as anatomical structures or other types of tissue structures, such as tumors) are noisy, and it has been found that traditional methods for segmenting such ultrasound images based on a CNN produce poor results. Therefore, various embodiments of the present invention provide a method and system for image processing based on a CNN, which seek to overcome or at least improve one or more problems associated with traditional methods and systems for image processing based on a CNN, particularly to enhance or improve the prediction ability (such as the accuracy of prediction results) associated with image processing based on a CNN (such as but not limited to image segmentation).
[0039] Figure 1 A schematic flowchart of a method 100 for image processing based on a CNN using at least one processor according to various embodiments of the present invention is depicted. The method 100 includes: receiving (at 102) an input image; performing (at 104) a plurality of feature extraction operations using a plurality of convolutional layers of the CNN based on the input image to respectively generate a plurality of output feature maps; and generating (at 106) an output image of the input image based on the plurality of output feature maps of the plurality of convolutional layers. In particular, for each of the plurality of feature extraction operations, performing the feature extraction operation using a convolutional layer includes: generating an output feature map of the convolutional layer based on the input feature map received by the convolutional layer and a plurality of weighted coordinate maps; generating a plurality of weighted coordinate maps based on the plurality of coordinate maps and a spatial attention map; and generating a spatial attention map based on the input feature map received by the convolutional layer for modifying the coordinate information of each of the plurality of coordinate maps to generate a plurality of weighted coordinate maps.
[0040] Accordingly, it has been advantageously found that the method 100 of image processing enhances or improves the prediction ability, particularly with respect to image segmentation, and more specifically, with respect to ultrasonic image segmentation. In particular, by performing the feature extraction operation using the corresponding convolutional layer in the above-described manner, not only can the relevant convolutional operation access the coordinate information (by using the coordinate map (extra coordinate channels)), but the associated convolutional operation can pay more attention (i.e., increased attention) to certain coordinates that may be beneficial to the feature extraction operation (by using the spatial attention map, which may also be simply referred to as the attention map). Thus, this increased attention (i.e., increased attention) is guided by the input feature map received by the convolutional layer through the spatial attention map derived from the input feature map. Accordingly, the associated convolutional operation not only knows its spatial position (e.g., in the Cartesian space), but also knows where to pay more attention through the spatial attention map. For example, through the spatial attention map, additional weights can be added to certain coordinates that may require more attention or focus, and the weights can be reduced to certain coordinates that may require less attention or focus, as guided by the input feature map (e.g., the more important parts of the input feature map can thus receive more attention in the feature extraction operation), resulting in the associated convolutional operation of the convolutional layer advantageously having attention coordinate guidance. Thus, this feature extraction operation using a convolutional layer with attention coordinate guidance can be referred to as attention coordinate-guided convolution (AC-Conv), and such a convolutional layer with attention coordinate guidance can be referred to as an AC-Conv layer. In this regard, through attention coordinate guidance, it has been advantageously found that the method 100 of image processing enhances or improves the prediction ability. As the method 100 of image processing and the corresponding system for image processing are described in more detail according to various embodiments and example embodiments of the present invention, these advantages or technical effects and / or other advantages or technical effects will become more apparent to those skilled in the art.
[0041] In various embodiments, the above generation of the spatial attention map includes: performing a first convolutional operation based on the input feature map received by the convolutional layer to generate a convolutional feature map; and applying an activation function based on the convolutional feature map to generate the spatial attention map.
[0042] In various embodiments, the activation function is a sigmoid activation function.
[0043] In various embodiments, the above generation of the plurality of weighted coordinate maps includes multiplying each of the plurality of coordinate maps by the spatial attention map so as to modify the coordinate information of each of the plurality of coordinate maps.
[0044] In various embodiments, a plurality of coordinate maps includes a first coordinate map and a second coordinate map. The first coordinate map includes coordinate information about a first dimension, and the second coordinate map includes coordinate information about a second dimension. The first dimension and the second dimension are two dimensions on which a first convolution operation is configured to be performed.
[0045] In various embodiments, generating the output feature map of the convolutional layer described above includes: concatenating the input feature map received by the convolutional layer and a plurality of weighted coordinate maps channel by channel to form a concatenated feature map; and performing a second convolution operation based on the concatenated feature map to generate the output feature map of the convolutional layer.
[0046] In various embodiments, the CNN includes a prediction sub-network, and the prediction sub-network includes at least one convolutional layer among a plurality of convolutional layers of the CNN. In this regard, method 100 further includes generating a set of predicted feature maps using the prediction sub-network based on an input image. Generating the set of predicted feature maps includes performing at least one feature extraction operation among a plurality of feature extraction operations using at least one convolutional layer of the prediction sub-network. In addition, a plurality of predicted feature maps in the set of predicted feature maps have different spatial resolution levels.
[0047] In various embodiments, the prediction sub-network has an encoder-decoder structure including a set of encoder blocks and a set of decoder blocks. The set of encoder blocks of the prediction sub-network includes a plurality of encoder blocks, and the set of decoder blocks of the prediction sub-network includes a plurality of decoder blocks. In this regard, method 100 further includes: for each of the plurality of encoder blocks of the prediction sub-network, generating a downsampled feature map using the encoder block based on the input feature map received by the encoder block; and for each of the plurality of decoder blocks of the prediction sub-network, generating an upsampled feature map using the decoder block based on the input feature map and the downsampled feature map generated by the encoder block corresponding to the decoder block received by the decoder block.
[0048] In various embodiments, generating a set of predicted feature maps using the prediction sub-network described above includes generating a plurality of predicted feature maps respectively based on a plurality of upsampled feature maps generated by a plurality of decoder blocks.
[0049] In various embodiments, generating a downsampled feature map using an encoder block of the prediction sub-network described above includes: extracting multi-scale features based on the input feature map received by the encoder block; and generating a downsampled feature map based on the multi-scale features extracted by the encoder block. In various embodiments, generating an upsampled feature map using a decoder block of the prediction sub-network described above includes: extracting multi-scale features based on the input feature map and the downsampled feature map generated by the encoder block corresponding to the decoder block received by the decoder block; and generating an upsampled feature map based on the multi-scale features extracted by the decoder block.
[0050] In various embodiments, each of the plurality of encoder blocks of the prediction sub-network includes at least one convolutional layer of the plurality of convolutional layers of the CNN, and generating the downsampled feature map using the encoder block of the prediction sub-network includes performing at least one feature extraction operation among the plurality of feature extraction operations using at least one convolutional layer of the encoder block. In various embodiments, each of the plurality of decoder blocks of the prediction sub-network includes at least one convolutional layer of the plurality of convolutional layers of the CNN, and generating the upsampled feature map using the decoder block of the prediction sub-network includes performing at least one feature extraction operation among the plurality of feature extraction operations using at least one convolutional layer of the decoder block.
[0051] In various embodiments, each convolutional layer of each of the plurality of encoder blocks of the prediction sub-network is one of the plurality of convolutional layers of the CNN. In various embodiments, each convolutional layer of each of the plurality of decoder blocks of the prediction sub-network is one of the plurality of convolutional layers of the CNN.
[0052] In various embodiments, each of the plurality of encoder blocks of the prediction sub-network is configured as a residual block. In various embodiments, each of the plurality of decoder blocks of the prediction sub-network is configured as a residual block.
[0053] In various embodiments, the CNN further includes a refinement sub-network, which includes at least one convolutional layer of the plurality of convolutional layers of the CNN. In this regard, method 100 further includes generating a set of refined feature maps using the refinement sub-network based on the fused feature map, and generating the set of refined feature maps includes performing at least one feature extraction operation among the plurality of feature extraction operations using at least one convolutional layer of the refinement sub-network. In addition, the plurality of refined feature maps of the set of refined feature maps have different spatial resolution levels.
[0054] In various embodiments, method 100 further includes connecting a set of predicted feature maps to generate a fused feature map.
[0055] In various embodiments, the refinement sub-network includes a plurality of refinement blocks, which are configured to generate a plurality of refined feature maps respectively, and each of the plurality of refinement blocks has an encoder-decoder structure, and the encoder-decoder structure includes a set of encoder blocks and a set of decoder blocks. The set of encoder blocks of the refinement sub-network includes a plurality of encoder blocks, and the set of decoder blocks of the refinement sub-network includes a plurality of decoder blocks. In this regard, method 100 further includes, for each of the plurality of refinement blocks: for each of the plurality of encoder blocks of the refinement block, generating a downsampled feature map using the encoder block based on the input feature map received by the encoder block; and for each of the plurality of decoder blocks of the refinement block, generating an upsampled feature map using the decoder block based on the input feature map and the downsampled feature map generated by the encoder block corresponding to the decoder block received by the decoder block.
[0056] In various embodiments, the multiple encoder-decoder structures of the multiple refinement blocks have different heights.
[0057] In various embodiments, the above-mentioned use of the refinement sub-network to generate a set of refined feature maps includes: for each of the multiple refinement blocks, generating the refined feature map of the refinement block based on the fused feature map received by the refinement block and the upsampled feature map generated by the first decoder block among the multiple decoder blocks of the refinement block.
[0058] In various embodiments, the above-mentioned use of the encoder block of the refinement block to generate the downsampled feature map includes: extracting multi-scale features based on the input feature map received by the encoder block; and generating the downsampled feature map based on the multi-scale features extracted by the encoder block. In various embodiments, the above-mentioned use of the decoder block of the refinement block to generate the upsampled feature map includes: extracting multi-scale features based on the input feature map and the downsampled feature map generated by the encoder block of the refinement block corresponding to the decoder block received by the decoder block; and generating the upsampled feature map based on the multi-scale features extracted by the decoder block.
[0059] In various embodiments, for each of the multiple refinement blocks: each of the multiple encoder blocks of the refinement block includes at least one convolutional layer among the multiple convolutional layers of the CNN, and the above-mentioned use of the encoder block of the refinement block to generate the downsampled feature map includes performing at least one feature extraction operation among the multiple feature extraction operations using at least one convolutional layer of the encoder block. In various embodiments, for each of the multiple refinement blocks: each of the multiple decoder blocks of the refinement block includes at least one convolutional layer among the multiple convolutional layers of the CNN, and the above-mentioned use of the decoder block of the refinement block to generate the upsampled feature map includes performing at least one feature extraction operation among the multiple feature extraction operations using at least one convolutional layer of the decoder block.
[0060] In various embodiments, each convolutional layer of each of the multiple encoder blocks of the refinement block is one of the multiple convolutional layers of the CNN. In various embodiments, each convolutional layer of each decoder block of the multiple decoder blocks of the refinement block is one of the multiple convolutional layers of the CNN.
[0061] In various embodiments, for each of the multiple refinement blocks, each of the multiple encoder blocks of the refinement block is configured as a residual block, and each of the multiple decoder blocks of the refinement block is configured as a residual block.
[0062] In various embodiments, the output image is generated based on a set of refined feature maps.
[0063] In various embodiments, the output image is generated based on the average value of a set of refined feature maps.
[0064] In various embodiments, the receiving (at 102) of the input image includes receiving a plurality of input images, each of the plurality of input images being a labeled image for training a CNN to obtain a trained CNN. In this regard, for each of the plurality of input images: a plurality of feature extraction operations are performed using a plurality of convolutional layers of the CNN based on the input image, respectively, to generate a plurality of output feature maps, respectively; and an output image of the input image is generated based on the plurality of output feature maps of the plurality of convolutional layers.
[0065] In various embodiments, the labeled image is a labeled ultrasound image including an organizational structure.
[0066] In various embodiments, the output image is the result of inferring the input image using the CNN.
[0067] In various embodiments, the input image is an ultrasound image including an organizational structure.
[0068] Figure 2 A schematic block diagram of a system 200 for CNN-based image processing according to various embodiments of the present invention is depicted, corresponding to the method 100 of image processing as described above with reference to Figure 1 the method 100 described above. The system 200 includes: a memory 202; and at least one processor 204 communicatively coupled to the memory 202, and the processor is configured to execute the method 100 of image processing described herein according to various embodiments of the present invention. Thus, in various embodiments, the at least one processor 204 is configured to: receive an input image; perform a plurality of feature extraction operations using a plurality of convolutional layers of the CNN based on the input image, respectively, to generate a plurality of output feature maps, respectively; and generate an output image of the input image based on the plurality of output feature maps of the plurality of convolutional layers. In particular, as described above, for each of the plurality of feature extraction operations, performing the feature extraction operation using the convolutional layer includes: generating an output feature map of the convolutional layer based on the input feature map received by the convolutional layer and a plurality of weighted coordinate maps; generating a plurality of weighted coordinate maps based on the plurality of coordinate maps and a spatial attention map; and generating a spatial attention map based on the input feature map received by the convolutional layer for modifying the coordinate information of each of the plurality of coordinate maps to generate a plurality of weighted coordinate maps.
[0069] Those skilled in the art will understand that the at least one processor 204 can be configured to perform various functions or operations through an instruction set (e.g., software module) of various functions or operations executable by the at least one processor 204. Thus, as Figure 2As shown, the system 200 may include an input image receiving module (or input image receiving circuit) 206 configured to receive an input image; a feature extraction module (or feature extraction circuit) 208 configured to perform a plurality of feature extraction operations using a plurality of convolutional layers of a CNN respectively on the input image to generate a plurality of output feature maps respectively; and an output image generating module (or output image generating circuit) 210 configured to generate an output image of the input image based on the plurality of output feature maps of the plurality of convolutional layers.
[0070] Those skilled in the art will understand that the above modules are not necessarily separate modules, and one or more modules may be implemented or implemented by a functional module (such as a circuit or a software program) as needed or appropriately without departing from the scope of the present invention. For example, two or more of the input image receiving module 206, the feature extraction module 208, and the output image generating module 210 may be implemented (e.g., compiled together) as an executable software program (e.g., a software application or simply referred to as an "app"), which may be stored in the memory 202 and may be executed by at least one processor 204 to perform various functions / operations described herein according to various embodiments of the present invention.
[0071] In various embodiments, the image processing system 200 corresponds to the image processing method 100 as described above according to various embodiments. Therefore, the various functions or operations configured to be executed by at least one processor 204 may correspond to the various steps or operations of the image processing method 100 as described above according to various embodiments, and thus, for the sake of clarity and conciseness, there is no need to repeat them with respect to the image processing system 200. In other words, the various embodiments described in the context of the method are similarly effective for the corresponding system, and vice versa. Figure 1
[0072] For example, in various embodiments, the memory 202 may store therein the input image receiving module 206, the feature extraction module 208, and / or the output image generating module 210, which respectively correspond to the various steps (or operations or functions) of the image processing method 100 as described herein according to various embodiments, and which may be executed by at least one processor 204 to perform the corresponding functions or operations described herein.
[0073] In various embodiments, according to various embodiments of the present invention, a method for segmenting tissue structures in an ultrasound image using a CNN is provided, which uses at least one processor. The method includes: performing the CNN-based image processing method 100 as described above according to various embodiments, whereby the input image is an ultrasound image including tissue structures; and the output image having the segmented tissue structures is the result of inferring the input image using a CNN.
[0074] In various embodiments, the CNN is trained according to various embodiments as described above. That is to say, the CNN is the trained CNN as described above.
[0075] In various embodiments, according to various embodiments, a system for using a CNN to segment tissue structures in an ultrasound image is provided, corresponding to the method for segmenting tissue structures in an ultrasound image according to various embodiments as described above. The system includes: a memory; and at least one processor communicatively coupled to the memory and configured to execute the method for segmenting tissue structures in an ultrasound image as described above. In various embodiments, the system for segmenting tissue structures in an ultrasound image may be the same as the system 200 for image processing, whereby the input image is an ultrasound image including tissue structures; the output image has the segmented tissue structures and is the result of inferring the input image using the CNN.
[0076] According to various embodiments of the present disclosure, a computing system, a controller, a microcontroller, or any other system providing processing capabilities may be provided. Such a system may include one or more processors and one or more computer-readable storage media. For example, the system 200 for image processing described above may include a processor (or controller) 204 and a computer-readable storage medium (or memory) 202, which are used, for example, to perform various processes as described herein. The memory or computer-readable storage medium used in various embodiments may be a volatile memory, such as DRAM (Dynamic Random Access Memory), or a non-volatile memory, such as PROM (Programmable Read-Only Memory), EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or flash memory, such as floating-gate memory, charge-trapping memory, MRAM (Magnetoresistive Random Access Memory), or PCRAM (Phase-Change Random Access Memory).
[0077] In various embodiments, "circuit" may be understood as any type of logical implementation entity, which may be a dedicated circuit or a processor executing software stored in a memory, firmware, or any combination thereof. Therefore, in one embodiment, "circuit" may be a hardwired logic circuit or a programmable logic circuit, such as a programmable processor, such as a microprocessor (e.g., a Complex Instruction Set Computer (CISC) processor or a Reduced Instruction Set Computer (RISC) processor). "Circuit" may also be a processor executing software, such as any kind of computer program, such as a computer program using virtual machine code (e.g., Java). According to various embodiments, any other kind of implementation of the corresponding function may also be understood as "circuit". Similarly, "module" may be a part of a system according to various embodiments and may include the "circuit" as described above, or may be understood as any kind of logically implemented entity.
[0078] Certain portions of the present disclosure are presented explicitly or implicitly in the form of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and functional or symbolic representations are the means by which those skilled in the data processing arts most effectively convey the substance of their work to others skilled in the art. An algorithm is generally conceived herein as a self-consistent sequence of steps leading to a desired result. These steps require physical manipulation of physical quantities, such as electrical, magnetic, or optical signals capable of being stored, transferred, combined, compared, and otherwise manipulated.
[0079] Unless otherwise specifically stated and as will be apparent from the following, throughout this specification, descriptions or discussions using terms such as "receive", "execute", "generate", "multiply", "connect", "extract", etc. refer to the actions and processes of a computer system or similar electronic device that manipulate and transform data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or other information storage, transmission, or display device.
[0080] This specification also discloses a system (e.g., which may also be embodied as an apparatus or device), such as an image processing system 200 for performing various operations / functions described herein. Such a system may be specifically constructed for the desired purpose or may include a general-purpose computer or other devices selectively activated or reconfigured by a computer program stored in the computer. The algorithms given herein are not inherently related to any particular computer or other device. According to the teachings herein, various general-purpose machines may be used with a computer program. Alternatively, it may be appropriate to construct more specialized devices to perform the various method steps.
[0081] Furthermore, this specification also at least implicitly discloses a computer program or software / functional modules, since it will be apparent to those skilled in the art that the various steps of the methods described herein can be implemented by computer code. The computer program is not intended to be limited to any particular programming language and its implementation. It should be understood that the teachings of the disclosure contained herein can be implemented using a variety of programming languages and their encodings. Furthermore, the computer program is not limited to any particular control flow. Without departing from the scope of the present invention, there are many other variations of the computer program, which can use different control flows. Those skilled in the art will realize that the various modules described herein (e.g., the input image receiving module 206, the feature extraction module 208, and / or the output image generating module 210) can be software modules implemented by a computer program or instruction set executable by a computer processor to perform the desired functions, or can be hardware modules of functional hardware units designed to perform the desired functions. It will also be understood that they can be implemented as a combination of hardware and software modules.
[0082] In addition, one or more steps of the computer programs / modules or methods described herein can be executed in parallel rather than sequentially. Such computer programs can be stored on any computer-readable medium. The computer-readable medium can include, for example, a magnetic or optical disk, a memory chip, or other storage devices suitable for interfacing with a general-purpose computer. When the computer program is loaded and executed on such a general-purpose computer, the computer program effectively results in an apparatus for implementing the method steps described herein.
[0083] In various embodiments, there is provided a computer program product embodied in one or more computer-readable storage media (non-transitory computer-readable storage media), the computer program product including instructions (e.g., input image receiving module 206, feature extraction module 208, and / or output image generation module 210) executable by one or more computer processors to perform a method 100 of image processing, as described herein with reference to Figure 1 According to various embodiments. Thus, the various computer programs or modules described herein can be stored in a computer program product receivable by a system, such as Figure 2 The system 200 for image processing shown, the computer programs or modules being executed by at least one processor 204 of the system 200 to perform various functions.
[0084] In various embodiments, there is provided a computer program product embodied in one or more computer-readable storage media (non-transitory computer-readable storage media), the computer program product including instructions executable by one or more computer processors to perform the method of segmenting tissue structures in an ultrasound image as described above according to various embodiments. Thus, the various computer programs or modules described herein can be stored in a computer program product receivable by a system (such as the system for segmenting tissue structures in an ultrasound image described above), and executed by at least one processor of the system to perform various functions.
[0085] The software or functional modules described herein can also be implemented as hardware modules. More specifically, in the hardware sense, a module is a functional hardware unit designed for other components or modules. For example, a module can be implemented using discrete electronic components, or it can form part of an entire electronic circuit such as an application-specific integrated circuit (ASIC). There are many other possibilities. Those skilled in the art will understand that the software or functional modules described herein can also be implemented as a combination of hardware and software modules.
[0086] In various embodiments, the system 200 for image processing can be implemented by any computer system (e.g., a desktop or portable computer system) including at least one processor and a memory, such as by way of example and not limitation Figure 3The computer system 300 schematically shown therein. Various methods / steps or functional modules may be implemented as software, such as a computer program executed within the computer system 300, and direct the computer system 300 (in particular, one or more of its processors) to perform the various functions or operations described herein according to various embodiments. The computer system 300 may include a computer module 302, an input module such as a keyboard and / or touch screen 304 and a mouse 306, and a plurality of output devices such as a display 308 and a printer 310. The computer module 302 may be connected to a computer network 312 via a suitable transceiver device 314 to allow access to, for example, the Internet or other network systems, such as a local area network (LAN) or a wide area network (WAN). The computer module 302 in this example may include a processor 318 for executing various instructions, a random access memory (RAM) 320, and a read only memory (ROM) 322. The computer module 302 may also include a plurality of input / output (I / O) interfaces, such as an I / O interface 324 to the display 308 and an I / O interface 326 to the keyboard 304. The components of the computer module 302 typically communicate via an interconnect bus 328 and in a manner known to those skilled in the relevant art.
[0087] Those skilled in the art should understand that the terms used herein are only for describing various embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0088] Any element or feature referred to herein by the names "first", "second", etc. does not limit the quantity or order of such element or feature unless otherwise stated or the context otherwise requires. For example, such naming may be used herein as a convenient way to distinguish between two or more elements or instances of an element. Thus, referring to a first and a second element does not necessarily mean that only two elements can be used, or that the first element must precede the second element. In addition, the phrase "at least one" in reference to a list of items refers to any single item in the list or any combination of two or more items in the list.
[0089] For the purpose of facilitating an easy understanding of the present invention and putting it into practice, various exemplary embodiments of the present invention will be described hereinafter by way of example only and not by way of limitation. However, those skilled in the art should understand that the present invention can be implemented in various different forms or configurations and should not be construed as being limited to the exemplary embodiments set forth hereinafter. On the contrary, these exemplary embodiments are provided to make the present disclosure thorough and complete and to fully convey the scope of the present invention to those skilled in the art.
[0090] Specifically, for a better understanding of the present invention and without limiting or losing generality, various exemplary embodiments of the present invention will now be described with respect to input images of ultrasound images and image processing for ultrasound image segmentation (i.e., a method of image processing based on CNN for segmenting tissue structures in ultrasound images). Although such a specific application (i.e., ultrasound image segmentation) may be preferred according to various exemplary embodiments, those skilled in the art should understand that the present invention is not limited to such a specific application, and the method of image processing can be implemented in other types of applications as needed or appropriately (e.g., for applications where the input image may be relatively noisy and / or the structures of interest generally have similar positions and / or shapes in the input image), such as but not limited to image classification.
[0091] Ultrasound image segmentation is a challenging task due to the presence of artifacts inherent in the modality, such as attenuation, shadowing, speckle noise, non-uniform texture, and blurred boundaries. In this regard, various exemplary embodiments provide a prediction-refinement attention network (which is a CNN) for segmenting soft tissue structures in ultrasound images, which may be referred to herein as ACU 2 E-Net or simply referred to as this CNN or model hereinafter. The prediction-refinement attention network includes: a prediction module or block (e.g., corresponding to the above-mentioned prediction sub-network according to various embodiments and may be referred to herein as ACU 2E-Net), which includes attention coordinate convolution (AC-Conv); and a multi-head residual refinement module or block (e.g., corresponding to the above-mentioned refinement sub-network according to various embodiments and may be referred to as MH-RRM or E-module herein), which includes multiple (e.g., three) parallel residual refinement modules or blocks (e.g., corresponding to the above-mentioned multiple refinement blocks according to various embodiments). In various example embodiments, AC-Conv is configured or designed to improve segmentation accuracy by perceiving the shape and position information of the target anatomical structure. By integrating residual refinement and integration strategies, it has been advantageously found that MH-RRM reduces segmentation bias and variance and avoids the multi-pass training and inference common in integration methods. According to various example embodiments, to demonstrate the effectiveness of the method for image processing based on this CNN in segmenting tissue structures in ultrasound images, a dataset of thyroid ultrasound scans was collected, and this CNN was evaluated against prior art segmentation methods. Comparison with prior art models shows that the performance of this CNN on transverse and sagittal thyroid images is competitive or improved. For example, ablation studies show that the AC-Conv and MH-RRM modules increase the segmentation Dice coefficient of the baseline model from 79.62% to 80.97% and 83.92%, while reducing the variance from 6.12% to 4.67% and 3.21%.
[0092] As described in the background art, ultrasonic images can be obtained from a handheld probe and thus are operator-dependent and affected by a large number of artifacts, such as severe speckle noise, shadows, and blurred boundaries. This increases the difficulty of segmenting the tissue structure of interest (e.g., anatomical structure) from adjacent tissues. Many traditional methods (e.g., active contours, graph cuts, superpixels, and deep models (e.g., fully convolutional network (FCN), U-Net, etc.)) have been proposed and applied to ultrasonic image segmentation. However, due to the noise characteristics of ultrasonic images, such traditional methods usually produce poor results. Although deep models have made great progress compared to traditional methods, accurately segmenting soft tissue structures from ultrasonic images remains a challenging task.
[0093] Regarding ultrasonic image segmentation, various example embodiments note that, different from general objects with different shapes and positions in natural image segmentation, the tissue structures (e.g., anatomical structures) in ultrasonic images have similar position and shape patterns. However, these geometric features are rarely used for segmenting deep models because they are difficult to represent and encode. Therefore, traditionally, how to utilize the specific geometric constraints of soft tissue structures in ultrasonic images remains a challenge. Another problem associated with the use of a single deep model to segment ultrasonic images is that due to blurred boundaries and textures, they usually produce results with high bias, and due to noise and inhomogeneity, they produce high variance.
[0094] Therefore, to overcome these challenges, various example embodiments provide the above-described attention-based prediction-refinement architecture (i.e., this CNN), which includes a prediction module and a multi-head residual refinement module (MH-RRM) constructed based on the above AC-Conv. This attention-based prediction-refinement architecture advantageously utilizes the anatomical location and shape constraints present in ultrasound images to reduce the bias and variance of the segmentation results while avoiding multi-pass training and inference. Thus, the contributions of this CNN include: (a) AC-Conv is configured to improve segmentation accuracy by perceiving geometric information (e.g., shape and location information) from ultrasound images; and / or (b) the prediction-refinement architecture with MH-RRM improves segmentation accuracy by integrating an integration strategy and a prediction-refinement strategy. As will be described below, the method of image processing for ultrasound image segmentation based on this CNN according to various example embodiments is tested on a dataset of thyroid ultrasound scans and achieves improved performance (e.g., accuracy) relative to traditional models.
[0095] CNN architecture
[0096] Figure 4A and 4B together depict an example network architecture of an example CNN 400 according to various example embodiments of the present invention. Also as described above, the example CNN 400 includes: a prediction module or block (ACU 2 -Net) 410( Figure 4A ) and an MH-RRM 450( Figure 4B ). In various example embodiments, the prediction module 410 can be configured based on the U 2 -Net disclosed in Qin et al., "U 2 -Net: Going Deeper with nested U-structure for salient object detection, Pattern Recognition", 106:107404, 2020 (which is referred to herein as the Qin reference, the content of which is incorporated herein by reference in its entirety for various purposes), by replacing each ordinary convolutional layer in the U 2 -Net with an AC-Conv layer described herein according to various example embodiments to form an attention coordinate-guided U 2 -Net (which can be referred to as ACU 2 -Net). In various example embodiments, the refinement module 450 includes a set of parallel variants of the prediction module (ACU 2 -Net) (e.g., in order to produce refined feature maps with different levels of spatial resolution). As an example, as Figure 4BAs shown, the refinement module 450 can be configured to have three refinement heads or blocks arranged in parallel (which are three ACUs for generating refinement feature maps with different spatial resolution levels 2 -Net variants) 454-1, 454-2, 454-3. These three refinement heads or blocks are respectively represented as ACU Figure 4B in 2 -Net-Ref7, ACU 2 -Net-Ref5 and ACU 2 -Net-Ref3 (for example, it looks like the character "E". Therefore, this example configuration of the refinement module 450 can also be referred to as the E module in this article). In the Figure 4A shown legend, the term AC-CBR represents AC-conv + BatchNorm + ReLU.
[0097] For illustration purposes only and not by way of limitation, Figure 5 a table (Table 1) is shown. Table 1 shows an example detailed configuration of the prediction module 410 and the refinement module 450 of the example CNN400 according to various example embodiments. The blank cells in Table 1 indicate that there is no such stage. In addition, "I", "M", and "O" represent the number of input channels (C in ), intermediate channels, and output channels (C out ) of each AC-RSU block (attention coordinate-guided residual U-shaped block). "En_i" and "De_j" represent the encoder and decoder levels respectively. The number "L" in "AC-RSU-L" represents the height of the AC-RSU block. Those skilled in the art will understand that the present invention is not limited to a CNN having the Figure 5 shown example detailed configuration (or parameters). This example detailed configuration (or parameters) is provided only for illustration purposes and not by way of limitation. Those skilled in the art will understand that the parameters of the CNN can vary or be modified as needed or suitable for various purposes, such as but not limited to the desired height of the encoder-decoder structure of the ACU 2 -Net, the desired different spatial resolution levels (and / or the desired number of different spatial resolution levels) of the generated prediction feature maps, the desired different spatial resolution levels (and / or the desired number of different spatial resolution levels) of the generated refinement feature maps, the desired number of layers in the encoder or decoder blocks, the desired number of channels in the encoder or decoder blocks, and so on.
[0098] The Qin reference discloses a deep network architecture (referred to as U 2 -Net) for salient object detection (SOD). U 2The network architecture of -Net is a two-level nested U-shaped structure. This network architecture has the following advantages: (1) Due to the mixture of receptive fields of different sizes in the residual U-shaped block (RSU block, which can be simply referred to as RSU), it can capture more context information from different scales; and (2) Due to the pooling operations used in these RSU blocks, it increases the depth of the entire architecture without significantly increasing the computational cost. This network architecture can train a deep network from scratch without using the backbone from the image classification task. In particular, U 2 -Net is a two-level nested U-shaped structure, which is designed for SOD without using any pre-trained backbone from image classification. It can be trained from scratch to achieve competitive performance. In addition, the network architecture allows the network to be deeper and obtain high resolution without significantly increasing memory and computational cost. This is achieved through the nested U-shaped structure, where RSU blocks that can extract multi-scale features within the level without reducing the resolution of the feature map are configured at the bottom layer; at the top layer, there is a U-Net-like structure (encoder-decoder structure), where each level is filled with an RSU block. The two-level configuration results in a nested U-shaped structure, and as Figure 4A shown in the example of the nested U-shaped structure (encoder-decoder structure) according to various exemplary embodiments, so as described above, U 2 each ordinary convolutional layer in -Net is replaced by the AC-Conv layer described herein according to various exemplary embodiments, thereby forming ACU 2 -Net410.
[0099] In summary, the multi-level depth feature integration method mainly focuses on developing better multi-level feature aggregation strategies. On the other hand, the methods in the category of multi-scale feature extraction aim to design new modules for extracting local and global information from the features obtained from the backbone network. In this regard, U 2 -Net or ACU 2 -Net410's network architecture is configured to directly extract multi-scale features level by level.
[0100] Residual U-shaped block (RSU) / Attention Coordinate Guided Residual U-shaped block (AC-RSU)
[0101] Local and global context information is important for salient object detection and other segmentation tasks. In modern CNN designs such as VGG, ResNet, DenseNet, etc., small convolutional filters of size 1×1 or 3×3 are the most commonly used components for feature extraction. They are popular because they require less storage space and are computationally efficient. For example, the output feature maps of the shallow layers only contain local features because the receptive fields of 1×1 or 3×3 filters are too small to capture global information. To obtain more global information from the shallow layers on high-resolution feature maps, the most straightforward method is to enlarge the receptive field. However, performing multiple dilated convolutions (especially in the early stages) on the input feature maps with the original resolution requires too much computational and memory resources. To reduce the computational cost, a parallel configuration can adopt a pyramid pooling module (PPM) that uses small kernel filters on the downsampled feature maps instead of dilated convolutions on the original-sized feature maps. But fusing features of different scales by direct upsampling and concatenation (or summation) may lead to the degradation of high-resolution features.
[0102] Therefore, as described in the Qin reference, an RSU block is provided to capture intra-level multi-scale features. By way of example only and not limitation, Figure 6 an example structure of the RSU-L(C in , M, C out ) block 600 is shown, where L is the number of layers in the encoder, C in , C out represent the input and output channels, and M represents the number of channels in the internal layers of the RSU block 600. Those skilled in the art will understand that the RSU-L block 600 is not limited to Figure 6 the specific dimensions (such as the number of layers L) shown, which are by way of example only and not limitation. Thus, the RSU block 600 includes three components:
[0103] (i) An input convolutional layer that converts the input feature map x(H×W×C in ) into an intermediate map with C out channels This is a normal convolutional layer for local feature extraction;
[0104] (ii) A U-Net-like symmetric encoder-decoder structure of height L that takes the intermediate feature map as input and learns to extract and encode multi-scale context information representing Figure 6The U-Net structure shown. A larger L results in deeper Residual U-shaped blocks (RSUs), more pooling operations, a larger receptive field, and richer local and global features. Configuring this parameter allows for the extraction of multi-scale features from input feature maps with arbitrary spatial resolution. Multi-scale features are extracted from the gradually downsampled feature maps and encoded into high-resolution feature maps through progressive upsampling, concatenation, and convolution. This process alleviates the loss of fine details caused by direct upsampling at large scales.
[0105] (iii) Residual connections that fuse local and multi-scale features through summation:
[0106] For better understanding, Figure 7A and 7B depicts the original Residual block 700 for comparison ( Figure 7A ) and the Residual U-shaped block (RSU) 720 ( Figure 7B ). The operations in the original Residual block 700 can be summarized as where, represents the desired mapping of the input feature x; represents the weight layer, which is a convolution operation in this setting. The main design difference between the RSU block 720 and the original Residual block 700 is that the RSU block 720 replaces the ordinary, single-stream convolution with a U-Net-like structure 600 and replaces the original features with local features transformed by the weight layer: where, represents the multi-layer U-shaped structure 600 as shown in Figure 6 . This difference between the RSU block 720 and the original Residual block 700 enables the network to directly extract features from multiple scales from each RSU block. Additionally, the computational overhead due to the U-shaped structure is small because most operations are applied to the downsampled feature maps.
[0107] In various exemplary embodiments, the AC-RSU block can be formed based on (e.g., the same as or similar to the above block 720) the above RSU block 720 (not limited to any specific size, such as the number of layers L, which can vary or be modified as needed or appropriately, whereby in various exemplary embodiments, each ordinary convolutional layer in the RSU block 720 is replaced with the AC-Conv layer described herein.
[0108] ACU 2 -Net architecture
[0109] According to various exemplary embodiments, an ACU n-Net, in which multiple U-Net-like structures are stacked in a nested manner. In particular, the exponential notation refers to nested U-shaped structures rather than cascaded stacks. In theory, the exponent n can be set to any positive integer to achieve single-stage or multi-stage nested U-shaped structures. However, architectures with too many nested levels will be too complex to implement and use in practical applications. For example, n can be set to 2 to form ACU 2 -Net. ACU 2 -Net has a two-stage nested U-shaped structure, Figure 4A depicts an example ACU that forms the prediction module 410 according to various example embodiments 2 -Net's schematic block diagram. The top layer is a U-shaped structure including multiple levels ( Figure 4A multiple cubes in). For example but not limited to 14 levels. Each level is filled with a configured AC-RSU block (bottom U-shaped structure). Therefore, the nested U-shaped structure can more effectively extract multi-scale features within the level and aggregate multi-level features between levels.
[0110] As Figure 4A shown, the prediction module (ACU 2 -Net) 410 has an encoder-decoder structure, which includes a set of encoder blocks 420 and a set of decoder blocks 430. By way of example only and not limitation, the prediction module 410 includes three parts: (1) a multi-level (e.g., seven-level) encoder structure 420; (2) a multi-level (e.g., seven-level) decoder structure 430; and (3) a feature map fusion module or block 440 coupled or connected to the decoder stage 430.
[0111] For the encoder stage 420, an example configuration of a set of encoder blocks 420 is as shown in Figure 5 Table 1 in. For the decoder stage 430, an example configuration of a set of decoder blocks 430 is also as shown in Figure 5 Table 1 in. As described above, "7", "6", "5", and "4" represent the height (L) of the AC-RSU block. For example, L can be configured according to the spatial resolution of the input feature map. For feature maps with larger height and width, a larger L can be used to obtain more large-scale information. For example, the resolution of the feature maps in En_6 and En_7 is relatively low, and further downsampling these feature maps will result in the loss of useful context. Therefore, in the En_6 and En_7 levels, AC-RSU-4F is used, where "F" indicates that the AC-RSU block is an extended version, in which, for example, pooling and upsampling operations are replaced by extended convolutions. In this case, all intermediate feature maps of AC-RSU-4F have the same resolution as their input feature maps.
[0112] For decoder stage 430, an example configuration of a set of decoder blocks (AC-RSU) is also as shown in Table 1 of Figure 5 . In various example embodiments, decoder stage 430 may have a similar or corresponding structure to its symmetric or corresponding encoder stage 420. For example, an extended version of AC-RSU-4F is also used for decoder blocks De_6 and De_7, which is similar to or corresponds to that used for symmetric or corresponding encoder blocks En_6 and En_7. As shown in Figure 4A , each decoder stage may be configured to take as input the concatenation of the upsampled feature map from its previous stage and the downsampled feature map from its symmetric or corresponding encoder stage.
[0113] In various example embodiments, prediction module 410 may be configured to generate a plurality of predicted feature maps based on the upsampled feature map generated by decoder stage 430. By way of example only and not limitation, in the example configuration shown in Figure 4A , seven predicted feature maps (e.g., side output saliency probability output maps) from decoder stages De_1, De_2, De_3, De_4, De_5, De_6, De_7 may be generated respectively based on a 3×3 convolutional layer and a sigmoid function Then, prediction module 410 may upsample the logarithm of the side output saliency map (the convolutional output before the sigmoid function) to the input image size, fuse it through a concatenation operation, and then generate a fused feature map (e.g., final saliency probability map) through a 1×1 convolutional layer and a sigmoid function
[0114] Thus, the configuration of ACU 2 -Net allows for a deep architecture that has rich multi-scale features and relatively low computational and storage costs. Additionally, in various example embodiments, since the ACU 2 -Net architecture is built on AC-RSU blocks without using any pre-trained backbones adapted from image classification, it is flexible and easy to adapt to different working environments with little performance loss.
[0115] Thus, in various example embodiments, prediction module 410 has an encoder-decoder structure that includes a set of encoder blocks (e.g., En_1 to En_7) 420 and a set of decoder blocks (e.g., De_1 to De_7) 430. As shown in Figure 4A , for each of the plurality of encoder blocks (e.g., En_1 to En_5) of the set of encoder blocks, based on the input feature map received by the encoder block, the encoder block may generate a downsampled feature map. Additionally, as shown in Figure 4AAs shown, for each of the multiple decoder blocks (e.g., De_1 to De_5) of the group of decoder blocks, an upsampled feature map is generated using the decoder block based on the input feature map and the downsampled feature map generated by the encoder block corresponding to the decoder block received by the decoder block.
[0116] In various example embodiments, multiple predicted feature maps are generated based on the multiple upsampled feature maps generated by the multiple decoder blocks, respectively.
[0117] Attention Coordinate Convolution (AC-Conv)
[0118] Various example embodiments note that soft tissue structures such as the thyroid gland in medical images seem to have predictable position and shape patterns, which can be used to assist the segmentation process. The Coordinate Convolution (CoordConv) as shown in Figure 8A has been disclosed to solve the coordinate transformation problem (see Liu et al., "An intriguing failing of convolution neural networks and the CoordConv solution", In NIPS, 9605-9616, 2018, referred to herein as the Liu reference, the content of which is incorporated herein by reference in its entirety for all purposes). In particular, Figure 8A shows a schematic block diagram of the original CoordConv layer 800. Specifically, given an input feature map M in (h×w×c)804, CoordConv can be described as M out =conv(cat(M in ,M i ,M j )),where M i 806 and M j 808 represent the row and column coordinate maps, respectively. However, various example embodiments of the present invention note that since the coordinate maps (M i ,M j ) connected to the features in different layers are almost constant, directly connecting them to the feature map M in in different layers may reduce the generalization ability of the network. This is because their corresponding convolution weights are responsible for synchronizing their value scales with the value scale of the feature map M in and extracting geometric information. To solve this problem, various example embodiments provide the Attention Coordinate Convolution (AC-Conv) 850 as shown in Figure 8B . Figure 8BDepicts a schematic block diagram of the AC-Conv850 according to various example embodiments of the present invention. The AC-Conv850 adds a spatial attention-like operation before connecting (channel-wise) the input feature map 854 and the coordinate maps 856', 858' (corresponding to the multiple weighted coordinate maps as described above according to various embodiments):
[0119] M out = conv(cat(M in , σ(conv(M in )) · cat(M i , M j )) (Equation 1)
[0120] where σ is the sigmoid function.
[0121] Thus, in various example embodiments, performing the feature extraction operation using the convolutional (AC-Conv) layer 850 includes: generating the output feature map 870 of the convolutional layer 850 based on the input feature map 854 and the multiple weighted coordinate maps 856', 858' received by the convolutional layer 850; generating the multiple weighted coordinate maps 856', 858' based on the multiple coordinate maps 856, 858 and the spatial attention map 860; and generating the spatial attention map 860 based on the input feature map 854 received by the convolutional layer 850 to modify the coordinate information of each of the multiple coordinate maps 856, 858 to generate the multiple weighted coordinate maps 856', 858'. In various example embodiments, generating the spatial attention map 860 includes performing a first convolutional operation 862 based on the input feature map 854 received by the convolutional layer 850 to generate a convolutional feature map; and applying an activation function 864 to the convolutional feature map to generate the spatial attention map 860. In various example embodiments, generating the multiple weighted coordinate maps 856', 858' includes multiplying each of the multiple coordinate maps 856, 858 by the spatial attention map 860 so as to modify the coordinate information of each of the multiple coordinate maps 856, 858. In various example embodiments, generating the output feature map 870 of the convolutional layer 850 includes: connecting (channel-wise) the input feature map 854 received by the convolutional layer 850 and the multiple weighted coordinate maps 856', 858' to form a connected feature map 866; and performing a second convolutional operation 868 based on the connected feature map 866 to generate the output feature map 870 of the convolutional layer 850.
[0122] The spatial attention-like operation plays two roles: i) as a synchronization layer to reduce M in and {M i , M j} scale difference between; ii) re-weighting the coordinates of each pixel, instead of using a constant coordinate map, to capture more important geometric information under the guidance of the attention map 860 derived from the current input feature map 854. For example, for two coordinates i, j, an i-coordinate map (or i-coordinate channel) 856 and a j-coordinate map (or j-coordinate channel) 858 may be provided. For example, the i-coordinate map 856 may be an h×w rank-1 matrix, the first row of which is filled with zeros (0), the second row is filled with ones (1), the third row is filled with twos (2), and so on. The j-coordinate map 858 may be the same or similar to the i-coordinate map 856, but with the above values filling the columns instead of the rows. As described above according to various example embodiments, the U-coordinate maps may be modified or adapted according to various example embodiments by replacing their convolutional layers with AC-Conv layers 850. 2 -Net to generate or construct an AC-RSU according to various example embodiments. For example, compared to RSU 720, AC-RSU is able to extract texture and geometric features from different receptive fields. In various example embodiments, the prediction module ACU 2 -Net410 and three sub-network ACUs in the refined E module 450 2 -Net-Ref7, ACU 2 -Net-Ref5 and ACU 2 -Net-Ref3 is built on AC-RSU.
[0123] Parallel Multi-Head Residual Refinement Module (MH-RRM)
[0124] To further improve accuracy, many traditional prediction-refinement models have been designed through cascaded sub-networks (cascaded refinement modules): c =F p (X), Recursively or incrementally refine the rough result, e.g. Figure 9A As shown. In theory, the final output is the most accurate and is therefore usually regarded as the final result. This cascaded refinement strategy can reduce the deviation of the segmentation results. However, various exemplary embodiments have found that in practice, due to low image quality and blurred boundaries, the use of such a network to segment soft tissue in ultrasound images often has large differences. Multi-model integration strategies can be used to reduce prediction bias and variance. However, various exemplary embodiments have found that the direct integration of multiple deep models requires a lot of computational and time costs. In order to solve these problems associated with traditional technologies, various exemplary embodiments embed the integration strategy into the refinement module. In particular, according to various exemplary embodiments of the present invention, there is provided such Figure 4BThe simple and effective parallel multi - head residual refinement module (MH - RRM) 450 shown. By way of example only and not limitation, the number of MH - RRM heads 454 - 1, 454 - 2, 454 - 3 according to various example embodiments (e.g., corresponding to the multiple refinement blocks described above according to various embodiments) is set to three. As Figure 4B shown. As described above, the three refinement heads or blocks 454 - 1, 454 - 2, 454 - 3 can each be formed based on the ACU 2 -Net, and the ACU 2 -Net is configured to generate refined feature maps with different levels of spatial resolution based on the fused feature map 444. In various example embodiments, the multiple refinement blocks 454 - 1, 454 - 2, 454 - 3 respectively generate multiple refined feature maps 464 - 1, 464 - 2, 464 - 3. Thus, in various example embodiments, the multiple refined feature maps 464 - 1, 464 - 2, 464 - 3 have different levels of spatial resolution.
[0125] In various example embodiments, each of the multiple refinement blocks 454 - 1, 454 - 2, 454 - 3 has an encoder - decoder structure, and the encoder - decoder structure includes multiple encoder blocks and multiple decoder blocks. For each refinement block and each of the multiple encoder blocks of the refinement block, as Figure 4B shown, the encoder block can generate a downsampled feature map based on the input feature map received by the encoder block. Additionally, for each refinement block and each of the multiple decoder blocks of the refinement block, as Figure 4B shown, based on the input feature map and the downsampled feature map generated by the encoder block corresponding to the decoder block received by the decoder block, the decoder block generates an upsampled feature map. In various example embodiments, the multiple encoder - decoder structures of the multiple subdivision blocks have different heights.
[0126] In various example embodiments, as Figure 4B shown, for each refinement block, the refined feature map of the refinement block can be generated based on the fused feature map 444 received by the refinement block and the upsampled feature map generated by the first decoder blocks 458 - 1, 458 - 2, 458 - 3 of the multiple decoder blocks of the refinement block. In various example embodiments, the output image of the example CNN400 is generated based on the average value of a set of refined feature maps 464 - 1, 464 - 2, 464 - 3.
[0127] Thus, in various example embodiments, given an input image X, the final segmentation result of the example CNN400 can be expressed as:
[0128]
[0129] Figure 9B shows the semantic workflow of the prediction-refinement architecture of an example CNN400 with the above parallel refinement modules. In Figure 9A and 9B the bold indicates the final prediction result.
[0130] Training and inference
[0131] During the training process, the three refinement outputs R (1) 464-1, R (2) 464-2 and R (3) 464-3 of the E module 450, together with the seven side outputs S (i) (i = {1, 2, 3, 4, 5, 6, 7}) and one fused output S fuse 444 from the prediction module 410, are monitored through independently computed losses, as Figure 4A and 4B shown. The entire model can be trained end-to-end with binary cross-entropy (BCE) loss:
[0132]
[0133] where, is the total loss, and are the losses corresponding to the side output, the fused output, and the refinement output. and are their corresponding weights to emphasize different outputs. In experiments performed according to various example embodiments, all the λ weights are set to 1.0. During the inference process, the average of R (1) 464-1, R (2) 464-2 and R (3) 464-3 is taken as the final prediction result (e.g., corresponding to the output image of the CNN as described above according to various embodiments).
[0134] Experiments
[0135] The thyroid is a butterfly-shaped organ located at the base of the neck just above the collarbone, with the left and right lobes connected by a narrow band of tissue called the isthmus in the middle (see Figure 10 ). In particular, Figure 10 depicts a schematic diagram of the thyroid and an ultrasound scanning protocol, as well as the corresponding ultrasound image with manually marked thyroid lobe coverage 1010. Figure 10 The dashed arrows in the top row of the image in Figure 10The bottom row of the images in [Figure 0] shows sample TRX (left) and SAG (right) images with manually marked thyroid lobe coverage 1010.
[0136] To diagnose thyroid abnormalities, clinicians can assess the size of the thyroid by manually segmenting it from the collected ultrasound scans. Only for illustrative purposes and not by way of limitation, the example CNN 400 is evaluated on the thyroid tissue segmentation problem as a case study.
[0137] Dataset
[0138] It seems that none of the existing public datasets are suitable for large-scale learning-based methods. To enable large-scale clinical applications, a comprehensive thyroid ultrasound segmentation dataset was collected with the approval of the health research ethics committees of the participating centers.
[0139] Regarding ultrasound scan collection, 777 ultrasound scans were retrospectively collected from 700 patients aged between 18 and 82 years who had thyroid ultrasound examinations at 12 different imaging centers. The scans were divided according to the scan direction of the ultrasound probe in the transverse (TRX) and sagittal (SAG) planes (e.g., see Figure 10 ). Thus, two partitions (the TRX set and the SAG set) are available. Based on the patient ID, each partition was randomly divided into three subsets for training, validation, and testing, so that no same patient would appear in two different subsets. Figure 11 [Table 2] is depicted, and Table 2 shows the volume and the number of corresponding slices (images) in each subset. In particular, Table 2 shows the number of TRX and SAG thyroid scans in the thyroid dataset, where "Volume #" and "Slice #" represent the volume and the number of corresponding marked images, respectively.
[0140] Regarding annotation or marking, the images in the dataset were manually marked by five experienced sonographers and verified by three radiologists. Given the relatively large total number of available images, every three or five slices of the ultrasound scans in the training set were marked to save marking time. However, for accurate volume assessment, the validation and test sets were marked slice by slice.
[0141] Regarding the implementation details, the example CNN400 is implemented using PyTorch. Specified training, validation, and test sets are used to evaluate the performance of the example CNN400. During training, the input images are first resized to 160×160×3 and then randomly cropped to 144×144×3. Online random horizontal and vertical flips are used to augment the dataset. The training batch size is set to 12. The model weights are initialized by default He uniform initialization (e.g., see He et al., "Delving deep into rectifiers: Surpassing human-level performance on imagenet classification", In Proceedings of the IEEE international conference on computer vision, 1026 - 1034, 2015). The learning rate of the Adam optimizer (e.g., see Kingma, "Adam: A method for stochastic optimization", arXiv preprint arXiv:1412.6980, 2014") is 1e-3 and there is no weight decay. The training loss converges after approximately 50,000 iterations, which takes about 24 hours. During testing, the input images are resized to 160×160×3 and fed into the example CNN. Bilinear interpolation is used for the downsampling and upsampling processes. Both the training and testing processes are carried out on a 12-core, 24-thread PC equipped with an AMD Ryzen Threadripper 2920x 4.3GHz CPU (128GB RAM) and an NVIDIA GTX 1080Ti GPU.
[0142] Regarding the evaluation metrics, two metrics are used to evaluate the overall performance of the method: Volumetric Dice (e.g., see Popovic et al., "Statistical validation metric for accuracy assessment in medical image segmentation", IJCARS, 2(2 - 4):169 - 181, 2007) and its standard deviation σ. The Dice coefficient is defined as:
[0143]
[0144] where P and G represent the predicted segmentation mask scan (h×w×c) and the ground truth mask scan (h×w×c), respectively. The standard deviation of the Dice coefficient is calculated as follows:
[0145]
[0146] where N is the number of test volumes, Dice μ represents the average volume Dice coefficient of the entire test set. In the experiments conducted, the average dice (Dice) and the standard deviation (σ) of each test set were reported.
[0147] Example CNN (ACU 2 E-Net) 400 was compared with 11 state-of-the-art (SOTA) models. The 11 state-of-the-art (SOTA) models included U-Net (Ronneberger et al., "U-net: Convolutional networks for biomedical image segmentation", In MICCAI, 234 - 241, 2015) and its five variants (including Res U-Net (e.g., see Xiao et al., "Weighted Res-UNet for high-quality retina vessel segmentation", In ITME, 327 - 331, 2018), Dense U-Net (e.g., see Guan et al., "Fully Dense U Net for 2-D Sparse Photoacoustic Tomography Artifact Removal", IEEE JBHI, 24(2): 568 - 576, 2019), Attention U-Net (e.g., see Oktay et al., "Attention u-net: Learning where to look for the pancreas", arXiv preprint arXiv:1804:03999, 2018), U-Net++ (e.g., see Zhou et al., "Unet++: Anested u-net architecture for medical image segmentation", In MICCAI-W, 3 - 11, 2018) and U 2 -Net (e.g., see Qin et al., "U 2-Net: Going Deeper with nested U-structure for salient object detection”, Pattern Recognition, 106:107404, 2020)), and five prediction-refinement models, including the stacked hourglass network (e.g., see Newell et al., “Stacked hourglass networks for human pose estimation”, In ECCV, 483-499, 2016), SRM (e.g., see Wang et al., “A stagewise refinement model for detecting salient objects in images”, In ICCV, 4019-4028, 2017), C-U-Net (e.g., see Tang et al., “Quantized densely connected u-nets for efficient landmark localization”, In ECCV, 339-354, 2018), R 3 -Net (Deng et al., “R3net: Recurrent residual refinement network for saliency detection”, In AAAI, 2018) and BASNet (Qin et al., “Basnet: Boundary-aware salient object detection”, In CVPR, 7479-7489, 2019).
[0148] Figure 12 depicts a table (Table 3), which shows a quantitative evaluation or comparison of the example CNN400 with other prior art segmentation models on the TRX and SAG test sets. The top of Table 3 includes a comparison with the classic U-Net and its variants (such as Attention U-Net), while the bottom of Table 3 shows a comparison with models involving a prediction-refinement strategy (such as R 3 -Net). It can be observed that the example CNN400 produces the highest DICE coefficient on both TRX and SAG images. Additionally, compared with the second-best model (BASNet) and other refinement module designs such as R 3 -Net, the parallel refinement module 450 increases the Dice coefficient by 2.55% and 1.22% respectively, and reduces the standard deviation σ by 31.99% and 7.51% respectively.
[0149] Figures 13A to 13L And FIGS. 14A through 14L show sample segmentation results on the TRX and SAG thyroid images. In particular, Figures 13A to 13L depicts a qualitative comparison of the ground truth results (dashed white lines) and the segmentation results (solid white lines) of different methods on a sampled TRX slice with a homogeneous thyroid, and Figures 14A to 14L depicts a qualitative comparison of the ground truth results (dashed white lines) and the segmentation results (solid white lines) of different methods on a sampled SAG slice with a heterogeneous thyroid. It can be seen that the exemplary CNN400 is able to produce improved (more accurate) segmentation results. Specifically, Figures 13A to 13L shows a homogeneous TRX thyroid lobe with severe flash noise and blurred boundaries. Res U-Net, U-Net++, SRM, C-U-Net, R 3 -Net, and BASNet failed to capture the accurate boundaries. Other models, such as U-Net, Dense U-Net, Attention U-Net, U 2 -Net, and the Stacked Hourglass Net were all unable to segment the slender area in the upper left of the thyroid. Figures 14A to 14L shows the segmentation results of a heterogeneous SAG view thyroid containing several complex nodules. Thus, it can be seen that the exemplary CNN400 produces relatively better results than other models.
[0150] To further evaluate the robustness of the exemplary CNN400, the success rate curves of the exemplary CNN400 and 11 other prior art models on the TRX image and the SAG image are plotted in Figure 15A and 15B respectively. The success rate is defined as the ratio of the number of scan predictions (coefficient higher than a specific dice threshold) to the total number of scans. A higher success rate means better performance, so the top curve (ACU 2 E-Net) is better than the 11 other prior art models being compared. Thus, it can be seen that the exemplary CNN400 significantly outperforms other models on both the TRX and SAG test sets.
[0151] To verify the effectiveness of the AC-Conv according to various exemplary embodiments, by replacing the adapted U 2- Ablation studies were conducted using ordinary convolutions (ordinary Conv) in - Net (LeCun et al., "Gradient-based learning applied to document recognition", Proceedings of IEEE, 86(11):2278-2324, 1998): SE-Conv (Hu et al., "Squeeze-and-excitation networks", In CVPR, 7132-7141, 2018), which explicitly models the interdependence of channels through its squeeze-and-excitation blocks, CBAM-Conv (Woo et al., "Cbam: Convolutional block attention module", In ECCV, 3-19, 2018), which refines the feature maps through its channel and spatial attention blocks, and CoordConv (Liu et al., "An intriguing failing of convolutional neural networks and the CoordConv solution", In NIPS, 9605-9616, 2018), which provides access to the input coordinates of the convolution itself by using coordinate channels and our AC-Conv. Figure 16 Figure shows a table (Table 4) depicting the ablation studies conducted on different convolutional blocks and refinement architectures. In Table 4, Ref7 is the abbreviation of ACU 2 - Net-Ref7. The experiments were conducted on the TRX thyroid test set. The results on the TRX test set are shown at the top of Table 4. It can be seen that ACU 2 - Net using AC-Conv gives the best results in terms of Dice coefficient and standard deviation. This further demonstrates that the combined strategy of jointly perceiving geometric and spatial information is more effective than individual methods based on spatial attention (CBAM) or coordinates (CoordConv).
[0152] To verify the performance of the MH-RRM (E-module), ablation studies were also conducted on different refinement configurations, including cascaded RRM Ref3(Ref5(Ref7))), parallel RRMs with three identical RRM averages (Refk, Refk, Refk) {k = 3, 5, 7}, and fused parallel RRM conv(Ref7, Ref5, Ref3), where the parallel refinement outputs were fused through a convolutional layer instead of averaging during inference. The bottom of Table 4 shows the ablation results on the RRM, which indicate that the cascaded RRM, parallel RRM with identical branches, and fused parallel RRM according to various example embodiments are all inferior to the MH-RRM.
[0153] Accordingly, various example embodiments advantageously provide an attention-based prediction-refinement network (ACU 2 E-Net) 400 for segmenting soft tissue structures in ultrasound images. In particular, the ACU 2 E-Net is based on (a) attention coordinate convolution (AC-conv) 850 and (b) parallel multi-head refinement module (MH-RRM) 450. The attention coordinate convolution 850 makes full use of the geometric information of the thyroid gland in the ultrasound image, and the parallel multi-head refinement module 450 refines the segmentation result by integrating an integration strategy with a residual refinement method.
[0154] Thorough ablation studies and comparisons with the prior art models described above demonstrate the effectiveness and robustness of the example CNN 400 without complicating the training and inference processes. Although the example CNN 400 has been described for the segmentation of thyroid tissue from ultrasound images, it should be understood that the example CNN 400, as well as the AC-Conv 850 and MH-RRM 450, are not limited to the application of segmenting thyroid tissue from ultrasound images, but can be applied as needed or appropriately to segment other types of tissues from ultrasound images, such as but not limited to the liver, spleen, and kidney, as well as tumors (e.g., hepatocellular carcinoma (HCC) in the liver or subcutaneous mass).
[0155] While the embodiments of the present invention have been specifically shown and described with reference to specific embodiments, those skilled in the art should understand that various changes in form and detail may be made without departing from the scope of the present invention as defined by the appended claims. Accordingly, the scope of the present invention is indicated by the appended claims, and all changes within the meaning and equivalent scope of the claims are included.
Claims
1. A method for image processing using at least one processor based on a convolutional neural network (CNN), the method comprises: receiving an input image; based on the input image, performing a plurality of feature extraction operations using a plurality of convolutional layers of the CNN to generate a plurality of output feature maps, wherein corresponding feature extraction operations of the plurality of feature extraction operations are performed by corresponding convolutional layers of the plurality of convolutional layers, and comprise: receiving, by the corresponding convolutional layer, a corresponding input feature map and a plurality of coordinate maps; generating, by the corresponding convolutional layer, a corresponding spatial attention map based on the corresponding input feature map; generating, by the corresponding convolutional layer, a plurality of weighted coordinate maps based on the plurality of coordinate maps and the corresponding spatial attention map; and outputting, by the corresponding convolutional layer, a corresponding output feature map of the corresponding convolutional layer based on the corresponding input feature map and the plurality of weighted coordinate maps; and generating an output image corresponding to the input image based on the plurality of output feature maps of the plurality of convolutional layers.
2. The method according to claim 1, wherein, generating, by the corresponding convolutional layer, the corresponding spatial attention map based on the corresponding input feature map comprises: performing a first convolution operation based on the corresponding input feature map received by the corresponding convolutional layer to generate a corresponding convolved feature map; and applying an activation function to the corresponding convolved feature map to generate the corresponding spatial attention map.
3. The method according to claim 2, wherein, the activation function is a sigmoid activation function.
4. The method according to claim 2 or claim 3, wherein, generating, by the corresponding convolutional layer, the plurality of weighted coordinate maps comprises multiplying each of the plurality of coordinate maps by the corresponding spatial attention map so as to modify the coordinate information of each of the plurality of coordinate maps.
5. The method according to any one of claims 2 to 4, wherein, the plurality of coordinate maps comprises a first coordinate map and a second coordinate map, the first coordinate map comprises coordinate information about a first dimension, the second coordinate map comprises coordinate information about a second dimension, and the first dimension and the second dimension are two dimensions on which the first convolution operation is configured to be performed.
6. The method according to any one of claims 1 to 5, wherein, outputting, by the corresponding convolutional layer, the corresponding output feature map of the corresponding convolutional layer comprises: connecting the corresponding input feature map received by the corresponding convolutional layer and the plurality of weighted coordinate maps channel by channel to form a corresponding connected feature map; and performing a second convolution operation based on the corresponding connected feature map to generate the corresponding output feature map of the corresponding convolutional layer.
7. The method according to any one of claims 1 to 6, wherein: the CNN comprises a prediction sub-network, the prediction sub-network comprises at least one convolutional layer of the plurality of convolutional layers of the CNN; and the method further comprises: generating a set of predicted feature maps using the prediction sub-network based on the input image comprises: Performing at least one of the multiple feature extraction operations using at least one convolutional layer of the prediction sub-network, wherein the set of predicted feature maps includes multiple predicted feature maps with different levels of spatial resolution.
8. The method according to claim 7, wherein: The prediction sub-network has an encoder-decoder structure including a set of encoder blocks and a set of decoder blocks, wherein the set of encoder blocks includes multiple first encoder blocks, the set of decoder blocks includes multiple first decoder blocks, each first encoder block of the multiple first encoder blocks corresponds to a respective first decoder block of the multiple first decoder blocks, and The method further includes: Based on the respective input feature maps received by the respective first encoder blocks, generating, by the respective first encoder of the multiple first encoder blocks, respective downsampled feature maps; and Based on the respective input feature maps and the respective downsampled feature maps generated by the respective first encoder blocks corresponding to the respective first decoder blocks, generating, by the respective first decoder blocks of the multiple first decoder blocks corresponding to the respective first encoder blocks, upsampled feature maps.
9. The method according to claim 8, wherein, Generating the set of predicted feature maps using the prediction sub-network includes generating the multiple predicted feature maps based on the multiple upsampled feature maps generated by the multiple first decoder blocks.
10. The method according to claim 8 or 9, wherein: For a respective first encoder block of the multiple first encoder blocks, generating the respective downsampled feature map includes: Extracting first multi-scale features based on the respective input feature map received by the respective first encoder block; and Generating the respective downsampled feature map based on the extracted first multi-scale features, and For a respective first decoder block of the multiple first decoder blocks, generating the respective upsampled feature map includes: Extracting second multi-scale features based on the respective input feature map and the respective downsampled feature map generated by the respective first encoder block corresponding to the respective first decoder block received by the decoder block; and Generating the respective upsampled feature map based on the multi-scale features extracted by the respective decoder block.
11. The method according to any one of claims 8 to 10, wherein: Each of the multiple first encoder blocks of the prediction sub-network includes at least one convolutional layer of the multiple convolutional layers of the CNN; and Generating the respective downsampled feature map by the respective first encoder block of the multiple first encoder blocks includes: Using at least one convolutional layer in the respective first encoder block to perform at least one of the multiple feature extraction operations, and Each of the multiple first decoder blocks of the prediction sub-network includes at least one convolutional layer of the multiple convolutional layers of the CNN, and Generating the corresponding upsampled feature maps by the respective first decoder blocks of the plurality of first decoder blocks includes: Performing at least one feature extraction operation among the plurality of feature extraction operations using at least one convolutional layer of the respective first decoder block.
12. The method according to claim 11, wherein, each convolutional layer of each of the plurality of first encoder blocks of the prediction sub-network is one of the plurality of convolutional layers of the CNN, and each convolutional layer of each of the plurality of first decoder blocks of the prediction sub-network is one of the plurality of convolutional layers of the CNN.
13. The method according to any one of claims 8 to 12, wherein, each of the plurality of first encoder blocks of the prediction sub-network is configured as a residual block, and each of the plurality of first decoder blocks of the prediction sub-network is configured as a residual block.
14. The method according to any one of claims 7 to 13, wherein, the CNN further includes a refinement sub-network, the refinement sub-network includes at least one convolutional layer of the plurality of convolutional layers of the CNN, the method further includes generating a set of refined feature maps using the refinement sub-network based on the fused feature maps, the generating including: Performing at least one feature extraction operation among the plurality of feature extraction operations using at least one convolutional layer of the refinement sub-network, wherein the set of refined feature maps includes a plurality of refined feature maps having different spatial resolution levels.
15. The method according to claim 14, further including connecting the set of predicted feature maps to generate the fused feature map.
16. The method according to claim 14 or 15, wherein, the refinement sub-network includes a plurality of refinement blocks configured to generate the plurality of refined feature maps, each of the plurality of refinement blocks having an encoder-decoder structure including a set of encoder blocks and a set of decoder blocks, wherein the set of encoder blocks includes a plurality of second encoder blocks, the set of decoder blocks includes a plurality of second decoder blocks, wherein a corresponding second encoder block among the plurality of second encoder blocks corresponds to a corresponding second decoder block among the plurality of decoder blocks, and the method further includes: for each refinement block of the plurality of refinement blocks: Generating corresponding downsampled feature maps by each of the plurality of second encoder blocks using the respective second encoder block based on the input feature map received by the respective second encoder block; and Generating corresponding upsampled feature maps by each of the plurality of second decoder blocks using the respective second decoder block based on the respective input feature map and the corresponding downsampled feature map generated by the respective second encoder block received by the respective second decoder block corresponding to the respective second decoder block.
17. The method according to claim 16, wherein, the plurality of refinement blocks include a plurality of encoder-decoder structures having different heights.
18. The method according to claim 16 or 17, wherein, the plurality of refinement blocks are configured to generate the plurality of refined feature maps by the following operations: For each refinement block of the plurality of refinement blocks, based on the fused feature map received by the corresponding refinement block and the corresponding upsampled feature map generated by the corresponding second decoder block among the plurality of second decoder blocks corresponding to the corresponding refinement block, generate the corresponding refined feature map of the plurality of refined feature maps.
19. The method according to any one of claims 16 to 18, wherein: For each second encoder block of the plurality of second encoder blocks, generating the corresponding downsampled feature map includes: extracting first multi-scale features based on the corresponding input feature map received by the corresponding second encoder block; and generating the corresponding downsampled feature map based on the extracted first multi-scale features extracted by the corresponding second encoder block, and For each second decoder block of the plurality of second decoder blocks, generating the corresponding upsampled feature map includes: extracting second multi-scale features based on the corresponding input feature map and the corresponding downsampled feature map generated by the corresponding second encoder block corresponding to the corresponding second decoder block received by the corresponding second decoder block; and generating the corresponding upsampled feature map based on the extracted multi-scale features extracted by the corresponding decoder block.
20. The method according to any one of claims 16 to 19, wherein, for the corresponding refinement block of the plurality of refinement blocks: each of the plurality of second encoder blocks corresponding to the corresponding refinement block includes at least one convolutional layer of the plurality of convolutional layers of the CNN; and using the corresponding second encoder block of the corresponding refinement block, generating the corresponding downsampled feature map by each second encoder block of the plurality of second encoder blocks includes: performing at least one feature extraction operation among the plurality of feature extraction operations using the at least one convolutional layer of the corresponding second encoder block, and each of the plurality of second decoder blocks corresponding to the corresponding refinement block includes at least one convolutional layer of the plurality of convolutional layers of the CNN; and using the corresponding second decoder block of the corresponding refinement block, generating the corresponding upsampled feature map by each second decoder block of the plurality of second decoder blocks includes: performing at least one feature extraction operation among the plurality of feature extraction operations using the at least one convolutional layer of the corresponding second decoder block.
21. The method according to claim 20, wherein: each convolutional layer of each of the plurality of second encoder blocks of the refinement block is one of the plurality of convolutional layers of the CNN, and each convolutional layer of each of the plurality of second decoder blocks of the refinement block is one of the plurality of convolutional layers of the CNN.
22. The method according to any one of claims 16 to 21, wherein, for each of the plurality of refinement blocks: Each of the plurality of second encoder blocks of the refinement block is configured as a residual block, and each of the plurality of second decoder blocks of the refinement block is configured as a residual block.
23. The method according to any one of claims 14 to 21, wherein, the output image is generated based on the set of refined feature maps.
24. The method according to claim 23, wherein, the output image is generated based on the average value of the set of refined feature maps.
25. The method according to any one of claims 1 to 24, wherein, receiving the input image includes receiving a plurality of input images, each of the plurality of input images being a labeled image for training the CNN to obtain a trained CNN, and the method further includes, for each of the plurality of input images: performing the plurality of feature extraction operations using the plurality of convolutional layers of the CNN to generate the plurality of output feature maps; and generating the output image corresponding to the input image based on the plurality of output feature maps of the plurality of convolutional layers.
26. The method according to claim 25, wherein, the labeled image is an ultrasound image including labels of tissue structures.
27. The method according to any one of claims 1 to 24, wherein, the output image is the result of inferring the input image using the CNN.
28. The method according to claim 27, wherein, the input image is an ultrasound image including tissue structures.
29. A system for image processing based on a convolutional neural network (CNN), the system comprising: a memory; and at least one processor communicatively coupled to the memory and configured to execute the method for image processing based on a convolutional neural network (CNN) using at least one processor according to any one of claims 1 to 28.
30. A computer program product embodied in one or more non-transitory computer-readable storage media, the computer program product including executable instructions that can be executed by at least one processor to perform the method for image processing based on a convolutional neural network (CNN) using at least one processor according to any one of claims 1 to 28.
31. A method for segmenting tissue structures in an ultrasound image using a convolutional neural network (CNN), the method using at least one processor, the method comprising: performing the method for image processing based on a convolutional neural network (CNN) using at least one processor according to any one of claims 1 to 24, wherein: the input image is the ultrasound image including the tissue structures; and the output image has the segmented tissue structures and is the result of inferring the input image using the CNN.
32. The method according to claim 31, wherein, the CNN is trained according to claim 25 or 26.
33. A system for segmenting tissue structures in an ultrasound image using a CNN, the system comprising: a memory; and At least one processor, the processor being communicatively coupled to the memory and configured to execute the method of segmenting tissue structures in an ultrasound image using a convolutional neural network CNN according to claim 31 or 32.
34. A computer program product embodied in one or more non-transitory computer-readable storage media, the computer program product including executable instructions that can be executed by at least one processor to perform the method of segmenting tissue structures in an ultrasound image using a convolutional neural network CNN according to claim 31 or 32.
Citation Information
Patent Citations
Feature image extraction method based on deformable convolutional layer and feature image extraction device thereof
CN107292319A
Image deblurring method based on multi-task CNN
CN110782399A