System and method for labeling ultrasound data
By training the convolutional neural network to process ultrasonic data, the problem of difficulty in labeling tissue boundaries in ultrasonic images is solved, and accurate labeling and tissue segmentation of multiple pixels is achieved.
Patent Information
- Application Number
- CN202080047975.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-12
- Filing Date
- 2020-06-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-06-12
AI Technical Summary
Precise labeling of tissue boundaries in ultrasound images is difficult, especially when adjacent tissues have similar acoustic-mechanical properties.
By training a convolutional neural network (CNN), ultrasonic data is used to segment tissues, including downsampling of radio frequency waveform data and ultrasonic images, and labeling of tissues based on the output of CNN.
Accurate marking of multiple pixels in ultrasound images is achieved, including identifying muscles, fascia, fat, transplanted fat and other tissues, improving the accuracy and efficiency of tissue segmentation.
Smart Images

Figure CN114026654B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 860,403, filed on June 12, 2019, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates generally to ultrasound image processing and, in non-limiting embodiments or aspects, to systems and methods for labeling ultrasound data. Background Art
[0004] Ultrasound has become an increasingly popular medical imaging technique. For example, ultrasound may be relatively low risk (e.g., relatively few potential side effects and / or the like), relatively inexpensive (e.g., compared to other types of medical images), and / or the like.
[0005] However, analysis of ultrasound (e.g., ultrasound images and / or the like) can be more challenging than many other medical imaging modalities because ultrasound pixel values can depend on the path through intervening tissues and the orientation and properties of reflective tissue interfaces. As a result, even experts with extensive anatomical knowledge can have difficulty drawing precise boundaries between tissue interfaces in ultrasound images, particularly when adjacent tissues have similar acoustic-mechanical properties. For example, in superficial subcutaneous tissue, fascia tissue may appear similar to adipose tissue in an ultrasound image. In addition, certain methods for identifying soft tissue in ultrasound images use algorithms designed to identify specific targets, such as blood vessels and prostate glands. Such methods are typically only accurate in limited circumstances and often cannot reliably distinguish between several types of tissue. Summary of the invention
[0006] According to a non-limiting embodiment or aspect, a method for labeling ultrasound data is provided, comprising: training a convolutional neural network (CNN) based on ultrasound data, wherein the ultrasound data comprises ultrasound waveform data (e.g., radio frequency (RF) waveform data); downsampling the RF input of each downsampling layer of a plurality of downsampling layers in the CNN, wherein the RF input comprises RF waveform data for ultrasound; and segmenting tissue in the ultrasound based on an output of the CNN.
[0007] In a non-limiting embodiment or aspect, the method further includes: downsampling an image input of each of a plurality of downsampling layers in the CNN, the image input comprising a plurality of pixels of the ultrasound. In a non-limiting embodiment or aspect, the image input and the radio frequency input are processed substantially simultaneously. In a non-limiting embodiment or aspect, segmenting tissue in the ultrasound comprises marking a plurality of pixels. In a non-limiting embodiment or aspect, the plurality of pixels comprises a majority of pixels in the ultrasound. In a non-limiting embodiment or aspect, segmenting tissue comprises identifying at least one of: muscle, fascia, fat, transplanted fat, skin, tendon, ligament, nerve, blood vessel, bone, cartilage, needle, surgical instrument, or any combination thereof.
[0008] In a non-limiting embodiment or aspect, the multiple downsampling layers include an ultrasound image encoding branch and multiple radio frequency encoding branches, each radio frequency encoding branch includes a respective kernel size different from other radio frequency encoding branches in the multiple radio frequency encoding branches, and the respective kernel size of each radio frequency encoding branch corresponds to a respective wavelength. Additionally or alternatively, each radio frequency encoding branch includes multiple convolution blocks, each convolution block includes a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer, and a second activation layer, and at least one convolution block in the multiple convolution blocks includes a maximum pooling layer. Additionally or alternatively, downsampling includes downsampling the RF input of each RF encoding branch of the multiple RF encoding branches in the CNN, and downsampling the image input of each ultrasound image encoding branch in the CNN, the image input including multiple pixels of the ultrasound. In a non-limiting embodiment or aspect, the method further includes concatenating the RF coding branch output of each RF coding branch and the ultrasound image coding branch output of the ultrasound image coding branch to provide a concatenated coding branch output, and / or upsampling the concatenated coding branch output using multiple upsampling layers in the CNN. In a non-limiting embodiment or aspect, the multiple upsampling layers include a decoding branch, which includes multiple upconvolution blocks. In addition or alternatively, the CNN further includes multiple residual connections, each residual connection connecting a respective convolution block in the multiple convolution blocks to a respective upconvolution block in the multiple upconvolution blocks, and the size of the upconvolution block corresponds to the respective convolution block.
[0009] According to a non-limiting embodiment or aspect, a system for labeling ultrasound data is provided, comprising at least one computing device programmed or configured to train a convolutional neural network (CNN) based on ultrasound data, wherein the ultrasound data comprises ultrasound waveform data (e.g., radio frequency (RF) waveform data); downsample the RF input of each downsampling layer of a plurality of downsampling layers in the CNN, wherein the RF input comprises RF waveform data for ultrasound; and segment tissue in the ultrasound based on the output of the CNN.
[0010] In a non-limiting embodiment or aspect, the computing device is further programmed or configured to downsample an image input to each of a plurality of downsampling layers in the CNN, the image input comprising a plurality of pixels of the ultrasound. In a non-limiting embodiment or aspect, the image input and the radio frequency input are processed substantially simultaneously. In a non-limiting embodiment or aspect, segmenting tissue in the ultrasound comprises marking a plurality of pixels. In a non-limiting embodiment or aspect, the plurality of pixels comprises a majority of pixels in the ultrasound. In a non-limiting embodiment or aspect, segmenting tissue comprises identifying at least one of: muscle, fascia, fat, transplanted fat, or any combination thereof.
[0011] In a non-limiting embodiment or aspect, the multiple downsampling layers include an ultrasound image encoding branch and multiple radio frequency encoding branches, each radio frequency encoding branch includes a respective kernel size different from other radio frequency encoding branches of the multiple radio frequency encoding branches, and the respective kernel size of each radio frequency encoding branch corresponds to a respective wavelength. Additionally or alternatively, each radio frequency encoding branch includes multiple convolution blocks, each convolution block includes a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer, and a second activation layer, and at least one convolution block of the multiple convolution blocks includes a maximum pooling layer. Additionally or alternatively, downsampling includes downsampling the RF input of each RF encoding branch of the multiple RF encoding branches in the CNN, and downsampling the image input of each ultrasound image encoding branch in the CNN, the image input including multiple pixels of the ultrasound. In a non-limiting embodiment or aspect, the computing device is further programmed or configured to concatenate the RF coding branch output of each RF coding branch and the ultrasound image coding branch output of the ultrasound image coding branch to provide a concatenated coding branch output, and / or to upsample the concatenated coding branch output for use in multiple upsampling layers in the CNN. In a non-limiting embodiment or aspect, the multiple upsampling layers include a decoding branch, which includes a plurality of upconvolution blocks. Additionally or alternatively, the CNN further includes a plurality of residual connections, each of which connects a respective convolution block in the plurality of convolution blocks to a respective upconvolution block in the plurality of upconvolution blocks, and the size of the upconvolution block corresponds to the respective convolution block.
[0012] According to a non-limiting embodiment or aspect, a method for labeling ultrasound data is provided, comprising: receiving an ultrasound image represented by a plurality of pixels; and segmenting the ultrasound image by labeling a majority of the plurality of pixels.
[0013] In a non-limiting embodiment or aspect, the majority of pixels are labeled as at least one of: muscle, fascia, fat, transplanted fat, or any combination thereof. In a non-limiting embodiment or aspect, the ultrasound image is segmented based on a convolutional neural network (CNN), further comprising training the CNN based on ultrasound data, wherein at least one input ultrasound image in the ultrasound data includes fuzzy overlapping labels for multiple pixels.
[0014] According to a non-limiting embodiment or aspect, a system for labeling ultrasound data is provided, comprising at least one computing device programmed or configured to: receive an ultrasound image represented by a plurality of pixels; and segment the ultrasound image by labeling a majority of the plurality of pixels.
[0015] In a non-limiting embodiment or aspect, the majority of pixels are labeled as at least one of: muscle, fascia, fat, transplanted fat, or any combination thereof. In a non-limiting embodiment or aspect, the ultrasound image is segmented based on a convolutional neural network (CNN), and the computing device is further programmed or configured to train the CNN based on ultrasound data, wherein at least one input ultrasound image in the ultrasound data includes fuzzy overlapping labels for multiple pixels.
[0016] According to a non-limiting embodiment or aspect, a method for labeling ultrasound data is provided, comprising: training an artificial neural network (ANN) based on ultrasound data, the ultrasound data comprising ultrasound waveform data; and segmenting or otherwise labeling tissue in an ultrasound image or video based on the output of the ANN.
[0017] In a non-limiting embodiment or aspect, the ANN includes at least one of a convolutional neural network (CNN), a capsule network, a probabilistic network, a recurrent network, a deep network, or any combination thereof. In a non-limiting embodiment or aspect, the ultrasound waveform data includes at least one of an ultrasound image, raw radio frequency (RF) waveform data, beamformed RF waveform data, an intermediate representation derived from RF waveform data, or any combination thereof. In a non-limiting embodiment or aspect, the ultrasound waveform data or its intermediate representation preserves frequency information. In a non-limiting embodiment or aspect, the method further includes at least one of the following: downsampling the RF input of each of the multiple downsampling layers in the ANN, the RF input including the RF waveform data for ultrasound; or downsampling the image input of each of the multiple downsampling layers in the ANN, the image input including multiple pixels of the ultrasound.
[0018] Additional embodiments or aspects are listed in the following numbered clauses:
[0019] Item 1. A method for labeling ultrasound data, comprising: training a convolutional neural network (CNN) based on ultrasound data, wherein the ultrasound data includes radio frequency (RF) waveform data; downsampling the RF input of each downsampling layer of a plurality of downsampling layers in the CNN, wherein the RF input includes RF waveform data for ultrasound; and segmenting tissue in the ultrasound based on the output of the CNN.
[0020] Clause 2. The method according to clause 1 further comprises: downsampling an image input to each of the multiple downsampling layers in the CNN, the image input comprising multiple pixels of the ultrasound.
[0021] Clause 3. A method according to any preceding clause, wherein the image input and the radio frequency input are processed substantially simultaneously.
[0022] Clause 4. The method of any preceding clause, wherein segmenting tissue in the ultrasound comprises marking a plurality of pixels.
[0023] Clause 5. The method of any preceding clause, wherein the plurality of pixels comprises a majority of pixels in the ultrasound.
[0024] Clause 6. The method of any of the preceding clauses, wherein segmenting the tissue comprises identifying at least one of: muscle, fascia, fat, transplanted fat, or any combination thereof.
[0025] Clause 7. A method according to any one of the preceding clauses, wherein the multiple downsampling layers include an ultrasound image coding branch and multiple RF coding branches, each RF coding branch includes a respective kernel size different from other RF coding branches in the multiple RF coding branches, and the respective kernel size of each RF coding branch corresponds to a respective wavelength, wherein each RF coding branch includes a plurality of convolution blocks, each convolution block includes a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer and a second activation layer, and at least one convolution block in the multiple convolution blocks includes a maximum pooling layer, and wherein downsampling includes downsampling the RF input of each RF coding branch of the multiple RF coding branches in the CNN and downsampling the image input of each ultrasound image coding branch in the CNN, and the image input includes multiple pixels of the ultrasound.
[0026] Clause 8. The method according to any of the preceding clauses further includes: concatenating the RF coding branch output of each RF coding branch and the ultrasound image coding branch output of the ultrasound image coding branch to provide a concatenated coding branch output; and upsampling the concatenated coding branch output using multiple upsampling layers in the CNN.
[0027] Clause 9. A method according to any of the preceding clauses, wherein the multiple upsampling layers include a decoding branch, the decoding branch includes a plurality of upconvolution blocks, wherein the CNN further includes a plurality of residual connections, each residual connection connecting a respective convolution block in the plurality of convolution blocks to a respective upconvolution block in the plurality of upconvolution blocks, the upconvolution blocks having a size corresponding to the respective convolution blocks.
[0028] Item 10. A system for labeling ultrasound data, comprising at least one computing device, the at least one computing device being programmed or configured to: train a convolutional neural network (CNN) based on ultrasound data, the ultrasound data comprising radio frequency (RF) waveform data; downsample the RF input of each downsampling layer of a plurality of downsampling layers in the CNN, the RF input comprising RF waveform data for ultrasound; and segment tissue in the ultrasound based on the output of the CNN.
[0029] Clause 11. A system according to clause 10, wherein the computing device is further programmed or configured to downsample an image input of each downsampling layer of a plurality of downsampling layers in the CNN, the image input comprising a plurality of pixels of the ultrasound.
[0030] Clause 12. The system of any of clauses 10 to 11, wherein the image input and the radio frequency input are processed substantially simultaneously.
[0031] Clause 13. The system according to any one of Clauses 10 to 12, wherein the tissue segmented in the ultrasound includes marking a plurality of pixels.
[0032] Clause 14. The system according to any one of Clauses 10 to 13, wherein the plurality of pixels includes most of the pixels in the ultrasound.
[0033] Clause 15. The system according to any one of Clauses 10 to 14, wherein segmenting the tissue includes identifying at least one of the following: muscle, fascia, fat, transplanted fat, or any combination thereof.
[0034] Clause 16. A method for marking ultrasound data, comprising: receiving an ultrasound image represented by a plurality of pixels; and segmenting the ultrasound image by marking most of the plurality of pixels.
[0035] Clause 17. The method according to Clause 16, wherein most of the pixels are marked as at least one of the following: muscle, fascia, fat, transplanted fat, or any combination thereof.
[0036] Clause 18. The system according to any one of Clauses 10 to 17, wherein the ultrasound image is segmented based on a convolutional neural network (CNN), further comprising training the CNN based on the ultrasound data, wherein at least one input ultrasound image in the ultrasound data includes blurred overlapping markings for a plurality of pixels.
[0037] Clause 19. A system for marking ultrasound data, comprising at least one computing device programmed or configured to: receive an ultrasound image represented by a plurality of pixels; and segment the ultrasound image by marking most of the plurality of pixels.
[0038] Clause 20. The system according to Clause 19, wherein most of the pixels are marked as at least one of the following: muscle, fascia, fat, transplanted fat, skin, tendon, ligament, nerve, blood vessel, bone, cartilage, needle, surgical instrument, or any combination thereof.
[0039] Clause 21. The system according to any one of Clauses 19 to 20, wherein the ultrasound image is segmented based on a convolutional neural network (CNN), and the computing device is further programmed or configured to train the CNN based on the ultrasound data, wherein at least one input ultrasound image in the ultrasound data includes blurred overlapping markings for a plurality of pixels.
[0040] Clause 22. A method for labeling ultrasound data, comprising: training an artificial neural network (ANN) based on ultrasound data, the ultrasound data comprising ultrasound waveform data; and segmenting or otherwise labeling tissue in ultrasound based on an output of the ANN.
[0041] Clause 23. The method of clause 22, wherein the ANN comprises at least one of a convolutional neural network (CNN), a capsule network, a probabilistic network, a recurrent network, a deep network, or any combination thereof.
[0042] Clause 24. A method according to any one of clauses 22 to 23, wherein the ultrasound waveform data comprises at least one of an ultrasound image, raw radio frequency (RF) waveform data, beamformed RF waveform data, an intermediate representation derived from the RF waveform data, or any combination thereof.
[0043] Clause 25. The method of any one of clauses 22 to 24, wherein the ultrasonic waveform data stores frequency information.
[0044] Clause 26. The method according to any one of clauses 22 to 25 further includes at least one of the following: downsampling the RF input of each of the multiple downsampling layers in the ANN, wherein the RF input includes RF waveform data for the ultrasound; or downsampling the image input of each of the multiple downsampling layers in the ANN, wherein the image input includes multiple pixels of the ultrasound.
[0045] These and other features and characteristics of the present disclosure, the methods of operation and functions of the related structural elements, and the combination and economy of manufacture of the parts will become more apparent after considering the following description and appended claims with reference to the accompanying drawings, all of which form a part of this specification, wherein like reference numerals designate corresponding parts in the various figures. It is to be expressly understood, however, that the drawings are for illustration and description purposes only and are not intended as a definition of the limits of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Further advantages and details will be explained in more detail below with reference to a non-limiting, exemplary embodiment shown in the accompanying drawings, in which:
[0047] Figure 1 A system for labeling ultrasound data according to a non-limiting embodiment or aspect is shown;
[0048] Figure 2 Exemplary components for a computing device related to non-limiting embodiments or aspects are shown;
[0049] Figure 3An exemplary implementation of a convolutional neural network according to a non-limiting embodiment or aspect is shown;
[0050] Figure 4 An exemplary implementation of a convolutional neural network according to a non-limiting embodiment or aspect is shown;
[0051] Figure 5 An exemplary implementation of a convolutional neural network according to a non-limiting embodiment or aspect is shown;
[0052] FIG. 6A to FIG. 6C Test data illustrating results of implementations according to non-limiting embodiments or aspects;
[0053] Figure 7 is a flow chart of a method for labeling ultrasound data according to a non-limiting embodiment or aspect; and
[0054] Figure 8 is a flow chart of a method for labeling ultrasound data according to a non-limiting embodiment or aspect. DETAILED DESCRIPTION
[0055] It should be understood that, except for the contrary situation clearly specified, embodiments of the present invention can bear various alternative changes and step sequences. It should also be understood that the specific devices and processes described in the following specification are only exemplary embodiments or aspects of the present disclosure. Therefore, the specific dimensions and other physical characteristics related to the embodiments or aspects disclosed herein should not be considered restrictive. Any aspects, parts, elements, structures, behaviors, steps, functions, instructions and / or the like used herein should not be understood as key or necessary unless such description is clearly made. In addition, as used herein, the articles "one" and "an" are intended to include one or more items and can be used interchangeably with "one or more" and "at least one". In addition, as used herein, the term "having" or similar terms are intended to be open terms. In addition, the phrase "based on" means "based at least in part on", unless otherwise explicitly stated.
[0056] As used herein, the term "computing device" may refer to one or more electronic devices configured to process data. In some embodiments, the computing device may include necessary components for receiving, processing, and outputting data, such as a processor, a display, a memory, an input device, a network interface, and / or the like. A computing device may be a mobile device. The computing device may also be a desktop computer or other form of non-mobile computer. In a non-limiting embodiment or aspect, the computing device may include an artificial intelligence accelerator, including an application-specific integrated circuit (ASIC) neural engine, such as Apple's "neural engine" or Google's Tensor processing unit. In a non-limiting embodiment or aspect, the computing device may be composed of multiple separate circuits representing each connection in a neural network, such that each circuit is configured to measure the input from each node in the neural network. In such an arrangement, logic gates and / or analog circuits may be used without the need for software, processors, or memory.
[0057] A non-limiting embodiment or aspect provides a system and method for segmenting ultrasound data by using ultrasound waveform data (e.g., radio frequency (RF) waveform data). In a non-limiting embodiment or aspect, a deep learning computer vision method is used to automatically identify and mark soft tissue visible in ultrasound. The non-limiting embodiment or aspect allows the distinction between muscle, fascia, fat, and transplanted fat. The non-limiting embodiment or aspect can be applied to plastic surgery operations (e.g., adding or removing fat) and obtaining STEM cells from a patient's fat, including for treating radiation damage from cancer treatment. Muscle, fat, and transplanted fat may look similar in ultrasound images, making automatic distinction very challenging. The non-limiting embodiment or aspect allows the ultrasound of superficial subcutaneous tissue (e.g., muscle, fascia, fat, transplanted fat, or any combination thereof) to be segmented by using deep learning (e.g., convolutional neural network (CNN)). The non-limiting embodiment or aspect allows the ultrasound to be segmented by marking most (e.g., all, substantially all, and / or similar) pixels in the ultrasound image without using background markers. A non-limiting embodiment or aspect enables modification of a CNN to process image pixels and RF waveform data simultaneously (e.g., for deep learning / CNN segmentation of ultrasound). In a non-limiting embodiment or aspect, a CNN is created that enables learning of RF convolution kernels. Such a configuration may involve processing RF waveform scales that are different from image pixel sampling in both the vertical (e.g., "axial" or RF time) and horizontal axes. For example, an encoder-decoder-decoder CNN architecture may be modified to have an image (e.g., ultrasound image) downsampling branch (e.g., column, path, or channel group) and at least one parallel RF downsampling branch with different kernels learned for each RF downsampling branch, and multi-channel convolution, which may take the form of a multi-column encoder, with late fusion between the ultrasound image branch and the RF branch. A non-limiting embodiment or aspect provides a system and method for segmenting ultrasound data by using multiple parallel RF encoding branches that can be incorporated into a CNN. Therefore, in a non-limiting embodiment or aspect, RF waveform data can be processed (e.g., downsampled, etc.) simultaneously with ultrasound image data to improve the accuracy and efficiency of segmenting the ultrasound. Non-limiting embodiments or aspects provide data padding, including a novel scheme for padding the deep end of RF waveform data. Thus, the same CNN can be used to process ultrasound data (e.g., ultrasound images, RF images, and / or the like) even if certain items of the data have different dimensions (e.g., imaging depth, size, and / or the like).
[0058] Non-limiting embodiments can be implemented as a software application for processing ultrasound data output by an ultrasound device. In other non-limiting embodiments, the system and method for marking ultrasound data can be directly incorporated into the ultrasound device as hardware and / or software.
[0059] Reference now Figure 1, a system 100 for marking ultrasound data according to a non-limiting embodiment or aspect is shown. The system 100 may include an ultrasound / RF system 102. For example, the ultrasound / RF system 102 may include an ultrasound device configured to physically capture ultrasound waveform data (e.g., RF waveform data). In a non-limiting embodiment or aspect, the ultrasound / RF system 102 may only save (e.g., store, communicate, and / or the like) certain data related to the RF waveform data (e.g., the amplitude envelope of the RF waveform and / or the like), which may be used to create a grayscale ultrasound image. For example, the original, per-element RF waveform may be combined into a beamformed RF waveform, and the envelope of the beamformed RF waveform may form the basis of a grayscale ultrasound image (e.g., for display on a screen and / or the like). Additionally or alternatively, the frequency content may be used by the ultrasound / RF system 102 to calculate a Doppler shift to measure velocity (e.g., which may be displayed in color). In a non-limiting embodiment or aspect, after certain data (such as envelope, Doppler shift and / or the like) is calculated (such as derived, determined and / or the like), the raw RF waveform data can be discarded. Additionally or alternatively, the ultrasound / RF system 102 can save the RF waveform data for additional analysis (e.g., storage, analysis and / or the like of RF waveform data). In a non-limiting embodiment or aspect, the ultrasound / RF system 102 may include an ultrasound device that captures and saves RF waveform data (e.g., beamformed RF waveform data, RF waveform data for each element, any other suitable representation of the RF waveform (e.g., saving frequency content), any combination thereof and / or the like). Additionally or alternatively, the RF waveform data may be used in real time for online analysis, may be stored for later analysis, any combination thereof and / or the like. In a non-limiting embodiment or aspect, the ultrasound / RF system 102 may include a portable ultrasound machine, such as a crystal linear array scanner and / or the like. For example, the ultrasound / RF system 102 may include a Clarius L7 portable ultrasound machine. In non-limiting embodiments or aspects, the ultrasound / RF system 102 may be used to obtain at least one ultrasound image 104a (e.g., a grayscale image and / or the like), RF waveform data 104b (e.g., at least one RF image, a set of RF waveforms and / or the like, which may correspond to 104a), any combination thereof, and / or the like from at least one patient. For example, a clinician may use the ultrasound / RF system 102 to obtain such images. Additionally or alternatively, the ultrasound / RF system 102 may output (e.g., communicate and / or the like) ultrasound data 104, which may include at least one ultrasound image 104a, RF waveform data 104b, any combination thereof, and / or the like.In non-limiting embodiments or aspects, the ultrasound / RF system 102 may include one or more devices capable of receiving information from and / or communicating information to a computing device 106 , a database 108 , and / or the like.
[0060] The computing device 106 may include one or more devices capable of receiving information from and / or communicating information to the ultrasound system / RF system 102, the database 108, and / or the like. In a non-limiting embodiment or aspect, the computing device 106 may implement at least one convolutional neural network (e.g., W-Net, U-Net, AU-Net, SegNet, any combination thereof, and / or the like), as described herein. In a non-limiting embodiment or aspect, the computing device 106 may receive ultrasound data 104 (e.g., ultrasound images 104a, RF waveform data 104b, any combination thereof, and / or the like) from the ultrasound / RF system 102. Additionally or alternatively, the computing device 106 may receive (e.g., retrieve, and / or the like) ultrasound data 104 (e.g., historical ultrasound data, which may include at least one ultrasound image 104a, RF waveform data 104b, at least one labeled ultrasound image 104c, any combination thereof, and / or the like, as described herein) from a database 108.
[0061] In a non-limiting embodiment or aspect, the computing device 106 may train a CNN based on the ultrasound data 104, as described herein. Additionally or alternatively, the computing device 106 may downsample the ultrasound data 104 (e.g., RF waveform data 104b, ultrasound image 104a, any combination thereof, and / or the like) using a CNN implemented by the computing device 106, as described herein. For example, the computing device 106 may downsample the RF input of each of the multiple downsampling layers in the CNN (e.g., RF waveform data 104b of ultrasound and / or the like), as described herein. Additionally or alternatively, the computing device 106 may downsample the image input of each of the multiple downsampling layers in the CNN (e.g., at least one ultrasound image 104a including multiple pixels of the ultrasound), as described herein. In non-limiting embodiments or aspects, image input (e.g., ultrasound image 104a and / or the like) and RF input (e.g., RF waveform data 104b) can be processed substantially simultaneously (e.g., via parallel branches of the CNN, as separate channels of the input image to the CNN, any combination thereof, and / or the like), as described herein. In non-limiting embodiments or aspects, the computing device 106 can segment tissue in the ultrasound based on the output of the CNN, as described herein. For example, segmenting tissue in the ultrasound can include marking a plurality of pixels (e.g., most pixels, all pixels, and / or the like) in the ultrasound, as described herein. Additionally or alternatively, segmenting tissue (e.g., marking pixels and / or the like) can include identifying at least one of: muscle, fascia, fat, transplanted fat, skin, tendon, ligament, nerve, blood vessel, bone, cartilage, needle, surgical instrument, any combination thereof, and / or the like, as described herein. In non-limiting embodiments or aspects, the computing device 106 may output segmented ultrasound data 110 (eg, segmented ultrasound images and / or the like), as described herein.
[0062] In non-limiting embodiments or aspects, the computing device 106 can be separate from the ultrasound / RF system 102. Additionally or alternatively, the computing device 106 can be incorporated (eg, fully, partially, and / or similarly incorporated) into the ultrasound / RF system 102.
[0063] The database 108 may include one or more devices capable of receiving information from and / or communicating information to the ultrasound / RF system 102, the computing device 106, and / or the like. In a non-limiting embodiment or aspect, the database 108 may store ultrasound data 104 (e.g., historical ultrasound data) from previous ultrasound / RF scans (e.g., RF scans performed by the ultrasound / RF system 102, other ultrasound and / or RF systems, and / or the like). For example, the (historical) ultrasound data 104 may include at least one ultrasound image 104a, RF waveform data 104b, at least one labeled ultrasound image 104c, any combination thereof, and / or the like, as described herein. In a non-limiting embodiment or aspect, a clinician may provide a label for the labeled ultrasound image 104c. Additionally or alternatively, such a labeled ultrasound image 104c may be used to train and / or test a CNN (e.g., to determine how accurate the segmented tissue based on the output of the CNN is with the label provided by the clinician, and / or the like), as described herein.
[0064] In non-limiting embodiments or aspects, database 108 can be separate from computing device 106. Additionally or alternatively, database 108 can be implemented by computing device 106 (eg, completely, partially, and / or similarly).
[0065] In non-limiting embodiments or aspects, the ultrasound / RF system 102, computing device 106, and database 108 may be implemented by a single device, a single system, and / or the like (eg, completely, partially, and / or similarly).
[0066] Reference now Figure 2 , which is a diagram of exemplary components of a computing device 900 for implementing and executing the systems and methods described herein, according to a non-limiting embodiment. In non-limiting embodiments or aspects, the device 900 may include additional components, fewer components, different components, or components that are different from the computing device 900. Figure 2Components of different arrangements shown. Device 900 may include a bus 902, a processor 904, a memory 906, a storage component 908, an input component 910, an output component 912, and a communication interface 914. Bus 902 may include components for allowing communication between components of device 900. In a non-limiting embodiment or aspect, processor 904 may be implemented with hardware, firmware, or a combination of hardware and software. For example, processor 904 may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc.), a microprocessor, a digital signal processor (DSP), and / or any processing component that can be programmed to perform a function (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.). Memory 906 may include a random access memory (RAM), a read-only memory (ROM), and / or other types of dynamic or static storage devices (e.g., flash memory, magnetic storage, optical storage, etc.), which store information and / or instructions used by processor 904.
[0067] Continue to refer Figure 2 , the storage component 908 can store information and / or software related to the operation and use of the device 900. For example, the storage component 908 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a solid-state disk, etc.) and / or other types of computer-readable media. The input component 910 may include a component for allowing the device 900 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, etc.). In addition or alternatively, the input component 910 may include a sensor for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.). The output component 912 may include a component for providing output information from the device 900 (e.g., a display, a speaker, one or more light-emitting diodes (LEDs), etc.). The communication interface 914 may include a transceiver-like component (e.g., a transceiver, a separate receiver and transmitter, etc.) that enables the device 900 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 914 may allow the device 900 to receive information from another device and / or provide information to another device. For example, the communication interface 914 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface, cellular network interface, and / or the like.
[0068] The device 900 can perform one or more processes described herein. The device 900 can perform these processes based on the processor 904 executing software instructions stored by a computer-readable medium (e.g., memory 906 and / or storage component 908). The computer-readable medium may include any non-temporary memory device. The memory device includes a storage space located inside a single physical storage device or a storage space distributed in multiple physical storage devices. The software instructions can be read into the memory 906 and / or storage component 908 from another computer-readable medium or from another device through the communication interface 914. When executed, the software instructions stored in the memory 906 and / or storage component 908 can cause the processor 904 to perform one or more processes described herein. In addition or alternatively, hard-wired circuits can be used in place of or in combination with software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any specific combination of hardware circuits and software. The term "programming or configuration" used herein refers to the arrangement of software, hardware circuits, or any combination thereof on one or more devices.
[0069] Reference now Figure 3 , according to non-limiting embodiments or aspects, an exemplary CNN 300 (e.g., W-Net CNN architecture) is shown. For example, CNN 300 can be implemented (e.g., fully, partially, and / or similarly) by computing device 106. Additionally or alternatively, CNN 300 can be implemented (e.g., fully, partially, and / or similarly) by at least one other computing device and / or directly in (e.g., digital and / or analog) circuits, separate from or including computing device 106.
[0070] like Figure 3 As shown, the CNN 300 may include a downsampling layer (e.g., an ultrasound image encoding branch 330, an RF encoding branch 340 (e.g., a first RF encoding branch 341, a second RF encoding branch 342, a third RF encoding branch 343 and / or a fourth RF encoding branch 344) and / or the like), a bottleneck portion 350, an upsampling layer (e.g., a decoding branch 360 and / or the like), any combination thereof and / or the like.
[0071] Continue to refer Figure 3, the ultrasound image encoding branch 330 may downsample the image input (e.g., ultrasound image 304a and / or the like). For example, the ultrasound image 304a may include a plurality of pixels of a grayscale image of the ultrasound. Additionally or alternatively, the ultrasound image 304a may be in color, may include a color Doppler overlay, any combination thereof, and / or the like. In a non-limiting embodiment or aspect, the ultrasound image encoding branch 330 may include a plurality of convolution blocks (e.g., a first convolution block 330a, a second convolution block 330b, a third convolution block 330c, and / or a fourth convolution block 330d). Each convolution block (e.g., 330a, 330b, 330c, and / or 330d) may include at least one convolution layer set 320. For example, each convolution block (e.g., 330a, 330b, 330c, and / or 330d) may include two convolution layer sets 320. In a non-limiting embodiment or aspect, each convolutional layer set 320 may include a convolutional layer, a batch normalization layer, an activation layer, any combination thereof, and / or the like. Additionally or alternatively, each convolutional block (e.g., 330a, 330b, 330c, and / or 330d) may include a maximum pooling layer 322. In a non-limiting embodiment or aspect, each convolutional layer set 320 of the first convolutional block 330a may have 16 feature maps. Additionally or alternatively, the size of each convolutional layer set 320 of the first convolutional block 330a may be based on the size of the input image (e.g., the ultrasound image 304a, which may have a size of 784x192 and / or similar) and / or the number of feature maps (e.g., 16). For example, the size of each convolutional layer set 320 of the first convolutional block 330a may be 784x192x16. In non-limiting embodiments or aspects, each convolutional layer set 320 of the second convolutional block 330b may have a greater number (e.g., twice) of feature maps than the first convolutional block 330a (e.g., 32 feature maps), and / or other dimensions of the second convolutional block 330b may be smaller than the dimensions of the first convolutional block 330a. For example, the dimensions of each convolutional layer set 320 of the second convolutional block 330b may be 392x96x32. Additionally or alternatively, each convolutional layer set 320 of the third convolutional block 330c may have more feature maps (e.g., 64 feature maps) than the convolutional layer set 320 of the second convolutional block 330b, and / or other dimensions of the third convolutional block 330c may be smaller than the dimensions of the second convolutional block 330b. For example, the dimensions of each convolutional layer set 320 of the third convolutional block 330c may be 196x48x64. Additionally or alternatively, each convolutional layer set 320 of the fourth convolutional block 330d may have a larger number (e.g., twice) of feature maps (e.g., 128 feature maps) than the third convolutional block 330c, and / or other dimensions of the fourth convolutional block 330d may be smaller than the dimensions of the third convolutional block 330c. For example, the dimensions of each convolutional layer set 320 of the fourth convolutional block 330d may be 98x24x128.In a non-limiting embodiment or aspect, the activation layer of each convolutional layer set 320 in the ultrasound image encoding branch 330 may include a rectified linear unit (ReLU) layer.
[0072] In a non-limiting embodiment or aspect, at least some of the ultrasound data items (e.g., ultrasound images 304a, RF waveform data 304b, labeled ultrasound images, and / or the like) may have a different size than other items. For example, at least some of the ultrasound data items may have a size of 592x192. In a non-limiting embodiment or aspect, the structure of the CNN 300 may be limited to a fixed-size input, and ultrasound data items having a size different from the fixed-size input (e.g., smaller in at least one dimension) may be zero-padded (e.g., at their bottom) to match the input size. Additionally or alternatively, in order to reduce (e.g., minimize and / or similarly) phase artifacts introduced when padding the RF waveform data (e.g., RF images), RF images having a size different from the fixed-size input (e.g., smaller in at least one dimension) may be mirrored and / or reflected at the last (e.g., deepest) zero crossing of each A-scan / waveform to avoid waveform discontinuities and fill in padding values. In non-limiting embodiments or aspects, an error metric (e.g., for training and / or testing) may treat padded regions as special purpose background in a segmentation task and / or exclude padded regions from a loss function (e.g., when training CNN 300).
[0073] In a non-limiting embodiment or aspect, the output of the first convolution block 330a can be provided as an input to the second convolution block 330b. Additionally or alternatively, the output of the second convolution block 330b can be provided as an input to the third convolution block 330c. Additionally or alternatively, the output of the third convolution block 330c can be provided as an input to the fourth convolution block 330d. Additionally or alternatively, the output of the fourth convolution block 330d can be provided as an input to the bottleneck portion 350.
[0074] exist Figure 3In the embodiment of the present invention, the RF encoding branch 340 may include multiple RF encoding branches, for example, a first RF encoding branch 341, a second RF encoding branch 342, a third RF encoding branch 343 and / or a fourth RF encoding branch 344. Additionally or alternatively, each RF encoding branch (e.g., 341, 342, 343 and / or 344) may downsample the RF input (e.g., RF waveform data 304b and / or similar data). In a non-limiting embodiment or aspect, each RF encoding branch (e.g., 341, 342, 343 and / or 344) may include a respective kernel size different from that of other RF encoding branches. Additionally or alternatively, at least some kernels may be shaped and / or oriented to analyze along a range of single or small groups of adjacent waveform values. For example, for a vertical A-scan waveform of a linear probe, a tall, thin rectangular kernel may be used. For example, the kernel size of the first RF encoding branch 341 may be 51x9, the kernel size of the second encoding branch 342 may be 21x5, the kernel size of the third encoding branch 343 may be 11x3, and the kernel size of the fourth encoding branch 344 may be 7x3.
[0075] In a non-limiting embodiment or aspect, the respective kernel size of each RF encoding branch (e.g., 341, 342, 343, and / or 344) can correspond to the respective wavelength (e.g., RF spectrum and / or the like). For example, due to the different kernel sizes, the RF encoding branch 340 can bin the RF waveform analysis into different frequency bands corresponding to the wavelength support of each branch, which can help segmentation (e.g., classification and / or the like). In a non-limiting embodiment or aspect, the weights of at least some of the convolution blocks of the RF encoding branches (e.g., 341, 342, 343, and / or 344) can be initialized with a local frequency analysis kernel (e.g., wavelet, vertical Gabor kernel, and / or the like), for example, to encourage CNN 300 to learn an appropriate Gabor kernel to better bin the RF input into the various frequency bands. For example, the initial Gabor filter may include spatial frequencies in the range of [0.1, 0.85], variances σx∈[3, 5, 10, 25] and σy∈[1, 2, 4], and such Gabor filters may have a frequency separation of 3 to 8 MHz (which may be within the range of standard clinical practice, e.g., for portable point-of-care ultrasound (POCUS)). In a non-limiting embodiment or aspect, the first two convolution blocks of each RF encoding branch may include kernels designed to maintain Gabor filters of specific sizes, such as sizes 7x3, 11x3, 21x5, and 51x9 (one for each branch), as described herein. For example, a 11x3 Gabor filter may be embedded in a convolution kernel of size 21x5 (e.g., removing two standard deviations instead of one). In a non-limiting embodiment or aspect, the kernel size (e.g., the kernel size described above) may be selected to allow connections 380 (e.g., residual connections, skip connections, and / or the like) to enter the decoding branch 360 (e.g., matching the output size of the ultrasound image encoding branch 330). In a non-limiting embodiment or aspect, as described herein, the RF encoding branch with 11x3, 21x5, and 51x9 kernels may not have a maximum pooling (e.g., downsampling) layer 322 in its fourth, fourth, and third convolutional blocks, respectively. For example, the omission of such a maximum pooling (e.g., downsampling) layer 322 may compensate for lost input image boundary pixels.
[0076] In a non-limiting embodiment or aspect, each RF encoding branch (e.g., 341, 342, 343, and / or 344) may include a plurality of convolution blocks, for example, a first convolution block (e.g., 341a, 342a, 343a, and / or 344a), a second convolution block (e.g., 341b, 342b, 343b, and / or 344b), a third convolution block (e.g., 341c, 342c, 343c, and / or 344c), and / or a fourth convolution block (e.g., 341d, 342d, 343d, and / or 344d). Each convolution block may include at least one convolution layer set 320 (e.g., two convolution layer sets 320), as described herein, and each convolution layer set 320 may include a convolution layer, a batch normalization layer, an activation layer, any combination thereof, and / or the like. Additionally or alternatively, at least some of the convolution blocks (e.g., each convolution block, a subset of the convolution blocks, and / or the like) may include a maximum pooling layer 322, as described herein. For example, the third convolution block 341c of the first RF encoding branch 341, the fourth convolution block 342d of the second RF encoding branch 342, and / or the fourth convolution block 343d of the third RF encoding branch 343 may not include a maximum pooling layer 322, and / or the other convolution blocks of each RF encoding branch (e.g., 341, 342, 343, and / or 344) may each include a maximum pooling layer 322. In a non-limiting embodiment or aspect, each convolutional layer set 320 of the first convolutional block (e.g., 341a, 342a, 343a, and / or 344a) can have 16 feature maps and / or the size of each convolutional layer set 320 of the first convolutional block (e.g., 341a, 342a, 343a, and / or 344a) can be 784x192x16, as described herein. Additionally or alternatively, each convolutional layer set 320 of the second convolutional block (e.g., 341b, 342b, 343b, and / or 344b) can be 392x96x32, as described herein. Additionally or alternatively, each convolutional layer set 320 of the third convolutional block (e.g., 341c, 342c, 343c, and / or 344c) can be 196x48x64, as described herein. Additionally or alternatively, the size of each convolutional layer set 320 of the fourth convolutional block (e.g., 341d, 342d, 343d, and / or 344d) can be 98x24x128, as described herein. In a non-limiting embodiment or aspect, the activation layer of each convolutional layer set 320 in the RF encoding branch 340 can include a ReLU layer.
[0077] In a non-limiting embodiment or aspect, the output of each respective first convolution block (e.g., 341a, 342a, 343a, and / or 344a) may be provided as an input to each respective second convolution block (e.g., 341b, 342b, 343b, and / or 344b). Additionally or alternatively, the output of each respective second convolution block (e.g., 341b, 342b, 343b, and / or 344b) may be provided as an input to each respective third convolution block (e.g., 341c, 342c, 343c, and / or 344c). Additionally or alternatively, the output of each respective third convolution block (e.g., 341c, 342c, 343c, and / or 344c) may be provided as an input to each respective fourth convolution block (e.g., 341d, 342d, 343d, and / or 344d). Additionally or alternatively, the output of each respective fourth convolutional block (e.g., 341d, 342d, 343d, and / or 344d) may be provided as input to the bottleneck portion 350.
[0078] like Figure 3 As shown, the bottleneck portion 350 may include at least one convolutional layer set 352 (e.g., two convolutional layer sets 352), as described herein. For example, each convolutional layer set 352 may include a convolutional layer, a batch normalization layer, an activation layer, any combination thereof, and / or the like, as described herein. In a non-limiting embodiment or aspect, the activation layer of each convolutional layer set 352 may include a ReLU layer. In a non-limiting embodiment or aspect, each convolutional layer set 352 may have a greater number (e.g., twice) of feature maps (e.g., 768 feature maps) than the fourth convolutional block (e.g., 330d, 341d, 342d, 343d, and / or 344d) of the ultrasound image encoding branch 330 and / or the RF encoding branch 340, and / or other sizes of each convolutional layer set 352 may be smaller than the size of the fourth convolutional block. For example, the size of each convolutional layer set 352 may be 49x12x768.
[0079] In a non-limiting embodiment or aspect, the output of the fourth convolution block 330d of the ultrasound image encoding branch 330 and the output of each respective fourth convolution block (e.g., 341d, 342d, 343d, and / or 344d) of the RF encoding branch 340 may be provided as input to the bottleneck portion 350. Additionally or alternatively, such outputs from the encoding branches may be combined (e.g., concatenated, aggregated, and / or the like) before being provided as input to the bottleneck portion 350.
[0080] In a non-limiting embodiment or aspect, the output of the bottleneck portion 350 may be provided as input to the decoding branch 360 .
[0081] Continue to refer Figure 3, the decoding branch 360 may upsample its input (e.g., output from the bottleneck portion 350 and / or the like). In a non-limiting embodiment or aspect, the decoding branch 360 may include a plurality of up-convolution blocks (e.g., a first up-convolution block 360a, a second up-convolution block 360b, a third up-convolution block 360c, and / or a fourth up-convolution block 360d). Each up-convolution block (e.g., 360a, 360b, 360c, and / or 360d) may include at least one up-convolution layer 362 (e.g., a transposed convolution layer and / or the like) and / or at least one convolution layer set 364. For example, each upper convolution block (e.g., 360a, 360b, 360c, and / or 360d) may include two upper convolution layers 362 and / or two convolution layer sets 364 (e.g., first upper convolution layer 362, first convolution layer set 364, second upper convolution layer 362, and second convolution layer set 364, in sequence). In a non-limiting embodiment or aspect, each convolution layer set 364 may include a convolution layer, a batch normalization layer, an activation layer, any combination thereof, and / or the like. In a non-limiting embodiment or aspect, each upper convolution layer 362 and / or each convolution layer set 364 of the first upper convolution block 360a may have 256 feature maps. Additionally or alternatively, the size of each upper convolution layer 362 and / or each convolution layer set 364 of the first upper convolution block 360a may be based on the size of its input (e.g., the output of the bottleneck portion 350 and / or the like), the size of the fourth convolution block of the ultrasound image encoding branch 330 and / or the RF encoding branch 340, and / or the number of feature maps (e.g., 256). For example, the size of each upper convolution layer 362 and / or each convolution layer set 364 of the first upper convolution block 360a may be 98x24x16. In a non-limiting embodiment or aspect, each upper convolution layer 362 and / or each convolution layer set 364 of the second upper convolution block 360b may have a smaller number (e.g., half) of feature maps (e.g., 128 feature maps) than that of the first upper convolution block 360a, and / or other sizes of the second upper convolution block 360b may be based on the size of the third convolution block of the ultrasound image encoding branch 330 and / or the RF encoding branch 340. For example, the size of each upper convolution layer 362 and / or each convolution layer set 364 of the second upper convolution block 360b may be 196x48x128. Additionally or alternatively, each upper convolution layer 362 and / or each convolution layer set 364 of the third upper convolution block 360c may have a smaller number (e.g., half) of feature maps (e.g., 64 feature maps) than that of the second upper convolution block 360b, and / or other sizes of the third upper convolution block 360c may be based on the size of the second convolution block of the ultrasound image encoding branch 330 and / or the RF encoding branch 340. For example, the size of each upper convolution layer 362 and / or each convolution layer set 364 of the third upper convolution block 360c may be 392x96x64.Additionally or alternatively, each upper convolution layer 362 and / or each convolution layer set 364 of the fourth upper convolution block 360d may have a smaller number (e.g., half) of feature maps (e.g., 32 feature maps) than the third upper convolution block 360c, and / or other sizes of the fourth upper convolution block 360d may be based on the size of the first convolution block of the ultrasound image encoding branch 330 and / or the RF encoding branch 340. For example, the size of each upper convolution layer 362 and / or each convolution layer set 364 of the fourth upper convolution block 360d may be 784x192x32. In a non-limiting embodiment or aspect, the activation layer of each convolution layer set 364 in the decoding branch 360 may include a ReLU layer.
[0082] In a non-limiting embodiment or aspect, the output of the first up-convolution block 360a may be provided as an input to the second up-convolution block 360b. Additionally or alternatively, the output of the second up-convolution block 360b may be provided as an input to the third up-convolution block 360c. Additionally or alternatively, the output of the third up-convolution block 360c may be provided as an input to the fourth up-convolution block 360d. Additionally or alternatively, the output of the fourth up-convolution block 360d may be provided as an input to the output layer set 370.
[0083] In a non-limiting embodiment or aspect, the output layer set 370 may include at least one convolution layer and / or at least one activation layer. In a non-limiting embodiment or aspect, the activation layer of the output layer set 370 may include a softmax layer. In a non-limiting embodiment or aspect, the size of the output layer set 370 may be based on the size of the fourth upper convolution block 360d of the decoding branch 360 and / or the size of the input ultrasound data (e.g., ultrasound image 304a and / or RF waveform data 304b). For example, the size of the output layer set 370 may be 784x192. In a non-limiting embodiment or aspect, the activation layer of the output layer set 370 may include a classification layer. For example, the activation layer (e.g., classification layer) may assign a classification index (e.g., an integer class label and / or the like) to each pixel of the ultrasound to provide semantic segmentation (e.g., a label map and / or the like). In a non-limiting embodiment or aspect, the class labels can be selected from the set {1, 2, 3, 4, 5}, where the following integers can correspond to the following types of tissue: (1) skin (e.g., epidermis / dermis); (2) fat; (3) adipose fascia / fascia; (4) muscle; and (5) muscle fascia.
[0084] In a non-limiting embodiment or aspect, the CNN 300 may include a plurality of connections 380 (e.g., skip, residual, feedforward, and / or similar connections). For example, each connection 380 may connect a respective convolution block (e.g., its output) of an encoding branch (e.g., ultrasound image encoding branch 330 and / or RF encoding branch 340) to a respective up-convolution block (e.g., its input) of a plurality of up-convolution blocks of a decoding branch 360, and the size of the up-convolution block may correspond to the size of the respective convolution block. In a non-limiting embodiment or aspect, the encoded feature data (e.g., the output of each convolution block) from such a residual connection 380 may be concatenated with the input of the respective up-convolution block. In a non-limiting embodiment or aspect, each convolution block may have a connection with a respective up-convolution block, and the respective up-convolution blocks may have corresponding (e.g., matching, compatible, and / or similar) sizes.
[0085] Reference now Figure 4 , an exemplary CNN 400 (e.g., a U-NetCNN architecture) is shown according to non-limiting embodiments or aspects. For example, CNN 400 can be implemented (e.g., completely, partially, and / or similarly) by computing device 106. Additionally or alternatively, CNN 400 can be implemented (e.g., completely, partially, and / or similarly) by at least one other computing device, separate from or including computing device 106.
[0086] like Figure 4 As shown, CNN 400 may include a downsampling layer (e.g., encoding branch 430 and / or the like), a bottleneck portion 450, an upsampling layer (e.g., decoding branch 460 and / or the like), any combination thereof, and / or the like.
[0087] Continue to refer Figure 4 , the encoding branch 430 can downsample the input ultrasound data 404. For example, the ultrasound data 404 can include at least one ultrasound image, radio frequency waveform data, any combination thereof, and / or the like, as described herein. In a non-limiting embodiment or aspect, at least one ultrasound image (e.g., a single-channel ultrasound image) can be combined (e.g., concatenated and / or the like) with the RF waveform data (e.g., a single-channel RF image) corresponding to the ultrasound image to form a multi-channel input image, for example, for use as the input ultrasound data 404.
[0088] In a non-limiting embodiment or aspect, the encoding branch 430 may include a plurality of convolution blocks (e.g., a first convolution block 430a, a second convolution block 430b, a third convolution block 430c, and / or a fourth convolution block 430d). Each convolution block (e.g., 430a, 430b, 430c, and / or 430d) may include at least one convolution layer set 420. For example, each convolution block (e.g., 430a, 430b, 430c, and / or 430d) may include two convolution layer sets 420. In a non-limiting embodiment or aspect, each convolution layer set 420 may include a convolution layer (e.g., a 3x3 convolution layer and / or the like), a batch normalization layer, an activation layer (e.g., a ReLU layer and / or the like), any combination thereof, and / or the like. Additionally or alternatively, each convolutional block (e.g., 430a, 430b, 430c, and / or 430d) may include a max pooling layer 422 (e.g., a 2x2 max pooling layer and / or the like). In a non-limiting embodiment or aspect, each convolutional layer set 420 of the first convolutional block 430a may have 32 feature maps. Additionally or alternatively, the size of each convolutional layer set 420 of the first convolutional block 430a may be based on the size of the input image (e.g., an ultrasound image and / or RF image, which may have a size of 784x192 and / or the like) and / or the number of feature maps (e.g., 64), as described herein. In a non-limiting embodiment or aspect, each convolution layer set 420 of the second convolution block 430b can have a greater number (e.g., twice) of feature maps (e.g., 64 feature maps) than the feature maps of the first convolution block 430a, and / or other dimensions of the second convolution block 430b can be smaller than the dimensions of the first convolution block 430a, as described herein. Additionally or alternatively, each convolution layer set 420 of the third convolution block 430c can have a greater number (e.g., twice) of feature maps (e.g., 128 feature maps) than the feature maps of the second convolution block 430b, and / or other dimensions of the third convolution block 430c can be smaller than the dimensions of the second convolution block 430b, as described herein. Additionally or alternatively, each convolutional layer set 420 of the fourth convolutional block 430d may have a greater number (e.g., twice) of feature maps (e.g., 256 feature maps) than feature maps of the third convolutional block 430c, and / or other dimensions of the fourth convolutional block 430d may be smaller than dimensions of the third convolutional block 430c, as described herein. In a non-limiting embodiment or aspect, the activation layer of each convolutional layer set 420 in the encoding branch 430 may include a ReLU layer.
[0089] In a non-limiting embodiment or aspect, the output of the first convolution block 430a can be provided as an input to the second convolution block 430b. Additionally or alternatively, the output of the second convolution block 430b can be provided as an input to the third convolution block 430c. Additionally or alternatively, the output of the third convolution block 430c can be provided as an input to the fourth convolution block 430d. Additionally or alternatively, the output of the fourth convolution block 430d can be provided as an input to the bottleneck portion 450.
[0090] like Figure 4 As shown, the bottleneck portion 450 may include at least one convolutional layer set 452 (e.g., two convolutional layer sets 452), as described herein. For example, each convolutional layer set 452 may be a convolutional layer (e.g., a 3x3 convolutional layer and / or the like), a batch normalization layer, an activation layer (e.g., a ReLU layer and / or the like), any combination thereof, and / or the like, as described herein. In a non-limiting embodiment or aspect, the activation layer of each convolutional layer set 452 may include a ReLU layer. In a non-limiting embodiment or aspect, each convolutional layer set 452 may have a greater number (e.g., twice) of feature maps (e.g., 512 feature maps) than the feature maps of the fourth convolutional block 403d, and / or other dimensions of each convolutional layer set 452 may be smaller than the dimensions of the fourth convolutional block 430d, as described herein.
[0091] In a non-limiting embodiment or aspect, the output of the fourth convolutional block 430d of the encoding branch 430 may be provided as an input to the bottleneck portion 450. Additionally or alternatively, the output of the bottleneck portion 450 may be provided as an input to the decoding branch 460.
[0092] Continue to refer Figure 4, the decoding branch 460 may upsample its input (e.g., the output from the bottleneck portion 450 and / or the like). In a non-limiting embodiment or aspect, the decoding branch 460 may include a plurality of up-convolution blocks (e.g., a first up-convolution block 460a, a second up-convolution block 460b, a third up-convolution block 460c, and / or a fourth up-convolution block 460d). Each up-convolution block (e.g., 460a, 460b, 460c, and / or 460d) may include at least one up-convolution layer 462 (e.g., a transposed convolution layer and / or the like) and / or at least one convolution layer set 464. For example, each up-convolution block (e.g., 460a, 460b, 460c, and / or 460d) may include one up-convolution layer 462 and / or two convolution layer sets 464 (e.g., an up-convolution layer 462, a first convolution layer set 464, and a second convolution layer set 464, in sequence). In a non-limiting embodiment or aspect, each convolution layer set 464 may include a convolution layer (e.g., a 3x3 convolution layer and / or the like), a batch normalization layer, an activation layer (e.g., a ReLU layer and / or the like), any combination thereof, and / or the like. In a non-limiting embodiment or aspect, each upper convolution layer 462 of the first upper convolution block 460a and / or each convolution layer set 464 may have 256 feature maps, and / or the size of each upper convolution layer 462 of the first upper convolution block 460a and / or each convolution layer set 464 may be based on the size of the fourth convolution block 430d of the encoding branch 430, as described herein. Additionally or alternatively, each upper convolution layer 462 and / or each convolution layer set 464 of the second upper convolution block 460b may have a smaller number (e.g., half) of feature maps (e.g., 128 feature maps) than the feature maps of the first upper convolution block 460a, and / or other sizes of the second upper convolution block 460b may be based on the size of the third convolution block 430c of the encoding branch 430. Additionally or alternatively, each upper convolution layer 462 and / or each convolution layer set 464 of the third upper convolution block 460c may have a smaller number (e.g., half) of feature maps (e.g., 64 feature maps) than the feature maps of the second upper convolution block 460b, and / or other sizes of the third upper convolution block 460c may be based on the size of the second convolution block 430b of the encoding branch 430. Additionally or alternatively, each up-convolution layer 462 and / or each convolution layer set 464 of the fourth up-convolution block 460d may have a smaller number (e.g., half) of feature maps (e.g., 32 feature maps) than the feature maps of the third up-convolution block 460c, and / or other sizes of the fourth up-convolution block 460d may be based on the size of the first convolution block 430a of the encoding branch 430.
[0093] In a non-limiting embodiment or aspect, the output of the first up-convolution block 460a may be provided as an input to the second up-convolution block 460b. Additionally or alternatively, the output of the second up-convolution block 460b may be provided as an input to the third up-convolution block 460c. Additionally or alternatively, the output of the third up-convolution block 460c may be provided as an input to the fourth up-convolution block 460d. Additionally or alternatively, the output of the fourth up-convolution block 460d may be provided as an input to the output layer set 470.
[0094] In a non-limiting embodiment or aspect, the output layer set 470 may include at least one convolution layer and / or at least one activation layer. For example, the output layer set 470 may include a 1x1 convolution layer and an activation layer (e.g., a softmax layer and / or the like). In a non-limiting embodiment or aspect, the size of the output layer set 470 may be based on the size of the fourth upper convolution block 460d of the decoding branch 460 and / or the size of the input ultrasound data (e.g., ultrasound image and / or RF image). For example, the size of the output layer set 470 may be 784x192. In a non-limiting embodiment or aspect, the activation layer of the output layer set 470 may include a classification layer. For example, the activation layer (e.g., a classification layer) may assign a classification index (e.g., an integer class label and / or the like) to each pixel of the ultrasound to provide semantic segmentation (e.g., a label map and / or the like).
[0095] In a non-limiting embodiment or aspect, the CNN 400 may include a plurality of feature forwarding connections 480 (e.g., skip connections, residual connections, and / or the like). For example, each residual connection 480 may connect a respective convolution block (e.g., its output) of the encoding branch 430 to a respective up-convolution block (e.g., its input) of the decoding branch 460, and its size may have a size corresponding to that of the respective convolution block. In a non-limiting embodiment or aspect, the encoded feature data (e.g., the output of the respective convolution block) from such a residual connection 480 may be concatenated with the input at the respective up-convolution block.
[0096] In a non-limiting embodiment or aspect, the classification output of the CNN 400 can be optimized during training. For example, a cross entropy loss can be used as an objective function for training the CNN 400 before the activation layer (e.g., the final softmax layer and / or the like) of the output layer set 470 to seek a large numerical separation between the maximum (e.g., final output) score for each pixel and the non-maximum responses for other classes (e.g., this can create a more robust, more general CNN).
[0097] In a non-limiting embodiment or aspect, CNN 400 may be similar to the CNN described in Ronneberger et al., “U-net: Convolutional Networks for Biomedical Image Segmentation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, 234-241 (2015), the disclosure of which is incorporated herein by reference in its entirety.
[0098] Reference now Figure 5 , an exemplary CNN 500 (e.g., a SegNetCNN architecture) is shown according to non-limiting embodiments or aspects. For example, CNN 500 can be implemented (e.g., completely, partially, and / or similarly) by computing device 106. Additionally or alternatively, CNN 500 can be implemented (e.g., completely, partially, and / or similarly) by at least one other computing device, separate from or including computing device 106.
[0099] like Figure 5 As shown, CNN 500 may include downsampling layers (e.g., encoding branch 530 and / or the like), upsampling layers (e.g., decoding branch 560 and / or the like), any combination thereof, and / or the like.
[0100] Continue to refer Figure 5 , the encoding branch 530 can downsample the input ultrasound data. For example, the ultrasound data can include at least one ultrasound image, RF waveform data, any combination thereof, and / or the like, as described herein. In a non-limiting embodiment or aspect, at least one ultrasound image (e.g., a single-channel ultrasound image) can be combined (e.g., concatenated and / or the like) with RF waveform data (e.g., a single-channel RF image) corresponding to the ultrasound image to form a multi-channel input image, for example, for use as input ultrasound data.
[0101] In a non-limiting embodiment or aspect, the encoding branch 530 may include a plurality of convolution blocks (e.g., a first convolution block 530a, a second convolution block 530b, a third convolution block 530c, a fourth convolution block 530d, and / or a fifth convolution block 530e). Each convolution block (e.g., 530a, 530b, 530c, 530d, and / or 530e) may include at least one convolution layer set 520. For example, each convolution block (e.g., 530a, 530b, 530c, and / or 530d) may include two or three convolution layer sets 520 (e.g., the first convolution block 530a and the second convolution block 530b may each include two convolution layer sets 520, and the third convolution block 530c, the fourth convolution block 530d, and the fifth convolution block 530e may each include three convolution layer sets 520). In non-limiting embodiments or aspects, each convolution layer set 520 may include a convolution layer, a batch normalization layer, an activation layer (e.g., a ReLU layer and / or the like), any combination thereof, and / or the like. Additionally or alternatively, each convolution block (e.g., 530a, 530b, 530c, 530d, and / or 530e) may include a pooling layer 522 (e.g., a max pooling layer and / or the like).
[0102] In a non-limiting embodiment or aspect, the output of the first convolution block 530a can be provided as an input to the second convolution block 530b. Additionally or alternatively, the output of the second convolution block 530b can be provided as an input to the third convolution block 530c. Additionally or alternatively, the output of the third convolution block 530c can be provided as an input to the fourth convolution block 530d. Additionally or alternatively, the output of the fourth convolution block 530d can be provided as an input to the fifth convolution block 530e. Additionally or alternatively, the output of the fifth convolution block 530e can be provided as an input to the decoding branch 560.
[0103] Continue to refer Figure 5, the decoding branch 560 can upsample its input (e.g., output from the encoding branch 530 and / or the like). In a non-limiting embodiment or aspect, the decoding branch 560 can include a plurality of upsampling blocks (e.g., a first upsampling block 560a, a second upsampling block 560b, a third upsampling block 560c, a fourth upsampling block 560d, and / or a fifth upsampling block 560e). Each upsampling block (e.g., 560a, 560b, 560c, 560d, and / or 560e) can include at least one upsampling layer 562 and / or at least one convolutional layer set 564. For example, each upsampling block (e.g., 560a, 560b, 560c, 560d, and / or 560e) may include one upsampling layer 562 and two or three convolutional layer sets 564 (e.g., the first upsampling block 560a, the second upsampling block 560b, and the third upsampling block 560c may each include three convolutional layer sets 564, and the fourth upsampling block 560d and the fifth upsampling block 560e may each include three convolutional layer sets 520). In a non-limiting embodiment or aspect, each upsampling layer may upsample its input using the maximum pooling index captured and stored in the encoding branch 530 step. In a non-limiting embodiment or aspect, each convolutional layer set 564 may include a convolutional layer, a batch normalization layer, an activation layer (e.g., a ReLU layer and / or the like), any combination thereof, and / or the like.
[0104] In a non-limiting embodiment or aspect, the output of the first upsampling block 560a may be provided as an input to the second upsampling block 560b. Additionally or alternatively, the output of the second upsampling block 560b may be provided as an input to the third upsampling block 560c. Additionally or alternatively, the output of the third upsampling block 560c may be provided as an input to the fourth upsampling block 560d. Additionally or alternatively, the output of the fourth upsampling block 560d may be provided as an input to the fifth upsampling block 560e. Additionally or alternatively, the output of the fifth upsampling block 560e may be provided as an input to the output layer set 570.
[0105] In a non-limiting embodiment or aspect, the output layer set 570 may include at least one activation layer. For example, the activation layer may include a softmax layer. In a non-limiting embodiment or aspect, the size of the output layer set 570 may be based on the size of the input ultrasound data (e.g., ultrasound image and / or RF image). For example, the size of the output layer set 570 may be 784x192. In a non-limiting embodiment or aspect, the activation layer of the output layer set 570 may include a classification layer. For example, the activation layer (e.g., classification layer) may assign a classification index (e.g., integer class label and / or the like) to each pixel of the ultrasound to provide semantic segmentation (e.g., label map and / or the like).
[0106] In a non-limiting embodiment or aspect, the CNN 500 may include a plurality of feature forwarding connections 580 (e.g., skip connections, residual connections, and / or similar connections). For example, each connection 580 may connect a respective convolution block (e.g., its output) of the encoding branch 530 to a respective upsampling block (e.g., its input) of the decoding branch 560, which may have a size corresponding to the size of the respective convolution block. In a non-limiting embodiment or aspect, the encoded feature data (e.g., the output of the respective convolution block) from such a residual connection 580 may be concatenated with the input at the respective upsampling block.
[0107] In a non-limiting embodiment or aspect, the CNN 500 can be similar to the CNN described in Badrinarayanan et al., “Segnet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation,” 39 IEEE Symposium on Pattern Analysis and Machine Intelligence, 2481-2495 (2017), the disclosure of which is incorporated herein by reference in its entirety.
[0108] Reference now Fig. 6A , shown in accordance with a non-limiting embodiment or aspect Figure 3Graphs of average loss and mean intersection over union (mIoU) metrics for various tissue classes of an exemplary CNN 300 versus training time. The first curve 691 shows an example of the value of the loss function averaged over tissue types as it generally improves (gets smaller) during the training of the CNN 300. The second curve 692 shows an example of the value of the mIoU metric averaged over tissue types as it generally improves (gets closer to 1.0) during the training of the CNN. The third curve 693 shows an example of the mIoU of a CNN 300 for segmenting skin tissue as it generally improves during training. The fourth curve 694 shows an example of the mIoU of a CNN 300 for segmenting fat fascia tissue as it generally improves during training. The fifth curve 695 shows an example of the mIoU of a CNN 300 for segmenting fat tissue as it generally improves during the initial training process (in this particular example, the fat mIoU 395 ends up getting a bit worse as the fat fascia / breast mIoU 394 improves). The sixth curve 696 shows one embodiment of the mIoU of the CNN 300 for segmenting muscle fascia tissue as it generally improves during training. The seventh curve 697 shows one embodiment of the mIoU of the CNN 300 for segmenting muscle tissue as it generally improves during initial training (in this particular embodiment, the muscle mIoU 397 ends up being somewhat worse as the muscle fascia mIoU 396 improves). Each of these embodiments is illustrative and individual training sessions will hopefully have different values / curves for each of them.
[0109] Reference now Figure 6B, shown are exemplary test input images, corresponding labeled images, and output segmented ultrasound images of various exemplary CNNs according to non-limiting embodiments or aspects. The first column 601 shows four exemplary ultrasound images (e.g., test input ultrasound images), and the second column 602 shows four exemplary RF images (e.g., test input RF images, using color maps to display both positive and negative waveform values), corresponding to ultrasound images, respectively. The third column 603 shows four exemplary labeled images (e.g., labeled by clinicians and / or similar personnel based on four exemplary ultrasound images and / or RF images, respectively). The fourth column 604 shows four exemplary output segmented ultrasound images from an exemplary CNN with a U-Net architecture based on four ultrasound images (e.g., without using four RF images), and the fifth column 605 shows four exemplary output segmented ultrasound images from an exemplary CNN with a U-Net architecture based on four ultrasound images and four RF images, respectively. The sixth column 606 shows four exemplary output segmented ultrasound images from an exemplary CNN with an attention U-Net (AU-Net) architecture (e.g., the same or similar CNN architecture as described in Oktay et al., "Attention U-Net: Studying Where to Find the Pancreas" (First Medical Imaging Conference on Deep Learning (MIDL) (2018)), the entire disclosure of which is incorporated herein by reference), and the seventh column 607 shows four exemplary output segmented ultrasound images from an exemplary CNN with an AU-Net architecture based on four ultrasound images and four RF images, respectively. The eighth column 608 shows four exemplary output segmented ultrasound images from an exemplary CNN with a W-Net architecture based on four ultrasound images and four RF images, respectively.
[0110] For illustrative purposes, Table 1 shows examples of pixel-level accuracy and mIoU for various exemplary CNNs (due to the stochastic nature of training, different training sessions can be expected to have slightly different and / or more divergent values).
[0111] Table 1
[0112]
[0113] Reference now Figure 6C, shown are exemplary test input images, corresponding labeled images, and output segmented ultrasound images of various exemplary CNNs according to non-limiting embodiments or aspects. The first row 611 shows five exemplary ultrasound images (e.g., test input ultrasound images), and the second row 612 shows five exemplary RF images (e.g., test input RF images, using color maps to display both positive and negative waveform values), corresponding to ultrasound images, respectively. The third row 613 shows five exemplary labeled images (e.g., labeled by clinicians and / or similar personnel based on five exemplary ultrasound images and / or RF images, respectively). The fourth row 614 shows five exemplary output segmented ultrasound images from an exemplary CNN of a U-Net architecture based on five ultrasound images (e.g., without using five RF images), and the fifth row 615 shows five exemplary output segmented ultrasound images from an exemplary CNN of a SegNet architecture based on five ultrasound images (e.g., without using five RF images). The sixth row 616 shows five exemplary output segmented ultrasound images from an exemplary CNN with a U-Net architecture based on five ultrasound images and five RF images, respectively, and the seventh row 617 shows five exemplary output segmented ultrasound images from an exemplary CNN with a SegNet architecture based on five ultrasound images and five RF images, respectively.
[0114] For illustrative purposes, Table 2 shows examples of pixel-level accuracy and mIoU for various exemplary CNNs (due to the stochastic nature of training, different training sessions can be expected to have slightly different and / or more divergent values).
[0115] Table 2
[0116]
[0117] Reference now Figure 7 , shown is a method 700 for marking ultrasound data according to a non-limiting embodiment. It can be understood that Figure 7 The order of steps shown in is for illustrative purposes only, and non-limiting embodiments may involve more steps, fewer steps, different steps, and / or a different order of steps. In addition, Figure 7The embodiments shown in the drawings relate to ultrasound data, but as explained herein, the systems and methods disclosed herein may be used in many other situations. In non-limiting embodiments or aspects, one or more steps of method 700 may be performed (e.g., in full, in part, and / or similarly) by computing device 106. In non-limiting embodiments or aspects, one or more steps of method 700 may be performed (e.g., in full, in part, and / or similarly) by another system, another device, another group of systems, or another group of devices, separate from or including computing device 106, such as ultrasound / RF system 102, at least one other computing device, and / or the like.
[0118] like Figure 7 As shown, at step 702, the method 700 may include training a CNN, as described herein. For example, the computing device 106 may train a CNN (e.g., W-Net, U-Net, AU-Net, SegNet, any combination thereof, and / or the like) based on ultrasound data 104, which may include at least one ultrasound image 104a, RF waveform data 104b (e.g., at least one RF image), any combination thereof, and / or the like. In a non-limiting embodiment or aspect, the computing device 106 may receive (e.g., retrieve, request, obtain, and / or the like) ultrasound data for training the CNN from at least one of the ultrasound / RF system 102 and / or the database 108. For example, the computing device 106 may train a CNN based on historical ultrasound data (e.g., historical ultrasound images 104a, RF waveform data 104b, labeled ultrasound images 104c, any combination thereof, and / or the like) from the database 108, as described herein.
[0119] like Figure 7 As shown, at step 704, the method 700 may include downsampling the RF input (e.g., RF waveform data 104b and / or the like) of each of the plurality of downsampling layers in the CNN, as described herein. For example, the computing device 106 may downsample the RF input of the downsampling layer (e.g., encoding branch, RF encoding branch, and / or the like, as described herein). In a non-limiting embodiment or aspect, the RF input may include RF waveform data 104b for ultrasound received from at least one of the ultrasound / RF system 102 and / or the database 108, as described herein.
[0120] like Figure 7As shown, at step 706, the method 700 may include downsampling the image input (e.g., ultrasound image 104a and / or the like) of each of the multiple downsampling layers in the CNN, as described herein. For example, the computing device 106 may downsample the image input of the downsampling layer (e.g., the encoding branch, the ultrasound image encoding branch, and / or the like, as described herein). In a non-limiting embodiment or aspect, the image input may include at least one ultrasound image 104a for ultrasound received from at least one of the ultrasound / RF system 102 and / or the database 108, as described herein. Additionally or alternatively, the image input may include multiple pixels of the ultrasound, as described herein.
[0121] In a non-limiting embodiment or aspect, as described herein, the image input and the RF input are processed substantially simultaneously.
[0122] like Figure 7 As shown, at step 706, method 700 may include segmenting the tissue in the ultrasound based on the output of the CNN, as described herein. For example, the computing device 106 may segment the tissue in the ultrasound (e.g., labeling the pixels identified therein) based on the output of the CNN, as described herein. In a non-limiting embodiment or aspect, segmenting the tissue in the ultrasound may include labeling a plurality of (e.g., most, all, and / or similar) pixels in the ultrasound, as described herein. In a non-limiting embodiment or aspect, segmenting the tissue may include identifying at least one of: muscle, fascia, fat fascia, muscle fascia, fat, transplanted fat, any combination thereof, and / or the like, as described herein.
[0123] Reference now Figure 8 , which is a method 800 for labeling ultrasound data according to a non-limiting embodiment. It is understood that Figure 8 The order of steps shown in is for illustrative purposes only, and non-limiting embodiments may involve more steps, fewer steps, different steps, and / or a different order of steps. Figure 8 The embodiments shown in relate to ultrasound data, but as explained herein, the systems and methods disclosed herein may be used in many other situations. In non-limiting embodiments or aspects, one or more steps of method 800 may be performed (e.g., completely, partially, and / or similarly) by computing device 106. In non-limiting embodiments or aspects, one or more steps of method 800 may be performed (e.g., completely, partially, and / or similarly) by another system, another device, another group of systems, or another group of devices, separate from or including computing device 106 (e.g., ultrasound / RF system 102, at least one other computing device, and / or the like).
[0124] like Figure 8 As shown, at step 802, method 800 may include receiving an ultrasound image represented by a plurality of pixels. For example, computing device 106 may receive ultrasound image 104a from at least one of ultrasound / RF system 102 and / or database 108, as described herein.
[0125] like Figure 8 As shown, at step 804, the method 800 may include segmenting the ultrasound image by marking a majority of the plurality of pixels. For example, the computing device 106 may segment the ultrasound image 104a by marking a majority of its pixels, as described herein. In a non-limiting embodiment or aspect, the majority of pixels may be labeled as at least one of: muscle, fascia, fat fascia, muscle fascia, fat, transplanted fat, any combination thereof, and / or the like, as described herein. In a non-limiting embodiment or aspect, the computing device 106 may segment the ultrasound image 104a based on a CNN (e.g., an output generated based on using the ultrasound image 104a as input), as described herein. In a non-limiting embodiment or aspect, the computing device 106 may train a CNN based on ultrasound data 104 (e.g., historical ultrasound data from a database 108, ultrasound images 104a received from an ultrasound / RF system 102, and / or the like), as described herein. Additionally or alternatively, at least one input ultrasound image in the ultrasound data 104 (eg, the labeled ultrasound image 104 c and / or the like) may include blurred overlapping labels for a plurality of pixels.
[0126] Although the embodiments have been described in detail for the purpose of illustration, it should be understood that such details are for that purpose only, and the present disclosure is not limited to the disclosed embodiments, but is intended to cover modifications and equivalent arrangements within the spirit and scope of the appended claims. For example, it should be understood that the present disclosure contemplates that, to the extent possible, one or more features of any embodiment can be combined with one or more features of any other embodiment.
Claims
1. A method for labeling ultrasound data, include: Training a convolutional neural network (CNN) based on ultrasound data, wherein the ultrasound data includes radio frequency (RF) waveform data; downsampling an RF input of each of a plurality of downsampling layers in the CNN, the RF input comprising RF waveform data for ultrasound; wherein the plurality of downsampling layers comprises an ultrasound image encoding branch and a plurality of radio frequency encoding branches, each radio frequency encoding branch comprises a respective kernel size different from other radio frequency encoding branches in the plurality of radio frequency encoding branches, the respective kernel size of each radio frequency encoding branch corresponds to a respective wavelength, Each RF coding branch includes a plurality of convolution blocks, each convolution block includes a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer and a second activation layer, and at least one convolution block in the plurality of convolution blocks includes a maximum pooling layer, and wherein downsampling comprises downsampling the RF input of each of the plurality of RF encoding branches in the CNN and downsampling the image input of each ultrasound image encoding branch in the CNN, the image input comprising a plurality of pixels of the ultrasound; as well as Tissue in ultrasound is segmented based on the output of the CNN.
2. The method according to claim 1, further comprising: include: An image input of each of a plurality of downsampling layers in the CNN is downsampled, the image input comprising a plurality of pixels of the ultrasound.
3. The method of claim 2, wherein the image input and the RF input are processed simultaneously. The method of claim 1 , wherein segmenting the tissue in the ultrasound comprises marking a plurality of pixels. The method of claim 4 , wherein the plurality of pixels comprises a majority of pixels in the ultrasound. The method of claim 1 , wherein segmenting tissue comprises identifying at least one of: muscle, fascia, fat, or any combination thereof.
7. The method of claim 6, wherein the fat comprises transplanted fat.
8. The method according to claim 1, further comprising: include: concatenating the RF encoding branch output of each RF encoding branch and the ultrasound image encoding branch output of the ultrasound image encoding branch to provide a concatenated encoding branch output; as well as The concatenated encoding branch outputs are upsampled using multiple upsampling layers in the CNN.
9. The method of claim 8, wherein the plurality of upsampling layers comprises a decoding branch, the decoding branch comprising a plurality of upconvolution blocks, The CNN further comprises a plurality of residual connections, each residual connection connecting a respective convolution block in the plurality of convolution blocks to a corresponding up-convolution block in the plurality of up-convolution blocks, the up-convolution blocks having a size corresponding to the respective convolution block.
10. A system for labeling ultrasound data, comprising at least one computing device programmed or configured to: Training a convolutional neural network (CNN) based on ultrasound data, the ultrasound data including radio frequency (RF) waveform data; downsampling an RF input of each of a plurality of downsampling layers in the CNN, the RF input comprising RF waveform data for ultrasound; wherein the plurality of downsampling layers comprises an ultrasound image encoding branch and a plurality of radio frequency encoding branches, each radio frequency encoding branch comprises a respective kernel size different from other radio frequency encoding branches in the plurality of radio frequency encoding branches, the respective kernel size of each radio frequency encoding branch corresponds to a respective wavelength, Each RF coding branch includes a plurality of convolution blocks, each convolution block includes a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer and a second activation layer, and at least one convolution block in the plurality of convolution blocks includes a maximum pooling layer, and wherein downsampling comprises downsampling the RF input of each of the plurality of RF encoding branches in the CNN and downsampling the image input of each ultrasound image encoding branch in the CNN, the image input comprising a plurality of pixels of the ultrasound; and Tissue in ultrasound is segmented based on the output of the CNN.
11. The system of claim 10, wherein the computing device is further programmed or configured to downsample an image input of each of a plurality of downsampling layers in the CNN, the image input comprising a plurality of pixels of the ultrasound.
12. The system of claim 10, wherein the image input and the radio frequency input are processed simultaneously.
13. The system of claim 10, wherein segmenting the tissue in the ultrasound comprises marking a plurality of pixels.
14. The system of claim 13, wherein the plurality of pixels comprises a majority of pixels in the ultrasound.
15. The system of claim 10, wherein segmenting tissue comprises identifying at least one of: muscle, fascia, fat, or any combination thereof.
16. The system of claim 15, wherein the fat comprises transplanted fat.
Citation Information
Patent Citations
Ultrasound imaging system with a neural network for image formation and tissue characterization
WO2018127498A1
Probability map-based ultrasound scanning
WO2018209193A1