Method and system for generating synthetic elastography images

A neural network converts B-mode ultrasound images into synthetic elastography images, addressing the computational and hardware requirements of SWE, enabling efficient and accurate elasticity mapping on conventional ultrasound systems.

JP7752532B2Active Publication Date: 2025-10-10KONINKLIJKE PHILIPS NV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021573494
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-20
Filing Date
2020-06-10
Publication Date
2025-10-10
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

Conventional ultrasound elastography techniques, such as shear wave elastography (SWE), are computationally demanding and require specialized hardware, making them unavailable in many ultrasound systems and sensitive to motion artifacts.

Method used

A method using a trained artificial neural network to generate synthetic elastography images from conventional B-mode ultrasound images, eliminating the need for specialized hardware and ultrafast imaging, by recognizing tissue texture information linked to mechanical properties and converting it into elasticity maps.

Benefits of technology

Enables the generation of SWE-like images at the frame rate of B-mode ultrasound, reducing computational intensity and allowing use on conventional scanners, with high correlation to SWE images and retention of diagnostic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752532000002
    Figure 0007752532000002
  • Figure 0007752532000003
    Figure 0007752532000003
  • Figure 0007752532000004
    Figure 0007752532000004
Patent Text Reader

Abstract

The present invention relates to a method for generating a synthetic elastography image, comprising the steps of: (a) receiving a B-mode ultrasound image of a region of interest; and (b) generating a synthetic elastography image of the region of interest by applying the trained artificial neural network to the B-mode ultrasound image. The present invention also relates to a method for training an artificial neural network useful for generating synthetic elastography images, and related computer programs and systems.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a computer-implemented method for generating synthetic elastography images, a method for training an artificial neural network useful in generating synthetic elastography images from B-mode ultrasound images, a method for training an artificial neural network or a second artificial neural network to provide a confidence map, and related computer programs and systems. [Background technology]

[0002] Tissue elasticity or stiffness is an important marker in biomedicine, for example, in cancer imaging or the evaluation of musculoskeletal pathologies. Whether a tissue is stiff or soft can provide important diagnostic information. For example, cancerous tumors are often stiffer than surrounding tissue, and liver stiffness typically indicates various disease states, including cirrhosis and hepatitis. To image tissue elastic properties, it is necessary to record how the tissue behaves when deformed. Ultrasound elastography typically uses the radiation force of focused ultrasound pulses to remotely create a "push" within the tissue. For example, shear wave elastography (SWE) is an advanced technique that enables local elasticity estimation and the generation of 2D elastograms. In SWE, a "push" is induced deep within the tissue by a sequence of acoustic radiation forces. The disturbances created by this "push" travel laterally (sideways) through the tissue as shear waves. Using ultrafast ultrasound imaging, the stiffness of intervening tissue can be inferred by observing how quickly the shear waves travel to different lateral positions, generating elastography images, also known as elastograms. However, to assess shear wave velocity, either stroboscopic techniques or ultrafast imaging using frame rates on the order of 1000 Hz must be employed. These recordings must then be processed to generate 2D elastographic maps.

[0003] SWE images are known to be sensitive to ultrasound probe pressure, motion artifacts, and region of interest (ROI) location, among other things. Furthermore, SWE requires specialized ultrasound (US) transducers capable of generating high acoustic radiation force "push" pulses immediately followed by high-frame-rate images. Therefore, SWE is not readily available in many ultrasound systems, which are computationally demanding and have low frame rates. Therefore, it would be advantageous to have a robust technique for acquiring elasticity images without requiring sophisticated acquisition strategies or specialized hardware components.

[0004] Conventional B-mode ultrasound produces 2D images with fast acquisition and image reconstruction times.

[0005] US 2018 / 0132830 A1 discloses a method for reconstructing displacement and strain maps of human tissue in a subject in vivo. The method includes applying a mechanical deformation to a target region, imaging the tissue while the deformation is applied to the target region, and measuring axial and lateral displacements and strains in the target region to thereby distinguish tissues with different stiffness and strain responses. In particular, this document discloses the reconstruction of displacement and strain images from conventional ultrasound B-mode images. Therefore, the described method requires the application of a mechanical "push," for example, performed by a user's hand.

[0006] In "A Deep Learning Framework for One-Sided Sound Speed ​​Inversion in Medical Ultrasound," M. Feigin et al., a paper submitted to IEEE on December 10, 2018, an alternative approach to shear wave imaging is used. This approach employs one-sided pressure wave velocity measurements from a conventional ultrasound probe. The method uses raw ultrasound channel data and generates a corresponding tissue sound speed map using a fully convolutional deep neural network. Summary of the Invention [Problem to be solved by the invention]

[0007] It is therefore an object of the present invention to provide a method (and associated computer program and system) that is faster, less computationally intensive, and does not require a specialized ultrasound probe. [Means for solving the problem]

[0008] This object is met or exceeded by a computer-implemented method for generating a synthetic elastography image as claimed in claim 1, a method for training an artificial neural network useful for generating a synthetic elastography image as claimed in claim 10, a method for training a second artificial neural network useful for providing a confidence map as claimed in claim 12, a computer program as claimed in claim 13, and a system as claimed in claim 14. Advantageous embodiments are set out in the dependent claims. Any feature, advantage, or alternative embodiment described herein in relation to the claimed method may also be applied to other claim categories, in particular the claimed system, computer program, and vice versa. In particular, the trained artificial neural network may be provided or improved by the training method as claimed. Furthermore, input / output data to the trained artificial neural network may comprise advantageous features and embodiments of the input / output training data, and vice versa.

[0009] According to a first aspect, the present invention provides a method for generating a synthetic elastography image, the method comprising the steps of: (a) receiving a B-mode ultrasound image of a region of interest; and (b) generating a synthetic elastography image of the region of interest by applying a trained artificial neural network to the B-mode ultrasound image. Preferably, the method is performed by a computing unit to which the B-mode ultrasound image is received via a first interface. Optionally, the output of the trained artificial neural network (NN), i.e., the synthetic elastography image, is output, e.g., via a second interface. The method may further comprise the step of displaying the synthetic elastography image, e.g., side-by-side or overlaid, with the B-mode ultrasound image that forms the input to the NN.

[0010] Thus, the present invention provides a method for generating elastograms from a single B-mode image, rather than from a series of B-mode images tracking mechanical deformation or from displacement maps estimated from raw ultrasound channel data (high-frequency data). Rather, any B-mode image can serve as input for the trained artificial neural network (NN) of the present invention. Unlike prior art methods, the present invention does not calculate displacement to estimate elasticity, but generates elastograms based on specific B-mode characteristics. The present invention recognizes that while B-mode ultrasound measures tissue echogenicity rather than elasticity, B-mode ultrasound images can also carry texture information linked to underlying tissue structure and mechanical properties. The NN is trained to recognize this and convert that information into an elasticity map (synthetic elasticity image) that is at least similar to elastography images that can currently be generated using ultrasound elastography techniques, such as shear wave elastography protocols.

[0011] Thus, the present invention provides a deep learning solution that enables the robust synthesis of SWE-like images from simple conventional B-mode acquisitions. The present invention thus allows for the generation of synthetic elastography images at the frame rate of B-mode ultrasound. Therefore, the present invention can be used to add or replace shear wave elastography in conventional scanners, mitigating the heavy system requirements and operator sensitivity still exhibited by conventional ultrasound elastography techniques such as SWE. Furthermore, the method of the present invention can be performed on ultrasound scanners that are not fully equipped for SWE itself. Another application could be retrospective elastography, in which synthetic elastography images are generated from previously recorded B-mode images. This could be used for further tissue typing at a later stage in the diagnostic process.

[0012] The computing unit executing the method can be any data processing unit, such as a graphics processing unit (GPU), a central processing unit (CPU), or other digital data processing device. Preferably, the method is performed by a computing unit that is part of or connected to an ultrasound scanner used for diagnostic imaging. In a useful embodiment, the output of the neural network, i.e., the composite elasticity image, is output via an interface of the computing unit and may be displayed to a user, for example, on a computer screen, a mobile device screen, a television set, or other display device. The method can also include acquiring a B-mode ultrasound image of the region of interest, which can be provided from a data storage device, either locally or remotely. Because the method of the present invention is very fast, it can be performed in real time immediately after acquisition of the B-mode ultrasound image, thereby enabling direct evaluation of the elastography image.

[0013] The B-mode ultrasound image may be any image in brightness mode, i.e., the intensity of the ultrasound echo is depicted as brightness. It may be acquired by any known method. The B-mode ultrasound image is preferably a two-dimensional (2D) image, for example, represented by a 2D matrix of grayscale values. This method can also be applied to a stack of 2D images that together cover a volume of interest. However, this method is also applicable to three-dimensional (3D) B-mode ultrasound images, or even one-dimensional (1D) images. The B-mode ultrasound image may be provided in any format, for example, the DICOM (Digital Imaging and Communications in Medicine) standard.

[0014] An artificial neural network (NN) is based on a collection of connected artificial neurons, also called nodes, where each connection (also called an edge) can transmit a signal from one node to another. Each artificial neuron that receives a signal can process it and forward it to further artificial neurons connected to it. In a useful embodiment, the artificial neurons of the NN of the present invention are arranged in layers. The input signal, i.e., the pixel values ​​of a B-mode ultrasound image, travels from the first layer, also called the input layer, to the last layer, the output layer.

[0015] In a useful embodiment, the NN is a feedforward network. The neural network preferably includes several layers, including a hidden layer, and is therefore preferably a deep network. In one embodiment, the NN is trained based on machine learning techniques, particularly deep learning, for example, by backpropagation. The training data used may be elastography images generated by any elastography technique, preferably ultrasound elastography. The synthetic elastography images generated by the NN are similar to the images generated by ultrasound elastography, which are used as output training data.

[0016] Although the NN is trained on SWE in a preferred embodiment, other ultrasound elastography images can also be used for training. For example, the NN may be trained on quasi-static elastography images. In this technique, external compression is applied to the tissue, and ultrasound images before and after compression are compared. The least deformed areas of the image are the most stiff, and the most deformed areas are the stiffest. Elastography images are images of relative strain (distortion). The NN may also be trained on acoustic radiation force impulse images (ARFI). This technique uses acoustic radiation force from a focused ultrasound beam to create a "push" within the tissue. The amount of tissue depression along the axis of the beam reflects the tissue's stiffness. By pushing at many different locations, a map of the tissue's stiffness is constructed. The NN may also be trained on elastography images obtained by SWE-based ultrasound shear imaging (SSI). This technique uses acoustic radiation force to induce a "push" within the tissue of interest, and tissue stiffness is calculated from how fast the resulting shear wave travels through the tissue. Local tissue velocity maps are obtained using special tracking techniques, providing a complete movie of shear wave propagation through tissue. In a further embodiment, the NN of the present invention may also be trained on magnetic resonance elastography images (MRE). In MRE, a mechanical vibrator is used on the surface of a patient's body, and an imaging acquisition sequence that measures the velocity of shear waves traveling into the patient's deeper tissues is used to infer tissue stiffness.

[0017] Whatever the training data used, the synthetic elastography images are abbreviated herein as sSWE images.

[0018] The trained artificial neural network according to the present invention may be provided in the form of a software program, but may also be implemented as hardware. Furthermore, the trained NN may be provided in the form of a trained function, which is not necessarily structured in exactly the same way as the trained neural network. For example, if a connection / edge has a weight of 0 after training, such a connection may be omitted when providing a trained function based on the trained NN.

[0019] The synthetic elastography images generated by the neural network of the present invention have been shown to be highly correlated with SWE images obtained by shear wave elastography imaging. Therefore, the method of the present invention allows for the generation of synthetic elastography images without the need for specialized hardware and with less computational effort than conventional shear wave elastography images. Because there is no ultrafast imaging scheme, no acoustic radiation force "push" pulses, and no sophisticated sequences are required to generate 2D elastograms, the generation of synthetic elastography images (sSWE) is much faster and less computationally intensive than SWE itself, and can be performed with conventional ultrasound probes that are compatible with B-mode ultrasound imaging but not SWE. The method of the present invention does not require a separate ultrasound probe. The synthetic elastography images generated by the neural network of the present invention can be used for diagnostic purposes. They retain diagnostic information in the form of elasticity values.

[0020] According to one embodiment, the input to the trained artificial neural network, i.e., a B-mode ultrasound image, has the same size and dimensions as the output of the trained NN, i.e., a synthetic elastography image. In this context, size refers to the size of the matrix of data points and the dimensionality of the input, e.g., 2D or 3D. In other words, the trained NN preferably acts as an end-to-end mapping function that converts a B-mode image into an elastography-like image of the same size and dimensions. Preferably, the input to the NN is a stack of 2D or 2D B-mode images, and the output is a stack of 2D or 2D elastography images of the same size. According to a preferred embodiment, the trained NN includes at least one convolutional layer. A convolutional layer applies a relatively small filter kernel across the entire input layer, so that neurons in the layer are connected to only a small region of the previous layer. This architecture ensures that the learned filter kernel generates the strongest response to spatially local input patterns. Each filter kernel is replicated across the entire input layer. In a useful embodiment of the present invention, the parameters of a convolutional layer have a small receive field, but contain a set of learnable filter kernels that can extend across the full depth of the input volume. During a forward pass through the convolutional layer, each filter kernel is convolved across the width and height of the input volume, computing the dot product between the filter kernel's entries and the input to generate a feature map for that filter. The full output volume of the convolutional layer is formed by stacking the feature maps of all filter kernels along the depth dimension. Thus, all entries in the output layer, including several feature maps or combinations of feature maps, can be interpreted as looking at a small region in the input and as the output of neurons that share parameters with neurons in the same feature map.

[0021] According to a preferred embodiment, the artificial neural network to be trained is a deep fully convolutional neural network. In this regard, fully convolutional means that there are no fully connected layers. Rather, the NN of the present invention exploits spatially local correlations by enforcing sparse local connectivity patterns between neurons in adjacent layers, with each neuron connected to only a small region of the input volume. A deep convolutional neural network includes at least two convolutional layers. In fact, the NN to be trained of the present invention preferably includes at least four convolutional layers, and more preferably eight or more convolutional layers. A further advantage of a fully convolutional neural network is that the NN to be trained can be applied to input images of any size. Because the filter kernel is replicated across the entire input layer or previous layer, the convolutional network can be applied to images of any size. The input to the NN is preferably two-dimensional, as this significantly reduces the number of training parameters in the NN. In a useful embodiment, the convolutional layer or layers of the NN include a convolutional filter kernel having a size of 3x3 pixels for 2D input images, or less preferably 2x2 or 4x4 pixels. For a 3D input image, the filter kernel is also 3D, e.g., 3x3x3 or 2x2x2 pixels.

[0022] In a useful embodiment, the depth of each convolutional layer is between 8 and 64, preferably between 16 and 32, and most preferably 32, meaning that 32 different feature maps generated by different filter kernels are part of each convolutional layer. When two convolutional layers follow each other, their depth preferably remains the same, and each of the 32 feature maps in the subsequent layer is generated by convolving each of the feature maps from the previous layer with a separately trained filter kernel and adding them together. Thus, a deep convolutional layer with a depth of 32 connected to a subsequent layer with a depth of 32 requires 32 x 32 filter kernels to train.

[0023] According to further useful embodiments, the artificial neural network to be trained comprises at least one unit with one to three, preferably two to three, most preferably two convolutional layers followed by a pooling or upsampling layer. For example, every two convolutional layers is followed by a pooling (downsampling) operation, preferably a 2x2 max-pooling operation, which reduces a kernel of four pixels to one by projecting only the highest value to the subsequent layer, thus resulting in a smaller size of the subsequent layer. This step allows the network to learn larger-scale features that are less sensitive to local variations. This may be performed in the first part of the NN, called the encoder. In some embodiments, there is a second part of the network, called the decoder, in which the max-pooling operation is replaced, for example, by nearest-neighbor upsampling, to increase the layer size again so that the output layer has the same size as the input layer.

[0024] In one embodiment, the artificial neural network to be trained has an encoder-decoder architecture. Preferably, the NN includes exactly one encoder section and exactly one decoder section. Thus, the size of the layers (e.g., the number of pixels processed therein) gradually decreases as one goes deeper into the network. The central layer is called the deep latent space. In a useful embodiment, the layers in the deep latent space are reduced in size relative to the input layer, e.g., by a factor of 8, from 4 to 16. The deep latent space may also include at least one convolutional layer. For example, it may consist of two convolutional layers.

[0025] According to one embodiment, the NN comprises an encoder section comprising at least one, preferably multiple, pooling or downsampling layers. The encoder section of the NN is followed by a decoder section comprising at least one, preferably multiple, upsampling layers. Preferably, the NN comprises exactly one encoder section comprising multiple pooling or downsampling layers and exactly one decoder section comprising multiple upsampling layers. Optionally, the deep latent space between the decoder and encoder sections comprises further layers, in particular, two convolutional layers. The pooling layer may be a max pooling layer, but may also use other functions such as average pooling or l2 norm pooling. The upsampling layer may use nearest neighbor upsampling, but may also use linear interpolation. Preferably, the decoder section is a mirrored version of the encoder section, with the downsampling layer replaced by an upsampling layer. Preferably, the encoder section comprises at least two, preferably three, max pooling layers with a 2x2 filter kernel and a stride of 2. If the stride is 2, the filter kernel moves two pixels at a time as it slides around the layer. Thereby, from each 2x2 square of neurons in the preceding convolutional layer, only the activity of the most active (i.e., highest value) neuron is used for further calculations. This allows the receptive field to grow deeper automatically without the need to reduce the size of the filter kernel. Furthermore, it is possible to build deeper networks, allowing for more complex behavior.

[0026] According to one embodiment, the artificial neural network to be trained comprises one or more layers between the decoder portion and the encoder portion. The layer or layers between the decoder portion and the encoder portion are referred to as a deep latent space. Optionally, there may be operations performed within the deep latent space. The deep latent space may comprise at least one convolutional layer. According to one embodiment, the deep latent space also comprises or consists of two convolutional layers. According to another embodiment, the deep latent space comprises at least one fully connected layer. In a preferred embodiment, the encoding portion of the network consists of a total of six convolutional layers and up to three pooling layers that map input images into the deep latent space.

[0027] Preferably, the decoder portion of the network may also be configured with a total of six convolutional layers and three upsampling layers. According to one embodiment, the NN comprises an encoder portion including multiple convolutional layers, each of which has one to three, preferably two to three, most preferably two convolutional layers followed by a pooling layer; and a decoder portion including multiple convolutional layers, each of which has one to three, preferably two to three, most preferably two convolutional layers followed by an upsampling layer. The convolutional layers of the decoder portion preferably have the same filter kernel size and stride as those of the encoder portion, e.g., a 3x3x3 filter kernel for 3D input images and a 2x2 filter kernel for 2D input images. The operations performed by the convolutional layers on the decoder portion may be described as deconvolution. Preferably, each of the two convolutional layers is followed by a pooling layer or an upsampling layer, respectively, in the encoder portion and the decoder portion. In one embodiment, the encoder section includes two to four, preferably three, such units, each consisting of one to three convolutional layers followed by a pooling layer. The decoder section preferably also includes two to four, preferably three, units, each consisting of one to three convolutional layers followed by an upsampling layer, since the decoder section is preferably a mirrored version of the encoder. According to a further embodiment, the NN includes at least one skip connection from a layer in the encoder section to an equal-sized layer in the decoder section. In other words, the NN preferably includes a direct "skip" connection from at least one encoder filter layer to its equal-sized decoder counterpart. By transferring the encoder layer output across the latent space and concatenating it with larger-scale model features during decoding, the NN allows for optimal combination of fine- and coarse-level information to generate higher-resolution elastography image estimates. In a useful embodiment, there is one skip connection from each unit containing one to three convolutional layers followed by a pooling layer in the encoder portion to the corresponding unit containing one to three convolutional layers followed by an upsampling layer in the encoder portion.The skip connections can be implemented by adding or concatenating the output of each layer on the encoder section to the output of the corresponding layer in the decoder section. For example, if the convolutional layers have a depth of 32 in the encoder section, they are concatenated to 32 feature maps in the corresponding layer in the decoder section, resulting in 64 feature maps. By using twice the number of filter kernels in the convolutional layers in the decoder section (e.g., 64 x 32), these 64 feature maps are mapped back to 32 feature maps in the subsequent layer. Thus, each feature map of the following convolutional layers in the decoder section can be influenced by inputs provided by the skip connections from the encoder section.

[0028] The stride with which the filter kernel is applied is preferably 1, i.e., the filter kernel is moved one pixel at a time. This leads to overlapping received fields between columns, and the grid size is not reduced between convolutional layers. Alternatively, the stride may be 2, and the filter kernel jumps two pixels at a time as it slides. This reduces the size of subsequent layers by a factor of 2 in each dimension; for example, a 2-dimensional grid would be 2x2. Thus, convolutional networks can be built without the need for downsampling layers.

[0029] In most embodiments of the present invention, the neural network includes at least one layer that includes an activation function, preferably a nonlinear activation function. For example, the results of each convolutional layer can be passed through the nonlinear activation function. Preferably, such an activation function propagates both positive and negative input values ​​with unbounded output values. "Unbounded" means that the output value of the activation function is not limited to any particular value (such as +1 or -1). Preferably, any value can be obtained in principle, thereby preserving the dynamic range of the input data. Preferably, the activation function of the present invention introduces nonlinearity while preserving negative signal components, minimizing the risk of vanishing gradients during training. In a useful embodiment, the activation function used in the convolutional layer is a leaky rectifier linear unit (LReLU), which can be written as follows: The α value is between 0.01 and 0.4, e.g., 0.1. Alternatively, a hyperbolic tangent activation may be used to preserve negative values. In other embodiments, the activation function is a rectified linear unit (ReLU), hyperbolic tangent, or sigmoid function. In a further embodiment, the activation function may be an inverse rectifier function that combines L2 normalization of the sampling direction with two rectified linear unit (ReLU) activations, thereby concatenating the positive and negative portions of the input. Preferably, the results of each convolutional layer are passed through such a nonlinear activation function. According to another embodiment of the present invention, the latent space includes a domain-adaptive structure or function that allows the sSWE generation to be transferred to other ultrasound machines and acquisitions, e.g., other types of ultrasound scanners and / or other elastographic acquisition techniques than the one on which the network is trained. This may be done, for example, by applying coefficients and / or shifts and / or other functions to each node in the latent space. Therefore, at least one layer in the latent space, i.e., the smallest layer between the encoder and decoder portions of the network, can be equipped with a domain adaptation function to align the feature maps to a different system. For example, if a NN is trained on SWE images, it can be further adapted to supersonic elastography imaging by shifting and / or scaling the nodes in the latent space, e.g., by adding a layer that applies such an adaptation function to each node. This can correspond to a transformation applied to the layer in the latent space. The required translation and scaling can be established by comparing the encoded latent spaces of two datasets in different domains. In a preferred embodiment, the mean and variance of the latent vectors are corrected.

[0030] According to another aspect, the present invention provides a method for training an artificial neural network useful for generating a synthetic elastography image from B-mode ultrasound images, comprising the steps of: (a) receiving input training data, i.e., at least one B-mode ultrasound image of a region of interest, the B-mode ultrasound image being acquired during an ultrasound examination of a human or animal subject; (b) receiving output training data, i.e., at least one ultrasound elastography image of the region of interest acquired by an ultrasound elastography technique during the same ultrasound examination; and (c) training the artificial neural network using the input training data and the output training data. The training method may be used to train an artificial neural network capable of generating a synthetic elastography image of the region of interest described herein. It is trained using B-mode ultrasound images as input training data and elastography images of the same region of interest acquired by any ultrasound elastography technique during the same ultrasound examination, or simultaneously with the B-mode ultrasound image. The elastography images may be shear wave elastography (SWE) images, but may also be acquired using any other ultrasound elastography technique. For example, it may be obtained by quasi-static elastography imaging, acoustic radiation force impulse imaging (ARFI), supersonic shear imaging (SSI), or transient elastography, or even magnetic resonance elastography (MRE). Depending on the training data used, the appearance of the output of the NN will differ, as the method of the present invention is primarily data-driven.

[0031] The training method described herein may be used to provide an artificial neural network useful for generating synthetic elastography images, i.e., for initially training the neural network. It may also be used to recalibrate an already trained network. Thus, it is possible to perform a combined SWE / SWE plan, performing real-time full-view synthetic shear wave elastography imaging (using the method of claim 1 and B-mode images as input) but occasionally calibrating the model with conventional elastography images, particularly conventional SWE images, during the scan. According to an alternative method for calibrating an artificial neural network configured to generate synthetic elastography images from B-mode ultrasound images, the method includes the steps of: acquiring at least one (conventional) ultrasound elastography image of the region of interest acquired by ultrasound elastography techniques during the same ultrasound examination as the B-mode ultrasound image; and recalibrating the trained artificial neural network based on the elastography image and the synthetic elastography image.

[0032] In one embodiment, the method further comprises estimating a confidence map of the synthetic elastography image (sSWE). Accordingly, the method further comprises applying a neural network (NN) or a second trained artificial neural network to the B-mode ultrasound image, the output of which is a confidence map comprising a plurality of confidence scores, each confidence score representing a confidence level for the value of a corresponding pixel in the synthetic elastography image. The method also includes the optional further step of providing and / or displaying the confidence map. This step thus allows for simultaneous estimation of the SWE confidence, which is used to identify low-confidence regions due to shear wave artifacts, such as signal voids in (pseudo) fluid lesions or B-mode artifacts such as shadowing or reverberation. The method may further comprise displaying the confidence map, for example, side-by-side with or overlaid with the sSWE.

[0033] Preferably, the NN or second NN is trained to provide a confidence map in a manner including the steps of: (a) receiving input training data, i.e., at least one synthetic elastography image generated by the method described herein, wherein the B-mode ultrasound images used to generate the synthetic elastography image are acquired during an ultrasound examination of a human or animal subject; (b) receiving output training data, i.e., at least one ultrasound elastography image of the region of interest acquired by an ultrasound elastography technique during the same ultrasound examination; and (c) training the artificial neural network using the input training data and the output training data.

[0034] In one embodiment, the confidence map includes not only the confidence of the NN or second NN in predicting SWE, but also the intrinsic confidence of the SWE acquisition itself (which is typically estimated during the SWE protocol based on fitting, noise, etc.). This confidence can also be trained using the training method embodiment described above. Preferably, both the sSWE image and the corresponding confidence map are generated by the same artificial NN using B-mode images as input. Alternatively, the confidence map can comprise a second NN, which can have the same architecture as the (first) NN useful for generating elastography images of the region of interest. Alternatively, the second NN can have a simpler architecture, e.g., a deep convolutional neural network, but with fewer layers than the first NN. For example, the second NN can have an encoder-decoder architecture with one or two units, each including one or three convolutional layers followed by a pooling or upsampling layer on the encoder and decoder portions.

[0035] Training of the first and second neural networks can be accomplished by backpropagation, where input training data is propagated through the NN using a given filter kernel. The output is compared to the output training data using an error or cost function, and the output is propagated back through the NN, thereby computing the gradient to find the filter kernel (and possibly other parameters, such as bias) that results in the smallest error. This can be done by adjusting the weights of the filter kernel to follow the negative gradient of the cost function.

[0036] In a useful embodiment, a neural network uses dropout layers during training, where certain nodes or filter kernels within the dropout layer are randomly selected and their values / weights are set to 0. For example, a dropout layer may have a predetermined percentage of dropout nodes, such that 30-80%, preferably 40-60%, of all nodes are dropouts and their values / weights are set to 0. For subsequent backpropagation of the training data, a different set of nodes in the dropout layer is set to 0. This generates noise during training, but has the advantage that the training converges to a useful minimum. Thus, the NN is better trainable. In a useful embodiment, each of two or three convolutional layers is followed by a dropout layer during training.

[0037] The present invention also relates to a computer program comprising instructions that, when executed by a computing unit, cause the computing unit to perform the method of the present invention. This applies to the method for generating a synthetic elastography image, the additional method step of applying a second trained artificial neural network to provide a confidence map including a plurality of confidence scores, and the method for training the first and second artificial neural networks. Alternatively, the neural network may be implemented as hardware, for example, with a fixed connection on a chip or other processing unit. The computing unit capable of executing the method of the present invention may be any processing unit, such as a central processing unit (CPU) or a graphics processing unit (GPU). The computing unit may be part of a computer, cloud, server, laptop, tablet computer, mobile phone, smartphone, or other mobile device. In particular, the computing unit may be part of an ultrasound imaging system. The ultrasound imaging system may also include a display device, such as a computer screen.

[0038] The invention also relates to a computer readable medium comprising instructions which, when executed by a computing unit, cause the computing unit to carry out the method according to the invention, in particular the method or training method according to any one of claims 1 to 9. Such a computer readable medium may be any digital storage medium, e.g. a hard disk, a server, a cloud or a computer, an optical or magnetic digital storage medium, a CDROM, an SSD card, an SD card, a DVD or a USB or other memory stick.

[0039] According to another aspect, the present invention also relates to a system for generating a synthetic elastography image, the system having: a) a first interface configured to receive a B-mode ultrasound image of a region of interest; b) a computing unit configured to apply a trained artificial neural network to the B-mode ultrasound image, thereby generating a synthetic elastography image of the region of interest; and c) a second interface configured to output the synthetic elastography image of the region of interest.

[0040] Preferably, the system is configured to implement the present invention to generate a synthetic elastography image. Such a system may be implemented on an ultrasound imaging system, for example, on one of its processing units, such as a GPU. However, it is also important that the B-mode ultrasound images acquired by the ultrasound imaging system are transferred, for example via the Internet, to another local or remote computing unit, from which the synthetic elastography image of the region of interest is returned to the ultrasound imaging system for display. Therefore, the second interface may be connected to a display device, such as a computer screen, touch screen, etc.

[0041] Furthermore, the present invention also relates to a system for training a first or second artificial neural network according to the training method described herein.

[0042] According to a further aspect, the present invention relates to an ultrasound imaging system comprising an ultrasound transducer configured to transmit and receive ultrasound signals and a computing unit configured to generate a B-mode ultrasound image from the received ultrasound signals, the computing unit also being configured to perform the method according to any of claims 1 to 9. Due to the low computational cost of the method, such a computing unit can be integrated into existing ultrasound systems. Useful embodiments of the present invention will now be described with reference to the accompanying drawings, in which similar elements or features are designated with the same reference numerals. [Brief explanation of the drawings]

[0043] [Figure 1] FIG. 1 is a schematic diagram of conventional B-mode ultrasound imaging. [Figure 2] FIG. 1 is a schematic diagram of conventional shear wave elastography. [Figure 3] FIG. 1 is a schematic diagram of a method for generating an sSWE image according to an embodiment of the present invention. [Figure 4] FIG. 1 is a schematic diagram of a deep convolutional neural network according to an embodiment of the present invention. [Figure 5] FIG. 2 is a more detailed schematic diagram of a unit of a neural network according to an embodiment of the invention, comprising two convolutional layers and a pooling layer; [Figure 6] 1A-1C are examples of B-mode, SWE, and sSWE images generated in accordance with an embodiment of the present invention. [Figure 7] 1 illustrates an ultrasound imaging system according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0044] 1 shows a schematic of the B-mode ultrasound process, in which an ultrasound probe 2, typically comprising an array of ultrasound transducers, transmits a series of ultrasound pulses 3, e.g., as compressional wavefronts, into a region of interest 4, typically within the human or animal body. By recording the echoes and applying appropriate signal processing, such as beamforming, a B-mode ultrasound image 5 of the region of interest is acquired. This can be done at high frame rates, especially for 2D images.

[0045] FIG. 2 shows a conventional SWE image. An ultrasound probe 2 generates a sequence of "pushing" acoustic radiation force pulses 6 into a region of interest 4. The "pushing" pulses create laterally traveling shear waves 8, which are recorded using ultrafast imaging by the ultrasound probe 2 and additional ultrasound transmission pulses 3. The recorded echoes are transferred to a computing unit 10, which processes the ultrafast image recordings and generates a 2D SWE image 12. The B-mode image 5 and SWE image 12 shown in FIGS. 1 and 2 are acquired from the same region of interest, here during a prostate examination of a human subject.

[0046] FIG. 3 is a schematic diagram of a method for generating a synthetic SWE image according to an embodiment of the present invention. First, a B-mode image 5 is generated in a conventional manner, as shown in FIG. 1. The B-mode image 5 is propagated through a trained artificial neural network 16 according to an embodiment of the present invention, implemented (by software or hardware) on a computing unit 102, which may be the computing unit of a commercially available ultrasound scanner. The result is a synthetic elastography image (sSWE) 18, which preferably has the same size and dimensions as the B-mode image, although the grid may be slightly coarser, as in a conventional SWE image. FIG. 4 shows an embodiment of the NN 16 according to the present invention. The input image is propagated forward through the NN 16 from left to right. The pixel size of each layer is noted at the right of each layer.

[0047] In this case, a 2D B-mode image 5 with an image size of 64 × 96 pixels is fed to the input layer 22. This is followed by two convolutional layers 24 with a depth of 32. Thus, 32 filter kernels are applied to the input layer 22, resulting in 32 feature maps, which form part of each convolutional layer 24. In a preferred embodiment, the network's convolutional layers 24, 24a, and 38 each contain a 32 or 32 × 32 two-dimensional 3 × 3 pixel convolutional filter kernel, the results of which are passed through a nonlinear activation function, specifically a leaky rectified linear unit. The first two convolutional layers 24 in the encoder section 30 are followed by a 2 × 2 max-pooling layer 26, which reduces the four-pixel kernel to 1 by projecting only the highest value to the corresponding node in the next layer, which is again the convolutional layer 24a. The two convolutional layers 24 and the max-pooling layer 26 together form a unit 28. This unit architecture is repeated with the next unit 28a, which consists of two convolutional layers 24a and a max pooling layer 26a. The pixel size of the layers shows a 2x2 reduction in size from each unit 28 to the next unit 28a. However, the depth (i.e., the number of feature maps included in each convolutional layer) remains the same at 32. In this embodiment, there are a total of three units 28, 28a, and 28b in the encoding portion 30 of the network. The pooling layer of the third unit 28b is followed by several layers in a deep latent space 34, with the grid / layer only having a size of 8x12x32 or 8x12x64. In this embodiment, the deep latent space consists of two convolutional layers followed by an upsampling layer. In another embodiment, unit 34 can also be counted as part of the decoding portion 32 of the network. Each unit 36 ​​in the decoder section contains two convolutional layers 38 followed by an upsampling layer 40, which projects each pixel / node in the preceding layer onto a 2x2 pixel in the subsequent layer by nearest neighbor upsampling. The decoder section 32 is thus a mirrored version of the encoder section, containing three units 36, each consisting of two convolutional layers followed by an upsampling layer, or in the case of the final unit 36a, an output activation layer 42.The output of the NN is a synthetic shear wave elastography image 18.

[0048] Additionally, the deep convolutional neural network (DCNN) 16 includes direct "skip" connections 44 from the encoder filter layers to their corresponding equal-sized decoders. In a useful embodiment, there is one skip connection from each unit 28, 28a, 28b in the encoder portion 30 to an equal-sized layer in the decoder portion 32.

[0049] FIG. 5 shows unit 28, the first unit of the encoding portion 30 of network 16, in more detail. Thus, a B-mode image 5, here, for illustrative purposes, is fed into input layer 22 and represented by a one-dimensional matrix containing 16 pixels. Input layer 22 is a convolutional layer that applies four different filter kernels K1, K2, and K3 and K4 (not shown), each having a size of three pixels, to input data 22. This therefore results in the next layer 24 having a depth of four, i.e., containing four feature maps 48a-48d, each of which is activated upon detecting a feature of some particular type at a corresponding spatial location in input layer 22. The next convolutional layer 24′ contains 4-16 filter kernels, each sweeping across one feature map in layer 24 and adding its result to one of the four feature maps in layer 24′. For example, filter kernel K4,1 sweeps the fourth feature map 48d of layer 24 and adds its result to the first feature map 49a of layer 24′. Filter kernel K3,1 convolves the third feature map 48c of layer 24 and adds it to the first feature map 49a of layer 24'. Thus, 4 × 4 = 16 filter kernels are trained during the training step. Layer 24' is also a fully convolutional layer, resulting in an output with a depth of 4, i.e., four feature maps. This output is fed to pooling layer 26, which reduces each kernel of two pixels to 1 by projecting only the highest value onto a small grid, denoted 50. In the NN framework, layer 50 may already be in the latent space or may be the first convolutional layer of the next unit 28.

[0050] An embodiment of the present invention was tested as follows: Fifty patients diagnosed with prostate cancer underwent transrectal SWE examinations at the Martini Clinic, University Hospital Hamburg-Eppendorf, Germany. An Aixplorer™ (SuperSonic Imagine, Aixen-Provence, France) with an SE12-3 ultrasound probe was used. For each patient, SWE images were obtained at the base, middle, and apical sections of the prostate. Regions of interest were selected to cover the entire prostate or a portion of the prostate. Given the first 40 patients in the training set, a fully convolutional deep neural network was trained to synthesize SWE images given corresponding B-mode (side-by-side view) images. Data augmentation was utilized to reduce the risk of overfitting and prevent artifacts that hinder training by only estimating the loss gradient from high-confidence SWE measurements. The method was tested on 30 image planes from the remaining 10 patients. The results are shown in Figure 6. We show that the NN can accurately map B-mode images to sSWE images with a pixel-wise mean absolute error of approximately 4.8 kPa in terms of Young's modulus. Qualitatively, tumor regions characterized by high stiffness were largely preserved (verified by histopathological examination). Figure 6 shows examples from five test patients, where the first column (a) shows B-mode ultrasound images, the second column (b) shows shear wave elastographic acquisitions, and the third column (c) shows the corresponding synthetic SWE images acquired by a method according to one embodiment of the present invention.

[0051] 7 is a schematic diagram of an ultrasound system 100 configured to perform the methods of the present invention, according to one embodiment of the present invention. The ultrasound system 100 includes a conventional ultrasound hardware unit 102 having a CPU 104, a GPU 106, and a digital storage medium 108, such as a hard disk or solid-state disk. Computer programs can be loaded into the hardware unit from a CD-ROM 110 or via the Internet 112. The hardware unit 102 is connected to a user interface 114 that includes a keyboard 116 and, optionally, a touchpad 118. The touchpad 118 may also function as a display device for displaying imaging parameters. The hardware unit 102 is connected to an ultrasound probe 120, which includes an array of ultrasound transducers 122, capable of acquiring B-mode ultrasound images, preferably in real time, from a subject or patient (not shown). The B-mode images 124 acquired by the ultrasound probe 120 and the sSWE images 18 generated by the method of the present invention executed by the CPU 104 and / or GPU are displayed on a screen 126, which may be any commercially available display unit, such as a screen, television, flat screen, projector, etc. Furthermore, there may be a connection to a remote computer or server 128, for example via the Internet 112. The method of the present invention may be executed by the CPU 104 or GPU 106 of the hardware unit 102, but may also be executed by a processor of the remote server 128.

[0052] The above discussion is intended to be merely illustrative of the present system and should not be construed as limiting the appended claims to any particular embodiment or group of embodiments. Accordingly, while the present system has been invented in particular detail with reference to exemplary embodiments, it will also be understood that those skilled in the art can devise numerous modifications and alternative embodiments without departing from the broader intended spirit and scope of the present system as set forth in the following claims. Accordingly, the specification and drawings are to be regarded in an illustrative manner and are not intended to limit the scope of the appended claims.

Claims

1. 1. A computer-implemented method for generating a synthetic elastography image, the method comprising: a) receiving a plurality of B-mode ultrasound images of a region of interest; b) generating a plurality of synthetic elastography images of the region of interest at a B-mode ultrasound frame rate by applying a trained artificial neural network to the plurality of B-mode ultrasound images; and the artificial neural network is trained based on machine learning using a plurality of elastography images generated by an elastography technique as training data; method.

2. 2. The method of claim 1, wherein the input to the trained artificial neural network, i.e., the plurality of B-mode ultrasound images, has the same size and dimensions as the output of the trained artificial neural network, i.e., the plurality of synthetic elastography images.

3. 3. The method of claim 1, wherein the trained artificial neural network has at least one convolutional layer, the one or more convolutional layers having a filter kernel with a size of 3x3 pixels.

4. 4. The method of claim 1, wherein the trained artificial neural network is a deep fully convolutional neural network.

5. 5. The method according to claim 1, wherein the trained artificial neural network comprises at least one unit having two convolutional layers followed by a pooling or upsampling layer.

6. 6. The method of claim 1, wherein the trained artificial neural network has an encoder-decoder architecture, the artificial neural network having one encoder part and one decoder part.

7. 7. The method of claim 1, wherein the trained artificial neural network comprises one or more layers in a deep latent space between the encoder and decoder portions.

8. 8. The method of claim 1, wherein the trained artificial neural network comprises an encoder portion having a plurality of convolutional layers, one to three convolutional layers each followed by a pooling layer and a decoder portion having a plurality of convolutional layers, and one to three convolutional layers each followed by an upsampling layer.

9. 9. The method of claim 7 or 8, wherein the trained artificial neural network has at least one skip connection from a layer in the encoder portion to an equally sized layer in the decoder portion.

10. 10. The method according to claim 1, wherein the trained artificial neural network comprises at least one layer including a non-linear activation function such as Leaky ReLU, ReLU, hyperbolic tangent, sigmoid, or anti-rectifier.

11. A method of training an artificial neural network useful for generating a plurality of synthetic elastography images from a plurality of B-mode ultrasound images at a B-mode ultrasound frame rate, comprising: (a) receiving input training data, i.e., at least one B-mode ultrasound image of a region of interest, the B-mode ultrasound image being acquired during an ultrasound examination of a human or animal subject; (b) receiving output training data, i.e., at least one ultrasound elastography image of said region of interest acquired by an ultrasound elastography technique during the same ultrasound examination; (c) training the artificial neural network by using the input training data and the output training data; A method comprising:

12. c) applying the trained artificial neural network or a second trained artificial neural network to the plurality of B-mode ultrasound images, wherein the output of the trained artificial neural network or the second trained artificial neural network is a confidence map having a plurality of confidence scores, each confidence score representing a confidence level of a value of a corresponding pixel in the plurality of synthetic elastography images.

11. The method according to claim 1, comprising:

13. 1. A method for training a second artificial neural network or said trained artificial neural network to provide a confidence map having a plurality of confidence scores, each confidence score representing a confidence level of a pixel value of a synthetic elastography image, said method comprising: (a) receiving input training data, i.e., at least one synthetic elastography image generated by the method of any one of claims 1 to 10, wherein the B-mode ultrasound images used to generate the synthetic elastography image are acquired during an ultrasound examination of a human or animal subject; (b) receiving output training data, i.e., at least one ultrasound elastography image of said region of interest acquired by an ultrasound elastography technique during the same ultrasound examination; (c) training the artificial neural network by using the input training data and the output training data; A method comprising:

14. A computer program comprising instructions which, when said program is executed by a computing unit, cause said computing unit to carry out a method according to any one of claims 1 to 13.

15. A system for generating multiple synthetic elastography images at a B-mode ultrasound frame rate, the system comprising: a) a first interface configured to receive a plurality of B-mode ultrasound images of a region of interest; b) a computing unit configured to apply the trained artificial neural network to the plurality of B-mode ultrasound images, thereby generating a plurality of synthetic elastography images of the region of interest; and c) a second interface configured to output the plurality of composite elastography images of the region of interest; and and the artificial neural network is trained based on machine learning using a plurality of elastography images generated by an elastography technique as training data; system.

Citation Information

Patent Citations

  • Medical image processing apparatus

    JP2019076541A

  • Automated cardiac volume segmentation

    JP2019504659A

  • Ultrasound imaging apparatus and method of controlling the same

    US20160113630A1

  • Method and apparatus to measure tissue displacement and strain

    US20180132830A1

  • Automated cardiac volume segmentation

    US20180259608A1