Adaptive Ultrasound Beamforming

A deep learning-based ultrasound beamforming method with convolutional layers and known operators addresses the limitations of traditional beamformers by reducing errors and improving image quality and throughput.

JP2025538734APending Publication Date: 2025-11-28UNIVERSITY OF LEEDS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025532907
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing ultrasound beamformers, such as delay-and-sum (DaS) beamformers, suffer from poor lateral resolution due to fixed apodization, while data-adaptive methods like minimum variance techniques are computationally complex and limited by non-stationary ultrasound conditions, making real-time implementation challenging.

Method used

A deep learning architecture incorporating convolutional layers and known operators is used to predict apodization weights, reducing trainable parameters and improving lateral resolution by artificially decorrelating signals, thus enhancing image quality.

Benefits of technology

The method reduces training time and computational complexity, leading to improved ultrasound image quality and throughput by optimizing apodization weights, even in non-stationary conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538734000022
    Figure 2025538734000022
  • Figure 2025538734000023
    Figure 2025538734000023
  • Figure 2025538734000024
    Figure 2025538734000024
Patent Text Reader

Abstract

Ultrasound beamforming is performed by a deep learning architecture incorporating convolutional layers and known operators. The network is trained to optimally combine multiple observations of complex signal data into a single complex signal. This is achieved by using convolutional layers with learnable parameters to embed operators that perform multiple weighted sums of subsets of the original observations and average them to produce a single complex signal. Known operators, including forward-backward operators, are optionally incorporated into this architecture.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to methods and systems for performing adaptive ultrasound beamforming. In particular, the present disclosure relates to a deep learning architecture for ultrasound beamforming consisting of one or more neural networks, where one or each network is trained to predict a set of apodization weights and perform apodization within the architecture. [Background technology]

[0002] Brightness-mode (B-mode) ultrasound image reconstruction is performed using the magnitude (brightness) of ultrasound echoes from insulating media. The echo signals are measured by an array of sensors known as elements. Each element converts pressure waves into an electrical signal, usually through the application of the piezoelectric effect. This electrical signal is typically amplified and then digitized at a specific sampling rate using an analog-to-digital converter (ADC). A process known as beamforming combines the data from these elements to estimate the magnitude of the echo from a point in the medium. This is achieved by applying spatial filtering, which results in a beam pattern with local maxima and minima known as lobes. The common goal of all beamformers is to achieve a beam with a narrow main lobe (i.e., the lobe radiating in the direction of interest) and low-level side lobes (i.e., undesired lobes outside the region of interest).

[0003] A narrow main lobe ensures sufficient lateral resolution of the system, which is achieved by applying focusing delays. Low side lobes ensure that interference from off-axis scattering is minimized. Strong scatterers in the side lobes can be mistaken for weak scatterers in the main lobe, limiting lateral resolution. Side lobes can be suppressed by applying per-element weights to the beamformer, a process known as apodization. Suppressing side lobes through apodization results in a broader main lobe. The most widely adopted beamformer is the delay-and-sum (DaS) beamformer with a fixed apodization weight. The poor lateral resolution of DaS beamformers limits their usefulness for anatomical imaging of small, deep structures.

[0004] To improve lateral resolution, data-adaptive beamformers, such as those using minimum variance (MV) techniques, can be employed. These beamformers calculate optimal apodization based on the statistics of the received data, thereby suppressing sidelobes in areas with strong interference and preserving them in areas without interference. However, such algorithms are inherently complex, and real-time implementation remains a major challenge.

[0005] Better estimates can be obtained by employing techniques such as subarray averaging, which uses multiple subarrays to estimate the covariance matrix, and temporal averaging, which uses multiple temporal samples to estimate the covariance matrix. Diagonal loading, which adds the coefficients of an identity matrix to the covariance matrix of a sample, and backward averaging, which reverses the imaging array and signal arrival angle (backward aperture) and then averages it with the original (forward aperture), can be applied to increase robustness in the presence of sound speed estimation errors.

[0006] Furthermore, the non-stationary nature of ultrasound limits the number of samples of received data that can be used to estimate the statistics needed to determine the optimal apodization, which ultimately limits the accuracy of the estimation.

[0007] Deep learning architectures can overcome some of the limitations of minimum variance algorithms. Their ability to generalize to the vast statistical range available in the datasets used for training may provide more accurate estimates of optimal apodization weights in the presence of noise or sound speed estimation errors. Deep learning architectures can be efficiently implemented on graphics processing units (GPUs) or tensor processing units (TPUs). These devices often incorporate hardware acceleration designed to perform common operations found in deep learning architectures, such as convolution operations. Deep learning architectures often consist of large, general-purpose data-driven models with many parameters. These models rely on large datasets that adequately represent the inverse problem and on networks large enough to generalize well under a variety of conditions. It is difficult to predict how these data-driven models will perform on unknown data. Furthermore, these networks can be inefficient in terms of computational complexity.

[0008] The reference to prior art herein is not an admission or implication that this prior art forms part of common general knowledge in any jurisdiction or worldwide, or that this prior art could reasonably be expected to be understood, considered relevant, and / or combined with other prior art by a person skilled in the art. Summary of the Invention [Problem to be solved by the invention]

[0009] It is an object of the present invention to ameliorate at least one or more of the above or other disadvantages of the prior art and / or to provide a useful alternative. [Means for solving the problem]

[0010] The present invention is a method and system as defined in the appended claims.

[0011] It will be understood that features and aspects of the present disclosure can be combined with other different aspects of the disclosure, as appropriate, and not just in the specific exemplary combinations described herein.

[0012] According to one aspect of the present disclosure, there is provided a method for generating beamformed ultrasound signals, the method including: receiving a plurality of first electrical signals from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasonic waves received from a location within the object; processing the first electrical signal, wherein said processing includes applying a predetermined delay profile; inputting the processed first electrical signal into a trained artificial neural network, wherein the trained artificial neural network includes a plurality of convolutional layers; Determining apodization weights for beamforming ultrasound signals in a trained artificial neural network; and outputting from the trained artificial neural network a beamformed signal representing scattered ultrasound waves received from the location within the object based on a plurality of second electrical signals, where each second electrical signal is based on a respective sum of products, where each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective apodization weight, and where each sum is based on being less than all of the processed first electrical signals.

[0013] By generating a beamformed signal representative of scattered ultrasound waves received from the location within the object based on a plurality of second electrical signals (where each second electrical signal is based on a respective sum of products, where each product is a product of a respective first input based on a respective first electrical signal and a respective second input based on a respective apodization weight, and where each sum is based on less than all of the first electrical signals), an advantage is obtained by reducing the effect of errors in the first electrical signals, which in turn can reduce the training time of the neural network and improve the quality of the ultrasound image data obtained during the inference stage of the artificial neural network.

[0014] The reduction in training time results from a reduction in the number of trainable parameters because the network incorporates known operators that replace what would normally be trainable layers. Furthermore, for a given input size, the disclosed network is arranged to predict fewer apodization weights than the input size, which allows for compression through the network, which can speed up inference in predicting apodization weights.

[0015] Signals are artificially decorrelated, allowing for improved image quality. An assumption of the minimum variance algorithm is that the desired signal is uncorrelated with the unwanted signals. This is not the case in ultrasound imaging, where interference is generally highly correlated. The covariance matrix can be calculated for several different sub-apertures. By averaging these, coherent signals can be artificially decorrelated, resulting in a better estimate of the covariance matrix. This can be used to avoid signal cancellation.

[0016] Furthermore, the artificial decorrelation of the present disclosure reduces the size of the covariance matrix compared to full-aperture and increases the number of observations: if the number of observations at a position is less than the number of elements in the subarray, the covariance matrix may not be invertible.

[0017] For example, reducing the effect of errors in the first electrical signal may allow for inversion of the covariance matrix used in minimum variance techniques to create training data for an artificial neural network, thereby enabling better estimation of the spatial covariance matrix to provide training data for the artificial neural network, which in turn leads to improved image quality. Predicting fewer apodization weights than the input size allows for smaller networks with fewer operations, improving throughput.

[0018] The plurality of second electrical signals may be based on respective plurality of first electrical signals received from the location within the object at different times.

[0019] This provides the advantage that the effect of errors in the first electrical signal can be further reduced.

[0020] In certain embodiments, processing the first electrical signal may further include amplifying the first electrical signal and digitizing the amplified first electrical signal, and optionally normalizing the digitized first electrical signal and applying a Hilbert transform to the normalized signal.

[0021] The trained artificial neural network may, in certain cases, further include at least one beamforming operator layer, which may optionally be a forward-backward block.

[0022] According to another aspect of the present disclosure, there is provided a method of providing a trained artificial neural network adapted to generate beamformed ultrasound signals, said method including: receiving input training data including a plurality of processed first electrical signals corresponding to a plurality of first electrical signals received from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective first electrical signal in response to ultrasonic waves received from a location within an object, and the processed first electrical signals correspond to the first electrical signals after application of a predetermined delay profile; receiving output training data including at least one of a plurality of output electrical signals and apodization weights determined by an adaptive beamforming method; and training an artificial neural network using the input training data and the output training data; wherein the artificial neural network is adapted to output a beamformed ultrasound signal representing scattered ultrasound received from the location within the object based on a plurality of second electrical signals, wherein each of the second electrical signals is based on a respective sum of products, wherein each of the products is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a predetermined apodization weight, and wherein each of the sums is based on being less than all of the processed first electrical signals.

[0023] In certain embodiments, the artificial neural network is further adapted to determine predicted apodization weights for beamforming the ultrasound signals to output a plurality of predicted second electrical signals and generate a predicted beamformed ultrasound signal based on the plurality of predicted second electrical signals, wherein training the artificial neural network includes minimizing a mean squared error combining the predicted apodization weights and the predicted beamformed ultrasound signal.

[0024] According to a further aspect of the present disclosure, there is provided a computer readable medium bearing instructions which, when executed by a processor, cause the processor to perform the method defined above.

[0025] According to a further aspect of the present disclosure, there is provided a system for beamforming of ultrasound signals, said system comprising: an input for receiving a plurality of first electrical signals from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasonic waves received from a location within the object; processor means for processing the first electrical signal and applying a trained artificial neural network to the processed first electrical signal, wherein the processor means: processing the first electrical signal by applying a predetermined delay profile; inputting the processed first electrical signal into a trained artificial neural network; determining apodization weights for beamforming the ultrasound signals; and outputting a beamformed ultrasound signal representing scattered ultrasound waves received from the location within the object based on a plurality of second electrical signals, wherein each of the second electrical signals is based on a respective sum of products, wherein each of the products is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective apodization weight, and wherein each of the sums is based on less than all of the processed first electrical signals. the processor means adapted to perform An output section for outputting the formed ultrasound beam signal.

[0026] The subject matter of the present disclosure will be more fully understood and appreciated from the following detailed description taken in conjunction with the drawings in which corresponding or like numerals or characters indicate corresponding or like components. Unless otherwise indicated, the drawings provide illustrative embodiments or aspects of the present disclosure and are not intended to limit the scope of the present disclosure. [Brief explanation of the drawings]

[0027] [Figure 1] FIG. 1 shows an exemplary data processing pipeline in which raw received sensor data is processed to obtain an analytic signal, a delay profile is applied, and adaptive beamforming of the delayed analytic signal is performed. [Figure 2] FIG. 2 shows an example of a virtual source delay model used to obtain the delay profile of a focused transmission. [Figure 3]FIG. 3 illustrates one embodiment of a network architecture for performing adaptive beamforming in accordance with the present disclosure. [Figure 4] FIG. 4 shows a comparison of beamformed images generated from beamformed data in accordance with the present disclosure, images generated using conventional minimum variance techniques, and images obtained via a beamformer with uniform apodization in accordance with prior art. [Figure 5A] FIG. 5A illustrates a schematic diagram of a system for ultrasound image reconstruction suitable for implementing the architecture of the present disclosure. [Figure 5B] FIG. 5B illustrates schematically another system for ultrasound image reconstruction suitable for implementing the architecture of the present disclosure. [Figure 6] FIG. 6 illustrates certain functional operations of an adaptive beamforming method according to the present disclosure. [Figure 7] FIG. 7 illustrates certain functional operations in training a network architecture used for adaptive beamforming in accordance with the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0028] Description of exemplary embodiments The detailed description set forth below in connection with the accompanying drawings is intended as a description of various exemplary embodiments of the present disclosure and is not intended to represent the only form in which the present disclosure may be practiced. It should be understood that the same or equivalent functions may be achieved by different embodiments intended to be encompassed within the scope of the present invention. Furthermore, terms such as "comprise," "comprising," "has," "contain," or other grammatical variations are intended to cover a non-exclusive inclusion, such that modules, circuits, device components, structures, and method steps that include a list of elements or steps may include not only those elements, but also other elements or steps not expressly listed or inherent in such module, circuit, device component, or step. An element or step preceded by "comprises...a" does not, without more constraints, exclude the existence of further (additional) identical elements or steps that include that element or step.

[0029] Instead of relying on the massive datasets and sufficiently large networks required by traditional deep learning architectures to perform data-adaptive beamforming, application-specific architectures can be developed that incorporate fewer parameters and prior knowledge. By incorporating known operators, these networks contain fewer parameters than architectures associated with purely data-driven approaches, and as a result, require fewer training samples. Incorporating known operators can also improve the predictive accuracy of the model and ensure a solution that generalizes well to unseen data.

[0030] The present invention provides a method for adaptive beamforming using neural networks that incorporate convolutional layers (layers that perform convolution or cross-correlation operations) and known operators to optimally combine multiple observations of real or complex data (from data recorded on different array elements) into a single real or complex signal, resulting in images with superior lateral resolution compared to delay-and-sum (DaS) beamformers with constant apodization.

[0031] Known operator layers are incorporated to perform subarray aperture shading and averaging. In the subarray aperture shading layer, a weighted sum is performed for each subarray of observations taken from the original input, with the weights provided by the output of the previous convolutional layer. In the subarray averaging layer, the average value of the data from the various subarray weighted sums is determined, resulting in a single complex signal.

[0032] The algorithmic constraints from the MV algorithm, subarray aperture shading and averaging, are embedded in the network. As a result, structural information present in the input is conveyed to the output through the final layer, which is designed to perform optimal combinations of the inputs. This constraint aids the network's ability to generalize to a variety of data outside the training set.

[0033] A series of transformations are applied to the inputs in a neural network to generate a desired output. The unknown parameters associated with this transformation are learned using a gradient descent algorithm, which is designed to minimize the error between the network output and ground truth obtained by other methods, such as a minimum variance algorithm. The network can be trained using a combination of simulated data (in-silico), data acquired from experimental phantoms (in-vitro), or data acquired from patients in a clinical setting (in-vivo).

[0034] In certain embodiments, the network utilizes as input multiple time delay samples along the axial dimension (defined as the dimension perpendicular to the plane of the ultrasound array) and determines multiple sets of apodization weights along the axial dimension, which are applied to the original observations in a subarray aperture shading layer of the network.

[0035] In certain embodiments, the network runs on sensor channel data (element space). Alternatively, or in addition, the network runs on multiple beamformed channel signals (beam space) obtained from a set of preliminary beamformers.

[0036] When an axisymmetric array is utilized for imaging, layers can be incorporated into the network that combine the input data (or slices of the input data) with inverted / reversed versions of the data across one or more dimensions of the layer's input tensor. These layers can be included to more accurately determine the optimal apodization function.

[0037] While particular examples are described in connection with two-dimensional (2D) focused cardiac ultrasound imaging, embodiments are not limited to echocardiography and may also be applied to imaging of other anatomical features. Embodiments are not limited to focused imaging and may also be applied to unfocused imaging. Embodiments are not limited to two-dimensional imaging and may also be applied to three-dimensional (3D) volumetric imaging. While pixel-based delay profiles are described herein, embodiments may alternatively or additionally utilize line-based delay profiles. In particular embodiments, beamformed data is acquired from a single emission or multiple emissions simultaneously.

[0038] For ease of understanding, the specific example shown in the figure is described and is not intended to be limiting in scope. While this embodiment operates in element space, other embodiments combine beam space observations. Furthermore, multiple embodiments can be coupled to combine data at multiple stages.

[0039] Consider a linear array of "M" elements (an array of sensors suitable for measuring ultrasonic echo signals). Each element is capable of converting the received ultrasonic signal from reflections in the medium (pressure waves scattered by objects in the medium) into an analog electrical signal. These "M" analog signals can be amplified and digitized (at a certain sampling rate) by analog-to-digital converters (ADCs) in an analog front end (AFE) to generate "M" digital signals, or observations, of a region of interest (ROI).

[0040] Figure 1 shows an example data processing pipeline, where digital signals (i.e., samples) corresponding to sensor data are normalized in step [1] and an analytic signal is obtained by a Hilbert transform in step [2]. In step [3], a virtual source delay profile is applied, and in step [4], the delay data is provided as input to an adaptive beamforming stage in accordance with the present disclosure.

[0041] A batch size of N is drawn, corresponding to a different delay profile for each line of pixels obtained from a single emission. Each input has dimensions corresponding to the axial dimension, the element dimension, and the real and imaginary components of the analytic signal.

[0042] In one example of this embodiment, a 64-element two-dimensional transthoracic cardiac array is used to obtain "R" axial samples from each of the 64 channels of recorded data ("R x 64"; Figure 1, step [1]). In other embodiments, arrays of different shapes can be used, such as curvilinear arrays, matrix arrays, or arrays required for intravascular imaging. These arrays can consist of any number of elements and provide any number of observations. In other embodiments, fewer or more samples can be recorded, which will depend on the sampling rate of the system, the axial length of the ROI, and the direction of transmission.

[0043] In step [1] of the embodiment of Figure 1, the digital signal (i.e., the samples) are normalized between -1 and 1, and then the analytic signal is obtained via a Hilbert transform (step [2]). In other cases, the samples can be normalized between 0 and 1 before the transform to obtain the analytic signal samples.

[0044] The delay profile used in step [3] of the embodiment is obtained from a virtual source model in which the focal position is modeled as a spherical virtual source. This virtual source model is described in: Kim, C., Yoon, C., Park, JH, Lee, Y., Kim, WH, Chang, JM, Choi, BI, Song, TK, and Yoo, YM (2013). Evaluation of Ultrasound Synthetic Aperture Imaging Using Bidirectional Pixel-Based Focusing: Preliminary Phantom and In Vivo Breast Study. IEEE Transactions on Biomedical Engineering, 60, 2716-272.

[0045] The delayed signal is input to the adaptive beamforming architecture in step [4].

[0046] The initial processing chain may differ from that described in connection with Figure 1; normalization (step [1]) may not be present, and an analytic signal may not be required, so that a delay profile is applied directly to the samples. However, delaying the data across the array in step [3] is a necessary aspect of signal processing at the input to the architecture to compensate for different path lengths of the different transducer elements.

[0047] Referring to elements 202, 204, and 206 in FIG. 2, the total two-way propagation time from the center of the array "κ" to a point "p" within the ROI and back to one element "e" can be written as:

number

[0048] In Figure 2, the angle at which the delay profile is applied can be approximated from the active aperture and focus position as follows:

number

[0049] From Figure 1, the output of the delay stage represents the input tensor of the architecture (Figure 1, step [3]). The dimensional shape of the tensor can be conventionally described as "N x H x W x C." For the input tensor, "N" represents the batch size (a set of inputs corresponding to different lateral locations within the medium, where lateral is defined as the dimension parallel to the plane of the transducer, i.e., the plane of the element array 210 in Figure 2), "H" represents the axial dimension (and consists of samples at multiple axial locations within the medium, where axial is defined as the dimension perpendicular to the plane of the ultrasound array, i.e., parallel to axis 212 in Figure 2), "W" represents the array element dimension (i.e., the number of elements in the array), and "C" represents the real and imaginary components of the analytic signal.

[0050] In the embodiment shown in step [3] of Figure 1, "N" different delay profiles can be obtained from a single ray corresponding to "N" different horizontal lines across the ROI. Each horizontal line in this embodiment consists of samples determined by the delay profile at 256 axial points "H=256".

[0051] In the illustrated example, a software pixel-based delay stage is used to obtain "N" horizontal lines of pixels directly from a single emission using a virtual source model. In other examples, unfocused emission (e.g., plane wave or diverging wave emission) can be utilized with a corresponding delay profile or a line-by-line delay profile of focused emission. Embodiments can use a parallel computing device, such as a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another dedicated processing device, such as a tensor processing unit (TPU), to beamform batches of data acquired simultaneously from a single emission or multiple emissions.

[0052] A deep learning architecture then combines the "M" observations into a single beamformed sample for all points within the ROI in accordance with the present disclosure. This architecture is used to perform the adaptive beamforming stage shown in step [4] of Figure 1. The network converts a batch of "N" horizontal lines of channel data into "N" horizontal lines of beamformed data, where "N" can be any number of lines.

[0053] As mentioned above, Figure 2 shows an example of a virtual source-delay model used to obtain the delay profile of a focused transmission. This shows how the propagation distance to a point within a region of interest (ROI) and back to the element of interest is approximated.

[0054] 3 illustrates one embodiment of a neural network architecture for performing adaptive beamforming according to the present disclosure. The network includes both convolutional layers and blocks of known operations. In the embodiment illustrated in FIG. 3, the layers are as shown in Table 1. [Table 1]

[0055] This embodiment requires as input the analytic signal from a set of element channels across the region of interest, such as that transmitted in step [3] of Figure 1. For the input tensor, "H" represents the axial dimension (H=256), "W" represents the array element dimension (W=64), and "C" represents the real and imaginary components of the analytic signal (C=2).

[0056] The first stage of the network uses a 2D convolutional layer (Layer 2) with a stride of (1,2) and 32 filter channels, followed by a max-pooling layer (Layer 3). This first reduces the element dimension to the size of the subarray (W=32). The network then finds a lower-dimensional representation of the data. Performing this compression in the first layer of the network reduces operations in subsequent layers, resulting in a computationally efficient implementation. In this embodiment, the stride was selected to reduce the subarray length to half the size of the transformer, but in other embodiments, different stride lengths can be selected to generate tensor shapes that fit other subarray lengths.

[0057] In this embodiment, the first known operator is in layer 4. Other embodiments may not include such a known operator. This layer takes as input the output from the max pooling layer and, subject to the constraint that its input C=2W, splits the input into two matrices of equal dimensions, HxWxW:

number

[0058] Having constructed a permutation matrix J of dimension W x W, this layer performs negation and matrix multiplication along the H dimension between the final two dimensions of the input tensor and the permutation matrix:

number

[0059] "B" is then composited with the original input:

number

[0060] As a result, this layer flips or inverts "A1" and "-A2" along the final two dimensions and combines them with the original input. The first convolutional layer is used to transform the input data for the Forward-Backward combination layer, utilizing a linear activation function as opposed to a Rectified Linear Unit (ReLU). This is done to preserve negative components of the output tensor.

[0061] Five convolutional layers (Layers 5 to 9) are used to transform the output data from the forward-backward combination layer (Layer 4), followed by an upsampling layer (Layer 10) and two further convolutional layers (Layers 11 and 12). The upsampling layer (Layer 10) transforms the data so that "H=256" is the same as the original input size and "W=32" is the same as the subarray length selected in this embodiment. The final two convolutional layers (Layers 11 and 12) transform the data so that "C=2" is the same as the original input, representing the real and imaginary components of the complex target weights. The final convolutional layer (Layer 12) again uses a linear activation function to preserve the negative components of these complex target weights.

[0062] The final layers of the network (layers 13 and 14) perform element-wise weighting of the original inputs by the weights predicted by the last convolutional layer (layer 12). This is achieved by splitting the original inputs into two HxW1x1 tensors representing the real and imaginary components of the analytic signal:

number

number

[0063] For subarray weighting and averaging, the constraint on the dimension of "W" is "W1>W2". In the embodiment shown in Figure 3, the delayed input tensor has shape "256x64x2" and the apodization weight tensor has shape "256x32x2".

[0064] The complex conjugate of the apodization weight can be determined by negating B2, which contains the imaginary part of the complex weight, and the partitioned data can be concatenated along the "C" dimension as follows:

number

[0065] Two depthwise convolutions, with the dimension of "H" being the depth dimension, are used with effective padding to perform apodization weighting of the delayed input data, one with the filter set to "B" and the other with the filter set to "B". * The outputs of these convolutions are concatenated along the C dimension to produce an output with shape H×(W1-W2+1)×2. In the embodiment shown in Figure 3, this corresponds to the output tensor (of layer 13), which has shape 256×33×2.

[0066] The final layer (layer 14) averages its inputs across the W dimension to obtain the final beamformed output with shape 256 × 1 × 2. This is achieved by utilizing an average pooling layer with pool size equal to (1, W1 - W2 + 1).

[0067] Combined, these two layers (i.e., layers 13 and 14) can be seen as embeddings of known operators:

number

[0068] Other embodiments may use different filter sizes or strides in the convolutional layers, may use a different number of convolutional layers, and may apply further compression and expansion through additional pooling and upsampling layers.

[0069] Thus, an embodiment of the network consists of a CNN and blocks containing known operators to predict a set of apodization weights. The CNN uses multiple observations from multiple axis samples to predict axis-dependent apodization weights. The convolutional layers in the CNN can use a rectified linear unit (ReLU) nonlinear activation function, as does the previous embodiment. To preserve the negative components of the output of some layers, which is necessary for some known operator blocks, other activation functions that preserve negative values, such as linear activation functions, can also be used.

[0070] Thus, rather than simply using a neural network to output apodization weights, the network architecture of the present disclosure embeds known beamforming operators to perform the apodization weighting. While in the above-described embodiment, the known weighting functions are implemented at the end of the neural network architecture, in particular embodiments, additional layers can be placed before, between, or after these known operator layers. Thus, the apodization weighting is considered an intrinsic part of the neural network architecture of the present disclosure. The resulting network architecture does not predict apodization weights, but rather optimally combines the input channel data. This optimal combination is output directly. In particular embodiments, the optimal combination of ground truth is determined using a minimum variance algorithm, but in other embodiments, it may be determined by a different algorithm.

[0071] In certain embodiments, the known operator incorporates subarray averaging (as in the embodiments described above). When subarray averaging is performed, the inputs to the known operator that performs the weighting are not of equal width. For example, the weights may be half the width of the contributing signals, and subarray averaging is performed. In such a case, all contributing RF signals are input to the neural network. In most scenarios, the contributing RF signals will be data recorded from all receiving elements of the transducer. The network can then predict a set of weights to apply to the subarrays. The output obtained from this known operator is averaged in another known operator layer to obtain an output.

[0072] In an array of many elements, such as a matrix array, the contributing RF signals may be a subset of the data recorded by the array. In such cases, multiple embodiments may be operated in conjunction to combine the data in multiple stages. In the two-stage case, the contributing RF signals may be combined in a first stage to produce a single output. Multiple outputs from the first stage, derived from different subsets of transducer elements, are input to a second stage where they are optimally combined to produce a single output.

[0073] If a minimum variance algorithm is implemented to determine the optimal combination of ground truths, subarray averaging can increase the robustness of the output, which is important in medical imaging equipment, at the expense of some loss of resolution.

[0074] As mentioned above, the architecture of the present disclosure introduces other known operators, including forward-backward blocks. These operators serve to reduce the error of the predicted output with respect to the ground truth of the MV algorithm. Furthermore, these other known operators also act to generalize the training process (which may then also improve inference speed) by reducing the number of learnable parameters and reducing the size of the network. Introducing known operators can help ensure that the architecture generalizes well to previously unseen data, a critical requirement for medical imaging.

[0075] Figure 4 shows a beamformed image ("predicted image") generated from data beamformed using a deep learning beamformer. An image obtained via a beamformer with uniform apodization ("boxcar image") is included for reference and comparison.

[0076] Figure 4 shows an image produced by the embodiment shown in Figure 3 after envelope detection and log compression have been performed. The network architecture in this embodiment was trained using simulated radio frequency (RF) ultrasound signals and corresponding ground truth ("target images") obtained via a minimum variance (MV) algorithm (called the Eigenspace Based Minimum Variance (EBMV) algorithm) in which weights are projected onto the signal subspace.

[0077] The minimum variance technique used to obtain the training data set in the above embodiment can be understood in more detail as follows:

[0078] Given an array of 'M' elements, the 'M' observations can be described by a column matrix:

number

[0079] As described in Synnevag, J.F., Austeng, A., and Holm, S. (2009), Benefits of minimum-variance beamforming in medical ultrasound imaging. IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control, 56, 1868-1879, a sample covariance matrix incorporating subarray averaging and time averaging can be estimated as follows:

number

[0080] In the formula, the notation (·) Hrepresents the conjugate transpose. The subarray length "L" was set to 32, representing half the aperture size of the phased array probe. The number of time samples "K" was chosen so that 9 time samples were used to estimate the covariance matrix. These values ​​were chosen to ensure that the sample covariance matrix was invertible and to preserve speckle statistics. To obtain a more accurate estimate of the covariance matrix, FB averaging was used as described in Asl, B.M. and Mahloojifar, A. (2011), Contrast enhancement and robustness improvement of adaptive ultrasound imaging using forward-backward minimum variance beamforming. IEEE Transactions on Ultrasonics, Ferroelectrics and Frequency Control, 58, 858-867. The FB average is calculated as follows:

number

[0081] where "J" is the exchange matrix, denoted (·) * denotes the conjugate. To improve the robustness of the beamformer in the presence of sound speed estimation errors, diagonal loading was applied. The weighting coefficients were automatically calculated using a shrinkage algorithm as described in Stoica, P., Jian Li, Xumin Zhu, and Guerci, J. (2008). On Using a Priori Knowledge in Space-Time Adaptive Processing. IEEE Transactions on Signal Processing, 56, 2598-2602. In this method, the shrinkage parameters "α" and "β" are calculated based on the received data, and the loaded covariance matrix is ​​written as follows:

number

[0082] The contraction parameters are determined as follows:

number

[0083] To minimize the variance of the beamformer output while maintaining unity gain at the focal point, a set of apodization weights was calculated using covariance matrix estimation according to Capon, J. (1969), High-resolution frequency-wavenumber spectrum analysis. Proceedings of the IEEE, 57, 1408-1418:

number

[0084] where "a" is the steering vector and is a vector of 1 when delays are pre-applied. To reduce noise and improve contrast, the MVDR weights are adjusted according to R d (t) into the signal subspace [see: Mehdizadeh, S., Austeng, A., Johansen, T.F. and Holm, S. (2012). Eigenspace Based Minimum Variance Beamforming Applied to Ultrasound Imaging of Acoustically Hard Tissues. IEEE Transactions on Medical Imaging, 31, 1912-1921]:

number

[0085] The signal and interference subspaces were determined automatically based on an estimate of the noise in the received signal. The eigenvalue threshold was set to the sample covariance matrix eigenvectors corresponding to eigenvalues ​​greater than a scalar described as:

number

[0086] During the ceremony, “s DAS (t) is the DaS beamformed signal, averaged over the array:

number

[0087] The output can be expressed as:

number

[0088] To train the network of the above embodiment, a suitable optimizer, for example the NAdam optimizer, can be used to optimize for mean squared error (MSE) loss. Training used batches of 16 horizontal lines with 256 axial samples each, and randomized the data fed to the network at each epoch. The network was trained for up to 1000 epochs on the simulated data. Early stopping was performed with a patience of 50. The predicted apodization weights, L W ” and the mean squared error (MSE) loss of the final beamformed output “L S The MSE loss of and a combined loss incorporating are used during training:

number

[0089] where "λ" is a user-defined parameter between 0 and 1. In the network trained in this example, λ=0.6.

[0090] FIG. 5A shows a schematic diagram of a system for ultrasound image reconstruction suitable for implementing the architecture of the present disclosure. In this system, imaging transmit sequences are programmed on a CPU. Arbitrary and flexible waveforms can be designed to probe the Route of Interference (ROI). Transmit parameters are communicated (via a high-speed data link) to a Field Programmable Gate Array (FPGA), which controls a series of high-voltage transmit circuits (Tx excitation) in parallel. These are used to excite the piezoelectric elements ("transducer elements") of the transducer and generate ultrasonic pressure waves. The high-voltage signal passes through a transmit / receive switch (TX / RX switch) to prevent damage to the receive electronics.

[0091] The transducer element receives the pressure waves scattered by the object (e.g., a "scatterer") and converts them into an (analog) electrical signal.

[0092] The TX / RX switch directs the received electrical signal to the receiver circuit (RX LNA / ADC), which includes a low-noise amplification (LNA) stage, a time-gain compensation amplification stage, and an analog-to-digital conversion (ADC) stage. The LNA stage is used to amplify the low-voltage electrical signal in parallel. This is followed by a time-gain compensation amplification stage, which is used to amplify the signal with increasing depth. In the ADC stage, the amplified signal is digitized in parallel by a high-speed analog-to-digital converter (ADC).

[0093] The FPGA is used to control the ADC and record the samples, which are transferred via a high-speed data interface to the GPU, where the data is normalized, the analytic signal is obtained via a Hilbert transform, and a delay transform is performed to take into account the different path lengths of the different transducer elements, e.g., as described in connection with Figure 1.

[0094] The adaptive beamforming stage is also performed on the GPU. The beamformed data from multiple transmits is then envelope detected, log-compressed, and displayed on a display.

[0095] While a single FPGA is shown in FIG. 5A, multiple FPGAs may be present to control different subdivisions of the transducer array. In such cases, a synchronization mechanism may be present to synchronize transmission and reception across these multiple subarrays. A GPU can then perform adaptive beamforming on the data recorded via the single FPGA. Additional GPUs can be used to perform adaptive beamforming in a second stage by taking the partially beamformed data generated by individual FPGAs / GPUs in the system before envelope detection and log compression are performed for display.

[0096] FIG. 5B shows a schematic diagram of another system for ultrasound image reconstruction suitable for implementing the architecture of the present disclosure. This system uses multiple GPUs for processing. In this case, data recorded across elements can be divided among a subset of the GPUs (e.g., three of the four available GPUs), with partial beamforming performed on each GPU. A fourth GPU can be used to perform a second stage of beamforming, combining data acquired in the previous stages. While multiple GPUs are shown in this case, the system can also include multiple CPUs and / or additional FPGAs for computing analytic signals in a similar manner.

[0097] FIG. 6 illustrates certain functional operations of an adaptive beamforming method according to the present disclosure.

[0098] In step 602, a plurality of first electrical signals are received from respective ultrasonic transducer elements, each adapted to generate one of the first electrical signals in response to ultrasonic waves received from a location within the object.

[0099] In step 604, the first electrical signal is processed for input to a trained artificial neural network. This processing includes applying a predetermined delay profile to compensate for different path lengths to each transducer element. The first electrical signal may also be processed by amplifying the first electrical signal and digitizing the amplified first electrical signal. In certain embodiments, processing the first electrical signal further includes normalizing the digitized first electrical signal and applying a Hilbert transform to the normalized signal.

[0100] In step 606, the processed first electrical signal is input to a trained artificial neural network, which may include multiple convolutional layers and optionally further include forward-backward blocks.

[0101] In step 608, the trained artificial neural network determines apodization weights for beamforming the ultrasound signals.

[0102] In step 610, the trained artificial neural network outputs a beamformed signal representative of scattered ultrasound waves received from the location within the object based on a plurality of second electrical signals, each of the second electrical signals based on a respective sum of products, where each of the products is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective apodization weight, and each of the respective sums is based on less than all of all of the processed first electrical signals.

[0103] FIG. 7 illustrates certain functional operations in training a network architecture used for adaptive beamforming in accordance with the present disclosure.

[0104] In step 702, input training data is received that includes a plurality of processed first electrical signals corresponding to a plurality of first electrical signals received from respective ultrasound transducer elements, each ultrasound transducer element adapted to generate a respective one of the first electrical signals in response to ultrasound waves received from a location within the object, the processed first electrical signals corresponding to the first electrical signals after application of a predetermined delay profile.

[0105] In step 704, output training data is received, which includes a plurality of output electrical signals and apodization weights determined by an adaptive beamforming method, such as the MV technique described above.

[0106] In step 706, the input training data and the output training data are used to train an artificial neural network.

[0107] The description of various embodiments of the present disclosure has been presented for purposes of illustration and illustration, but is not intended to be exhaustive or to limit the invention to the precise form disclosed. Those skilled in the art will recognize that changes could be made to the above-described embodiments without departing from the broad inventive concept thereof.

[0108] Further particular and preferred aspects of the invention are set out in the accompanying independent and dependent claims. It will be appreciated that features of the dependent claims may be combined with features of the independent claims in combinations other than those explicitly set out in the claims.

Claims

1. 1. A method for generating beamformed ultrasound signals, the method comprising: receiving a plurality of first electrical signals from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasonic waves received from a location within the object; processing the first electrical signal, wherein said processing includes applying a predetermined delay profile; inputting the processed first electrical signal into a trained artificial neural network, wherein the trained artificial neural network includes a plurality of convolutional layers; determining apodization weights for beamforming ultrasound signals in a trained artificial neural network; and outputting, from the trained artificial neural network, beamformed signals representative of scattered ultrasound waves received from the location within the object based on a plurality of second electrical signals, where each second electrical signal is based on a respective sum of products, where each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective apodization weight, and where each sum is based on less than all of the processed first electrical signals. The method comprising:

2. The method of claim 1 , wherein the plurality of second electrical signals are based on respective plurality of the first electrical signals received from the location within the object at different times.

3. The method of claim 1 or 2, wherein processing the first electrical signal further comprises amplifying the first electrical signal and digitizing the amplified first electrical signal.

4. 4. The method of claim 3, wherein processing the first electrical signal further comprises normalizing the digitized first electrical signal and applying a Hilbert transform to the normalized signal.

5. The method of any one of claims 1 to 4, wherein the trained artificial neural network further comprises at least one beamforming operator layer.

6. The method of claim 5 , wherein the at least one beamforming operator layer is a forward-backward block.

7. 1. A method of providing a trained artificial neural network adapted to generate beamformed ultrasound signals, the method comprising: receiving input training data including a plurality of processed first electrical signals corresponding to a plurality of first electrical signals received from respective ultrasound transducer elements, wherein each of the ultrasound transducer elements is adapted to generate a respective first electrical signal in response to ultrasound waves received from a location within the object, and the processed first electrical signals correspond to the first electrical signals after application of a predetermined delay profile; receiving output training data including at least one of a plurality of output electrical signals and apodization weights determined by an adaptive beamforming method; and Training an artificial neural network by using input training data and output training data Including, wherein the artificial neural network includes a plurality of convolutional layers and at least one beamforming operator layer; and The method further includes: an artificial neural network adapted to output a beamformed ultrasound signal representing scattered ultrasound received from the location within the object based on a plurality of second electrical signals, wherein each of the second electrical signals is based on a respective sum of products, wherein each of the products is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a predetermined apodization weight, and wherein each of the sums is based on being less than all of the processed first electrical signals.

8. The artificial neural network is further adapted to determine predicted apodization weights for beamforming the ultrasound signal, output a plurality of predicted second electrical signals, and generate a predicted beamformed ultrasound signal based on the plurality of predicted second electrical signals; and 8. The method of claim 7, wherein training the artificial neural network comprises minimizing a mean squared error combining the predicted apodization weights and the predicted beamformed ultrasound signal.

9. A computer program comprising instructions that cause a computer to carry out the method of any one of claims 1 to 8 when the program is executed by a processor.

10. 1. A system for beamforming ultrasound signals, the system comprising: an input for receiving a plurality of first electrical signals from respective ultrasonic transducer elements, wherein each ultrasonic transducer element is adapted to generate a respective first electrical signal in response to ultrasonic waves received from a location within the object; processor means for processing the first electrical signal and applying a trained artificial neural network to the processed first electrical signal, wherein the neural network includes a plurality of convolutional layers and at least one beamforming operator layer, wherein the processor means: processing the first electrical signal by applying a predetermined delay profile; inputting the processed first electrical signal into a trained artificial neural network; determining apodization weights for beamforming the ultrasound signals; outputting a beamformed signal representing scattered ultrasound waves received from the location within the object based on a plurality of second electrical signals, wherein each of the second electrical signals is based on a respective sum of products, wherein each of the products is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective apodization weight, and wherein each of the sums is based on being less than all of the processed first electrical signals; the processor means adapted to perform an output section for outputting a beamformed ultrasound signal; The system comprising: