Adaptive ultrasound beamforming

A deep learning-based ultrasound beamforming method using CNNs and known operators addresses the limitations of existing techniques by predicting fewer cut-off weights, enhancing resolution and reducing training time while improving image quality and generalization.

CN120322696APending Publication Date: 2025-07-15UNIVERSITY OF LEEDS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084317.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-06
Filing Date
2023-12-07
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the lateral resolution of the ultrasonic beamformer is insufficient, making it difficult to achieve high-quality image reconstruction in the presence of interference, and deep learning architectures have limitations in computing complexity and generalization capabilities.

Method used

Using a neural network architecture containing convolutional layers and known operators, we use the predicted toe weights to combine with known operators to perform adaptive beam formation to reduce the impact of errors and improve image quality.

Benefits of technology

The lateral resolution and image quality of ultrasound images are improved, training time and computational complexity are reduced, and generalization ability of unseen data is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005437836230000131
    Figure BDA0005437836230000131
  • Figure BDA0005437836230000132
    Figure BDA0005437836230000132
  • Figure BDA0005437836230000161
    Figure BDA0005437836230000161
Patent Text Reader

Abstract

Ultrasound beamforming is performed by a deep learning architecture that combines convolutional layers and known operators. The network is trained to combine multiple observations of complex signal data into a single complex signal in an optimal manner. This is accomplished using convolutional layers with learnable parameters and by embedding operators to perform a multi-weight addition operation of a subset of the original observations and average these to produce a single complex signal. Known operators, including forward-backward operators, are optionally included in the architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and system for performing adaptive ultrasonic beamforming. Specifically, the present disclosure relates to a deep learning architecture for ultrasonic beamforming composed of one or more neural networks. The network or each network is trained to predict a set of apodization weights and perform apodization within the architecture. Background Art

[0002] Brightness mode (B-mode) ultrasonic image reconstruction is performed using the amplitude (brightness) of ultrasonic echoes from a medium irradiated by sound waves. The echo signals are measured with an array of sensors called array elements. Typically by applying the piezoelectric effect, each array element converts a pressure wave into an electrical signal. This electrical signal is usually amplified and then digitized using an analog-to-digital converter (ADC) at a specified sampling rate. Through a process called beamforming, data from these array elements are combined to estimate the amplitude of the echo from a point in the medium. This is achievable by applying spatial filtering, which produces a beam pattern with local maxima (called lobes) and minima. A common goal of all beamformers is to achieve a beam with a narrow main lobe (i.e., the lobe emitted in the direction of interest) and a low sidelobe level (i.e., the unwanted lobes outside the region of interest).

[0003] The narrow main lobe ensures sufficient lateral resolution of the system, which is achieved by applying focusing delays. The low sidelobe level ensures that interference from off-axis scatter is minimized. Strong scatterers in the sidelobes may be misinterpreted as weak scatterers in the main lobe, which may limit the lateral resolution. Sidelobes can be suppressed by applying element-by-element weights in the beamformer, which is called apodization. When the sidelobes are suppressed by apodization, the main lobe becomes wider. The most widely adopted beamformer is the Delay and Sum (DaS) beamformer with fixed apodization weights. The insufficient lateral resolution of the DaS beamformer limits its usefulness for anatomical imaging of small deep structures.

[0004] Data adaptive beamformers such as those utilizing minimum variance (MV) techniques can be employed to improve the lateral resolution. In such beamformers, the optimal apodization is calculated based on the statistics of the received data. Thus, in regions containing strong interference, the sidelobes are suppressed, while in regions where there is no strong interference, the sidelobes are retained. However, this algorithm is inherently complex, and achieving real-time processing remains a significant challenge.

[0005] Techniques such as subarray averaging (using multiple subarrays in covariance matrix estimation) and temporal averaging (using multiple time samples in covariance matrix estimation) can be employed to obtain better estimates. Diagonal loading (where coefficients of the identity matrix are added to the sample covariance matrix) and forward-backward averaging (where the imaging array and the angle of arrival of the signal (backward aperture) are flipped and then averaged with the original values (forward aperture)) can be applied to increase robustness in the presence of sound speed estimation errors.

[0006] Furthermore, the non-stationary nature of ultrasound limits the number of input data samples available for estimating the statistics required to determine the optimal apodization. This ultimately limits the accuracy of the estimation.

[0007] Deep learning architectures can overcome some limitations of the minimum variance algorithm. The ability of deep learning architectures to generalize over the large number of available statistics in the training dataset may provide a more accurate estimate of the optimal apodization weights in the presence of noise or sound speed estimation errors. Deep learning architectures can be efficiently implemented on a graphics processing unit (GPU) or a tensor processing unit (TPU). These devices typically include hardware acceleration and are designed to perform common operations in deep learning architectures, such as convolutional operations. Deep learning architectures generally consist of large, general-purpose, data-driven models with many parameters. Such models rely on a massive dataset that can fully characterize the inverse problem, as well as a network that is large enough and can achieve good generalization under a range of conditions. It is difficult to predict the performance of these data-driven models on unseen data. In addition, these networks may be inefficient in terms of computational complexity.

[0008] Any reference to prior art in this specification is not an admission or implication that such prior art forms part of the common general knowledge in any jurisdiction or globally, or that such prior art can be reasonably expected to be understood, regarded as relevant / or combined with other prior art by a person skilled in the art.

[0009] An object of the present invention is to at least ameliorate one or more of the above or other disadvantages of the prior art and / or provide a useful alternative. Summary of the Invention

[0010] The present invention is a method and system as defined in the appended claims.

[0011] It should be understood that the features and aspects of the present disclosure can be appropriately combined with other different aspects of the present disclosure, not just in the specific illustrative combinations described herein.

[0012] According to one aspect of the present disclosure, there is provided a method of generating a beamformed ultrasound signal, the method comprising:

[0013] Receiving a plurality of first electrical signals from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasound received from a location in an object;

[0014] Processing the first electrical signals, the processing including applying a predetermined delay profile;

[0015] Inputting the processed first electrical signals into a trained artificial neural network, the trained artificial neural network including a plurality of convolutional layers;

[0016] Determining in the trained artificial neural network tapering weights for beamforming the ultrasonic signals; and

[0017] Outputting from the trained artificial neural network a beamformed signal, the beamformed signal being based on a plurality of second electrical signals representing scattered ultrasound received from the location in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective one of the tapering weights, and each sum is based on less than all of the processed first electrical signals.

[0018] By generating a beamformed signal that is based on a plurality of second electrical signals representing scattered ultrasound received from the location in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective one of the first electrical signals and a respective second input based on a respective one of the tapering weights, and each sum is based on less than all of the first electrical signals, this provides the advantage of reducing the impact of errors in the first electrical signals. This in turn enables a reduction in the training time of the neural network and an improvement in the quality of the ultrasonic image data obtained during the inference phase of the artificial neural network.

[0019] The reduction in training time is due to a reduction in the number of trainable parameters because the network incorporates known operators to replace layers that would otherwise be trainable. Additionally, for a given input size, the network of the present disclosure is arranged to predict fewer tapering weights than the input size, which allows for compression through the network, which can accelerate inference in the prediction of the tapering weights.

[0020] The improvement in image quality is made possible because the signals are artificially decorrelated. The assumption in the minimum variance algorithm is that the desired signal is uncorrelated with the undesired signals. This is not the case for ultrasonic imaging where the interference is typically highly correlated. The covariance matrix can be calculated for several different sub-apertures. These can be artificially decorrelated by averaging to obtain a better estimate of the covariance matrix. It can be used to avoid signal cancellation.

[0021] In addition, compared with the full aperture, the artificial decorrelation of the present disclosure reduces the size of the covariance matrix and increases the number of observations. If the number of observations at this location is less than the number of array elements of the subarray, the covariance matrix may not be invertible.

[0022] For example, the reduction of the influence of errors in the first electrical signal enables the covariance matrix used in the minimum variance technique for creating learning data for an artificial neural network to be invertible, thereby enabling better estimation of the spatial covariance matrix to provide learning data for the artificial neural network, which in turn leads to improved image quality. The prediction of tapering weights fewer than the input size results in a smaller network with fewer operations, thereby increasing throughput.

[0023] The plurality of second electrical signals may be based on the respective plurality of first electrical signals received from the location in the object at different times.

[0024] This provides the advantage of further reducing the influence of errors in the first electrical signal.

[0025] In some embodiments, processing the first electrical signal may further include amplifying the first electrical signal and digitizing the amplified first electrical signal, and optionally, normalizing the digitized first electrical signal and applying a Hilbert transform to the normalized signal.

[0026] In some cases, the trained artificial neural network may further include at least one beamforming operator layer. The at least one beamforming operator layer may optionally be a forward-backward block.

[0027] According to another aspect of the present disclosure, a method of providing a trained artificial neural network adapted to generate beamformed ultrasound signals is provided, the method comprising:

[0028] Receiving input training data, the input training data including a plurality of processed first electrical signals corresponding to a plurality of first electrical signals received from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasound received from a location in an object, and wherein the processed first electrical signals correspond to the first electrical signals after applying a predetermined delay profile;

[0029] Receiving output training data, the output training data including tapering weights and at least one of a plurality of output electrical signals determined by an adaptive beamforming method; and

[0030] Training an artificial neural network using the input training data and the output training data;

[0031] Wherein, the artificial neural network is adapted to output a beamformed ultrasound signal, the beamformed ultrasound signal being based on a plurality of second electrical signals representing scattered ultrasound received from the position in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a predetermined apodization weight, and each sum is based on less than all of the processed first electrical signals.

[0032] In some embodiments, the artificial neural network is further adapted to determine a predicted apodization weight for beamforming the ultrasound signal, output a plurality of predicted second electrical signals, and generate a predicted beamformed ultrasound signal based on the plurality of predicted second electrical signals. Here, training the artificial neural network includes minimizing a combined mean square error of the predicted apodization weight and the predicted beamformed ultrasound signal.

[0033] According to another aspect of the present disclosure, there is provided a computer-readable medium carrying instructions which, when executed by a processor, cause the processor to perform the method as defined above.

[0034] According to another aspect of the present disclosure, there is provided a system for beamforming an ultrasound signal, the system comprising:

[0035] An input section for receiving a plurality of first electrical signals from respective ultrasound transducer elements, wherein each of the ultrasound transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasound received from a position in the object;

[0036] A processor device for processing the first electrical signals and for applying a trained artificial neural network to the processed first electrical signals, wherein the processor device is adapted to:

[0037] Process the first electrical signals by applying a predetermined delay profile;

[0038] Input the processed first electrical signals into the trained artificial neural network;

[0039] Determine an apodization weight for beamforming the ultrasound signal; and

[0040] Output a beamformed ultrasound signal, the beamformed signal being based on a plurality of second electrical signals representing scattered ultrasound received from the position in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective one of the apodization weights, and each sum is based on less than all of the processed first electrical signals.

[0041] An output unit for outputting the beamformed ultrasonic signal. Description of the Drawings

[0042] The subject matter of the present disclosure will be more fully understood and appreciated from the following detailed description in conjunction with the accompanying drawings, in which corresponding or similar numbers or characters indicate corresponding or similar components. Unless otherwise specified, the drawings provide exemplary embodiments or aspects of the present disclosure and do not limit the scope of the present disclosure. In the drawings:

[0043] Figure 1 An exemplary data processing flow is shown, in which the received raw sensor data is processed to obtain an analysis signal, a delay profile is applied, and adaptive beamforming is performed on the delayed analysis signal.

[0044] Figure 2 An example of a virtual source delay model for obtaining a delay profile for focused transmission is shown.

[0045] Figure 3 An embodiment of a network architecture for performing adaptive beamforming according to the present disclosure is shown.

[0046] Figure 4 A comparison between a beamformed image generated from data beamformed according to the present disclosure, an image generated using a conventional minimum variance technique, and an image obtained by a beamformer using uniform apodization in the prior art is shown.

[0047] Figure 5A A system for ultrasonic image reconstruction suitable for implementing the architecture of the present disclosure is schematically shown.

[0048] Figure 5B An alternative system for ultrasonic image reconstruction suitable for implementing the architecture of the present disclosure is schematically shown.

[0049] Figure 6 Certain functional operations of an adaptive beamforming method according to the present disclosure are shown.

[0050] Figure 7 Certain functional operations in the training of a network architecture used in adaptive beamforming according to the present disclosure are shown. Detailed Embodiments

[0051] The detailed description set forth below in connection with the accompanying drawings is intended as a description of various illustrative embodiments of the present disclosure and is not intended to represent the only forms in which the present disclosure may be practiced. It is to be understood that the same or equivalent functions may be implemented by different embodiments that are intended to be included within the scope of the present invention. Additionally, terms such as "comprises," "comprising," "has," "contains," or any other grammatical variations thereof are intended to cover non-exclusive inclusion, such that a module, circuit, device component, structure, and method step that includes a list of elements or steps does not include only those elements but may include other elements or steps not expressly listed or inherent to such module, circuit, device component, or step. Without further limitation, an element or step prefaced with "comprising...a" does not exclude the presence of additional identical elements or steps that include the recited element or step.

[0052] When traditional deep learning architectures perform data-adaptive beamforming, they need to rely on a massive dataset and a large enough network. As an alternative, a dedicated architecture with fewer parameters and incorporating prior knowledge can be adopted. By combining known operators, these networks have fewer parameters than pure data-driven architectures, and thus require fewer training samples. The combination of known operators can also improve the accuracy of model prediction and ensure that the solution has good generalization ability for unseen data.

[0053] The present invention provides a method for adaptive beamforming using a neural network. The network includes convolutional layers (i.e., layers that perform convolution or cross-correlation operations) and known operators to optimally combine observations of multiple real or complex signal data (from data recorded by different array elements) into a single real or complex signal. Compared with a delay-and-sum (DaS) beamformer with fixed apodization, this provides an image with excellent lateral resolution.

[0054] A known operator layer is incorporated to perform subarray aperture weighting and averaging processing. In the subarray aperture weighting layer, for each subarray of the observations obtained from the original input, a weighted sum is performed, and the weights are provided by the output of the previous convolutional layer. In the subarray averaging layer, a single complex signal is generated by determining the average value of the weighted sum data of each subarray.

[0055] The algorithmic constraints from the MV algorithm, subarray aperture weighting and averaging, are embedded in the network. Thus, the structural information in the input data is passed to the output through the final layers, which are specially designed to optimally combine the input information. This constraint helps to improve the generalization ability of the network for various data outside the training set.

[0056] Inside the neural network, a series of transformations are applied to the input to produce the desired output. The unknown parameters associated with this transformation are learned through a gradient descent algorithm, which is designed to minimize the error between the network output and the true values obtained by some other method (e.g., the minimum variance algorithm). The network can be trained by combining simulated data (computer simulations), data obtained from experimental phantoms (in vitro), or data obtained from patients in a clinical setting (in vivo).

[0057] In some embodiments, the network uses multiple time delay samples along the axis (defined as the dimension perpendicular to the surface of the ultrasound array) as input to determine multiple sets of apodization weights along the axis. The set of apodization weights is applied to the original observations in the sub-array aperture weighting layer of the network.

[0058] In some embodiments, the network processes sensor channel data (element space). Alternatively or additionally, the network processes multiple beamformed channel signals (beam space) obtained from a set of pre-beamformers.

[0059] When using an axially symmetric array for imaging, layers can be incorporated into the network that combine the input data (or a slice of the input data) with the reversed / flipped version of the data in one or more dimensions of the layer input tensor. This structure is referred to as the forward-backward module. Incorporating these layers can more accurately determine the optimal apodization function.

[0060] Although the specific examples discussed herein relate to 2-dimensional (2D) focused cardiac ultrasound imaging, the embodiments are not limited to echocardiography and can also be applied to imaging other anatomical features. The embodiments are not limited to focused imaging and can also be applied to non-focused imaging. The embodiments are not limited to 2D imaging and can also be applied to 3-dimensional (3D) volume imaging. Although pixel-based delay profiles are discussed herein, the embodiments can alternatively or additionally utilize line-by-line delay profiles. In some embodiments, beamformed data is acquired from a single transmit or simultaneously from multiple transmits.

[0061] For ease of understanding, specific examples shown in the drawings will be described. This is not intended to limit the scope covered in any way. Although this embodiment operates in element space, other embodiments combine beam space observations. Additionally, multiple embodiments can operate in concert to combine data at multiple stages.

[0062] Consider a linear array of M elements (i.e., a sensor array suitable for measuring ultrasonic echo signals). Each element can convert an ultrasonic signal received from reflections within the medium (i.e., a pressure wave scattered by an object in the medium) into an analog electrical signal. These M analog signals can be amplified and digitized (at a specified sampling rate) by an analog-to-digital converter (ADC) in an analog front end (AFE) to produce M digital signals or observations of the region of interest (ROI).

[0063] Figure 1 A possible data processing flow is shown, where the digital signals (i.e., samples) corresponding to the sensor data are normalized (step [1]) before obtaining the analytic signal via a Hilbert transform (step [2]). According to the present disclosure, a virtual source delay profile is applied (step [3]), and the delayed data is provided as input to an adaptive beamforming stage (step [4]).

[0064] The case of a batch size of N is shown in the figure, which corresponds to different delay profiles for each pixel row obtained from a single transmission. Each input has an axial dimension, an element dimension, and dimensions corresponding to the real and imaginary parts of the analytic signal.

[0065] In an example of this embodiment, a 2D transthoracic cardiac array consisting of 64 elements is used, and R axial samples are obtained for each channel from the recorded data of 64 channels R×64.( Figure 1 , step [1]). Other embodiments may use arrays of different geometries, such as curved arrays, matrix arrays, or those required for intravascular imaging. These arrays can consist of any number of elements and provide any number of observations. In other embodiments, fewer or more samples may be recorded, which will depend on the sampling rate of the system, the axial length of the ROI, and the transmission direction.

[0066] In Figure 1 step [1] of the embodiment in, the digital signals (i.e., samples) are normalized between -1 and 1, and the analytic signal is obtained via a Hilbert transform (step [2]). In other cases, the samples can be normalized between 0 and 1 before the transform to obtain analytic signal samples.

[0067] The delay profile used in step [3] of the embodiment is obtained from a virtual source model, where the focus is modeled as a spherical virtual source. This virtual source model is described in the following article: Kim, C., Yoon, C., Park, J.H., Lee, Y., Kim, W.H., Chang, J.M., Choi, B.I., Song, T.K. & Yoo, Y.M. (2013), "Evaluation of Ultrasound Synthetic Aperture Imaging Using Bidirectional Pixel-Based Focusing: Preliminary Phantom and In Vivo Breast Study", (IEEE Transactions on Biomedical Engineering, Vol. 60, pp. 2716 - 272).

[0068] The delayed signals are input into an adaptive beamforming architecture, step [4],

[0069] The initial processing chain can be different from the chain described with respect to Figure 1 There can be no normalization (phase [1]), and there can be no need to analyze the signal, such that the delay profile can be directly applied to the sample. However, to compensate for the path length differences of different transducer elements, delaying the data on the array (step [3]) is a necessary part of the input signal processing of this architecture.

[0070] Referring to Figure 2 , for elements 202, 204, and 206, the total time of the bidirectional propagation from the center of the array, k, to a point p in the ROI and back to element e can be described as:

[0071]

[0072] where c is the speed of sound in the medium and the transmit focus is selected at f. Using these propagation times and knowing the sampling frequency, a set of delay profiles can be pre-determined and used to collect appropriate samples.

[0073] In Figure 2 , the angles at which the delay profile is applied can be approximately calculated from the effective aperture and the focus as follows:

[0074]

[0075] where (x0,0) and (x M ,0) correspond to the positions of the first and last elements in the effective aperture ( Figure 2 ). Beyond these limiting angles, the pixels are masked.

[0076] From Figure 1 the output during the delay phase represents the input tensor of the architecture( Figure 1 , step [3]). The tensor dimension shape can be conventionally described as N×H×W×C. For the input tensor, N represents the batch size (a set of inputs at different lateral positions in the medium, where the lateral is defined as the dimension parallel to the transducer surface, i.e., Figure 2 the plane 210 of the element array in Figure 2 ), H represents the axial dimension (including samples at multiple axial positions in the medium, where the axial is defined as the dimension perpendicular to the ultrasonic array surface, i.e., parallel to

[0077] the axis 212 in Figure 1 ), W represents the array element dimension (i.e., the number of elements in the array), and C represents the real and imaginary parts of the analyzed signal.

[0078] In the embodiment shown in

[0079] step [3], N different delay profiles corresponding to N different lateral lines on the ROI can be obtained from a single transmission. In this embodiment, each lateral line consists of samples determined by the delay profile at 256 axial points, H = 256. Figure 1 In the example shown, a pixel-based software delay phase is used to directly obtain N lateral pixel lines from a single transmission through a virtual source model. Other examples may utilize non-focused transmissions (such as plane wave or diverging wave transmissions) with corresponding delay profiles for focused transmissions or line-by-line delay profiles. Embodiments can use parallel computing devices such as graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), or tensor processing units (TPUs) to perform beamforming on batch data collected from a single transmission or multiple simultaneous transmissions.

[0080] As described above, Figure 2 shows an example of a virtual source delay model for obtaining the delay profile for a focused transmission. It shows how to approximately calculate the propagation distance to a point in the region of interest (ROI) and back to the element of interest.

[0081] Figure 3 shows an embodiment of an application network architecture for performing adaptive beamforming according to the present disclosure. The network includes convolutional layers and blocks composed of known operations. In Figure 3 the embodiment shown, the layers are listed in Table 1.

[0082]

[0083] Table 1

[0084] This embodiment requires the analysis signals from a set of array element channels over the region of interest as input, e.g., the analysis signals transmitted at step [3] in Figure 1 For the input tensor, H represents the axial dimension (H = 256), W represents the array element dimension (W = 64), and C represents the real and imaginary parts of the analysis signal (C = 2).

[0085] The first stage of the network uses a 2D convolutional layer (layer 2) with a stride of (1, 2) and 32 filter channels, followed by a max pooling layer (layer 3). This first reduces the element dimension to the size of the sub-array (W = 32). Then, it forces the network to find a low-dimensional representation of the data. By performing this compression in the first layer of the network, the operations in subsequent layers are reduced, enabling a computationally efficient implementation. Although in this embodiment, the stride is chosen to reduce the sub-array length to half of the transducer size, in other embodiments, different stride lengths can be chosen to generate a tensor shape that matches other sub-array lengths.

[0086] In this embodiment, the first known operator appears in layer 4. In other embodiments, such a known operator may not be included. This layer takes the output of the max pooling layer as input and, subject to the constraint that its input C = 2W, divides the input into two H×W×W matrices of equal size:

[0087]

[0088] Construct a swap matrix W×w with dimension J. This layer performs a negation operation and a matrix multiplication along the H dimension between the last two dimensions of its input tensor and the swap matrix:

[0089]

[0090] Subsequently, B is composed with the original input:

[0091] O = A + B;

[0092] Thus, this layer reverses or flips A1 and -A2 along the last two dimensions and combines them with the original input. The first convolutional layer for transforming the input data for the forward-backward combination layer uses a linear activation function instead of a rectified linear unit (ReLU). This is done to preserve the negative part in the output tensor.

[0093] Five convolutional layers (layers 5 to 9) are used to transform the output data from the forward-backward combination layer (layer 4), followed by an upsampling layer (layer 10) and two additional convolutional layers (layers 11 and 12). The upsampling layer (layer 10) transforms the data such that H = 256 is the same as the original input size and W = 32 is the same as the subarray length selected in this embodiment. The last two convolutional layers (layers 11 and 12) transform the data such that C = 2 is the same as the original input and represents the real and imaginary parts of the complex target weights. Again, a linear activation function is used for the final convolutional layer (layer 12) to retain the negative parts of these complex target weights.

[0094] The final layers of the network (layers 13 and 14) perform element-wise weighting of the original input using the weights predicted by the final convolutional layer (layer 12). This is achieved by dividing the original input into two H×W1×1 tensors representing the real and imaginary parts of the analysis signal:

[0095]

[0096] And similarly, the predicted weights represented as B are divided into two H×W2×1 inputs:

[0097]

[0098] For subarray weighting and averaging, the constraint on the W dimension is W1 > W2. For Figure 3 the embodiment shown, the shape of the delay input tensor is 256×64×2 and the shape of the windowing weight tensor is 256×32×2.

[0099] The complex conjugate of the windowing weight can be determined by negating B2 (containing the imaginary part of the complex weight), and the partitioned data can be concatenated along the C dimension as:

[0100]

[0101] Two depthwise convolutions (where the H dimension is the depth dimension) are used with valid padding to perform windowing weights on the delay input data, with the filter of one set to B and the filter of the other set to B * . The outputs of these convolutions are concatenated along the C dimension to produce an output with shape H×(W1 - W2 + 1)×2. For Figure 3 the embodiment shown, this corresponds to an output tensor with shape 256×33×2 (for layer 13).

[0102] The final layer (layer 14) obtains the average of its input across the W dimension to obtain the final beamforming output with shape 256×1×2. This is achieved by using an average pooling layer with a pool size equal to (1, W1 - W2 + 1).

[0103] Combined, these two layers (i.e., layer 13 and layer 14) can be considered to embed known operators:

[0104]

[0105] where L = 32 is the subarray length, M = 64 is the number of array elements, s(z) is the computed analytic signal as a function of the axial position, w(z) H is the conjugate transpose of the weight as a function of the axial position, and, Y l (z) is the observed subarray, the original input of the architecture, as a function of the axial position.

[0106] In other embodiments, different filter sizes or strides in the convolutional layer, different numbers of convolutional layers can be utilized, and additional compression and expansion can be applied by additional pooling and upsampling layers.

[0107] Thus, an embodiment of the network consists of a CNN and a block containing known operators for predicting a set of apodization weights. The CNN uses multiple observations from multiple axial samples to predict axially correlated apodization weights. The convolutional layers in the CNN can utilize the rectified linear unit (ReLU) non-linear activation function, as they did in the previous embodiments. To retain the negative part of some layer outputs, which is necessary for some known operator blocks, they can also utilize other activation functions that retain negative values, such as the linear activation function.

[0108] Thus, the network architecture of the present disclosure embeds known beamforming operators to perform apodization weights, rather than simply using a neural network to output apodization weights. Although in the embodiments discussed above, the known weight function is implemented at the end of the neural network architecture, in certain embodiments, additional layers can be placed before these known operator layers, interleaved with these known operator layers, or placed after these known operator layers. Thus, the apodization weights are considered an inherent part of the neural network architecture of the present disclosure. The resulting network architecture does not predict apodization weights, but rather combines the input channel data in an optimal manner. This optimal combination is directly output. In certain embodiments, the minimum variance algorithm is used to determine the optimal combination of true values, while in other embodiments, it can be determined by another algorithm.

[0109] In some embodiments, the known operator includes subarray averaging (as in the above embodiments). In the case of performing subarray averaging, the inputs to the weighted known operator do not have equal widths. For example, the weights can be half the width of the contributing signals, and subarray averaging is performed. In this case, all contributing RF signals become inputs to the neural network. In most cases, the contributing RF signals should be the data recorded from each receiving element of the transducer. Then, the network can predict a set of weights to be applied to the subarray. Then, the resulting output from this known operator is averaged in another known operator layer to give an output.

[0110] In an array consisting of many elements, such as a matrix array, the contributing RF signals can be a subset of the data recorded on the array. In this case, multiple embodiments can combine operations to combine data in multiple stages. In the case of consisting of two stages, the contributing RF signals can be combined to produce a single output in the first stage. Multiple outputs from the first stage obtained from different subsets of the transducer elements are input into the second stage and optimally combined to produce a single output.

[0111] In the case of deploying the minimum variance algorithm to determine the optimal combination of true values, subarray averaging can increase the robustness of the output, which is important in medical imaging devices, but at the cost of sacrificing some resolution.

[0112] As described above, the architecture of the present disclosure introduces other known operators including forward-backward blocks. These operators are used to reduce the error in the predicted output relative to the true value of the MV algorithm. In addition, these other known operators are also used to regularize the training process by reducing the number of learnable parameters and the size of the network, which in turn can also improve the inference speed. The introduction of known operators may help ensure that the architecture generalizes well to previously unseen data, which is a key requirement for medical imaging.

[0113] Figure 4 A beamformed image (“predicted image”) generated from data beamformed using a deep learning beamformer is shown. An image (“boxcar image”) obtained by a beamformer with uniform apodization is included for reference and comparison.

[0114] Figure 4 Shown after performing envelope detection and log compression, by Figure 3 The image generated by the embodiment shown in. The network architecture in this embodiment is trained using simulated radio frequency (RF) ultrasound signals and the corresponding true values (“target image”) obtained via the minimum variance (MV) algorithm, where the weights are projected onto the signal subspace, called the eigen-based minimum variance (EBMV) algorithm.

[0115] The minimum variance technique for obtaining the training data set of the above embodiments can be more specifically understood as follows:

[0116] Considering an array of M array elements, M observations can be described by a column matrix:

[0117]

[0118] where, y m (t) = s m (t) + i m (t) + n m (t) is the analysis signal of each array element m at time t and contains the desired signal (s), interference (i), and noise (n) components.

[0119] As described in Synnevag, J.F., Austeng, A. & Holm, S. (2009), "Benefits of minimum-variance beamforming in medical ultrasound imaging", (IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control, Vol. 56, pp. 1868 - 1879), the sample covariance matrix combining subarray averaging and time averaging can be estimated according to the following:

[0120]

[0121] where, the symbol (·) H denotes conjugate transpose. The subarray length L is set to 32, which represents half of the aperture size of the phased array probe. The number of time samples K is selected such that 9 time samples are used in the covariance matrix estimation. These values are chosen to ensure that the sample covariance matrix is invertible and to preserve the speckle statistics. FB averaging is used to obtain a more accurate estimate of the covariance matrix. As described in Asl, B.M. & Mahloojifar, A. (2011). "Contrast enhancement and robustness improvement of adaptive ultrasound imaging using forward-backward minimum variance beamforming", (IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control, Vol. 58, pp. 858 - 867). The FB averaging is calculated as follows:

[0122]

[0123] where, J is the exchange matrix, the symbol (·) *Denotes conjugate. To increase the robustness of the beamformer in the presence of sound speed estimation errors, diagonal loading is applied. A shrinkage algorithm is used to automatically calculate the loading coefficients, as described in Stoica, P., Jian Li, Xumin Zhu, and Guerci, J. (2008), "On Using a Priori Knowledge in Space-Time Adaptive Processing", (IEEE Transactions on Signal Processing, Vol. 56, pp. 2598-2602). In this method, the shrinkage parameters α and β are calculated based on the received data, and the loaded covariance matrix is described as:

[0124]

[0125] The shrinkage parameters are determined as:

[0126]

[0127] where I is the identity matrix, and the parameters ν and ρ are:

[0128]

[0129] where N = (2K + 1)(M - L + 1) and Y l (q) = Y l (t - k).

[0130] To minimize the variance of the beamformer output while maintaining unit gain at the focus, a set of tapered weights is calculated using a covariance matrix estimate from: Capon, J, (1969), "High-resolution frequency-wavenumber spectrum analysis", (Proceedings of the IEEE, Vol. 57, pp. 1408-1418):

[0131]

[0132] where a is the steering vector, which is a vector of 1s if delays are applied beforehand. To reduce noise and improve contrast, the MVDR weights are projected onto the signal subspace, R d(t). [See Mehdizadeh, S., Austeng, A., Johansen, T.F., and Holm, S., (2012), "Eigenspace Based Minimum Variance Beamforming Applied to Ultrasound Imaging of Acoustically Hard Tissues", (IEEE Transactions on Medical Imaging, Vol. 31, pp. 1912 - 1921)], according to:

[0133]

[0134] Based on the estimation of the received signal noise, the signal - plus - interference subspace is automatically determined. The eigenvalue threshold is set for the eigenvectors corresponding to the eigenvalues in the sample covariance matrix greater than the scalar described as follows:

[0135]

[0136] where, s DAS (t) is the DaS beamforming signal; the average value of the entire array is:

[0137]

[0138] The output can be described as:

[0139]

[0140] To train the network of the above - mentioned embodiments, a suitable optimizer (e.g., NAdam optimizer) can be used to optimize the mean - squared error (MSE) loss. During training, batches containing 16 lateral lines, each with 256 axial samples, are used, and the data provided to the network in each epoch is randomized. The network is trained for up to 1000 epochs on simulated data. An early - stopping strategy is implemented with a patience value set to 50. A combined loss is used during training, which includes the mean - squared error (MSE) loss of predicting the apodization weights L W and the MSE loss of the final beamforming output L S as follows:

[0141] L total = λL S +(1 - λ)L W ;

[0142] where, λ is a user - defined parameter between 0 and 1. For the network trained in this example, λ = 0.6.

[0143] Figure 5AA system for ultrasonic image reconstruction is schematically shown that is suitable for implementing the architecture of the present disclosure. In this system, the imaging transmit sequence is programmed on a CPU. Any arbitrary and flexible waveform can be designed to probe the RoI. The transmit parameters (via a high-speed data link) are transmitted to a field-programmable gate array (FPGA), which controls a series of high-voltage transmit circuits (Tx excitation) in parallel. These are used to excite the piezoelectric elements of the transducer ("transducer elements") and generate ultrasonic pressure waves. The high-voltage signal passes through a transmit / receive switch (TX / RX switch) to prevent damage to the receiving electronics.

[0144] At the transducer elements, the pressure waves scattered by the object (e.g., "scatterers") are received and converted into (analog) electrical signals.

[0145] The TX / RX switch directs the received electrical signal into a receiver circuit (RX LNA / ADC), which includes a low-noise amplification (LNA) stage, a time-gain compensation amplification stage, and an analog-to-digital conversion (ADC) stage. The LNA stage is used to amplify the low-voltage electrical signals in parallel. This is followed by a time-gain compensation amplification stage, which is used to amplify the signals for increasing depths. In the ADC stage, the amplified signals are digitized in parallel by a high-speed analog-to-digital converter (ADC).

[0146] The FPGA is used to control the ADC and record the samples. These can be transmitted via a high-speed data interface to a GPU, where the data is normalized, the analytical signal is obtained via a Hilbert transform, and a delay transform is performed to account for the different path lengths of the different transducer elements, e.g., as described with respect to Figure 1 as described.

[0147] An adaptive beamforming stage is also performed on the GPU. Then, the beamformed data from multiple transmissions can be subjected to envelope detection, logarithmic compression, and displayed on a monitor.

[0148] Although a single FPGA is shown in Figure 5A there can be multiple FPGAs to control different sub-parts of the transducer array. In this case, a synchronization mechanism can be employed to achieve synchronization of transmission and reception across multiple sub-arrays. Then, the GPU can perform adaptive beamforming on the data recorded via a single FPGA. An additional GPU can be used, in a second stage, to perform adaptive beamforming for display by acquiring the partial beamformed data generated by each individual FPGA / GPU in the system before performing envelope detection and logarithmic compression.

[0149] Figure 5BAn alternative system for ultrasonic image reconstruction, which is schematically shown and suitable for implementing the architecture of the present disclosure, is presented. In this system, multiple GPUs are used to perform processing. In this case, the data recorded across the array elements can be split among a subset of the GPUs (e.g., three out of four available GPUs), and partial beamforming can be performed on each GPU. The 4th GPU can be used to perform the second stage of beamforming, in which the data obtained from the previous stage is combined. Although multiple GPUs are shown in this case, the system can also include multiple CPUs and / or additional FPGAs in a similar manner for computational analysis of signals.

[0150] Figure 6 Certain functional operations of an adaptive beamforming method according to the present disclosure are shown.

[0151] In step 602, a plurality of first electrical signals are received from corresponding ultrasonic transducer elements. Each ultrasonic transducer element is adapted to generate a corresponding one of the first electrical signals in response to ultrasonic waves received from a location in the object.

[0152] In step 604, the first electrical signals are processed for input into a trained artificial neural network. This processing includes applying a predetermined delay profile to compensate for different path lengths to the corresponding transducer elements. The first electrical signals can also be processed by amplifying the first electrical signals and digitizing the amplified first electrical signals. In certain embodiments, the processing of the first electrical signals further includes normalizing the digitized first electrical signals and applying a Hilbert transform to the normalized signals.

[0153] In step 606, the processed first electrical signals are input into a trained artificial neural network. The network includes a plurality of convolutional layers and optionally may also include forward-backward blocks.

[0154] In step 608, the trained artificial neural network determines the apodization weights for beamforming ultrasonic signals.

[0155] In step 610, the trained artificial neural network outputs beamformed signals, which are based on a plurality of second electrical signals and represent the scattered ultrasonic waves received from the location in the object. Each of the second electrical signals is based on a corresponding product sum, where each product is the product of a corresponding first input based on the corresponding processed first electrical signal and a corresponding second input based on the corresponding apodization weight, and each sum is based on less than all of the processed first electrical signals.

[0156] Figure 7 Certain functional operations in the training of a network architecture used in adaptive beamforming according to the present disclosure are shown.

[0157] In step 702, input training data is received, the input training data including a plurality of processed first electrical signals corresponding to a plurality of first electrical signals received from respective ultrasonic transducer elements. Each ultrasonic transducer element is adapted to generate a respective one of the first electrical signals in response to ultrasound received from a location in the object, and the processed first electrical signals correspond to the first electrical signals after application of a predetermined delay profile.

[0158] In step 704, output training data is received. The output training data includes apodization weights and a plurality of output electrical signals determined using an adaptive beamforming method such as the MV technique outlined above.

[0159] In step 706, an artificial neural network is trained using the input training data and the output training data.

[0160] The description of the various embodiments of the present disclosure has been presented for purposes of illustration and example, but is not intended to be exhaustive or to limit the invention to the disclosed forms. Those skilled in the art will appreciate that changes may be made to the above embodiments without departing from its broad inventive concept.

[0161] Other specific and preferred aspects of the invention are set forth in the appended independent and dependent claims. It should be understood that the features of the dependent claims may be combined with the features of the independent claims in combinations different from those explicitly set forth in the claims.

Claims

1. A method for generating a beamformed ultrasound signal, the method comprising: Receiving a plurality of first electrical signals from respective ultrasound transducer elements, wherein each of the ultrasound transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasound received from a location in an object; Processing the first electrical signals, the processing including applying a predetermined delay profile; Inputting the processed first electrical signals into a trained artificial neural network, the trained artificial neural network including a plurality of convolutional layers; Determining in the trained artificial neural network tapering weights for beamforming the ultrasound signal; And Outputting from the trained artificial neural network a beamformed signal, the beamformed signal being based on a plurality of second electrical signals representing scattered ultrasound received from the location in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective one of the tapering weights, and each sum is based on less than all of the processed first electrical signals.

2. The method according to claim 1, wherein, The plurality of second electrical signals are based on respective ones of the plurality of first electrical signals received from the location in the object at different times.

3. The method according to claim 1 or 2, wherein Processing the first electrical signals further includes amplifying the first electrical signals and digitizing the amplified first electrical signals.

4. The method according to claim 3, wherein, Processing the first electrical signals further includes normalizing the digitized first electrical signals and applying a Hilbert transform to the normalized signals.

5. The method according to any one of claims 1 to 4, wherein, The trained artificial neural network further includes at least one beamforming operator layer.

6. The method according to claim 5, wherein The at least one beamforming operator layer is a forward-backward block.

7. A method for providing a trained artificial neural network adapted to generate a beamformed ultrasound signal, the method comprising: Receiving input training data, the input training data including a plurality of processed first electrical signals corresponding to a plurality of first electrical signals received from respective ultrasound transducer elements, wherein each of the ultrasound transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasound received from a location in an object, and wherein the processed first electrical signals correspond to the first electrical signals after applying a predetermined delay profile; Receiving output training data, the output training data including tapering weights and at least one of a plurality of output electrical signals determined by an adaptive beamforming method; and Training the artificial neural network using the input training data and the output training data; Wherein the artificial neural network includes a plurality of convolutional layers and at least one beamforming operator layer; and Wherein the artificial neural network is adapted to output a beamformed ultrasound signal, the beamformed ultrasound signal being based on a plurality of second electrical signals representing scattered ultrasound received from the location in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a predetermined tapering weight, and each sum is based on less than all of the processed first electrical signals.

8. The method according to claim 7, wherein The artificial neural network is also adapted to determine tapering weights for predicting beamforming of the ultrasonic signal, output a plurality of predicted second electrical signals, and generate a predicted beamformed ultrasonic signal based on the plurality of predicted second electrical signals; and wherein training the artificial neural network includes minimizing a combined mean square error of the predicted tapering weights and the predicted beamformed ultrasonic signal.

9. A computer program comprising instructions that, when executed by a processor, cause the computer to perform the method according to any one of the preceding claims.

10. A system for beamforming an ultrasonic signal, the system comprising: an input section for receiving a plurality of first electrical signals from respective ultrasonic transducer elements, wherein each of the ultrasonic transducer elements is adapted to generate a respective one of the first electrical signals in response to ultrasonic waves received from a location in an object; a processor device for processing the first electrical signals and for applying a trained artificial neural network to the processed first electrical signals, the neural network comprising a plurality of convolutional layers and at least one beamforming operator layer, wherein the processor device is adapted to: process the first electrical signals by applying a predetermined delay profile; input the processed first electrical signals into the trained artificial neural network; determine tapering weights for beamforming the ultrasonic signal; output a beamforming signal, the beamforming signal being based on a plurality of second electrical signals representing scattered ultrasonic waves received from the location in the object, wherein each of the second electrical signals is based on a respective product sum, wherein each product is a product of a respective first input based on a respective processed first electrical signal and a respective second input based on a respective one of the tapering weights, and each sum is based on less than all of the processed first electrical signals. an output section for outputting the beamformed ultrasonic signal.