Sonography device

The sonography device with a flexible ultrasound transducer array and machine learning-enhanced spatial data combination addresses image quality and user variability issues, providing a wide-view, accurate three-dimensional image.

WO2026039873A1PCT designated stage Publication Date: 2026-02-26COMMONWEALTH SCI & IND RES ORG +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2025/050916
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-23
Filing Date
2025-08-21
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Ultrasound devices face challenges such as low image quality, limited depth of field, artifacts, noise, and user variability due to the difficulty in correctly setting the orientation of the ultrasound probe, which is influenced by anatomical features and individual body characteristics.

Method used

A sonography device with a flexible array of multiple ultrasound transducers that tracks spatial positions and combines reflection data using a controller, applying machine learning models for accurate spatial transformation and beamforming to create a wide-view three-dimensional image.

Benefits of technology

The device provides a wide-view, robust image that compensates for changes in array placement, offering high accuracy and visibility of features otherwise obscured, with improved image quality and reduced user dependency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050916_26022026_PF_FP_ABST
    Figure AU2025050916_26022026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to a sonography device comprising a flexible array of multiple ultrasound transducers, each of the multiple ultrasound transducers being configured to generate an ultrasound signal and to sense a reflection of the ultrasound signal to generate reflection data comprising overlapping reflection data that covers a spatially overlapping field of view with other ultrasound transducers; a spatial measurement system configured to determine a relative spatial position related to the multiple ultrasound transducers; and a controller configured to spatially combine the reflection data based on the relative spatial position and based on the overlapping reflection data.
Need to check novelty before this filing date? Find Prior Art

Description

"Sonography device"Cross-Reference to Related Applications

[0001] The present application claims priority from Australian Provisional Patent Application No 2024902640 filed on 23 August 2024, the contents of which are incorporated herein by reference in their entirety.Technical Field

[0002] This disclosure relates to a device for sonography, such as ultrasound imaging.Background

[0003] Ultrasound devices are based on the principle of using high-frequency sound waves to create images of internal structures of the body. The sound waves are emitted by a device called a transducer, which is placed on the skin over the area of interest. The transducer also receives the reflected sound waves, which are processed by a computer to form an image on a screen.

[0004] The quality and resolution of the ultrasound image depend on several factors, such as the frequency and shape of the sound waves, the number and arrangement of the transducers, the speed and attenuation of the sound waves in different tissues, and the processing and display methods of the computer. The frequency of the sound waves determines the penetration depth and the resolution of the image. Higher frequencies can provide higher resolution, but they are more easily absorbed by the tissues and have lower penetration depth. Lower frequencies can penetrate deeper, but they have lower resolution and may miss small details. The optimal frequency for ultrasound imaging depends on the type and location of the tissue being scanned.

[0005] The speed and attenuation of the sound waves in different tissues affect the contrast and brightness of the image. The speed of the sound waves depends on the density and elasticity of the tissue. Different tissues have different speeds, which causethe sound waves to bend and change direction when they cross the boundaries between the tissues. This phenomenon is called refraction, and it can cause errors and distortions in the image. The attenuation of the sound waves depends on the absorption and scattering of the tissue. Different tissues have different attenuation coefficients, which cause the sound waves to lose energy and amplitude when they travel through the tissues. This phenomenon is called attenuation, and it can affect the signal-to-noise ratio and the dynamic range of the image.

[0006] The received sound waves are processed by a computer to generate an image. The processing and display methods of the computer affect the quality and interpretation of the image. The computer receives the signals from the transducers, which contain information about the amplitude, frequency, phase, and direction of the sound waves. The computer performs various operations on these signals, such as filtering, amplification, compression, enhancement, and reconstruction, to produce an image that is suitable for visualization and analysis. The computer can also apply different algorithms and techniques to improve the resolution, contrast, and accuracy of the image, such as harmonic images and speckle reduction.

[0007] Ultrasound devices can be used for medical diagnosis and treatment, as they are non-invasive, safe, relatively inexpensive, and portable. They can provide real-time and dynamic images of various organs and tissues, such as the heart, blood vessels, liver, kidney, brain, bone, and muscle. They can also detect and measure various physiological parameters, such as blood flow, pressure, velocity, and elasticity. However, ultrasound devices also have some limitations and challenges, such as low image quality, limited depth of field, artifacts, noise, and user variability.

[0008] Further, the operation of ultrasound devices typically requires significant training by the user to acquire the required skills. The main reason is that the suitability of the acquired image depends on the orientation of the ultrasound probe, which can be difficult to set correctly for the given circumstances, such as, anatomical feature to be imaged or individual body characteristics. Other factors may include limited field ofview, image dependency on probe pressure applied, and physiological variations of human body.

[0009] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each of the appended claims.

[0010] Throughout this specification the word "comprise", or variations such as "comprises" or "comprising", will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.Summary

[0011] This disclosure provides an improved sonography device that comprises a flexible array of multiple ultrasound transducers. The array is tracked in space to process the signals from the array. Further, the signals from the array are registered with each other, or more generally the information from the array is combined, based on overlapping image regions. This can create a wide-view three-dimensional image that is robust against changes in the placement of the ultrasound array.

[0012] A sonography device comprises a flexible array of multiple ultrasound transducers, each of the multiple ultrasound transducers being configured to generate an ultrasound signal and to sense a reflection of the ultrasound signal to generate reflection data comprising overlapping reflection data that covers a spatially overlapping field of view with other ultrasound transducers; a spatial measurement system configured to determine a relative spatial position related to the multiple ultrasound transducers; and a controller configured to spatially combine the reflection data based on the relative spatial position and based on the overlapping reflection data.

[0013] It is an advantage that the controller combines the reflection data based on the overlapping reflection data. This way, the accuracy of combining the reflection data is improved. This improved accuracy enables pixel-accurate imaging, which is difficult with only sensor based position measurements. In this way, the device compensates for a change in relative position caused by flex of the array accurately. As a result, the device can be made relatively large and adjust to the surface of the body while retaining a high accuracy in the spatial combination of reflection data. Therefore, the device can provide a wide view and is more robust against variations in the placement of the device on the body. Further, the wide view produces features that may otherwise not be visible.

[0014] In some embodiments, the flexible array comprises multiple rigid sub-arrays, each of the rigid sub-arrays comprising multiple transducers and the rigid sub-arrays are flexibly connected to form the flexible array.

[0015] In some embodiments, the spatial measurement system is configured to determine the relative spatial position of each of the rigid sub-arrays.

[0016] In some embodiments, the controller is configured to combine multiple visual representations generated by respective sub-arrays.

[0017] In some embodiments, combining the multiple visual representations comprises determining a spatial transformation to the multiple visual representations to register the multiple visual representations.

[0018] It is an advantage that the device is able to create a wide view anatomical representation of the imaged anatomy.

[0019] In some embodiments, determining the spatial transformation comprises applying a trained machine learning model to the multiple visual representations, the trained machine learning model comprising outputs indicative of the spatial transformation.

[0020] In some embodiments, the controller is further configured to modify generating the ultrasound signal based on the relative spatial position of the multiple ultrasound transducers.

[0021] In some embodiments, modifying the ultrasound signal comprises applying a trained machine learning model to the reflection data of a previous iteration.

[0022] A method for processing ultrasound signals comprises receiving reflection data indicative of a sensed reflection of an ultrasound signal by multiple ultrasound transducers of a flexible array of ultrasound transducers, the reflection data comprising overlapping reflection data that covers a spatially overlapping field of view with other ultrasound transducers; receiving position data indicative of a relative spatial position of multiple ultrasound transducers; and combining the reflection data spatially based on the position data and based on the overlapping reflection data.

[0023] In some embodiments, the method further comprises determining, based on the position data and based on the overlapping reflection data, signal parameters of the ultrasound signals for controlling the ultrasound transducers.

[0024] In some embodiments, the method further comprises receiving multiple visual representations comprising a visual representation for each of multiple rigid sub-arrays of the flexible array; and registering the multiple visual representations to determine a three-dimensional ultrasound representation.

[0025] In some embodiments, registering the multiple visual representations comprises performing a coarse registration using the position data; and performing a fine registration by applying a trained machine learning model to the reflection data.

[0026] In some embodiments, the machine learning model comprises outputs indicative of respective spatial transformations of the multiple visual representations.

[0027] In some embodiments, the trained machine learning model is a CNN or a transformer model.

[0028] In some embodiments, using the transformer model comprises applying the transformer model on a first visual representation and an adjacent second visual representation of the multiple visual representations and the transformer model comprises an output indicative of spatial transformations of the second visual representation that register the second visual representation with the first visual representation.

[0029] In some embodiments, applying the transformer model comprises splitting the first visual representation and the second visual representation into contiguous subsets, projecting the subsets into an embedding space and applying self-attention between parts of the first visual representation and between parts of the second visual representation.

[0030] In some embodiments, the method further comprises training the machine learning model using training data comprising labels indicative of a spatial transformation that registers a first visual representation with a second visual representation by reducing an error between the labels and the outputs of the machine learning model.

[0031] In some embodiments, the method further comprises training the machine learning model to reduce a similarity loss value applied between a first visual representation and a second visual representation .

[0032] In some embodiments, the loss value is defined by a global loss function to align more than two partially overlapping visual representations simultaneously.

[0033] In some embodiments, the method further comprises determining parameters of generating the ultrasound signal based on the position data.

[0034] In some embodiments, the method is performed in real-time to complete the step of combining the reflection data before receiving further reflection data.

[0035] In some embodiments, the method comprises determining, from the combined reflection data, a three-dimensional ultrasound representation comprising a three- dimensional image comprising intensity values for multiple voxels in a three dimensional space.

[0036] In some embodiments, the method further comprises, determining from the combined reflection data, a segmentation of three-dimensional objects in a volume covered by the reflection data .

[0037] In some embodiments, determining the segmentation is further based on automated measurements.

[0038] In some embodiments, the method comprises, classifying the combined reflection data into one or more diagnoses.

[0039] Software, when executed by a computer, causes the computer to perform the above method.Brief Description of Drawings

[0040] Examples will now be described with reference to the following drawings:

[0041] Figure 1 illustrates a sonography device.

[0042] Figure 2 illustrates the field of view of transducers on an object.

[0043] Figure 3 illustrates a method for processing ultrasound signals.

[0044] Figure 4 is a hardware technology schematic of a sonography device.

[0045] Figure 5a illustrates 2D CMUT arrays and Figure 5b illustrates ID CMUT arrays embedded on a flexible patch. Note that while only sixteen arrays are shown in the figures (a) and (b), the patch is completely covered with arrays located in close proximity.

[0046] Figures 6a, 6b, 6c illustrate optimiser method vs Unsupervised and Supervised Transformer Al algorithms for rigid registration of partially overlapping volumes.Description of Embodiments

[0047] Figure 1 illustrates a sonography device 100 which can be used to image body parts of humans animals or non-body parts, such as machinery, ground or others. Sonography device 100 comprises a flexible array 101 of multiple ultrasound transducers, such as first transducer 102 and second transducer 103. Each of the multiple ultrasound transducer of array 101 are configured to generate an ultrasound signal 104 and to sense a reflection 105 of the ultrasound signal off an object 106, such as an human organ. It is noted that each transducer may comprise multiple elements, such as signal sources or sensors. While the elements may operate as both, source and sensor, they may equally operate only as source or only as sensor.

[0048] Device 100 receives the reflection and converts it to a digital signal, such as by analog-digital (A / D) conversion of a voltage indicative of a pressure or pressure change in individual signal elements, to generate reflection data. Since the transducers, such as 102 and 103, are located relatively close to each other with respect to their field of view, the reflection data comprises overlapping reflection data that covers a spatially overlapping field of view. In other words, some points of object 106 are within the field of view of first transducer 102 as well as within the field of view of second transducer 103.

[0049] Device 100 further comprises a spatial measurement system 107 configured to determine a relative spatial position related to the multiple ultrasound transducers and acontroller 108 configured to spatially combine the reflection data based on the relative spatial position and based on the overlapping reflection data.

[0050] As can be seen in the example of this figure, there are image points indicated at 201, which might be imaged only by the first transducer 102. Similarly, there are image points indicated at 202, which are imaged only by the second transducer. It is noted, however, that in other examples there are no image points that are imaged by only a single transducer. Between those points, there are image points indicated at 203, which are imaged by both the first transducer 102 and the second transducer 103. Controller 108 combines the reflection data from first transducer 102 with the reflection data from second transducer 103 based on the overlapping reflection data representing the points at 203. Combining refers to the process of bringing together two or more pieces of data to form a unified whole. This can involve merging, blending, or integrating distinct parts so that they function as one entity or produce a collective result. The collective result may be a stitched image, combined field of view or combined set of parameters. In some cases, the overlapping areas are merged to have a single representation of the overlapping area formed from two or more datasets. There are a number of different ways of combining the reflection data based on the overlapping reflection data. Effectively, controller 108 uses the overlapping reflection data to combine the reflection data. More specifically, controller uses first reflection data from first transducer 102 and second reflection data from second transducer 103, where both transducers have an overlapping field of view. Controller 108 then uses the first and second reflection data to combine the first with the second reflection data. That is, controller 108 combines the first and second reflection data into combined reflection data that represents the combined view of the first and second transducer. Combining the reflection data may be achieved by applying a function to the first and second reflection data or by iteratively optimising the combination parameters or the parameters of a spatial transformation applied to one or both reflection data. For example, controller 108 optimises spatial transformation parameters, such as rotation and translation to the reflection data of the second transducer 103 until the overlapping points 203 are aligned optimally. Further, controller 108 may apply a machine learning model to the data to obtain the spatial transformation parameters. In some examples,controller 108 first performs a course alignment using the relative spatial position and then a fine alignment using overlapping image points to thereby combine the reflection data.

[0051] It is noted that the term “transducer” may refer to a single signal source and sensor or to multiple signal sources and sensors that operate as one transducer. In the latter case, it is possible to perform beamforming by applying delays to some signals of the multiple signal sources. The multiple signal sources of one transducer may be connected rigidly, such as on a rigid plate as illustrated in Figure 1 considering that each transducer (rectangle) comprises multiple signal sources and sensors. In this sense, each transducer 102, 103 may also be referred to as a rigid sub-array. The multiple subarrays are connected flexibly to create the flexible array 101.

[0052] The multiple signal sources and sensors of one sub-array can be arranged in a line to generate a 2-dimensional image or in a plane to generate a 3-dimensional image. Those images are then combined using overlapping image data. The relative location related to the multiple ultrasound transducers may therefore comprise the relative location of each source and sensor, or the relative location of each transducer comprising multiple sources and sensors. Relative location may refer to a difference in x, y, z direction between first transducer 102 and second transducer 103. The relative location may also comprise an angle in space (comprising three individual angle values) between the first transducer 102 and the second transducer 103.Spatial measurement

[0053] As set out above, spatial measurement system 107 determines a relative spatial position related to the multiple ultrasound transducers. In one example, the spatial measurement system collects high-resolution shape information using a fibre optic shape sensor (FOSS). The FOSS operates on the principle of optical frequency domain reflectometry (OFDR), which involves sweeping a laser across a range of optical frequencies to create interference fringes. These fringes are then transformed into complex reflection coefficients, enabling the detection of local strain changes in the optical core. The integration of FOSS into soft actuators provides a capability to detectbending, twisting, and axial strain with high precision. The sensor's design may feature a multicore fibre with a central core and three outer cores that spiral around it, which is advantageous for determining the bend radius and direction. The sensor's may be able to provide submillimeter and sub-degree resolution measurements of the array’s shape, body twist, and position in three-dimensional space.

[0054] In that sense, the spatial measurement system 107 may comprise a fibre optic tracking system, which may include a light source, an optical fibre with a sensing component, a detection unit, and a calculation unit. The optical fibre is attached to the array transducers, and the sensing component is designed to modify the light signals based on the movement of the object. The detection unit receives these modified signals and the calculation unit interprets them to determine the transducers’ spatial information in six degrees of freedom, which includes its position and orientation in three-dimensional space. The spatial measurement system 107 may use a monolithic, multicore FOSS to capture the shape, body twist, and array position of a transducer in 3D space with high precision.

[0055] The spatial measurement system 107 may be configured so that it measures the location of each of the transducers 102, 103, etc. Since those transducers are rigid, a single set of location information is sufficient to derive the location of every source / sensor on the transducer, or to assign that location to an image generated by that transducer.Combining data

[0056] As mentioned above, controller 108 is configured to combine multiple visual representations generated by respective sub-arrays. That is, the sub-arrays provide the reflection signals and there may be a separate controller on each of the sub-arrays to calculate an image from the sensed reflections 105, which may also include beamforming to direct the signal 104 at a specific direction to improve image quality. In other examples, the raw signal data is processed by controller 108 to calculate a visual representation for each sub-array and then combine the visual representations based on spatial location and overlapping image data. The visual representations maybe one-dimensional (i.e. a line comprising a list of line values), two-dimensional (i.e. images comprising pixels), or three-dimensional (volumes comprising voxels). Controller 108 may also perform beamforming.

[0057] Controller 108 may combine the multiple visual representations (lines, images, volumes) by determining a spatial transformation to the multiple visual representations to register the multiple visual representations. The spatial transformation may comprise a translation in one, two or three dimensions and may comprise a rotation about one, two or three axes. This way, the device is able to create a wide view representation of the imaged anatomy. Controller 108 may determine the spatial transformation by applying a trained machine learning model to the multiple visual representations as further described below. The trained machine learning model may comprise outputs indicative of the spatial transformation, such as one output for each translation axis and one output for each rotation axis.

[0058] It is noted that the data that is being combined may comprise related to brightness mode (B-mode) ultrasound, which is an imaging technique that displays a two-dimensional cross-sectional image of tissue. It works by sending ultrasound waves into the body and detecting the echoes that bounce back from internal structures. The intensity of each echo is represented as a brightness on the screen, creating a grayscale image that shows the internal anatomy. It is noted, however, that the spatial transformations can be applied to data generated by the transducers in other modalities, such as Doppler modalities (colour, power and spectral), elastography and others.

[0059] Doppler ultrasound is based on the Doppler effect of a change in signal frequency depending on the speed of the imaged object. The Doppler modalities of ultrasound can be used to assess and visualise blood flow within the body. Colour Doppler displays the direction and velocity of blood flow by overlaying colour-coded information onto a conventional grayscale ultrasound image, allowing clinicians to quickly identify flow patterns and potential abnormalities. Power Doppler is more sensitive to the presence of blood flow, especially in small vessels or in regions where flow is weak, and it displays the intensity of flow without indicating direction, resultingin a more detailed depiction of vascularity. Spectral Doppler provides a graphical representation of blood flow velocities over time, typically in the form of a waveform, which enables the measurement of peak velocities and the analysis of flow characteristics within a specific vessel.

[0060] These modalities can be near-simultaneously acquired from the same position as the B-mode image, meaning that the different modalities are implicitly registered. This means controller 108 can map the additional modality information using the same spatial transformations (i.e. controller 108 propagates images acquired using the additional modalities to the same three-dimension space using the rigid transformations determined from the B-mode images). As a result controller 108 generates a combined three-dimensional representation from these additional ultrasound derived modalities.Beamforming

[0061] Beamforming is a technique for image formation from a raw signal. As such, beam forming can be used to enhance the resolution and contrast of ultrasound images by applying weights to the signals received by each element of an array transducer. The weights can be adjusted to steer or focus the beam in a desired direction, or to suppress noise or interference from other directions. Beamforming can be performed in analog or digital domain, or in a combination of both. Analog beam forming uses analog delay lines and summing circuits to combine the signals from different elements. Digital beam forming converts the signals to digital samples and applies digital filters and delays before summing them. Digital beam forming offers more flexibility and accuracy, but also requires more computational power and memory.

[0062] One challenge of beamforming is to account for the relative spatial position of the multiple ultrasound transducers that form the array. The position of each transducer affects the phase and amplitude of the received signals, and hence the resulting image quality. To compensate for the position variations, the controller 108 may modify the weights or delays applied to each signal based on the estimated position of each transducer. The position estimation can be done using sensors such as optical fibres attached to the transducers, or by using image registration techniques that align theimages from different transducers based on their overlapping regions as disclosed herein. Alternatively, the controller may use a machine learning approach to learn the optimal weights or delays for each transducer based on the reflection data of a previous iteration. The machine learning model can be trained on a large dataset of ultrasound images and corresponding transducer positions, and can adapt to the changes in the position dynamically.

[0063] In one example, each transducer 102, 103 performs beamforming individually and independently from the other transducers. In that case, each transducer 102, 103 may generate a separate image (ID, 2D or 3D) that are then combined, which may also be referred to as mosaicking or stitching. This approach has the advantage that beamforming is simplified because the multiple signal sources and sensors are rigidly connected, so their relative spatial location is not required for beamforming. The beamforming for each transducer sub-array may be performed by a processor specific for each transducer and independently of other transducers or by central controller 108.

[0064] In another example, the controller 108 controls the entire array 101 as one single array of individual signal sources and sensors. In that sense, controller 108 calculates signal characteristics, such as delays, for each signal source or for multiple signals sources from different transducers. Those signal sources from multiple different transducers now operate together to generate a wavefront, which means they are synchronised to generate ultrasound signals at a precise point in time relative to the other signal sources. The same applies to the sensors, which are controlled temporally by controller 108 across different transducers or sub-arrays. In this example, controller 108 uses the relative location data to calculate the signals and process the reflection data. In simple terms, if the spatial measurement indicates that a particular signal source is further away from a desired focal point, controller 108 reduces the delay for that signal source and reduces the delay for that sensor. In that sense, controller 108 modifies generating the ultrasound signal based on the relative spatial position of the multiple ultrasound transducers. This may be achieved by applying a trained machine learning model to the reflection data of a previous iteration. That is, the machine learning model is trained to predict the delays and other signal parameters based on thelast reflection data measurement. This provides a way for using the reflection data to fine-tune the beamforming signal control in order to improve on the spatial measurements from the spatial measurement system 107.Method

[0065] Figure 3 illustrates a method 300 for processing ultrasound signals, which may be performed by controller 108. It is noted that controller may be any computing device and as such, may be embedded into the array 101 or integrated with array 101 into a single device. In other examples, controller is a separate device, such as a computer that is connected to array 101 via one or more wires that communicate the reflection data and control data or via wireless communication. Controller 108 may comprise one processor or multiple processors and may be implemented using any computing architecture, including one or more central processing units (CPU), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), graphic processing units (GPU) and others. Method 300 may be implemented by computer readable code, e.g. compiled C++ source code or others, stored on a non-transitory computer-readable medium that, when executed, causes controller 108 to execute method 300.

[0066] Controller 108 receives 301 reflection data indicative of sensed reflection 105 of ultrasound signal 104 by multiple ultrasound transducers 102, 103 of flexible array 101 of ultrasound transducers. The reflection data comprises overlapping reflection data that covers a spatially overlapping field of view 203 with other ultrasound transducers as described above. Controller 108 further receives 302 position data indicative of a relative spatial position of multiple ultrasound transducers from spatial measurement system 107. Then, the controller 108 combines 303 the reflection data spatially based on the position data and based on the overlapping reflection data as described herein. In some examples, controller 108 creates a mosaic by stitching together image data from the different transducers using the spatial measurement for coarse registration and the image data itself for fine registration. In other examples, controller 108 combines the raw measurement data using the spatial measurement for coarse registration and the raw measurement data itself for fine registration. The fineregistration may involve the use of a trained machine learning model is disclosed herein.

[0067] Once the reflection data is combined, such as by registering the multiple images from the multiple transducers, controller 108 determines, from the combined reflection data, a three-dimensional ultrasound representation. Such a three- dimensional ultrasound representation may comprise a three-dimensional image comprising intensity values for multiple voxels in a three dimensional space. Further, controller 108 may determine from the combined reflection data, a segmentation of three-dimensional objects in a volume covered by the reflection data. The segmentation of three-dimensional objects may be achieved by applying a segmentation algorithm on a two-dimensional image or a three dimensional volume representing the reflection data from the entire array 101. In other examples, however, the segmentation may also be achieved by applying a segmentation algorithm on the combined reflection data or on an intermediate output calculated by the controller 108.

[0068] The controller 108 may further classify the combined reflection data into one or more diagnoses. For example, the controller 108 may apply a machine learning model that has been trained on training samples labelled with the one or more diagnoses. Controller 108 may apply the model to the combined reflection data, to a calculated two- or three-dimensional image or an intermediate format. Further details on various algorithms and the like are provided further below.

[0069] As also described above, controller 108 may determine, based on the position data and based on the overlapping reflection data, signal parameters of the ultrasound signals for controlling the ultrasound transducers to effectively perform beamforming to improve image quality..Hardware

[0070] In an example, the building blocks of the system hardware of device 100 comprise: 1) Imaging Probe Components (also described as array 101), focusing on elements directly related to acquiring ultrasound signals from the patient's body andensuring optimal imaging conditions; 2) Control and Signal Processing Components (described as controller 108), encompassing elements responsible for processing and managing the acquired signals, controlling the ultrasound system's operation, and facilitating data transmission and communication. Array 101 might be an array of capacitive micromachined ultrasonic transducers (CMUT), Piezoelectric Micromachined Ultrasound Transducer (PMUT) or similar technology.

[0071] Figure 4 is a hardware technology schematic, including Imaging Probe Components, such as array 101 (i.e., CMUT arrays, coupling pad, tracking devices) and Control and Signal Processing Components 108 (i.e., pulser-receiver system, I / O interface, display, compute and storage, channel management system, power supply).Imaging Probe Components

[0072] CMUTs are micro-electromechanical-systems (MEMS) integrated on complementary metal-oxide-semiconductors (CMOS), capable of sending and receiving acoustic signals in the ultrasonic range and can thus act as a replacement to the standard piezoelectric transducers typically found in ultrasound probes.

[0073] CMUTs are composed of a flexible top plate, a shallow gap, and a fixed bottom plate. In the gap, a large direct current voltage bias is applied that deflects the top plate. An alternate current voltage, then, modulates the vibration of the top plate and, hence, the production of the acoustic waves. Reversely, in the receive mode, the signals from the scanned region are received by the top plate whose vibrations are converted into small voltage variations which are used to produce the images in the same way as for traditional technologies.

[0074] Moreover, CMUTs enhance the cost-effectiveness of ultrasound technology while offering a broader frequency bandwidth compared to conventional methods. This enables the use of a single probe for whole-body imaging. Notably, their compact dimensions make them suitable for miniaturized systems, facilitating ultra-portable and wireless solutions. CMUT drum sizes can vary from 20 pm up to 400 pm, significantly smaller than standard technology typically ranging from fractions of a millimetre toseveral millimetres in size and can thus be used to create miniaturised ultrasound probes containing thousands of vibrating drums.

[0075] Device 100 may comprise multiple rigid CMUT-based sub-arrays distributed on a flexible embedding, capable of securely housing the CMUT probes. These subarrays can be positioned on the anatomical area to be scanned to generate wide-view imaging. CMUTs can be arranged in different configurations on each sub-array. When the transducers are aligned and controlled in a single line, they form a ID sub-array, which can be used to generate 2D images as illustrated in Figure 5a. Conversely, when the transducers are arranged in a grid pattern, they form a 2D sub-array, enabling the generation of 3D volumes as illustrated in Figure 5b.

[0076] The size of each sub-array can span from a few millimetres to a few centimetres in height and width. The smallest size is constrained by the need to accurately track each array, while the largest size may depend on the curvature of the scanned surface, as each array is smaller than the body's profile approximated curvature, to maintain continuous coupling with the skin. The CMUT arrays are strategically positioned in close proximity to ensure comprehensive ultrasound coverage of the entire area, facilitating the utilization of intrinsic and overlapping image information for accurate image formation.Coupling pads

[0077] The flexible wearable array concept includes solutions to ensure proper acoustic coupling with the scanned surface (the skin of the patient or the surface of a phantom, for example). One approach uses water-based gels which me be unreliable as it usually relies on an operator continuously re-distributing it on the surface of the ultrasound probe head to create a homogeneous layer without air bubbles. The CMUT arrays in the embedding may be covered with a semi-solid coupling pad, such as silicone, hydrogel and polymer gel pads, to ensure effective coupling between the arrays and the skin without the operator.Tracking system

[0078] As multiple sensors, grouped in separate sub-arrays, are distributed on the skin surface on a flexible array, ensuring the relative localization of the flexible CMUT arrays is useful for generating consistent and coherent anatomical view of the scanned area. Moreover, the scanned surface may experience changes during imaging (i.e. flex), such as muscle contractions or breathing movements. To address these dynamic conditions, tracking systems can be employed to continuously monitor the position and orientation of the CMUT arrays in real-time.

[0079] Different types of tracking systems can be used to track an ultrasound probe in space, including electromagnetic tracking systems (such as the Northern Digital Inc (NDI) Aurora system (https: / / www.ndigital.com / electromagnetic-tracking- technology / aurora / )) or optical tracking solutions. Other examples involve the use of optical fiber tracking, where 3D position sensing of each CMUT array can be achieved by measuring strain along the cores of a specialized multicore fiber connected to the arrays as described above. This can be accomplished using various techniques. For instance, submillimeter and sub-degree resolution 3D tracking of a soft actuator using Optical Frequency Domain Reflectometry (OFDR) techniques can monitor the distributed strain of a multicore fiber. Depending on the number and size of CMUT sub-arrays embedded in device 100, different configurations of optical fibers can be used. For example, one could consider a grid-like distribution of optical fibers intersecting at the centre of each CMUT array to enhance their localisation accuracy.

[0080] The information on the relative location of the CMUT arrays can be then used in different ways for beamforming and / or tomographic imaging / volumes creation as elucidated further below.Control and Signal Processing Components Pulser-receiver system

[0081] An advanced multi-channel pulser-receiver system (e.g. Verasonics Vantage systems or ULA-OP) can be connected to the CMUT-based probes distribution and used to send / receive the signal to / from the probes for generating ultrasound images.The pulser-receiver system may be controlled via Matlab or C++ interface and enables full customisation of signal transmission and reception.

[0082] Each CMUT array can function independently as an ultrasound probe, driven by the pulser-receiver system, and utilizing standard beamforming techniques for image creation. This approach offers an efficient method for imaging, where each array operates autonomously to generate ultrasound images. Furthermore, quasi-real-time image acquisition can be achieved by minimizing delays in each acquisition.

[0083] In more complex scenarios, a single array can be designated for pulsing while the remaining arrays are used for receiving, operating cyclically. This rotational process enables the utilization of different arrays as if they were part of a unified system. Alternatively, channels from different arrays can be utilized to either pulse or receive signals, tailored to specific criteria, such as maximizing image quality or enhancing specific imaging features.

[0084] Additionally, it is worth noting that once optimal parameters for signal transmission and reception have been established, the pulser-receiver system can be miniaturised, as for example in imaging systems. This optimization process allows for more compact and efficient system designs without compromising performance.Channel Management System

[0085] Given the limited number of channels that can be simultaneously controlled by some pulser-receiver systems (typically between 32 and 256, depending on the model), and the desire to control multiple CMUT arrays within the system, an efficient channel allocation system is advantageous to address this constraint effectively.

[0086] One approach to achieve efficient channel allocation is using multiplexing / demultiplexing strategies, such as using Microchips HV2802 Low Harmonic Distortion, 32-Channel single pole single throw (SPST), High-Voltage Analog Switch. A multiplexing / demultiplexing system can be designed to enhance the number of channels available, allowing for efficient utilization of available resources.By multiplexing signals from multiple CMUT arrays onto a smaller number of channels for transmission and then demultiplexing them upon reception, the system can accommodate the control of a larger number of arrays than would be possible with the pulser-receiver system alone.Power supply and connections

[0087] CMUT-based probes can use a large direct current voltage bias (typically in the range of 80-110V) to collapse the flexible top plate in CMUT arrays’ drums in order to achieve better signal generation. This voltage bias can either be integrated into the pulser-receiver system or supplied externally by a dedicated power supply connected to the CMUT probes. The connectivity between the pulser-receiver system, the CMUT arrays on the flexible substrate, and the channel management system can be established through various means, including high-quality cables designed for ultrasound transmission / reception and / or wireless options such as Wi-Fi and Bluetooth technology.Prototype Versions

[0088] Multiple prototypes were developed as part of an iterative process aimed at advancing the technology. Each prototype represents a step forward in addressing specific challenges and incorporating advancements in the hardware system. The prototypes were tested using CMUT arrays provided by Philips Research although custom-made arrays can be utilized.

[0089] The CMUT array connected to the system may consist of 4032 CMUT drums, arranged along 96 lines, with each line containing 42 CMUT drums. A dedicated MATLAB code was developed to control and acquire B-mode ultrasound images from this array. The code encompasses various specifications of the CMUT array, such as the size of each drum, their spatial arrangement, probe frequency, and mapping between drum positions and connector pins, as well as imaging parameters such as beamforming type and transmission / reception delays.

[0090] The prototypes were used on an ultrasound-compatible leg phantom model and the CMUT array tomographic scan showed good correspondence in comparison to state-of-the-art 3D ultrasound probe scan. Other examples relate to ultrasound images of the internal jugular vein and the carotid artery obtained from an ultrasoundcompatible neck phantom and a volunteer. This outlines one of the potential applications of this solution, which involves automating the detection of thrombosis for astronauts during space missions.

[0091] In another example, 10 CMUT sub-arrays are connected to the pulser-receiver system. To achieve this, a distribution board was developed and employed to efficiently distribute the channels provided by the pulser-receiver system across all 10 CMUT arrays. The distribution board was specifically designed to receive as input 96 channels from the pulser-receiver system and output 96 channels multiplied by the number of arrays to connect. These output channels were soldered to the corresponding 10 CMUT printed circuit boards (PCB). For instance, channel 1 from the pulser-receiver system was transmitted to channel 1 of all 10 CMUT PCB Boards via the distribution board.

[0092] This solution implies that whenever a signal is passed through a channel at the input of the distribution board, the same signal is passed to the corresponding outputs of the distribution board, meaning that the corresponding channels of all the 10 CMUT arrays would be operating simultaneously. However, simultaneous operation of multiple arrays is not desired, as the signal returned to the pulser-receiver system would be a composite of signals from all 10 arrays. To overcome this issue, a switching mechanism can be integrated into the system. The switch can be strategically placed between the bias pin of each CMUT PCB Board and the power supply responsible for providing the required bias for probe activation. Manual selection or selection via a dedicated application enabled to designate specific probes using the switch, ensuring that only the selected probe received bias and was activated for imaging.

[0093] In another approach, switches are directly applied to the data cables to select which CMUT array to operate. This may utilize commercially available switches (High-Definition Multimedia Interface (HDMI) Switcher Splitter Hub 2 in 1 In Outultra high definition (UHD) 4K Bi Direction Hub High-bandwidth Digital Content Protection (HDCP) 3D) capable of managing a group of 8 channels as input, with the flexibility to designate 2 groups of 8 channels as output.Data processing

[0094] Method 300 can be implemented as software designed to process the signals acquired by the multiple CMUT arrays and deliver the final image interpretation to a user, such as a sonographer or physician. This process can be divided into three distinct phases:• Phase 1. Beamforming, which involves the transmission and reception of ultrasound signals to form images / volumes from the collected signals.• Phase 2. Tomography or Wide-View (Mosaicking) Generation, where the ultrasound information gathered by the CMUT arrays is integrated into a single coherent visualization. This includes: o Phase 2. a. Combining the individual images / volumes seamlessly into a wide-view or tomographic image / volume. o Phase 2.b. Compounding this information to ensure smooth and continuous visualization across the entire scanned area.• Phase 3. Automated Image Interpretation, which provides the user with image analysis and can include: o Segmentation of anatomical structures. o Tracking of specific features or regions. o Diagnostics of potential pathologies. o Measurement of relevant parameters, tailored to the specific application.

[0095] It is worth noting that Phases 1-3 can be implemented either sequentially or using algorithms that combine different phases for improved efficiency. These options are elucidated below.Option 1Phase 1. Beamforming.

[0096] Each CMUT sub-array can be treated as an individual probe (i.e. transducer), and beamforming techniques can be applied to generate B-mode 2D images when using ID CMUT sub-arrays, and 3D volumes when using 2D CMUT sub-arrays. This approach has been demonstrated with our ID sub-arrays to create 2D B-mode images using a beamforming technique. For 2D CMUT sub-arrays, especially those with smaller form factors, a phased array approach is used to enhance both image quality and the field of view. This method ensures that adjacent arrays can produce partially overlapping views, which are large enough to include meaningful features for combining views.

[0097] An alternative to classic beamforming techniques is to employ artificial intelligence (Al)-based beamforming. This involves training deep learning models, such as convolutional neural networks (CNNs), on datasets with raw ultrasound data as input and beamformed images as output.

[0098] In one example, an adaptive beamforming based on minimum variance (ABF- MV) with deep neural network (DNN) is used to improve the image performance and to speed up the beamforming process of ultrafast ultrasound imaging. In particular, a DNN, with a combination architecture of fully-connected network (FCN) and convolutional autoencoder (CAE), can be trained with channel radio-frequency (RF) data as input while minimum variance (MV) beamformed data as ground truth.

[0099] In another example, to perform beamforming and speckle reduction using neural networks, a controller 108 initiates the process by preparing the data. It simulates ultrasound channel data using tools like Field II Pro, creating realistic ultrasound signals. Real photographic images are used as ground truth echogenicity maps, ensuring that the simulated data closely resembles real-world scenarios. This step is useful for training the neural network effectively, as it provides a reliable reference for the network to learn from. The controller 108 then designs a deep convolutional neural network (CNN) with multiple layers and filters. This network processes the demodulated analytic signals from the transducer elements, which contain both amplitude and phase information used for accurate beamforming and speckle reduction.The architecture of the CNN is tailored to handle the specific characteristics of ultrasound data, such as its high dimensionality and the presence of noise. During the training phase, the controller 108 uses the simulated channel signals and their corresponding ground truth maps. It employs loss functions like fl, f2, and multi-scale structural similarity measure (MS-SSIM), modified to suit ultrasound imaging. These loss functions help the network learn to minimize the differences between the predicted and actual echogenicity maps. The training process requires a large dataset and significant computational resources, as the network needs to learn the intricate patterns and structures present in the ultrasound data. Finally, the controller 108 evaluates the performance of the trained network on various datasets, including simulation, phantom, and in vivo data. This step ensures that the network generalizes well to different types of data and performs consistently across various scenarios. The controller 108 compares the network's performance against traditional speckle reduction techniques to assess its effectiveness. This comprehensive evaluation helps in fine-tuning the network and improving its accuracy and robustness in real-world applications.

[0100] Al-based beamforming has the potential to capture complex patterns in the data, offering improved image quality, higher resolution and speed, and artefact reduction compared to traditional methods.Phase 2. Tomography or Wide- View (Mosaicking) Generation.

[0101] Phase 2. a. The images or the volumes generated in Phase 1 are combined to create a coherent and comprehensive view of the anatomy of interest in real-time based on the relative spatial position and based on the overlapping reflection data. The 3D spatial information of the arrays, gathered through the spatial measurement system 107 (e.g., tracking system) is used to coarsely register their corresponding generated images / volumes in a common coordinate system. Starting from these initial conditions, the images / volumes are then rigidly registered using artificial intelligence methods. A rigid registration approach is feasible because the CMUT arrays collect image information simultaneously and in real-time, capturing a common anatomical configuration for each timestamp.

[0102] There may be two algorithms (a supervised and an unsupervised model) using a Swin transformer backbone structure for rigid registration, based on a modified algorithm architecture. This type of algorithm has the ability to capture complex relationships and dependencies within data. Transformers excel at handling long-range interactions, which is beneficial for aligning and matching features in images. These algorithms offer real-time performance, enabling real-time comprehensive tomographic anatomical visualization. Real-time in this context means that the step of combining the reflection data is completed before further reflection data is received.

[0103] The algorithm takes as input two partially overlapping volumes at a time, identifying one as the fixed volume and the other as the moving volume. It rigidly registers the moving volume to the fixed volume, matching their overlapping area using common learned features. During training, the supervised model is optimised by minimising the difference between the translational and rotational parameters inferred by the algorithm and the provided ground-truth parameters to align the moving volume with the fixed volume. The network structure is illustrated in Figure 6b.

[0104] The unsupervised method, on the other hand, minimises the loss function, e.g., the normalised cross-correlation loss (a measure of image similarity) over the overlapping sub-volumes after applying the inferred translational and rotational parameters to the moving volume. The corresponding network architecture is illustrated in Figure 6c. Further, the registration may involve registering multiple volumes simultaneously by minimizing a global loss function.

[0105] For both structures of Figure 6b and 6c, the machine learning model comprises outputs 620, 630 indicative of respective spatial transformations (xyz translation, xyz rotation) of the multiple visual representations, e.g. the input images. The machine learning models may comprise a CNN while the models in Figures 6b and 6c comprise transformer models. More particularly, the transformer model is applied on a first image 621 / 631 and an adjacent second image 622 / 623. The transformer model calculates an output indicative of spatial transformations of the second image that register the second image with the first image.

[0106] The transformer model may also split the first image 621 / 631 and the second image 622 / 632 into contiguous subsets indicated at 623 / 633, projecting the subsets into an embedding space and applying self-attention between parts of the first image and between parts of the second image. To that end, the machine learning model comprises a linear projection block 624 and multiple blocks each having a Swin Transformer block 625 and a patch merging block 626. The first Swin transformer block may have dimensions of M / 4 x W / 4 x L / 4 x C, th second Swin transformer block M / 8 x W / 8 x L / 8 x 2C, the third Swin transformer block M / 16 x W / 16 x L / 16 x 4C and the last Swin transformer block may have M / 32 x W / 32 x L / 32 x 8C.

[0107] Training the machine learning model may comprise supervised training where the training data comprises labels indicative of a spatial transformation that registers a first image with a second image. In that case, controller 108 reduces an error between the labels and the outputs of the machine learning model, as depicted in Figure 6b. In another example, training the machine learning model comprises unsupervised training. In that case, controller 108 reduces a similarity loss value applied between a first image and a second image as shown in Figure 6c.

[0108] This problem may also be approached using other optimisation methods. These methods iteratively find the translational and rotational parameters for the moving volume by optimizing a loss function expressing the similarity between the two volumes (such as the normalized cross-correlation loss) as shown in Figure 6a. However, these methods may be not real-time (thus not desirable for some applications).

[0109] We evaluated and compared the Al methods with the optimisation method using a benchmark synthetic knee dataset collected using a state-of-the-art ultrasound system and volumetric probe (Epiq 7 ultrasound system and VL13-5 Philips probe), encompassing 240 registration cases from 16 volunteers. Each volume was split into two partially overlapping volumes, defined as the moving and fixed volumes The moving volume was then perturbed from its initial position by random translations and rotations within the range of + / - 3mm and + / - 3° along all directions. These pairs ofmoving and fixed volumes were then used to train and test the Al algorithms and test the traditional optimisation method, with the original volumes serving as the groundtruth registrations.

[0110] A 2-fold cross-validation was performed, with 10 volunteers per fold in the training set and 6 volunteers per fold in the test set. The training datasets included a total of 800 and 723 volume pairs per fold (including the original volumes plus data augmentation) and 85 volume pairs in each testing set. The results obtained (reported in Table 1) show the mean difference between translations / rotations inferred by the algorithm and ground-truth parameters. The transformers successfully registered all volumes, providing results comparable to ground-truth solutions in all cases. The supervised method showed slightly higher performance compared to the unsupervised method, which was expected since the supervised method is directly trained to minimize the differences from the ground-truth parameters. The optimisation approach achieved comparable results but failed in 9 cases out of 170.

[0111] Table 1 : Performance comparison between the optimisation method vs Supervised and Unsupervised Transformer Al algorithms for rigid registration of partially overlapping volumes. The three methods were tested on a total of 140 volume couples (85 from each test set). The mean difference in translation and rotation between the different algorithms and ground-truths are reported.Traditional optimisation Supervised Unsupervised method Transformer TransformerTest sets Mean difference Mean difference Mean difference translation andtranslation and translation androtations between rotations between rotations between algorithm and ground- algorithm and ground- algorithm and groundtruth truth truth[mm / deg] [mm / deg] [mm / deg]Test set 1 (6 0.32 0.21 0.14 0.16 0.15 0.14 0.23, 0.15, 0.16, patients, 850 16 0 24 0 16o .140.16 0.14 0.24, 0.37, 0.94 volume coupleTest set 2 {6 0.27 0.24 0.14 0.140.15 0.14 0.16, 0.10, 0.17, patients, 850 13 0 290.15 0.130.15 0.14 0.22, 0.40, 0.70 volume couple

[0112] It is to be noted that similar registration approaches could be developed with different backbones, such as convolutional neural networks and / or a combination of convolutional neural networks and transformer architectures as performed by controller 108. In this sense, controller 108 integrates deep learning models into iterative registration methods in a process referred to as deep iterative registration. This involves replacing intensity-based similarity metrics with deep similarity metrics derived from convolutional neural networks (CNNs), allowing the controller 108 to capture underlying image features more effectively. The controller 108 may further employ supervised registration methods. These methods involve training deep learning models using labelled reference data, which can include artificially generated deformation vector fields (DVFs) or biomechanical constraints. These methods have demonstrated high accuracy in various applications, including multi -modality and large motion registrations. In the case of weakly supervised registration, the controller 108 may use organ contours or segmentation labels, offering a more accessible and reliable training process. The controller 108 may perform unsupervised registration, which eliminates the need for labelled training data. Instead, it relies on loss functions that measure the similarity between the warped and fixed images. This category includes methods that use pre-trained CNNs for feature extraction and generative adversarial networks (GANs) for registration. The controller 108 may also perform the integration of deep similarity metrics into unsupervised models and the combination of whole image-based and patch-based training strategies. By following these steps, the controller 108 can effectively perform deep learning-based three-dimensional medical image registration.

[0113] Controller 108 may apply transformation models that can be rigid, affine, or deformable. Rigid registration allows only translations and rotations, while affine registration includes scaling and shearing. Deformable registration uses a displacement field to account for local differences between images. The similarity cost function quantifies how closely aligned the transformed moving image and fixed images are, with common functions including sum of square differences (SSD), cross-correlation (CC), and mutual information (MI). Controller 108 uses machine learning models to detect corresponding features, predict intensity similarity, and directly learn transformations. Techniques like spatial transformer networks (STNs) enhance CNNsby allowing explicit spatial invariance. VoxelMorph incorporates STNs to predict voxel-wise displacement fields, improving registration accuracy and efficiency.

[0114] While the disclosed algorithms are applied to partially overlapping volumes, a similar approach can be adapted for 2D images generated by ID arrays, by modifying the algorithm structure. With ID arrays, this registration process may lead to real-time tomographic reconstruction of 2D surfaces of the internal anatomy. However, when using arrays with small form factors, the distance between different surfaces can become relatively small, allowing the combination of these surfaces to reconstruct dynamic 3D volumes, similar to the output achieved with multiple 2D arrays. Interpolation techniques could be considered to retrieve the missing anatomical information. In particular, processor 108 may perform Neural Radiance Fields (NeRFs) algorithms, which utilizes a weighted multilayer perceptron to rebuild the volumetric density and color of a 3D scene.

[0115] NeRFs are a type of deep learning model used to synthesize novel views of complex 3D scenes by interpolating the visual information from multiple perspectives. By learning the 3D representation of an object from sparse input views, NeRFs can generate high-quality, continuous volumetric reconstructions. Applying NeRFs to ultrasound imaging could enhance the quality and continuity of the reconstructed volumes, even in regions with sparse data or occlusions.

[0116] Controller 108 begins by collecting posed multi -view images of a static scene. These images serve as the training set for the Neural Radiance Fields. The controller then uses multi-layer perceptrons (MLPs) to represent two volumes: an opacity field denoted as c, and a radiance field denoted as c. The opacity field represents a soft shape and is a function of 3D position, while the radiance field represents view-dependent surface texture and is parameterized by both 3D position and viewing direction. The next step for Controller 108 is to train these fields. This involves minimizing the discrepancy between the observed images (ground truth) and the predicted images rendered from c and c at the same viewpoints. Controller 108 accomplishes this optimization using stochastic gradient descent. The controller then ray -traces theimplicit volumes o and c to render each pixel of the predicted images. For each ray, its colour is determined by integrating the product of the opacity and radiance fields along the ray, adjusted by an exponential term that accounts for the accumulated opacity along the ray's path. To enhance the ability to synthesize sharper images, controller 108 employs a positional encoding that maps the 3D coordinates and viewing directions to their Fourier features. This encoding helps the network capture high-frequency details. When applying NeRF to 360° captures of unbounded scenes, controller 108 faces a significant challenge in the parameterization of space. To address this, the controller partitions the scene into an inner unit sphere and an outer volume represented by an inverted sphere. The inner volume contains the foreground and all cameras, while the outer volume contains the remainder of the environment. These two volumes are modelled with separate NeRFs, and their outputs are composited to render the final image. This approach allows controller 108 to achieve improved quantitative and qualitative results in challenging scenarios involving large-scale, unbounded 3D scenes.

[0117] Phase 2.b. Once the volumes / images from the CMUT arrays are properly aligned, the pixel / voxel intensity values for the overlapping part of the images / volumes are defined. Some approaches select the maximum or mean intensity values across each pair of overlapping sub-volumes / images. However, these methods can propagate errors from the registration process and are not ideal for ultrasound imaging, which may be affected by artefacts, shadows, and intensity variations between overlapping volumes. To overcome these limitations, generative models such as generative adversarial networks (GANs) or diffusion models could be employed to refine the ultrasound visualization by incorporating data from complementary imaging modalities like magnetic resonance imaging (MRI), computed tomography (CT), or cone beam CT (CBCT) or anatomical structures segmentations from these modalities. These modalities offer comprehensive anatomical coverage, which is beneficial for constructing a tomographic ultrasound image.

[0118] The controller 108 initiates the process by selecting the appropriate generative model based on the specific application. If the task asks for a good generalization with limited training data, the controller 108 opts for Locality -based Statistical ShapeModels (LSSMs). For applications asking for high specificity and realistic image generation, the controller 108 chooses deep-learning models like Diffeomorphic Autoencoders (DAEs) and Autoencoder Generative Adversarial Networks (AE-GANs). The controller 108 then prepares the input data. For shape modelling, it uses pointbased representations of segmentation labels. For deep-learning methods, the controller 108 inputs one-hot-encoded labels directly. For appearance modelling, the controller 108 uses intensity image volumes, with Statistical Appearance Models (SAMs) considering shape distortions and appearances separately. Next, the controller 108 trains the selected model on various training set sizes. It ensures that the model is learning the correct representations and patterns from the data. The controller 108 closely monitors the training process to avoid overfitting or underfitting. Once the model is trained, the controller 108 assesses its performance based on generalization ability, specificity, and likeness. For shape modelling, the controller 108 checks if the model can generate new shapes that are similar to the ones in the training data. For appearance modelling, it verifies if the model can generate diverse, sharp, and realistic images. The latent space of SSM-based models is more interpretable and compact, producing continuous and smooth interpolation results. On the other hand, deeplearning models have higher expressiveness but use larger training datasets and are harder to interpret. The choice of generative model may depend on the specific application and the trade-offs between interpretability, computational resources, and the quality of generated samples.

[0119] The disclosed approach merges partial fields of view obtained from different directions by the CMUT arrays. By integrating whole-anatomy data from complementary modalities, these generative models can interpolate missing information and correct for inconsistencies, significantly enhancing the quality and accuracy of the assembled tomographic visualization.

[0120] Alternatively, generative networks could be used to convert the tomographic ultrasound volume to another, more user-friendly modality, such as CT or MRI. This may comprise converting 3D ultrasound volumes of a spine model to MRI using an unsupervised cycle GAN.

[0121] In other words, controller 108 uses a deep learning approach with generative adversarial networks (GANs), specifically CycleGAN, to conduct unsupervised image- to-image translation. This produces spatially aligned ultrasound (US) and MRI volumes corresponding to their respective input volumes of the anatomical region. Controller 108 works with a dataset consisting of one MRI volume and US volumes. Controller 108 adjusts the original US voxel spacing through bilinear interpolation to match the MRI voxel spacing, ensuring a consistent aspect ratio. The CycleGAN architecture, implemented by controller 108, maps one image domain to another using two sets of GANs that work together. The structure of the generator is based on an encoder / decoder with a U-net structure, and the discriminator uses a PatchGAN discriminator. The least squares loss is used for the GAN loss, and the cycle consistency and identity loss are measured by the sum of the LI distance. During training, controller 108 feeds the generators with real US and MRI volumes with data augmentation applied, generating new volumes of the opposite imaging modality. The discriminators evaluate the quality of these generated volumes by comparing them with the corresponding real volumes. Controller 108 optimizes hyperparameters using Bayesian optimization based on the performance of the generated volumes and the training losses. This method could be further expanded and applied to the tomographic ultrasound images generated in this phase.Phase 3. Automated Image Interpretation

[0122] Once a high-quality tomographic volume reconstruction is generated from Phase 2, various Al algorithms can be utilized for the final image interpretation, including convolutional neural networks (CNNs), vision transformers (ViTs), GANs, and combinations thereof, including Supervised learning, Unsupervised learning, Reinforcement learning, Semi-supervised learning, Self-supervised learning, Multiinstance learning, Statistical inference, Inductive learning, Deductive inference, Transductive learning, Multi-task learning, Active learning, Online learning, Transfer learning, or Ensemble learning. The architectures may involve deep neural networks (DNN), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Deep conventional-extreme learning machine (DC-ELM), Deep Boltzmann Machine(DBM), Deep Belief Networks (DBN), Deep Autoencoders (DAN), Deep stacking networks (DSN), Long short-term memory / Gated Recurrent Unit (LSTM / GRU)Option 2Phase 1. Beamforming. + Phase 2. Tomography or Wide- View (Mosaicking) Generation (Phase 2. a.)

[0123] The CMUT arrays embedded in the solution function as a unified ultrasound system, with all arrays contributing to the formation of the ultrasound tomographic view. Thus, Phase 1 and Phase 2.a are performed concurrently. To achieve this, the tracking information gathered is provided in real-time as input to the beamforming algorithm, since the relative positions of the arrays (and individual elements within the arrays) should be known to reconstruct the tomographic visualization.

[0124] While for Option 1 the tracking system provides a coarse localisation prior to image registration, in Option 2, submillimeter accuracy of the tracking system is desirable. Standard beamforming techniques can be applied by selecting individual CMUT arrays or small groups of CMUTs within the arrays for transmission and reception. For example, a single array can be designated for pulsing while the remaining arrays receive signals, operating cyclically. Alternatively, channels from different arrays can be utilized to either pulse or receive signals, following specific patterns.Phase 2.b. Similar to Phase 2.b. in Option 1.Phase 3. Automated Image Interpretation. Similar to Phase 3. in Option 1.Option 3Phase 1. Beamforming + Phase 2. Tomography or Wide- View (Mosaicking) Generation (Phase 2. a.)

[0125] As in Option 2, the CMUT arrays are utilized as a unified probe, and the tracking information is provided as real-time input into the beamforming algorithm. However, in this approach, the beamforming method is based on artificial intelligence (Al).

[0126] Deep learning models such as convolutional neural networks (CNNs) can be trained using a dataset where both the input (raw ultrasound data) and output (corresponding tomographic images) are known. This approach allows the model to learn the mapping between raw ultrasound data and desired tomographic images. Ground-truth tomographic visualizations can be obtained through many ways, such as manual annotation, validated through complementary imaging modality (e.g., MRI, CT); generating the tomographic views as described in Phase 2 (2. a. or 2.b2) of Option 1 and Option 2. Alternatively, these ground-truth datasets could be generated experimentally using approaches such as full waveform inversion(Guasch et al., 2020) (Ulrich et al., 2023)(Robins et al., 2023)

[0127] Full waveform inversion (FWI) is a technique that reconstructs the subsurface structure by iteratively minimizing the difference between observed and simulated waveforms. It involves solving the full wave equation to update the model of the medium iteratively, leading to high-resolution images of the subsurface. FWI has shown potential in medical imaging research for providing detailed reconstructions of anatomical structures by utilizing high-frequency ultrasound data.

[0128] Some FWI applications utilize piezoelectric elements arranged on a rigid ring, where one transducer is fired while all the others record the received signals cyclically. While this method can serve as a basis for generating ground-truth data, the approach disclosed herein introduces other elements to this approach by leveraging the capabilities of CMUT arrays. Unlike current systems, our CMUT arrays can be directly placed on the anatomy, eliminating the need for a rigid structure and allowing for more flexible and accurate positioning. Additionally, CMUT arrays offer a much larger bandwidth than piezoelectric transducers (PZT), possibly facilitating improved anatomical reconstruction.

[0129] To perform the FWI algorithm, controller 108 begins with an initial model, which is an estimate of the target of interest. This model is composed of many parameters. The goal of controller 108 is to minimize the misfit function, which is defined as half the sum of the squares of the differences between the experimental data and an equivalent simulated dataset generated using the initial model. Mathematically, this is expressed as, where p and d represent the predicted and observed data, respectively. The predicted data are generated by controller 108 solving the wave equation using the initial model.

[0130] To compute the gradient of the misfit function with respect to the model parameters, controller 108 employs the adjoint-state method. This involves solving the wave equation forward in time to generate the predicted data, then injecting the residual data (the difference between the predicted and observed data) into the model at the receiver positions. Controller 108 then solves the wave equation backward in time using the injected residual data as a virtual source. The resulting wavefield is scaled using the differential of the wave equation operator with respect to the model parameters. Controller 108 then finds the zero lag of the cross-correlation of this wavefield in time with the original forward wavefield at every point in the model. This process yields a gradient vector oriented to point in the direction of maximum increase of the misfit function at the current position in the solution space. Controller 108 uses the gradient vector to iteratively update the model parameters, moving successively downhill on the hyper-surface defined by the misfit function to arrive close to the model that best predicts the observed data. The direction of steepest descent is typically preconditioned by controller 108 to speed convergence, often using spatial preconditioning to compensate for illumination variation within the model. This iterative process continues until the model converges to a solution that minimizes the misfit function, thereby producing an accurate representation of the target of interest. The computational cost of numerically solving the governing wave equation in threedimensions for many sources restricts solutions to iterated local gradient-descent methods, making the FWI algorithm highly effective for high-resolution imaging.Phase 2.b. Similar to Phase 2.b. in Option 1.Phase 3. Automated Image Interpretation. Similar to Phase 3. in Option 1.Option 4Phase 1. Beamforming + Phase 2. Tomography or Wide- View (Mosaicking) Generation (Phase 2. a.)

[0131] Option 4 introduces the utilization of the 'coherent wave plane compounding' approach (C. Wang et al., n.d.)in conjunction with tracked CMUT arrays as a unified probe for ultrasound imaging (as in Option 2 and Option 3). In this approach, all transducer elements emit plane waves simultaneously at a specific angle, and the process is repeated for several different angles. The system then combines the received signals from these multiple angles to produce a single high-quality image with improved resolution and contrast.

[0132] Coherent plane-wave compounding (CPWC) is a technique used to enhance the image quality in ultrasound plane-wave imaging (PWI). The process begins by transmitting a plane-wave into the medium, which allows for very high frame rates. However, due to the lack of transmitting focus, the resulting images may suffer from poor resolution and contrast. To address this, CPWC coherently combines multiple steered plane-wave signals. When the i-th plane-wave is steered at an angle a , the time to reach a focal point p(x,z) is given by tt= (zcosa. + xsinai) / c0, where c0is the speed of sound in the medium. The time for the echo signal to propagate back to a receiving element at xRis tR= sqrl((x - x,, )2+ z2) / c0. The beamformed output of PWI at point p is then calculated as p(i) =+ Z ) , where M is the number of elements in the transducer, h; Ris the signal received by the element at xR, andrepresents the receive apodization for the signal received by the element at xR.

[0133] The output of CPWC at point p is expressed as Scpwc= (1 / , where N is the total number of angles (i.e., the number of plane waves), anddenotes the angular apodization. In this method,is typically a Tukey window with a 25% taper, and \ (a ) is set to 1. To further enhance the image quality, CPWC can be weighted by the generalized coherence factor (GCF) and short-lag spatial coherence (SLSC). GCF is defined as the ratio of the spectral energy within a prespecified low-frequency region to the total energy. For point p(x,z), a spatial signal sequence Px,z = [ / ?(!), j>(2),..., j>(7V)] is assumed as the discrete sequence. The GCF can be calculatedthe spectrum of the sequence Px z, and K is the number of points in the discrete spectrum. The GCF-weighted CPWC is then given by SGCF(j») = GCF(p)Scpwc.

[0134] For SLSC, the normalized spatial coherence at lag 1 is defined asusually selected as one wavelength. The SLSC value is the summation of the corresponding spatial coherence function over the first L lags, given by SLSC p) = S^R(I). In this study, the parameter Q, where Q = L / N x 100%, is set to about 20%. The CPWC weighted by SLSC is expressed as SSLSC(p) = SLSC p)Scpwc. This detailed process ensures that CPWC can significantly improve the resolution and contrast of ultrasound images by effectively combining multiple steered plane-wave signals and applying adaptive weighting factors like GCF and SLSC.

[0135] In adapting this approach, the tracked CMUT arrays serve as the unified probe, with each array comprising individual transducer elements. By coordinating the emission of plane waves from these arrays at different angles and coherently summing the received echoes, controller 108 can perform the 'coherent plane wave compounding' technique to enhance image quality. The tracking information provided in real-time is crucial for accurately aligning and combining the signals from different angles.

[0136] Moreover, the system's capability can accommodate dynamically changing positions of the CMUT arrays, following the body's profile. This flexibility allows the system to dynamically adapt the 'coherent plane wave' strategy as the arrays adjust to fit the anatomical contours, ensuring continuous optimization of the imaging process regardless of changes in patient positioning or anatomical variations.

[0137] Phase 2.b. Similar to Phase 2.b. in Option 1.

[0138] Phase 3. Automated Image Interpretation. Similar to Phase 3. in Option 1.Option 5

[0139] Phase 1. Beamforming + Phase 2. Tomography or Wide-View (Mosaicking) Generation (Phase 2. a.) (Phase 2.b.) + Phase 3. Automated Image Interpretation

[0140] Option 5 involves a direct pathway from raw ultrasound data (such as radiofrequency (RF) signals or IQ-demodulated channel data), along with tracking information from the CMUT arrays to the final segmentation output without the intermediate step of beamforming or tomographic imaging generation.

[0141] During Phase 7, the system acquires RF signals from the CMUT arrays, along with tracking data. Instead of applying beamforming techniques, the system forwards the raw data directly to an Al algorithm. This Al algorithm is specifically trained to process this type of data, leveraging the tracking information to accurately interpret the spatial context of the signals.

[0142] The Al algorithm, trained on datasets comprising raw data and corresponding segmentations of structures of interest, performs real-time segmentation directly from the raw data. It dynamically adjusts its segmentation strategy based on the spatial orientation and movement of the CMUT arrays, as indicated by the tracking information. Further, the segmentation may be based on additional, measurements, which may include automated measurements. For example, the automated measurements may include a volume estimate, measurement of velocity of movingobjects, such as blood flow, measurement of stiffness of tissue, and other automatic measurements.

[0143] Phase 2, which typically involves tomography or wide-view generation, is skipped entirely in this approach. Instead, the Al algorithm directly outputs segmented structures of interest, delineating boundaries and identifying specific anatomical features within the ultrasound images.

[0144] In Phase 3, the system can further analyse the segmented structures, detecting abnormalities, measuring relevant parameters, and generating clinical reports — all directly from the raw data and without the need for intermediate tomographic images.

[0145] By circumventing beamforming and tomography steps, Option 5 streamlines the ultrasound imaging process. This approach harnesses the power of Al to perform efficient and automated output directly from raw signals, marking a significant advancement in ultrasound imaging technology. A similar strategy can be applied to single-probe systems for segmenting structures within phantoms and the adaptation of this approach to a dynamic multi-array system provides significant advantages.

[0146] In particular, this process begins by formulating the problem for unfocused input channel data. The aim is to generate a DNN beamformed image and a segmentation map prediction using downsampled IQ channel data as input. The network architecture is constructed on the U-Net architecture. This network has a single encoder and two decoders to produce both the DNN image and segmentation. The input data is downsampled and normalized to ensure stable DNN training. Enhanced beamformed images are used during network training to improve the identification of anechoic targets. The network can be trained using a combination of LILoss and dice similarity coefficient (DSC) Loss to jointly produce the DNN image and segmentation.

[0147] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The presentembodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

CLAIMS:

1. A sonography device comprising: a flexible array of multiple ultrasound transducers, each of the multiple ultrasound transducers being configured to generate an ultrasound signal and to sense a reflection of the ultrasound signal to generate reflection data comprising overlapping reflection data that covers a spatially overlapping field of view with other ultrasound transducers; a spatial measurement system configured to determine a relative spatial position related to the multiple ultrasound transducers; and a controller configured to spatially combine the reflection data based on the relative spatial position and based on the overlapping reflection data.

2. The device of claim 1, wherein the flexible array comprises multiple rigid subarrays, each of the rigid sub-arrays comprising multiple transducers and the rigid subarrays are flexibly connected to form the flexible array.

3. The device of claim 2, wherein the spatial measurement system is configured to determine the relative spatial position of each of the rigid sub-arrays.

4. The device of any one of claims 2 and 3, wherein the controller is configured to combine multiple visual representations generated by respective sub-arrays.

5. The device of claim 4, wherein combining the multiple visual representations comprises determining a spatial transformation to the multiple visual representations to register the multiple visual representations.

6. The device of claim 5, wherein determining the spatial transformation comprises applying a trained machine learning model to the multiple visual representations, the trained machine learning model comprising outputs indicative of the spatial transformation.

7. The device of any one of the preceding claims, wherein the controller is further configured to modify generating the ultrasound signal based on the relative spatial position of the multiple ultrasound transducers.

8. The device of claim 7, wherein modifying the ultrasound signal comprises applying a trained machine learning model to the reflection data of a previous iteration.

9. A method for processing ultrasound signals, the method comprising: receiving reflection data indicative of a sensed reflection of an ultrasound signal by multiple ultrasound transducers of a flexible array of ultrasound transducers, the reflection data comprising overlapping reflection data that covers a spatially overlapping field of view with other ultrasound transducers; receiving position data indicative of a relative spatial position of multiple ultrasound transducers; and combining the reflection data spatially based on the position data and based on the overlapping reflection data.

10. The method of claim 9, wherein the method further comprises determining, based on the position data and based on the overlapping reflection data, signal parameters of the ultrasound signals for controlling the ultrasound transducers.

11. The method of claim 9 or 10, wherein the method further comprises: receiving multiple visual representations comprising a visual representation for each of multiple rigid sub-arrays of the flexible array; and registering the multiple visual representations to determine a three- dimensional ultrasound representation.

12. The method of claim 11, wherein registering the multiple visual representations comprises: performing a coarse registration using the position data; and performing a fine registration by applying a trained machine learning model to the reflection data.

13. The method of claim 12, wherein the machine learning model comprises outputs indicative of respective spatial transformations of the multiple visual representations.

14. The method of claim 12 or 13, wherein the trained machine learning model is a CNN or a transformer model.

15. The method of claim 14, wherein using the transformer model comprises applying the transformer model on a first visual representation and an adjacent second visual representation of the multiple visual representations and the transformer model comprises an output indicative of spatial transformations of the second visual representation that register the second visual representation with the first visual representation.

16. The method of claim 15, wherein applying the transformer model comprises splitting the first visual representation and the second visual representation into contiguous subsets, projecting the subsets into an embedding space and applying selfattention between parts of the first visual representation and between parts of the second visual representation.

17. The method of any one of claims 12 to 16, wherein the method further comprises training the machine learning model using training data comprising labels indicative of a spatial transformation that registers a first visual representation with a second visual representation by reducing an error between the labels and the outputs of the machine learning model.

18. The method of any one of claims 12 to 17, wherein the method further comprises training the machine learning model to reduce a similarity loss value applied between a first visual representation and a second visual representation.

19. The method of claim 18, wherein the loss value is defined by a global loss function to align more than two partially overlapping visual representations simultaneously.

20. The method of any one of claim 11 to 19, wherein the method further comprises determining parameters of generating the ultrasound signal based on the position data.

21. The method of any one of claims 11 to 20, wherein the method is performed in real-time to complete the step of combining the reflection data before receiving further reflection data.22 The method of any one of claims 11 to 21, wherein the method further comprises determining, from the combined reflection data, a three-dimensional ultrasound representation comprising a three-dimensional image comprising intensity values for multiple voxels in a three dimensional space.

23. The method of any one of claims 11 to 22, wherein the method further comprises, determining from the combined reflection data, a segmentation of three- dimensional objects in a volume covered by the reflection data.

24. The method of claim 23, wherein determining the segmentation is further based on automated measurements.

25. The method of any one of claims 11 to 24, wherein the method comprises, classifying the combined reflection data into one or more diagnoses.

26. Software that, when executed by a computer, causes the computer to perform the method of any one of claims 11 to 25.

Citation Information

Patent Citations

  • Ultrasound system

    EP3881770A1

  • Foldable 2-d CMUT-on-CMOS arrays

    US20170119348A1

  • Methods and Apparatus for Imaging with Conformable Ultrasound Patch

    US20200121281A1

  • Flexible ultrasound transducer array

    US7878977B2

  • Flexible ultrasound array for measuring a curved object with scattering element

    WO2023018332A1