Rotating ultrasound transducer for volumetric estimation

The rotating ultrasonic transducer and INR algorithm address the challenges of 3D ultrasound imaging by mapping 2D ultrasound images to 3D coordinates, achieving efficient and accurate volumetric estimation with reduced complexity, as shown by high accuracy in simulated and in vivo applications.

WO2026062646A1PCT designated stage Publication Date: 2026-03-26RAMOT AT TEL AVIV UNIVERSITY LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Conventional ultrasound imaging methods face challenges in generating accurate 3D representations from 2D images due to the need for precise positional data, high computational complexity, and reliance on additional sensors, which limits accessibility and accuracy in applications like follicle size assessment during IVF procedures.

Method used

A system utilizing a rotating ultrasonic transducer and an implicit neural representation (INR) algorithm maps intensity values of 2D ultrasound images to absolute 3D coordinates, enabling efficient and accurate 3D image reconstruction and volumetric estimation.

Benefits of technology

The system achieves enhanced volumetric sampling and estimation with improved accuracy and reduced computational complexity, allowing for reliable 3D image generation from 2D slices, as demonstrated by a mean accuracy of 93% in simulated data and a mean error of 6.3% in vivo tumor volume measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050812_26032026_PF_FP_ABST
    Figure IL2025050812_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are methods and systems, for volumetric ultrasound imaging using rotational ultrasound transducer for sampling and estimating volume of objects, based on implicit volume representation thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ROTATING ULTRASOUND TRANSDUCER FOR VOLUMETRIC ESTIMATION

[0002] TECHNICAL FIELD

[0003] The present disclosure generally relates to systems and methods for volumetric ultrasound imaging using rotational ultrasound for estimating volume of objects in a target region.

[0004] BACKGROUND OF THE INVENTION

[0005] Ultrasound is a safe, and accessible imaging modality with temporal characteristics. Implicit volume representation is used for creating continuous functions that map positional coordinates to image density. This technique is used to define and manipulate 3D shapes and surfaces within, inter alia, scientific visualization, by defining volumes and surfaces using mathematical functions. For example, implicit volume representation can be used in reconstructing surfaces from 3D imaging data (such as CT or MRI scans) for visualization, surgical planning, and 3D printing. However, implication thereof in ID or 2D ultrasound is limited due to the need for precise positional data.

[0006] In conventional ultrasound imaging, a one-dimensional (ID) transducer array is held in contact with the patient and emits sound waves into the body. The reflected echoes are recorded and beamformed to produce two-dimensional (2D) images, representing tissue structure along the lateral (array width) and axial (propagation) dimensions.

[0007] Because 2D images are obtained from three-dimensional (3D) anatomy using a ID transducer, their appearance depends strongly on the transducer’s position and orientation (pose). As a result, obtaining clinically useful images requires a highly skilled operator with knowledge of the relevant anatomy. Three-dimensional ultrasound can mitigate this challenge by reducing the operator-dependent degrees of freedom, as a wider variety of probe positions and orientations can still capture the target anatomy.

[0008] One approach to volumetric ultrasound is to employ a two-dimensional (2D) matrix array transducer, which can capture echoes across both the lateral and elevational dimensions to produce a 3D volume directly. However, the number of transducer elements in a 2D matrix array increases quadratically with array size. While 2D array volumetric scanners are commercially available, they face significant technical challenges at the high frequencies and frame rates typical of ultrasound. In particular, they require high data transfer bandwidth and substantial real-time beamforming computation, which increases cost and complexity. These constraints limit their accessibility and offset some of the natural advantages of ultrasound imaging.

[0009] An alternative approach is to create 3D volumes by registering sequential 2D B-mode slices together, provided their poses are known. This approach leverages the abundant availability of 2D B-mode images from conventional ultrasound. However, because most 2D ultrasound images are acquired freehand without pose tracking, this strategy typically requires additional sensors such as inertial measurement units (IMUs), optical sensors, or cameras to estimate the position and orientation of each image. Alternatively, mechanical software-guided motion can be used to control the transducer and provide known positions for each acquired 2D slice, simplifying the reconstruction of a 3D volume.

[0010] Still, 2D B-mode images are subject to the partial volume effect because their elevational slice thickness can blur boundaries of obliquely intersected structures, complicating 3D reconstruction. Volumes generated from 2D array transducers can perform 3D focusing to mitigate this, but 2D-to-3D registration approaches lack this capability. Post-processing geometric constraints can be applied to reduce these effects, but they add complexity and may not fully recover the original object geometry.

[0011] Even when accurate 2D slice positions are known, converting them into a 3D representation is non-trivial. Common explicit 3D formats, such as voxel grids, polygon meshes and point clouds, are discrete and not always scalable. For example, doubling the resolution in each axis of a voxel grid requires approximately eight times more memory. Interpolating the intensities of missing voxels is particularly challenging in 3D because of the large number of potential reference points.

[0012] Neural radiance fields (NeRF) have been proposed to overcome these challenges by treating the volume as a continuous function and training a neural network to represent it. The network learns to map 3D coordinates and viewing directions to the expected intensity, allowing new views to be rendered from any pose. This approach aligns well with 2D-to-3D problems, since it models the underlying relationship between 3D structure and its 2D projections. In medical imaging, and particularly ultrasound, implicit neural representation (INR) extends this concept by training the network to directly predict image intensity values from 3D coordinates, without simulating the rendering process. INR-based algorithms can address fundamental challenges in 3D reconstruction, including filling in missing voxels and resampling volumes at arbitrary resolution, without relying on traditional interpolation methods. However, as with NeRF, INR approaches require accurate spatial coordinates for each input slice to achieve high-fidelity volumetric reconstruction.

[0013] A further challenge is that even when the target object is present within the field-of- view, it is not trivially separable from background tissue. Often, the goal of reconstruction is to obtain the 3D shape of a specific object within the imaged region. In recent years, deep learning has significantly advanced ultrasound image segmentation performance, and segmentation has become an integral component of many 2D-to-3D imaging pipelines. When segmentation masks of the 2D slices are available, INR or NeRF models can be trained to learn both the B-mode intensity and the segmentation information, enabling the generation of continuous 3D representations of the target object.

[0014] During IVF procedures patients are required to undergo numerous ultrasound exams to assess follicle size. They are required to report to a clinic frequently because it takes considerable experience to assess follicle size. This is because the follicle is a spheroid, but the ultrasound only shows a single plane through it. Thus, the ultrasound exam must be executed carefully to both detect each follicle and to measure the maximal possible diameter is extracted for each follicle.

[0015] Thus, there is a need for systems and methods for generating 3D images based on 2D ultrasound images, to facilitate volumetric estimation of objects in the images, in a reliable, cost effective and efficient manner.

[0016] SUMMARY OF THE INVENTION

[0017] According to embodiments, there are provided herein advantageous systems and methods for generating 3D images based on 2D ultrasound images, to enable volumetric estimation of objects in the reconstructed images.

[0018] According to some embodiments, the systems and methods disclosed herein make use of images obtained by a rotating ultrasonic transducer and a trained algorithm, which is configured to map intensity values of the 2D US images to the absolute 3D coordinates (i.e., x, y and z coordinates), to thereby enable reconstruction of 3D images and further allow the identification and / or volumetric estimation of objects in the image. In some embodiments, the systems and methods disclosed herein are accurate, cost effective, efficient, and ready to implement in various settings.

[0019] According to some embodiments, the systems and methods disclosed herein enhance volumetric ultrasound imaging using ID arrays and motorized rotation of the ultrasonic transducer. By leveraging precise rotation coordinates with the ultrasound data, the generated implicit volume representations improve volumetric sampling and estimation, enabling the generation of a complete 3D image from 2D image slices.

[0020] According to some embodiments, the systems and methods disclosed herein can be used in various volumetric ultrasound applications with static anatomy, including, for example, but not limited to, tumor detection and follicle analysis.

[0021] According to some embodiments, there is provided herein an advantageous pipeline using a ID array mounted on a programmable motor for precise volume scanning and implicit neural representations for continuous 3D reconstruction. As exemplified herein, the implicit neural network’s ability to sample the volume at arbitrary position and resolution was compared to standard interpolation algorithms, achieving enhanced performance boost (over 7.9) for such tasks while maintaining accuracy in locating objects in the volume. Accordingly, there is provided herein a volumetric reconstruction pipeline that was tested on simulated data with a mean accuracy of 93% using only 36 B-mode images. The disclosed algorithm was also evaluated in-vivo to measure the volume of tumors in mice, resulting in a mean error of 6.3%. According to some embodiments, as exemplified herein, implicit neural representations can reduce the amount of data necessary to recreate volumes from 2D slices and replace interpolation-based methods for sampling data from the volume to enable interactive analysis.

[0022] According to some embodiments, as detailed herein, the methods and systems disclosed herein, make use of computer-implemented methods including trained machine learning algorithms, for generating 3D images based on 2D ultrasound slices, enabling estimation of volume of various objects in the image.

[0023] According to some embodiments, the systems disclosed herein utilize a rotating transducer to sample multiple planes of an object, such as a spheroid, quickly. Afterwards, deep-learning algorithms are used for segmentation and reconstruction of the object dimension, diameter, volume, and the like.

[0024] According to some embodiments, the disclosed algorithms include deep neural networks referred to herein as implicit neural representation (INR), which is trained to map intensity values corresponding to a 3D coordinate, using 2D slices which are imaged / acquired at a plurality of plane waves, using a rotating US transducer.

[0025] According to some embodiments, the systems and methods disclosed herein can be used, for example, for estimating the volume of follicles or cysts (for example, in the ovary), for example, for in-vitro fertilization (IVF) procedures., for measuring volume, estimating weight of objects, such as, for example, in fetal imaging.

[0026] According to some embodiments, the systems and methods disclosed herein may be used to train a deep learning algorithm to extract cyst segmentation from 2D slices for volumetric reconstruction. In some embodiments, the segmentation may be facilitated using one or more segmentation tools, and may be performed automatically and / or semi -manually. To this aim, a phantom volume with randomly placed cysts of sizes in the range of 4-15mm was imaged. With a sample size of only 72 frames spaced 2.5° apart, the entire volume was reconstructed.

[0027] In some embodiments, the INR may learn / train to represent the relationship between position (x, y, z) and intensity while applying segmentation in a manner that is useful for recreating the full 3d structure.

[0028] According to some embodiments, as exemplified hereinbelow, in vivo, the systems and methods method were tested in mice to reconstruct 3D breast cancer tumors.

[0029] According to some embodiments, several advantages of implicit volume representation are demonstrated herein. First, 2D slices of the volume were oversampled at arbitrary resolution with high image quality compared to other methods. Next, 2D views of the volume that were not originally acquired in imaging were sampled, along a new axis or at new rotation angles. Finally, the cysts were segmented and their volume was estimated with a mean error of 5.6%. According to some embodiments, there is thus provided a method for generating 3D images of a target region, based on 2D ultrasound images and estimating a volume of an object in the target region, the method includes: acquiring a plurality of images, using a rotary imaging transducer, wherein at least some of the images are obtained at different angles; segmenting each of the obtained images; apply on the images a trained algorithm, to thereby reconstruct a 3D image of the target region, and provide an estimation of volume of one or more objects in the target region; wherein the algorithm is trained on a data set comprising images obtained at different angles and correlated with absolute coordinates of each obtained image.

[0030] According to some embodiments, the algorithm includes a deep neural network.

[0031] According to some embodiments, the segmenting may include applying a deep learning model configured to generate segmentation masks from the acquired 2D ultrasound images.

[0032] According to some embodiments, the deep learning model may include a transformerbased model, a convolutional neural network, or a combination thereof, optionally including a MedSAM-based architecture.

[0033] According to some embodiments, the trained algorithm is configured to produce intensity values for arbitrary 3D spatial coordinates not present in the original 2D ultrasound images.

[0034] According to some embodiments, the trained algorithm may reconstruct the 3D volume using fewer than 40 acquired 2D slices and achieves a volumetric estimation accuracy of at least about 90%.

[0035] According to some embodiments, the trained algorithm may be trained using a dataset comprising 2D ultrasound slices paired with corresponding segmentation masks and absolute spatial coordinates of each slice.

[0036] According to some embodiments, the transducer is configured to acquire 2D B-mode images. According to some embodiments, the transducer is configured to rotate about a single rotation axis.

[0037] According to some embodiments, the rotation is at linearly or non-linearly spaced intervals.

[0038] According to some embodiments, the angles are incremented in the range of about -45 to +45 degrees. In some embodiments, the rotation may be at 1.25 degrees, 2.5 degrees, 5 degrees, etc., intervals.

[0039] According to some embodiments, the objects may include ovarian follicles and / or cysts.

[0040] According to some embodiments, the object(s) may include a tumor.

[0041] According to some embodiments, the objects may include an internal organ or tissue.

[0042] According to some embodiments, the method may be used for identifying and / or estimating follicle and / or cysts volume.

[0043] According to some embodiments, the use may be during / for an IVF procedure.

[0044] According to some embodiments, the use may be for determining tumor size and / or monitoring tumor / cancer progression.

[0045] According to some embodiments, there is provided a system for estimating the volume of an object in a target region, the system includes: a rotating transducer configured to obtain a plurality of 2D images of a target region, wherein at least some of the images are obtained at different angles; and a processing unit functionally and / or physically associated therewith, and configured to: process and segment each obtained image; apply, on the images, a deep learning algorithm, to thereby create an INR for reconstruction of a 3D image of the target region and provide an estimation of volume of objects in the target region; wherein the algorithm is trained on a data set which includes images obtained at different angles and correlated with absolute coordinates (x, y, z) of each image.

[0046] According to some embodiments, the transducer may include: a linear array, phased array, vaginal transducer, convex array, therapeutic array, or any combinations thereof.

[0047] According to some embodiments, the system may include a motor and / or a motorized platform, configured to rotate the transducer.

[0048] According to some embodiments, the motor is functionally or physically associated with the transducer.

[0049] According to some embodiments, the rotation of the transducer is continuous, or at spaced intervals.

[0050] According to some embodiments, the rotation of the transducer is at a predetermined speed.

[0051] According to some embodiments, the rotation of the transducer is at a constant speed.

[0052] According to some embodiments, the system may further include a user interface, a monitor, a controller, a communication unit, or any combination thereof.

[0053] According to some embodiments, there is provided a non-transitory computer readable medium storing computer program instructions for executing the method for generating 3D images of a target region, based on 2D ultrasound images and estimating a volume of an object in the target region, as disclosed herein.

[0054] Certain embodiments of the present disclosure may include some, all, or none of the above advantages. One or more technical advantages may be readily apparent to those skilled in the art from the figures, descriptions and claims included herein.

[0055] In addition to the exemplary aspects and embodiments described above, further aspects and embodiments will become apparent by reference to the figures and by study of the following detailed descriptions. BRIEF DESCRIPTION OF THE FIGURES

[0056] The invention will now be described in relation to certain examples and embodiments with reference to the following illustrative figures.

[0057] Fig. 1A shows - an outline of a framework for generating or producing 3D Ultrasound (US) images based on 2D US images obtained by a rotating transducer, to identify and estimate volume of objects, by applying a trained deep neural network algorithm (INR); shown is the data acquisition step, which includes obtaining a plurality of images by a rotating US transducer; the training step which includes training the INR to map (correlate) the 3D coordinates of each image to the intensity levels, and segmentation of the image; and inference step of producing 3D images of objects using the trained INR, and estimating the volume of the object based thereon;

[0058] Fig. IB shows an illustration of a method for implicit volume representation, according to some embodiments. In the data acquisition phase, the transducer is rotated to acquire images of the volume, and segmentation masks are extracted from each frame. During training, an INR learns to represent the volume as a mapping from image coordinates to B-mode and segmentation masks. The trained INR can freely sample the volume and provide 3D reconstruction of the object.

[0059] Figs. 2A-E - Simulation results of spherical objects with varying contrast levels. Fig. 2A- Target volume representation with voxel-wise sound speed used to create contrast targets and the reconstructed volume. Fig. 2B- Reconstructed INR image presented as 3D B-mode image; Fig. 2C- Examples of a single frame from the set, at a rotation angle of 0°, of Original B-mode; Fig. 2D- Examples of a single frame from the set, at a rotation angle of 0°, of INR reconstructed image; Fig. 2E- Box-and-whisker plot with bars spanning the range of sphere contrast among original and reconstructed frame in the dataset from minimum contrast (bottom) to maximum (top). B-mode images displayed at 0.21mm x 0.32mm resolution (axial x lateral);

[0060] Figs. 3A-E - Sampling with arbitrary coordinates. Algorithm comparison of NN and INR interpolation on up-sampling new views of the rotation data along the x (depth) axis. Figs. 3 A-B- Sample outputs from each algorithm for NN (Fig. 3A), and INR (Fig. 3B) - evaluated at a depth of 4cm. Figs. C-E show Comparison metrics between NN and INR, showing the mean (bar height) and standard deviation (error bar) for each metric: Fig. 3C- Single slice processing time; Fig. 3D- dice coefficient; and Fig. 3E- loU. B-mode images displayed at 0.21mm x 0.32mm resolution (axial x lateral);

[0061] Figs. 4A-G - Effect of angle resolution on volume estimation. Fig. 4A- Ground truth model of individually labelled cysts. Figs. 4B-4E- 3D segmentations produced by the INR network trained at increasingly finer rotational resolution: Fig. 4B- 10°; Fig. 4C- 5°; Fig. 4D- 2.5°; Fig. 4E- 1.5°, Fig. 4F- Volume of each cyst in cm', Fig. 4G- Dice score, loU, and voxel count error in % for each cyst (bars, color-coded according to their labels in a-e) and the mean across all cysts (dotted black line) for each rotation resolution. Cysts are color-coded according to the legend visible in Fig. 4G;

[0062] Figs. 5A-E- Experimental results of volume estimation in a water bead phantom. Figs. 5A-5B- Frames from the dataset at different rotation angles: Fig. 5A- 0=35° and Fig. 5B- 0=140° highlight the need for 3D to understand the composition of the volume, as anechoic water beads enter and exit the frame during scanning; Fig. 5C - 3D B-mode image produced by an INR trained on this dataset; Fig. 5D- Labelled meshes recovered from the INR volume; Fig. 5E- Estimated volume and manual annotation. B-mode images displayed at 0.18mm x 0.15mm resolution (axial x lateral);

[0063] Figs. 6A-J- In-vivo tumor volume estimation results. Fig. 6A- Schematic illustration of the experimental setup. Figs. 6B-6C -Sample images from Tumor #1 at 9=0° (Fig. 6B) and 0=120° (Fig. 6C). Fig. 6D- 3D volume image reconstructed by an INR trained on Tumor #1. Fig. 6E- Mesh of Tumor #1 isolated from the surrounding volume; Fig. 6F- Estimated volume by analysis of the tumor mesh produced by INR compared to manual measurement - a table showing the calculated volume (cm3) of the tumors, based on the INR, compared to the respective measured volumes thereof; Figs. 6G-6 J - Similar to Figs. 6B and Fig. 6E, on Tumor #2; Figs. 6G-6H- sample B-mode images at (g) 9=0° (Fig. 6G) and 9=120° (Fig. 6H). Fig. 61- 3D volume image; and Fig. 6J- reconstructed mesh of Tumor #2. B-mode images displayed at 0.3mm x 0. 15mm resolution (axial x lateral).

[0064] DETAILED DESCRIPTION

[0065] In the following description, various aspects of the disclosure will be described. For the purpose of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the different aspects of the disclosure. However, it will also be apparent to one skilled in the art that the disclosure may be practiced without specific details being presented herein. Furthermore, well-known features may be omitted or simplified in order not to obscure the disclosure.

[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0067] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise, “a” and “an” are used herein to refer to one or more than one (i.e., to at least one) of the stated object, unless the context clearly dictates otherwise. By way of example, “a treatment” means one or more treatments.

[0068] As used herein, the term "about" when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20% or in some instances ±10%, or in some instances ±5%, or in some instances ±1%, or in some instances ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods.

[0069] The term “may” refer to an optional or possible, approach or possibility, but not a requirement. The term “can” refer to a permissible or plausible, approach or possibility, but not a requirement.

[0070] As used herein, "optional" or "optionally" means that the subsequently described event or circumstance does or does not occur, and that the description includes instances where said event or circumstance occurs and instances where it does not.

[0071] As used herein, the term “plurality” is directed to include more than one. In some embodiments, the term plurality includes at least two.

[0072] As used herein, the term “comprising” is synonymous with the terms "including," "containing," or "characterized by," and is inclusive or open-ended i.e. does not exclude additional, unrecited elements. According to some embodiments, the term comprising may be replaced with the term “consisting of’ which excludes any element, step, or ingredient not specified in the claim. According to some embodiments, the term comprising may be replaced with the term “consisting essentially of’ which limits the scope of a claim to the specified materials or steps "and those that do not materially affect the basic and novel characteristics of the claimed invention.

[0073] According to some embodiments, provided herein are ultrasonic systems and methods, including a computer-implemented method for implicit volume representation of volumetric sampling estimation and generation of 3D images from 2D slices.

[0074] In some embodiments, these methods and systems are useful for identifying and estimating the volume, size, diameter of objects (such as tissues and organs) for various biological applications. For example, the objects may be essentially spheroid objects, such as, follicles or cysts. In some embodiments, the objects are tumors (such as, for example, breast cancer). In some embodiments, the methods and systems can be used for measuring volume and estimate weight, such as in fetal imaging. In some embodiments, the methods and systems can be used for measuring various abdominal organs, such as, liver, spleen, kidney, Thyroid, and the like. Each possibility is a separate embodiment. According to some embodiments, the systems and methods disclosed herein can be used for musculoskeletal ultrasound applications.

[0075] According to some embodiments, the systems and methods disclosed herein may be used to determine and / or estimate static volumes.

[0076] According to some embodiments, the systems and methods disclosed herein can be used for follicle tracking during / for IVF procedures. In some embodiments, systems and methods can be used to assess / estimate / determine diameter of cysts or follicles. In some embodiments, the system may be configured to monitor ovarian follicles, including estimating follicle number and total ovarian volume, and to track follicular changes across multiple imaging sessions during in-vitro fertilization (IVF) treatment cycles.

[0077] According to some embodiments, the systems and methods disclosed herein can be used for tumor volume estimation and tracking. In some embodiments, the systems and methods can be used to assess / estimate / determine diameter, size and / or volume of a tumor.

[0078] In some embodiments, the system may be configured to monitor the growth of tumors over time or to provide volumetric guidance for needle-based interventions, including but not limited to biopsies, ablations, and localized drug delivery. According to some embodiments, the systems and methods disclosed herein enable a complete pipeline of volumetric analysis, including, image acquisition, segmentation, labeling and volume estimation.

[0079] Reference is made to Fig. 1A, which shows an outline of the framework for generating or producing 3D (US) images based on 2D US images obtained by a rotating transducer, to identify and estimate volume of objects, according to some embodiments. As shown in Fig. 1A, a rotating transducer is used to obtain / acquire a plurality of US images of a target region, wherein the images are obtained at different angles. As detailed herein, the rotation is facilitated along only a single rotation axis and may be performed at a linearly or non-linearly spaced intervals. At this data acquisition stage, the transducer acquires B-mode images, while rotating. Each image is processed individually and a segmentation mask is produced. The segmentation may be performed by any method known in the art, including, for example, MedSAM model (Nature Communications volume 15, 654 (2024)). Next, at the training step, deep neural network is trained to replicate the B-mode images while given the x-y-z absolute coordinates of each image. In some embodiments, the neural network is referred to herein as an implicit neural representation (INR), since the neural network is trained (learns) to represent the volume implicitly. At the inference step, once the INR is trained, it can be sampled arbitrarily (i.e., at various desired angles or slices). For volume estimation (for example, for follicle or cysts tracking), part or the entire 3D grid of desired points (including points that were not available in the original dataset, but can be inferred), may be sampled to reproduce the 3D object (follicles or cysts), based on the network output. In some embodiments, various volume estimation methods can be applied on the reproduced 3D object, including, for example, creating a mesh wrapping the 3D segmented object.

[0080] According to some embodiments, a deep neural network may be trained to map the intensity value in the image, corresponding to a 3D coordinate. Once the implicit volume network is trained, it can be sampled arbitrarily by providing the desired coordinates. In some embodiments, deep neural network may be based, for example on a NeRF framework (Mildenhall, Ben, et al. "Nerf: Representing scenes as neural radiance fields for view synthesis." Communications of the ACM 65.1 (2021): 99-106).

[0081] Thus, according to some embodiments, there is provided a 2D-to-3D pipeline where 2D images are acquired while advanced INR algorithm is applied to recreate 3D structures therefrom. Such methods are advantageous as it is much more cost and time effective to use 2D imaging, whereas direct 3D imaging requires complex hardware and is often slower to process.

[0082] According to some embodiments, using a rotating transducer enables the use of INR to create volumetric 3D images in a 2D-to-3D pipeline while providing notable advantages of INR. First, INR can sample 3D scenes at arbitrary view or resolution faster than interpolation methods, thereby reaching greater performance boost as the dataset size increases, without compromising image quality metrics like contrast or damaging segmentation masks. In the specific case of volume estimation, as demonstrated herein, the ability to reconstruct target objects with this setup in imaging scenarios that can be difficult for classic 2D-to-3D algorithms, using as little as 18 frames to reconstruct the volume.

[0083] According to some embodiments, using a mechanically operated transducer together with INR can enable existing 2D ultrasound applications to natively extend into 3D without dedicated 3D ultrasound transducers.

[0084] According to some embodiments, the US transducer is configured to acquire 2D B- mode images. In some embodiments, the US may be applied at a frequency of about 1-5 MHz. In some embodiments, the US is at a frequency of about 3.5 MHz. In some embodiments, the transducer is configured to rotate about the axial imaging direction and accordingly can acquire images from a 3D field-of-view, autonomously.

[0085] In some embodiments, the transducer may operate at a center frequency between about 1 MHz and 15 MHz, such as between 2 MHz and 10 MHz, or between 3 MHz and 5 MHz. In specific implementations for ovarian follicle or tumor imaging, the frequency may be about 3.5 MHz.

[0086] In some embodiments, the transducer may include between about 64 and 512 elements, with an element pitch between about 0.1 mm and 0.5 mm, such as about 0.22 mm.

[0087] In some embodiments, the elevation aperture of the transducer may be between about 5 mm and 20 mm, such as about 13.5 mm.

[0088] In some embodiments, the transducer may be rotated about a central axis through an angular range between about 90° and 360°, such as between 120° and 180°, or approximately 180°. In some embodiments, 2D images may be acquired at rotation intervals between about 0.5° and 10°, such as between 1° and 5°, or about 1.25°.

[0089] In some embodiments, the system may acquire between about 12 and 300 frames per volume, such as between 30 and 100 frames, or about 36 to 144 frames, depending on the rotation step and field of view.

[0090] In some embodiments, each 2D frame may be beamformed to an image having between about 128 x 128 and 512 x 512 pixels, such as about 256 x 256 pixels.

[0091] In some embodiments, the reconstruction algorithm may produce a complete volumetric image within between about 1 second and 2 minutes, such as within about 3.8 seconds for a 2563voxel grid.

[0092] In some embodiments, the reconstructed 3D volume may be sampled at a voxel spacing between about 0.05 mm and 1.0 mm, such as between 0.1 mm and 0.5 mm.

[0093] In some embodiments, the transducer may be rotated at an angular velocity between about 1 revolution per minute (rpm) and 60 rpm, such as between 5 rpm and 30 rpm, or about 10 rpm. The rotation speed may be selected based on the desired frame rate and rotation interval, such that sufficient frames are acquired for reconstruction while maintaining stable coupling with the tissue.

[0094] In some embodiments, the ultrasound system may acquire two-dimensional images at a frame rate between about 50 frames per second and 5000 frames per second, such as between about 500 and 2000 frames per second, depending on the number of plane waves used for compounding.

[0095] In some embodiments, a motor element configured to rotate an ultrasound transducer about the axial imaging direction is used, in order to rapidly and accurately analyze the volumetric sizes of objects, such as cysts, follicles or tumors.

[0096] According to some embodiments, the rotation is performed along a single rotation axis, at linearly spaced intervals. In some embodiments, the intervals are non-linear intervals.

[0097] According to some embodiments, the rotation angles may be in the range of about -180 to +180 degrees, -90 to +90 degrees, -60 to +60 degrees, -45 to +45 degrees etc. In some embodiments, the increments in angle measurements may be in the range of -45 to +45 degrees, for example, -10 to +10 degrees, -5 to +5 degrees, etc. Each possibility is a separate embodiment.

[0098] According to some embodiments, the methods and systems disclosed herein may be utilized for follicle tracking, which is a crucial part of IVF procedures that require great expertise to assess the diameter of cysts. According to some embodiments, the methods and systems disclosed herein may be used to instantly sample spheroid objects (such as follicles or cysts) in 3D, to provide a maximal diameter thereof immediately. In some embodiments, such a procedure can advantageously be performed quickly and autonomously without great expertise.

[0099] According to some embodiments, there is provided herein an ultrasound system for estimating the volume of an object in a target region, the system includes: a rotating transducer configured to obtain a plurality of 2D images of a target region, wherein at least some of the images are obtained at different angles; and a processing unit functionally and / or physically associated therewith, and configured to: process and segment each obtained image; apply, on the images, a trained algorithm, to thereby reconstruct a 3D image of the target region and provide an estimation of volume of objects in the target region; wherein the algorithm is trained on a data set which includes images obtained at different angles and correlated with absolute coordinates (x, y, z) of each image.

[0100] According to some embodiments, the system may further include a motor and / or a motorized platform, configured to rotate the transducer, continuously or at spaced intervals. In some embodiments, the rotation is at a predetermined speed. In some embodiments, the rotation is at constant speed. In some embodiments, the transducer is integrally formed with a motor. In some embodiments, the transducer is functionally and / or physically associated with the motor.

[0101] In some embodiments, the transducer may include any type of suitable transducer, including, for example, nut not limited to: linear array, phased array, vaginal transducer, convex array, therapeutic array, and the like. Each possibility is a separate embodiment. According to some embodiments, the transducer may include a one-dimensional linear array, a phased array, or a curved array. The transducer may be operated in any imaging mode suitable for generating two-dimensional B-mode images, including but not limited to planewave or focused beam transmission. The rotating mechanism may be mechanically coupled to any such transducer type.

[0102] In some embodiments, the transducer may be moved along a non-circular path, including but not limited to a helical trajectory, an oscillatory sweeping motion, or a combined rotational and translational path. The system may track the transducer’s position along such paths to assign spatial coordinates to the acquired images. The term “rotation” as used herein may encompass these equivalent scanning motions.

[0103] According to some embodiments, the system may further include one or more posetracking sensors coupled to the transducer, such as an inertial measurement unit (IMU), optical markers, or electromagnetic trackers. Pose information from such sensors may be used to determine the spatial coordinates of each acquired 2D image, either in place of or in addition to mechanical rotation.

[0104] According to some embodiments, the reconstruction algorithm may be configured to operate in real time or near real time, such that volumetric images are generated concurrently with the acquisition of the 2D ultrasound images. Such embodiments may be useful for interventional procedures or intraoperative guidance.

[0105] According to some embodiments, the INR network may be trained to jointly predict voxel intensity and voxel segmentation labels from 3D spatial coordinates. This multitask configuration may enable simultaneous reconstruction of tissue structure and identification of a target object within the reconstructed volume.

[0106] In some embodiments, each two-dimensional image may be formed from between about 1 and 15 compounded plane waves, such as about 5 plane waves.

[0107] In some embodiments, the implicit neural representation network may include between about 5 and 20 layers and between about 64 and 1024 neurons per layer.

[0108] In some embodiments, the network may be trained for between about 100 and 5000 epochs, such as between about 300 and 1000 epochs. In some embodiments, the positional encoding applied to the three-dimensional input coordinates may include between about 5 and 30 frequency components, such as about 15 frequency components.

[0109] In some embodiments, the training loss may comprise a structural similarity index (SSIM), binary cross-entropy loss, Dice coefficient loss, or any combination thereof, optionally with weighting between about 0.1 and 0.9 for each component.

[0110] In some embodiments, the reconstructed volume may include between about 643and 10243voxels, such as between about 1283and 5123voxels.

[0111] In some embodiments, the reconstructed volume may be converted into a three- dimensional surface mesh having between about 1,000 and 1,000,000 vertices, such as between about 10,000 and 100,000 vertices.

[0112] In some embodiments, the implicit neural representation may operate by parameterizing the volume as a continuous function which maps spatial coordinates to intensity values. The network may be configured to approximate this function such that sampling any coordinate within the spatial bounds produces a corresponding ultrasound intensity value.

[0113] In some embodiments, the INR may be configured to receive as input a set of spatial coordinates and to output corresponding intensity values for those coordinates without requiring an explicit voxel grid. The network may be queried at any arbitrary spatial point to generate local intensity values.

[0114] In some embodiments, the INR may reconstruct intermediate intensity values by learning a smooth function over space such that nearby coordinates have similar predicted values. The network may leverage its learned continuous gradient field to interpolate intensities between sparse measured points.

[0115] In some embodiments, the INR may represent both a continuous intensity field and a continuous segmentation field, such that for each spatial coordinate the network outputs a pair (I,S) representing the predicted intensity and the predicted segmentation probability.

[0116] In some embodiments, the INR may be trained by sampling coordinates from the known 2D ultrasound slices and associating each sampled coordinate with the corresponding intensity value from the slice. The loss may be computed as the difference between predicted and true intensities at those sampled coordinates.

[0117] In some embodiments, once trained, the INR may be used to generate a volumetric representation by querying a regular 3D grid of coordinates and assembling the predicted values into a volume. The voxel resolution of the generated volume may be selected independently of the resolution of the acquired images.

[0118] In some embodiments, the INR may be queried along virtual rays passing through the volume to synthesize simulated two-dimensional ultrasound images at arbitrary viewing angles, enabling virtual re-slicing of the reconstructed volume.

[0119] In certain embodiments, the INR may store only the network parameters and not an explicit volume, such that the memory required is substantially lower than the memory needed to store a voxel grid of equivalent resolution.

[0120] In some embodiments, the spatial coordinates input to the INR may be normalized to a fixed coordinate frame, such as rescaling all coordinates into a unit cube or spherical coordinate system, to stabilize training and inference.

[0121] In some embodiments, the implicit neural representation may be configured to preserve the relative contrast of target objects within the ultrasound data. The network may maintain contrast deviations below about 0.2 dB, enabling accurate reconstruction of low-contrast targets for downstream volumetric analysis.

[0122] In some embodiments, the INR model may represent the volumetric data in a compressed manner, such that the total size of the network parameters is smaller than the cumulative size of the acquired images. For example, the trained model may be approximately three times smaller than the original dataset while retaining equivalent volumetric information.

[0123] In some embodiments, the trained INR may be configured to generate two-dimensional images from arbitrary viewpoints by sampling the volume at a grid of spatial coordinates to produce a synthetic B-mode image. The computational complexity of such sampling may be independent of the size of the original dataset and may be implemented as a parallelizable GPU operation. In some embodiments, the inference time of the INR may not increase with the number of images used for training. The computational complexity may be substantially constant with respect to dataset size, providing an advantage over nearest-neighbor interpolation or k-d treebased approaches whose complexity increases as the dataset grows.

[0124] In some embodiments, the INR may be trained on datasets comprising objects with diverse shapes, sizes, and acoustic properties, including simulated phantoms, in-silico phantoms, and in-vivo biological tissues. The INR may reconstruct arbitrary geometries including irregularly shaped tumors or vascular structures.

[0125] In some embodiments, the reconstruction error of the INR may be substantially uncorrelated with the distance of a voxel from the rotation axis, while the effective voxel size may increase with radial distance due to the cylindrical geometry of the acquisition. In such embodiments, it may be advantageous to center the target object on the rotation axis to maximize spatial resolution.

[0126] In some embodiments, the INR may reconstruct three-dimensional volumes accurately using as few as about 12 to 40 acquired frames, achieving volumetric accuracy of at least 90 percent. In other embodiments, the rotation angle between frames may be decreased to increase accuracy when scan time is less critical.

[0127] In some embodiments, the imaging system may be configured to minimize the number of frames needed for reconstruction while maintaining accuracy by centering the object near the rotation axis and selecting the finest feasible rotation angle.

[0128] In some embodiments, the system may accurately reconstruct scenes containing multiple similar-looking objects entering and exiting the field-of-view. The known transducer coordinates provided by the mechanical rotation may enable reconstruction without reliance on feature-based registration methods.

[0129] In some embodiments, the INR-based reconstruction may be used for volumetric estimation of biological structures, including but not limited to, ovarian follicles, cysts, tumors, vascular structures such as the aorta, and neural structures such as the spinal cord.

[0130] In some embodiments, the volumetric accuracy of the reconstructed volume may depend on the quality of the segmentation masks used during training. In some embodiments, the segmentation module may comprise an automated model trained specifically for the target tissue type. Automated segmentation may improve accuracy compared to generic segmentation models when applied to ultrasound data of a specific anatomical region.

[0131] In some embodiments, the segmentation model may be configured to maintain accurate masks when an object is gradually exiting the field-of-view, thereby reducing geometric distortions in the reconstructed surface.

[0132] In some embodiments, the key mechanical feature may be the provision of known spatial coordinates of each acquired image. The transducer may be mounted on a robotic or mechanically guided apparatus configured to move along trajectories other than pure rotation, such as linear, helical, or combined paths, while still providing accurate pose data to the INR.

[0133] According to some embodiments, the system may further include a user interface, a monitor, a controller, a communication unit, and the like.

[0134] In some embodiments, the processor is remotely based. In some embodiments, the processor is physically associated with the system.

[0135] In some embodiments, the ultrasound data may be transmitted to a remote server for reconstruction, and the resulting volumetric data returned to the acquisition device for display. Such configurations may enable the use of lightweight imaging hardware in conjunction with centralized computing resources.

[0136] The following examples are presented in order to more fully illustrate some embodiments of the invention. They should in no way be construed, however, as limiting the broad scope of the invention. One skilled in the art can readily devise many variations and modifications of the principles disclosed herein without departing from the scope of the invention.

[0137] EXAMPLES

[0138] Materials and methods of Ultrasound Volumes

[0139] In classical computer vision, natural images describe a 2D projection of a 3D scene, and the NeRF algorithm learns a 3D function that implicitly describes the available image data, each of which is a projection of a volume from a different viewpoint. However, in ultrasound, the image plane directly intersects with the volume, and B-mode intensity of a pixel is directly representative of the volume at its coordinate. Optionally, the network can be trained to predict a binary segmentation mask of the volume. The resulting INR model learns to approximate a continuous function F that can be sampled at 3D positions to retrieve B-mode intensity I alongside the segmentation mask S:

[0140] F(x,y, z) = (I,S) (1)

[0141] Once an INR model is fitted to the volume, the continuous implicit representation for various tasks is used. By up-sampling the coordinates (x, y, z), the network output will be an up-sampled version of the original image. The (x, y, z) vectors can be selected in ways that were not originally present in the dataset, to synthesize new views of the data that the transducer did not necessarily image. Such tasks, normally completed with various interpolation schemes that require complex data structures to process in 3D, are simple function evaluations to an INR. Each instance of the INR model is separately trained on a specific dataset, in a one-to- one relationship between a single dataset and INR model.

[0142] Imaging setup

[0143] The rotating transducer setup is based on a 28.2 mm-wide, 128-element phased array with a pitch of 0.22 mm, elevation aperture of 13.5 mm, and center frequency of 3.5 MHz (IP 104, Sonic Concepts, Sonic Concepts, Bothell, WA, USA)’ assembled on a motorized rotary (RTY-IP100, Sonic Concepts) and controlled by a programmable ultrasound system (Vantage 256, Verasonics Inc., Redmond, WA, USA). This rotating transducer setup is used across all experiments. In the case of the in-silico and in-vivo experiments, the diameter of objects is manually annotated on the B-mode images.

[0144] The fast volumetric ultrasound scan includes sparse 2D slices, each incorporating several plane wave acquisitions alongside the corresponding absolute coordinates. The obtained frames are spaced apart, at linear intervals (for example, of about 1.25 or 2.5 degrees).

[0145] Simulated Data

[0146] In the first set of experiments, this transducer was simulated using a Python wrapper for k-Wave. The transducer’s operating frequency was set to 3 MHz and a field-of-view of 5.5 cm in the axial direction and 8.2 cm in the lateral direction was created using a virtual grid of size 280x412x88. Six spheres with random radius in the range of [4, 15] mm were randomly placed in a volume of soft-tissue. The volume was sampled at 1.25° rotation intervals such that the entire 180° of the volume was scanned in 144 frames. Five plane waves were simulated at each rotation angle, each steered at angles linearly spaced along [-5°, 5°]. Raw RF data was stored for post-processing. Two experiments were created by modifying the acoustic parameters of the spheres. In the first experiment spheres were simulated to be weak contrast targets in the range of [-2, 8] dB with respect to the background to test the INR’s accuracy in representation of intensity. A sound speed of 1580 m / s was used in dark spheres and 1480 m / s in bright targets. Sphere brightness was induced by varying the scattering coefficient between 0.2 in the brightest target to 0 in the darkest. The background received a scattering coefficient of 0.01 and sound speed of 1540 m / s to resemble soft tissue. In the second experiment aimed at estimating volume quality, the same volume was used but the targets were homogenous and anechoic.

[0147] In the phantom and in-vivo studies, the transducer setup was submerged in a water tank filled with degassed water. The transducer was located at the bottom of the tank so that the positive axial direction pointed upwards to the imaging target located above it (as illustrated in Figs. 1A-B). The transducer was operated by the programmable ultrasound system using an imaging script written in MATLAB (Mathworks, Natick, MA). The same imaging sequence described in the simulation acquisition was used to capture plane wave RF data and store it for offline processing. In both studies, the volume was sampled at 1.25° rotation intervals to produce a dataset of 144 frames in approximately 1 minute, which was deemed fast enough for the purposes.

[0148] Tissue-Mimicking Phantom

[0149] A custom 9 x 6 x 3 cm3rectangular mold was filled with 250 mL of degassed water with 2.5 g agarose powder (Thermo Fisher Scientific Inc, Waltham, MA) and 1.5 g of Silicon carbide (Sigma-Aldrich, St. Louis, MO) which acted as scatterers. Scatterers were prevented from sinking to the bottom of the phantom by submerging a magnetic rod in the solution and inducing rotation of the rod with a magnetic field until the solution was brought to room temperature. Next, the solution was poured into the phantom mold, and while solidifying, 3 water gel beads (UPC: 786194299504, Made Top), soaked in water for 30 minutes with slight variation between them to create size differences, were placed together at the same depth of 3.5 cm from the transducer but spaced evenly in a triangle shape so that the rotating transducer aligned to the center of the triangle could not image all three beads simultaneously. The transducer was placed in a degassed water tank with the phantom submerged in the bath directly above it so that by rotating the transducer different views of the beads could be captured. An imaging field-of-view of 4.7 cm in the axial direction and 3.8 cm in the lateral direction was acquired at each rotation angle. Gel beads presumably continued to grow throughout imaging as they are highly water absorbent, so their diameters were measured by annotating the B-mode images acquired from this experiment, resulting in diameters of 4.8 mm, 5.1 mm, and 6.9 mm. Due to the perfectly spherical shape of the beads, ground truth sphere volume was calculated using the sphere volume equation based on these diameters and compared to the volume of the convex hull bounding the INR’s segmentation prediction.

[0150] In-vivo Experiments

[0151] The estimation method was tested in vivo on five tumor-bearing mice. Animal-related procedures were conducted in accordance with the guidelines provided by the Institutional Animal Research Ethical Committee. Met-1 mouse breast carcinoma cells were cultured in Dulbecco modified Eagle medium (DMEM, high glucose, supplemented with 10% v / v fetal bovine serum, 1% v / v penicillin-streptomycin and 0.11 g / L sodium pyruvate) at 37 °C in a humidified 5% CO2 incubator until about 85% confluency on the day of the injection. Cells were collected by dissociation with TrypLE Express and resuspended at 1 x 106cells in 25 pL PBS+ / + for bilateral subcutaneous injection into #4 inguinal mammary fat pad of injected female FVB / NHanHsd mice (Envigo, Jerusalem, Israel). Before imaging, anesthesia was induced with 2% isoflurane in ambient air (180 mL / min) and the area in proximity of the tumor was shaved. Fur was removed with a depilatory cream to improve coupling along with ultrasound gel. Mice were placed on their side on an agar spacer prepared similarly to the one in the phantom study, at the top of the filled water tank. Unlike the phantom experiments, the imaging needed to pass through the entire depth of the water tank to reach the target, and as a result the axial field-of-view was enlarged to 7.7 cm. The lateral field-of-view remained the same at 3.8 cm.

[0152] Images were collected at 35 days after cell injection, and the tumor size was measured in each direction using the B-mode images from the 3.5 MHz rotating transducer. As the tumors were visibly elliptical to the eye, the ground-truth tumor volume was calculated as the volume of an ellipse given the measured diameters in each direction, and compared to the volume of the convex hull bounding the INR segmentation prediction. The various datasets used are summarized in Table 1 below.

[0153] Table 1. Summary of datasets used to showcase the properties of INR in 2D-to-3D ultrasound

[0154] Data Processing and INR Implementation

[0155] After data collection for a given volume, images were beamformed with a custom CUDA library to 2562resolution and a dynamic range of 20 dB for simulated data and 40 dB for phantom and in-vivo data. Segmentation masks were created semi-automatically using MedSAM to isolate the interest region from the background in each rotated image. An INR was trained using PyTorch on an NVIDIA RTX 3080 GPU (NVIDIA Corporation, Santa Clara, CA) with 10GB of GPU memory. Input coordinates to INR models employed positional encoding and sinusoidal representation network (SIREN) activations to encourage the models to learn high-frequency information. Each 3D coordinate Xi in (x, y, z) was encoded at a logarithmic harmonic scale:

[0156] This was the input to the model which was composed of a 10 layer multi-layer perceptron. Each layer had a width of 256, with skip connections every other layer and SIREN activations. Since positional encoding guides the network to finer resolution of volume representation, N=15 was empirically used for the positional encoding, at which point the output images created by the INR visibly showed speckle patterns similar to those in the original images. This gave the network input a shape of 93. The final layer produced an output of size 2 for each input position. Thus, a Bx3 batch of positions samples the network to produce a Bx2 vector of intensity and segmentation mask at each of the positions (Fig. IB).

[0157] During training of an INR on a particular volume, images, segmentation masks, and their corresponding coordinates were sampled from the available rotation angles, and the INR model was trained to replicate the image and binary mask for a given set of coordinates. The structural similarity index measure (SSIM) was used as the loss metric for image reconstruction (Limage, Fig. IB):

[0158] Where px, ox are the local mean and standard deviation of the image x and oxy is the local covariance of images x, y.

[0159] For binary segmentation, the loss between the prediction x and the ground truth mask y used binary cross-entropy together with the dice score and penalized the INR pixel-wise for false positives (LBCE, Fig. IB):

[0160] The Adam optimizer was used for training for 500 epochs per volume, sped up using grokfast to achieve peak performance in minimal training epochs. The learning rate was initially set at 1x10-3 with a cosine annealing schedule that reached a minimum of 1 xlO-6. After training, the entire volume was sampled at once as a voxel grid at the desired resolution, and objects in the volume were labelled using cc3d before being converted to a triangle mesh wrapping the convex hull of the voxels. Once trained, INR was compared to voxel-wise nearest neighbor to assess its capabilities in 3D imaging.

[0161] Image Similarity Metrics

[0162] In experiments using simulated data, the INR can be compared to the true model or original image data from which the data is simulated to evaluate the model’s ability to capture the scene accurately. In the contrast simulation, the contrast between the ithcontrast target 0 < i < N and the background in dB is compared between original images and INR-generated using the formula: where p is the mean intensity in a particular region in dB. In the case of volume estimation, the binary segmentation mask predicted by INR for the ithobject yi, is compared to the true segmentation mask yi using the dice coefficient and intersection-over-union (loU):

[0163] In addition, the final volumetric error in % was calculated, which indicates the discrepancy in estimation of a target volume Vi compared to the actual volume Vi:

[0164] The volume estimation metrics described in Equations (6-8) are all in the range [0, 1], It is desirable to maximize the dice coefficient and loU while minimizing the relative volume error.

[0165] Results:

[0166] Volumetric Imaging Simulation

[0167] First, the ability of INR to represent images with diverse contrast was demonstrated. A simulated dataset of low-contrast spheres embedded in tissue-mimicking phantom was used to train an INR on the full dataset that includes 2D images of the sampled volume at 1.25° rotation intervals such that the entire 180° of the volume was scanned in 144 frames. The volume include 6 spheres: three anechoic and three echogenic with random sphere sizes (4-15 mm in diameter) distributed randomly in the volume (as shown in Figs. 2A-E). The volume itself is displayed with each voxel colored according to its sound speed (Fig. 2A). Each 2D image captured a cross section of the volume, hence, only 1-3 spheres were observed in individual frames at a time (Fig. 2C). The trained INR reconstruction was compared to individual frames (Figs. 2C-D), to validate the accuracy of INR reconstruction. By sampling the INR at coordinates corresponding to the entire volume at once, the volume image was created (Fig. 2B) The INR reconstruction for each sphere was evaluated on a per-frame basis by analyzing each sphere contrast for each rotation angle in the dataset as a box-and-whisker plot (Fig. 2E), where the minimum, median, and maximum of contrast are shown across all frames containing each of the spheres as the lower bar, middle bar, and top bar for each box. During rotation of the transducer, sphere size and position determined how many of the 144 total frames contained a slice of each sphere. As a result, each of the target objects appears for a different number of frames in the dataset. The average deviation of contrast between INR-generated images and originals across all spheres and frames was 0.17 dB. A two-sided t-test suggested that the contrast of the original frames was not statistically significantly different from the INR frames with a p-value of 0.95.

[0168] Ability of the trained algorithm to reconstruct 3D objects based on 2D images

[0169] To assess the capability of INR for arbitrary sampling, the INR network is sampled in slices along the axial direction, perpendicular to the image plane used during data acquisition at xl6 the original image resolution (40962) and compared to nearest neighbor interpolation (NN). A sample slice from 4 cm depth shows the B-mode and segmentation mask results of NN interpolation (Fig. 3A) and INR (Fig. 3B). The dice coefficient and loU scores for each slice are calculated relative to the phantom model used to simulate the data (Figs. 3D-3E). A one sided t-test was performed to evaluate whether the INR outscored NN in dice coefficient or loll, and whether INR performed the calculation significantly faster. Although the p-value is insignificant for both dice coefficient (0.41) and loU (0.43), INR completes the calculation x7.9 faster (3.8 seconds for INR and 30.4 seconds for NN) thanks to its simplicity and reliance on GPU parallelization (Fig. 3C) resulting in a p-value of 0 when comparing the processing time for a single 40962frame. This means that INR can recreate images and segmentation masks significantly faster than NN without a decrease in segmentation metrics.

[0170] Ability of the trained algorithm for volume estimation of anechoic cysts

[0171] Next, a volume estimation setup, in which the same volume model is used but the cysts are anechoic was used. The INR’s volume estimation accuracy was analyzed based on the resolution of the rotation angle used in the volume sweep (Figs. 4A-G), compared to the digital model from which the data is simulated (Fig. 4A), as well as the INR reconstruction with progressively finer rotation angles (Figs. 4B-E). The resulting volumetric estimation is summarized in Fig. 4F. The dice coefficient, loU, and voxel count error for each sphere in the dataset are shown, as well as the mean across spheres (dotted line) at each rotation angle (Fig. 4G). All three metrics showcase the relationship between rotation angle resolution and the size of the target volume, where the smallest sphere with a volume of 0.53 cm3is notably underestimated when the rotation angle resolution is inadequate. Still, the larger volumes are reconstructed even with a rotation angle resolution of 10°- this corresponds to a dataset of just 18 frames spanning the entire volume. Although the volumetric accuracy continues to improve with finer rotation angles, an average accuracy of 93% is achieved across cysts using 5° resolution.

[0172] In addition, the spatial distribution of INR errors were tested, to determine the effect of the volume being sampled more densely near the rotation axis. The trained subsampled INRs (Figs. 4B-D) were tested on the full dataset of 144 frames sampled at 1.25° intervals. INR outputs of intensity and binary masks were compared to the true data sampled at each angle. For image outputs, the mean-squared error of the pixel -wise normalized images was calculated, and for binary segmentation masks the percentage of incorrectly segmented pixels. Finally, the spearman correlation between the errors with the radial distance from the rotation axis were calculated. This gave a mean correlation of -0.03 for image intensity errors and -0.05 for binary segmentation errors, suggesting that INR error is not monotonically distributed with the radial distance despite having more data points to train from near the rotation axis;

[0173] Volume Estimation of a Real-World Multi-Object Phantom

[0174] A 3D agarose-based phantom containing anechoic water beads embedded within agarose was used to test the method experimentally. The phantom contained 3 beads located triangularly at the same depth of 3.5 cm. In each 2D ultrasound image, the beads either entered or exited the frame but never appeared simultaneously during rotation of the transducer. Examples of individual 2D images at angles of 35° and 140° from the dataset of 144 are shown (Figs. 5A-B). After training, the resulting INR is sampled at the entire volume to produce the 3D intensity image (Fig. 5C). The 3D segmentation of each bead was generated and visualized as a triangle mesh (Fig. 5D), from which the volume of each bead was calculated. The recovered mesh volume was compared to the annotated volume of the beads (Fig. 5E), achieving a mean error of 5.5± 2.5%.

[0175] Volume Estimation of In-Vivo Tumors

[0176] In-vivo scans were performed on five tumor-bearing mice to reconstruct their 3D volumetric structure and calculate tumor volumes. Imaging was performed with the mice lying on their side on top of an agar spacer above the water tank (Fig. 6A). For the representative two tumors in Figs. 6A-J, a pair of 2D ultrasound images at angles of 0° and 120° are presented (Figs. 6B, 6C, 6G, 6H). These images highlight the view-dependent morphology of the tumors. After training an INR on each of these datasets, the 3D volume images are reconstructed (Figs. 6D and 61), and tumor meshes are recovered (Figs. 6E and 6J). Finally, the tumor volumes encompassed by the meshes are compared to the measurements, exhibiting an error of 6.8± 1.5% (Fig. 6F).

[0177] The experiment was repeated with five different mice. Table 2 compares the INR’s approximated volume for each tumor with the elliptical approximation annotated from the B- mode images.

[0178] Table 2. Tumor volume estimation in vivo using manual physical measurements and volumes estimated from the tumor mesh produced by INR across five mice.

[0179] The results presented herein indicate that the trained algorithm can successfully be utilized to generate 3D images based on the obtained 2D images, and facilitate determination of volume of objects in the target region being imaged.

[0180] Further, the results demonstrate the ability of the algorithm to successfully arbitrary sample coordinates and create 2D image slices at various angels that were not sampled.

[0181] The results further demonstrate that the algorithm is indifferent to image contrast, as demonstrated by the ability of the algorithm to reproduce the original image contrast. While certain embodiments of the invention have been illustrated and described, it will be clear that the invention is not limited to the embodiments described herein. Numerous modifications, changes, variations, substitutions and equivalents will be apparent to those skilled in the art without departing from the spirit and scope of the present invention as described by the claims, which follow.

Claims

CLAIMSWhat we claim is1. A method for generating 3D images of a target region, based on 2D ultrasound images and estimating a volume of an object in the target region, the method comprising: acquiring a plurality of images, using a rotary imaging transducer, wherein at least some of the images are obtained at different angles; segmenting each of the obtained images; applying on the images a trained algorithm, to thereby reconstruct a 3D image of the target region, and provide an estimation of volume of one or more objects in the target region; wherein the algorithm is trained on a data set comprising images obtained at different angles and correlated with absolute coordinates of each obtained image.

2. The method of claim 1, wherein the algorithm comprises a deep neural network.

3. The method according to any one of claims 1-2, wherein the transducer is configured to acquire 2D B-mode images.

4. The method according to any one of claims 1-3, wherein the transducer is configured to rotate about a single rotation axis.

5. The method according to any one of claims 1-4, wherein the rotation is at linearly or non- linearly spaced intervals.

6. The method according to any one of claims 1-5, wherein the angles are incremented in the range of about -45 to +45 degrees.

7. The method according to any one of claims 1-6, wherein the objects comprise ovarian follicles and / or cysts.

8. The method according to any one of claims 1-7, wherein the object(s) comprise a tumor.

9. According to some embodiments, wherein the object(s) comprise an internal organ or tissue.

10. The method according to any one of claims 1-9 wherein the method is used for identifying and / or estimating follicle and / or cysts volume.

11. The method according to claim 10, for use during or for an IVF procedure.

12. The method according to any one of claims 1-11, for estimating a volume of a tumor.

13. A system for estimating the volume of an object in a target region, the system comprising: a rotating transducer configured to obtain a plurality of 2D images of a target region, wherein at least some of the images are obtained at different angles; and a processing unit functionally and / or physically associated therewith, and configured to: process and segment each obtained image; apply, on the images, a deep learning algorithm, to thereby create an INR for reconstruction of a 3D image of the target region and provide an estimation of volume of objects in the target region; wherein the algorithm is trained on a data set which includes images obtained at different angles and correlated with absolute coordinates (x, y, z) of each image.

14. The system according to claim 13, wherein the transducer comprises: a linear array, phased array, vaginal transducer, convex array, therapeutic array, or any combinations thereof.

15. The system according to any one of claims 13-14, comprising a motor and / or a motorized platform, configured to rotate the transducer.

16. The system according to claim 15, wherein the motor is functionally or physically associated with the transducer.

17. The system according to any one of claims 13-16, wherein the rotation of the transducer is continuous or at spaced intervals.

18. The system according to any one of claims 13-17, wherein the rotation of the transducer is at a predetermined speed.

19. The system according to any one of claims 13-18, wherein the rotation of the transducer is at a constant speed.

20. The system according to any one of claims 13-19, further comprising a user interface, a monitor, a controller, a communication unit, or any combination thereof.

21. A non-transitory computer readable medium storing computer program instructions for executing the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Three-dimensional ultrasonic diagnosis device for gynecological diseases

    CN111436972A

  • Volume acquisition method for object in ultrasonic image and related ultrasonic system

    TW202216075A

  • Method and system for defining cut lines to generate a 3D fetal representation

    US20220047241A1