Surgical guidance with composite ultrasound imaging

By compounding ultrasound data in the nerve field and using 3D deformation field to deform the preoperative data in real time, the artifacts and synchronization problems of 3D ultrasound imaging guidance during the operation are solved, and high-precision and real-time image guidance effect is achieved.

CN119970225APending Publication Date: 2025-05-13SIEMENS MEDICAL SOLUTIONS USA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411590004.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-12
Filing Date
2024-11-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When using 3D ultrasound imaging to guide images during surgery, there is a problem of confusion between artifacts in rendering and operators for image guidance, and due to the pressure introduced by the ultrasound probe, the organ volume deforms, it is difficult to provide a view synchronized with preoperative imaging.

Method used

The ultrasound data is compounded into 3D representations through the nerve field, and the preoperative data is deformed in real time using the 3D deformation field to achieve a synchronous view of ultrasound and preoperative imaging.

Benefits of technology

This enables accurate 3D ultrasound imaging during surgery, reduces artifacts in rendering, improves operator clarity on image guidance, and allows real-time synchronous display of intraoperative and preoperative data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119970225A_ABST
    Figure CN119970225A_ABST
Patent Text Reader

Abstract

Surgical guidance with composite ultrasound imaging. In a method for surgical guidance with compound ultrasound imaging, a neural field uses both probe tracking and ultrasound imaging to compound ultrasound data into three dimensions while aligning component fields of view. Accurate compositing is provided, and the compositing can be operated in real time. In another method for visualization with preoperative data, modeling from ultrasound (e.g., a neural field) is used to generate a 3D deformation field, which is then applied to the preoperative data. The 3D deformation field may be applied to the rendering rather than the pre-operative volume dataset. The 3D deformation can be used in real time, allowing for synchronized 3D ultrasound and 3D preoperative imaging.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This patent document claims the benefit under 35 U.S.C. §119(e) of the filing date of U.S. Provisional Patent Application Serial No. 63 / 597,445, filed on November 9, 2023, which is incorporated herein by reference. Background Art

[0003] The present embodiments relate to ultrasound-based image guidance for surgical procedures (e.g., laparoscopy). For example, anatomical structures are displayed from intraoperative internal tissue imaging. Image-based surgical guidance typically involves displaying images from 3D data derived from preoperative imaging. Such data augmentation during invasive surgery can employ tissue surface reconstruction algorithms to support the use of geometry-based measurements and improved depth cues during visualization. The accuracy of such data augmentation depends largely on the type and quality of patient imaging and its alignment with the real world. Dynamic neural radiation fields (NeRFs) have been used for deformable tissue reconstruction from stereo laparoscopic cameras, which handle complex surgical scenarios better than surface-based methods. The separated neural radiation fields reconstructed during surgery can be further fused in a single global field.

[0004] Three-dimensional (3D) ultrasound facilitates more complete visualization of organs and internal structures (e.g., vascular trees in the liver), however, the performance of conventional methods for compounding 2D ultrasound into 3D volumes does not scale well to realistic intraoperative imaging, resulting in artifacts in renderings and operator confusion regarding image guidance. Furthermore, because intraoperative ultrasound introduces volumetric deformation of organs due to the required pressure from the ultrasound probe, it can be challenging to provide views that are synchronized in real time with preoperative imaging and planning data.

[0005] Deformable or non-rigid registration between imaging modalities is used in a variety of clinical applications. Manual segmentation of the prostate in both preoperative computed tomography (CT) and ultrasound can be used for registration, but this does not provide real-time performance and is susceptible to segmentation variability. Deep learning-based methods have only achieved real-time performance for 2D ultrasound to 3D CT registration. In another approach, the segmented volume data is converted into a surface (grid) and deformation physics is simulated on the grid data. Synthetic ultrasound images are generated from the grid to produce a deformed ultrasound data set for use in a surgical simulator. This method requires segmentation and deformation of ultrasound data, which does not provide image guidance from preoperative imaging. Summary of the invention

[0006] Systems, methods, and non-transitory computer-readable media are provided for surgical guidance using composite ultrasound imaging. In one method, a neural field uses both probe tracking and ultrasound imaging to composite ultrasound data into three dimensions while aligning the component fields of view. Accurate composites are provided, and the composites can operate in real time. In another method for visualization using preoperative data, modeling from ultrasound (e.g., a neural field) is used to generate a 3D deformation field, which is then applied to the preoperative data. The 3D deformation field can be applied to the rendering instead of the preoperative volume data set. The 3D deformation can be used in real time, allowing synchronized 3D ultrasound and 3D preoperative imaging.

[0007] In a first aspect, a method for surgical guidance using composite ultrasound imaging by an ultrasound system is provided. The ultrasound system scans tissue of a patient, thereby generating a two-dimensional (2D) representation. The position of the 2D representation is tracked. The 2D representation is composited by inputting the position and the 2D representation into a neural field. The neural field is trained using joint optimization of a pose and parameters of the neural field. The composite provides a 3D representation of the patient. A first image is rendered and displayed based on the 3D representation.

[0008] In a second aspect, a medical system for ultrasound compounding is provided. An ultrasound probe tracker is configured to track an ultrasound probe during acquisition of ultrasound data. A memory is configured to store a machine learning neural network formed as a neural field. An image processor is configured to use the neural field to compound the ultrasound data into a volume representation. The compounded ultrasound data has refined position information from the ultrasound probe tracker. A display is configured to display an image from the volume representation.

[0009] In a third aspect, a method for surgical guidance using composite ultrasound imaging by an ultrasound system is provided. The ultrasound system scans tissue of a patient, thereby generating a 2D representation. Pressure corresponding to the 2D representation is tracked. A model generates a 3D deformation field based on the pressure and the 2D representation. The 3D deformation field is used to render and display a preoperative image in real time with the scan.

[0010] The illustrative embodiments listed below summarize other features or aspects. Any one or more of the aspects described above or in the illustrative embodiments may be used alone or in combination with other in the illustrative embodiments, features or aspects. Any aspect or feature of one of the methods, systems, or computer-readable media may be used in other in the methods, systems, or computer-readable media. These and other aspects, features, and advantages will become apparent from the following detailed description of the preferred embodiments, which will be read in conjunction with the accompanying drawings. The present invention is defined by the following claims, and nothing in this section should be considered a limitation on those claims. Other aspects and advantages of the present invention are discussed below in conjunction with the preferred embodiments, and may be claimed later, independently or in combination. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Components and drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the embodiments. Moreover, in the figures, like reference numerals designate corresponding parts throughout the different views.

[0012] Figure 1 is a flow chart of one implementation of a method for 3D compounding in ultrasound;

[0013] Figure 2 Example ultrasound images from scanned and rendered images are shown;

[0014] Figure 3 An example of compositing a 2D image into a 3D representation is illustrated;

[0015] Figure 4 An example volume rendered segmented image is illustrated;

[0016] Figure 5 is a flow chart of one implementation of a method for real-time deformation based modeling of 3D ultrasound composite preoperative imaging;

[0017] Figure 6 An example 3D deformation is illustrated;

[0018] Figure 7 illustrates an example visualization using 3D deformation when rendering a preoperative image;

[0019] Figure 8 is a block diagram of an implementation of a medical system for 3D compounding in ultrasound;

[0020] Fig. 9 illustrates 3D compounding using deformation for image guidance in surgery; and

[0021] Fig.10 An example neural network is illustrated. DETAILED DESCRIPTION

[0022] In one method, surgical image guidance using neural 3D composite ultrasound improves the anatomical field of view during interventional surgery compared to 2D ultrasound imaging. In another method, 3D composite ultrasound and preoperative (e.g., CT planning) data are provided in real time, such as based on real-time deformable volume rendering. The methods can be used together, such as for composite neural fields to provide 3D deformations for deforming preoperative imaging to reflect the current state of tissue during surgery. In the example used in this article, a method for laparoscopy in the liver is described. For example, liver surgery guidance is provided using deformation-aware ultrasound compounding and visualization. The method or a combination of methods can be used in other interventions (e.g., open surgery or minimally invasive surgery) and / or for other organs (e.g., kidneys).

[0023] In a deformation method, 3D tissue deformations are modeled during intraoperative compounding of 2D ultrasound images. The resulting deformation field improves the alignment of preoperative images (e.g., CT volume data) with ultrasound for real-time image guidance during surgical procedures. Model-based deformations from ultrasound compounding algorithms are used to generate 3D deformation fields that can be directly applied during volume rendering from preoperative data. Techniques for direct visualization of deformation displacements and time-based visualization of tissue changes are provided. The continuously updated deformation field is applied to real-time volume and surface visualization.

[0024] In one implementation, the end-to-end system enables rendering of deformed pre-operative data sets based on ultrasound captured in real-time during surgery, allowing for use during surgery. For example, the architecture provides a real-time system for rendering from ultrasound compounding to deformed data with a 128×128×128 volume compounding update rate of 10 Hz and a rendering update rate of 60 Hz.

[0025] For the neural field approach, the neural field models the 3D scene. Instead of explicitly rendering primitives, a coordinate-based neural network is used in structural reconstruction and / or visualization. The weights in the neural network implicitly represent the voxel data. Ultrasound compounding and anatomical structure reconstruction are based on the neural field. The neural network is initialized from 2D ultrasound slices using meta-learning techniques. The neural network is trained based on pose input (e.g., tracked orientation) to jointly refine the physical probe tracking data with compounding the 2D ultrasound data into a 3D volumetric dataset. The neural field representation is refined using joint optimization of image pose and network weights.

[0026] Structural visualization (e.g., 3D segmentation) can be based on neural signed distance fields from neural fields. Rather than using 3D segmentation, neural fields can learn 3D signed distance functions directly from 2D markers or segmentation of 2D ultrasound images.

[0027] In other approaches, the deformation field is learned during training. The neural field is trained to output the deformation field along with or instead of the 3D compound ultrasound data. The deformation field is directly learned during training (e.g., by training in conjunction with differentiable deformable volume rendering).

[0028] In yet another approach, view optimization is based on differentiable direct volume rendering. In traditional image synthesis, differentiable rendering models the explicit relationship between rendering parameters and the resulting image. Image space derivatives are obtained with respect to the rendering parameters, which can be used in a variety of gradient-based optimization methods to solve the inverse rendering problem or to compute the loss of training machine learning models directly in the space of rendered images.

[0029] Figure 1Compounding involving ultrasound imaging. Compounding results from a scanned volume together based on the position of the field of view during the scan to form a volumetric representation. Figure 5 Involving generating a deformation field so that a pre-operative volumetric dataset is rendered in a manner that at least partially reflects the effect of pressure from a transducer on an organ of interest. A model is used for deformation, such as for compounding.

[0030] Figure 1 is a flow chart of one implementation of a method for surgical guidance using composite ultrasound imaging by an ultrasound system. A neural field is composited in response to input of ultrasound data and probe or field of view tracking information. The neural field is machine trained by joint optimization of parameters of the neural field and probe or field of view position, the neural field accurately refines position information and fills in missing information. The neural field representing the 3D scene is trained during image acquisition. In the context of ultrasound composite, the network is initialized at the beginning of the first ultrasound probe sweep. The network can be pre-trained with surgery-specific and / or patient-specific data to accelerate training.

[0031] Figure 1 The method is performed by a medical system such as Figure 8 A medical scanner, ultrasound system or medical system or another medical system. For example, an ultrasound scanner uses a probe (transducer) to acquire ultrasound data. A sensor and / or image processor tracks the probe or field of view. An image processor compounds (e.g., aligns and assembles) ultrasound data from scans at different times. An image processor and / or a graphics processing unit renders an image based on the composite ultrasound data. A display displays the image. Other devices may perform Figure 1 Any action in the action.

[0032] The method is performed in the order shown (e.g., from top to bottom or numerical order), but other orders may be used. For example, the compounding in action 110 may occur interleaved with or between repetitions of actions 100 and 102. Actions 100 and 102 may be performed simultaneously.

[0033] Additional, different, or fewer actions may be provided. For example, actions 120 and 130 are not performed (such as in the case of composite generation of deformation fields for preoperative data without rendering and displaying ultrasound data). As another example, actions for configuring scanning, scanning by another modality (e.g., CT or magnetic resonance (MR)), and / or use of output (e.g., measurements) are performed.

[0034] In act 100, an ultrasound system scans tissue of a patient. The patient is scanned. The scanning produces a plurality of 2D representations, such as images formatted in a display domain (e.g., scan conversion) or detected ultrasound data formatted in a scan domain rather than a display domain. In other embodiments, the 2D representations are retrieved from a memory or transferred over a computer network.

[0035] A 2D representation represents any part of a patient, such as representing an organ, head or torso of a patient. For example, the 2D representation is from a scan of a patient's liver. Figure 2 An example is illustrated. During surgery, a probe 202 is positioned against an organ (eg, liver). A 2D scan is performed, resulting in a 2D image 210. Image 220 shows the image as slabs at different orientations.

[0036] Different 2D representations represent the field of view at different locations within the organ of interest. For example, the transducer probe is translated, panned, and / or rotated while acquiring a series of 2D representations. The volume of the organ is scanned, thereby generating multiple 2D representations. Figure 2 In the example of , the probe 202 is translated along the organ without rotation or almost without rotation. At different positions, images 230 are acquired. A collection of such images from different probe positions and / or fields of view within the organ are aligned and assembled (ie, composited) to form a volume representation 240.

[0037] Figure 2 The example shows intraoperative scanning or scanning during open surgery. In alternative embodiments, the probe 202 is placed on the patient's skin to scan internal organs, or the probe is scanned from within the body (e.g., transesophageal or intracardiac echocardiography). In yet other alternatives, a matrix array is used to scan the patient's volume directly with little or no expected movement of the probe.

[0038] Scanning occurs when the transducer probe is moved relative to the patient to scan the volume. During scanning, the probe 202 is pressed against the tissue and may move along the tissue. This causes the tissue and adjacent tissues that are subjected to pressure to distort or deform.

[0039] The 2D representation produced by the scan is the ultrasound data or information derived therefrom. For example, different sample positions from the scan or pixel positions in the imaging represent ultrasound intensity (e.g., B-mode information). Doppler (e.g., velocity, variance, and / or power) from the ultrasound echo may be provided alternatively or also.

[0040] In another implementation, ultrasound data is used for segmentation. The 2D representation is segmented to identify locations corresponding to anatomical structures or objects of interest. Locations or data are marked, such as to identify blood vessels and / or tumors in the liver. Segmentation identifies boundaries and / or regions corresponding to the objects. Pixels or scan locations corresponding to the objects are identified as segmentations.

[0041] Segmentation is performed using any process or function, such as an intensity threshold using a low-pass filter. In one implementation, segmentation is performed using full width, half maximum (FWHM), or intensity thresholds with various standard deviations. Alternatively, the user (e.g., a radiologist) or an image processor manually segments. In another approach, a machine learning model (such as an encoder-decoder based neural network) outputs one or more segmentations in response to an input of a 2D representation. Training can be applied to output a segmented image to an image, U-Net, or encoder-decoder network in response to an input of a spatial 2D representation. The machine training model (segmenter) generates feature values ​​in a hidden layer in response to an input of a medical image, and uses the feature values ​​to output a segmentation.

[0042] In act 102, the position of the probe and / or field of view during the scan is tracked. The positions of the 2D representations relative to each other and / or another reference are determined. The probe or field of view is tracked in 3D space.

[0043] Any now known or later developed tracking may be used. For example, cameras, probe detection, electromagnetic sensing, and / or data processing (e.g., correlation of 2D representations) track position. In one implementation, fiducial markers on the probe are detected from a laparoscopic video feed (e.g., probe 202 is tracked from camera image 200, such as Figure 2 ). Machine learning can be used for detection and tracking. Stereo and / or depth cameras can improve tracking performance. In another implementation, the probe is detected from intraoperative imaging by other (non-ultrasound) modalities (such as MR or fluoroscopy). In yet another implementation, the probe pose is detected via optical, electromagnetic, inertial, and / or other types of tracking. As another implementation, ultrasound data from different 2D representations are correlated to determine displacement, thereby tracking the field of view.

[0044] The image processor composites the 2D representations into a 3D representation in act 110. The 2D representations are aligned and assembled (eg, stacked) to form the 3D representation. Figure 3 An example of a 2D representation as aligned according to tracking data is shown. A plurality of slices or 2D representations 310 are stacked to form a 3D representation 300 .

[0045] exist Figure 3In the example of , the 3D representation may have gaps or irregularities due to changes in velocity and / or pressure of the probe translation. The irregularities may be caused by or made worse by inaccuracies in the probe tracking.

[0046] In order to refine the tracking or 2D representation position and / or fill any gaps to provide a 3D representation on a regular grid, the image processor applies a neural field for compounding. The tracking position from action 102 and the 2D representation from action 100 are input into the neural field. The neural field is a neural network arranged in or having a coordinate-based architecture, such as a 128×128×128 voxel-based arrangement of a neural network. Different nodes (e.g., weights or activation functions) represent different positions in three dimensions, thereby providing a coordinate-based neural network for neural ultrasound compounding. The network architecture can use any position encoding or activation function that allows the network to fit high-frequency signals, such as a sine (e.g., SIREN), multi-resolution hashing, or another position encoding. The coordinate domain can be discretized using an irregular or hierarchical grid to handle the varying complexity of volume data due to sparse acquisition. Other neural networks or machine learning models can be used to refine the position and / or compound. For example, an encoder or transformer outputs a position in response to the input of a position and a 2D representation or multiple positions and 2D representations.

[0047] The neural field is trained to output a 3D representation in response to inputs of a 2D representation and a tracked position. The machine learning model may have been previously trained using training data of patient and / or procedure specific images. For a given patient, the neural field is further trained by optimizing using new inputs. The model (e.g., neural field) is formed by a schema (e.g., coordinate distribution) that defines learnable parameters. For any pre-training, the training data includes a sample of many inputs formed from a database of patient scans and the corresponding ground truth (e.g., correct output).

[0048] Meta-learning can be used to provide domain- and application-specific initialization of the network (e.g., different pre-trained 2D representations can be used depending on the target organ, disease, and / or surgery from the same patient or multiple patients). Different neural fields are provided for each of the different domains. Alternatively, neural fields are trained to composite multiple or all domains.

[0049] Any training data used for pre-training is collected from medical records (images from the same patient or a group of patients undergoing the same or similar surgery (based on similarity of metrics such as organ field of view or pathology)). Alternatively and / or additionally, simulations can be used to create the training data. Simulations can use physical objects, such as phantoms, and / or be based on computer modeling, such as using physical models.

[0050] The neural field as initialized (e.g., from pre-training or initial imaging / tracking) is trained for compounding for a given patient. The training uses joint optimization of the pose of the training images and the parameters (i.e., learnable parameters) of the neural field. A 2D representation (e.g., image) of the 2D ultrasound probe from tracking and its pose in 3D space is used as training input for the neural field, which maps 3D positions to ultrasound intensities (e.g., Figure 3 ). In the optimization used for training (e.g., Adam), the pose of the ultrasound image is jointly optimized together with the parameters of the neural field. Ultrasound probe tracking when applying the trained neural field is used only for pose initialization, where the neural field compounds and refines the pose. This can reduce the impact of small tracking errors and tissue deformations during scanning. Alternatively, the image pose can be trained without initialization to accommodate systems without probe tracking capabilities.

[0051] The neural field is trained using renderings. The loss used in the optimization is between one or more renderings from a 3D representation output by the neural field and one or more of the input 2D images. The rendering is from a 3D representation to a 2D representation, such as extracting one or more slices or volume rendering a view. The loss is based on or a comparison of the output one or more renderings to one of the input training images. The larger the gap, the larger the loss. The loss is used to change the value of one or more of the learnable parameters and the estimated pose in the joint optimization. The optimization continues until the loss is minimized for each sample set of the training data.

[0052] Any rendering may be used, such as multi-plane image generation, volume rendering using ray casting or ray tracing, surface rendering, or other rendering. The rendering generates views that are associable or comparable to one or more of the 2D representations that are input. In one implementation, differentiable rendering is used for 3D view optimization. Different renderings may be generated using different settings, each of which is compared to the input 2D representation. Alternatively, differentiable rendering is performed to maximize or minimize (e.g., optimize) a property of the rendered view (e.g., visual entropy or occlusion). The optimized views are then compared to calculate the loss for machine training.

[0053] In order to train the machine learning model, a machine learning model arrangement (architecture) is defined. Any machine learning model known now or developed later can be used. For example, an image-to-image network, U-Net, DenseNet, ResNet, or encoder-decoder network is used. Downsampling layers, convolutional layers, pooling layers, discard layers, skip connections, upsampling layers, and / or other neural network layers can be used. In one implementation, a coordinate-based neural field is used. Any architecture that receives image (spatial) information to output spatial (e.g., image) information can be used. The definition is implemented by configuration or programming of learning. The number of layers or units, the type of learning, and other characteristics of the model are controlled by the programmer or user. In other embodiments, one or more aspects (e.g., the number of nodes, the number of layers or units, or the type of learning) are defined and selected by the machine during learning. Training data (including many samples of input data and corresponding outputs) is used for training. The relationship between input and output is machine-learned.

[0054] The image processor or another processor machine trains the model (e.g., neural field). The training learns the values ​​of the learnable parameters (e.g., weights, connections, filter kernels, values ​​in activation functions, and / or other learnable parameters that define the architecture). Deep or another machine learning may be used. The weights, connections, filter kernels, and / or other parameters are the features being learned. Using the training data, the values ​​of the learnable parameters of the model are adjusted and tested to determine the values ​​that result in the best estimate of the output given the input. Training is performed using Adam or another optimization.

[0055] During training, the loss is minimized. The loss function is an L1, L2, or other error function between the network output or renderings from the network output and the ground truth of the training data or input data. Other losses can be used. Using optimization, different values ​​of the learnable parameters are tested to minimize the loss (or maximize the reward).

[0056] The input to the network is a 3D position, and the output is a vector of parameters used by the volume renderer (attenuation, reflection, scattering at the 3D position). Intensity or other metrics can be output, which can be converted into parameters used by the volume rendering. The vector provides information about the position, so different vectors are provided for different positions. The mapping from input to output uses a multi-layer perceptron network architecture (MLP, fully connected neurons, non-linear activations such as ReLU). Based on the input image and its pose, the 3D position and final intensity of each pixel are known. The neural field is then trained to minimize the difference between the intensity in the training image and the intensity produced by a differentiable volume rendering with fields such as attenuation, reflection, scattering, etc. A photometric loss function is used, but multiple metrics can be used to compare the training data images with images synthesized from the neural field, such as the structural similarity index (SSIM). Once the network (neural field) is trained, the neural field can be used in the algorithm of the visualization system instead of the voxel data (for example, for rendering).

[0057] Neural field compounding is used for this scan or surgery of the patient. Neural fields can be compounded rapidly, allowing the generation of 3D representations in real time.

[0058] In one implementation, a neural field or other model is trained to provide a 3D representation as voxels representing ultrasound intensity or velocity distributed in three dimensions. Input is 2D ultrasound with high in-plane resolution to output a 3D representation with high 3D resolution for B-mode or color mode ultrasound imaging.

[0059] In another implementation, a neural field or other model is trained to provide a 3D representation as a 3D segmentation of one or more objects of interest. In training, the input may be a 2D representation or a segmentation from a 2D representation and tracking, and the output is a 3D representation as a 3D segmentation.

[0060] For example, a 2D representation is segmented into a signed distance field. Each pixel or sample position is assigned a value representing the distance to the nearest surface or position of an object of interest. The 3D representation is a signed distance field or function of the same object but in three dimensions. Compounding is compounding into a 3D isolation mask, such as for rendering. A neural signed distance field can be trained directly from object segmentation on a 2D ultrasound image, as an alternative to performing segmentation on a voxel grid calculated from inference of a fully trained ultrasound neural field. For example, detected vessel contours are detected in a 2D ultrasound image, with the 3D image pose refined from a joint optimization, resulting in a 3D point cloud that can then be used to infer a 3D signed distance function. Other segmentation representations besides signed distance fields or functions may be used.

[0061] In another implementation, a neural field, neural network, machine learning model, or another model is trained and used to output a deformation field. The scan may deform the tissue due to probe pressure. With the input tracking (e.g., pressure) and the 2D representation, the 3D representation can be a field of values ​​representing the magnitude and / or direction (e.g., vector) of the deformation of each voxel. The deformation field is based on real-time scanning with ultrasound scanning or imaging (e.g., less than 0.1, 0.5, 1, or 3 seconds delay between scanning and generating the deformation field). The neural field or another model is trained or programmed using differentiable deformable volume rendering to calculate losses. The model is trained to output a deformation field.

[0062] Ultrasound compounding can utilize a 3D matrix probe. Since volume scanning of a matrix probe (e.g., a 2D array for 3D scanning) takes time, the sample position may shift due to movement of the patient and / or the sonographer. Compounding through neural fields or other models can refine the position information to eliminate distortions caused by motion. In other implementations, x-ray, CT, MR, or other non-ultrasound imaging modalities are used instead of ultrasound. A differentiable direct volume renderer is used to train voxel representations from x-ray or other modality projection images.

[0063] In action 120, an image processor and / or a graphics processing unit renders an image based on the 3D representation output by the model. Surface rendering, ray casting, ray tracing, or another type of rendering from the volume data set to the 2D image domain may be used. One or more renderings based on the 3D representation may be used for augmented reality, virtual reality, and / or display on a screen. The 3D representation is rendered as pixels or positions in two dimensions.

[0064] Rendering generates an image. The image can be updated as the scan is updated. For real-time viewing, the image is updated to reflect the current state currently represented by the ultrasound scan. For example, a B-mode image representing a volume is rendered based on the ultrasound intensity of the 3D representation.

[0065] In one implementation, rendering uses voxel classification to depict a 3D representation of an object (segmentation) through a transfer function. For example, rendering blood vessels and / or tumors in the liver. The 3D representation is of voxels that are labeled as belonging to blood vessels and / or tumors. The rendered image shows the object without showing other objects or reducing the emphasis on other objects. Figure 2 An example rendering of a compound ultrasound with a transfer function emphasizing liver vessels 220 at the beginning of the ultrasound sweep and liver vessels 240 at the end of the ultrasound sweep is shown. The 3D rendering is aligned with the direction of the ultrasound probe sweep. Figure 4 Another example rendering 400 of views from different viewing directions relative to the volume of blood vessels in the liver is shown.

[0066] In one implementation, the image processor or graphics processing unit samples the neural fields directly for rendering. The composite neural field representation of the ultrasound data is in a regular grid or can be converted to a regular grid by inferring the voxel positions and visualized with conventional volume rendering techniques. In this implementation, the neural fields are sampled directly, leveraging a compact and data-adaptive neural network representation for memory-efficient processing. The same applies to neural signed distance fields trained directly from structures in the input image - conventional surface extraction can be performed on a voxel grid reconstruction of the neural field SDF, or the neural field SDF can be sampled directly in a ray casting or sphere tracing algorithm.

[0067] In another implementation, rendering uses differentiable rendering. Rendering is repeated in optimization to maximize or minimize the characteristics of the rendered image. One or more variables used for rendering are changed to produce a desired view from the 3D representation. The visualization system can use a differentiable renderer to calculate an optimized view of the segmented structure, for example, when displaying enhanced guided visual effects on a separate display in an operating room. The differentiable renderer can operate on the surface of the segmented object, directly on the volume data, on selected objects in the scene (such as the liver vascular tree), and / or on the entire 3D scene. Typically, the renderer calculates the image space derivatives of various rendering parameters (such as camera orientation), and then uses gradient descent or other optimization techniques to calculate the changes in (one or more) rendering parameters that maximize the objective function (such as visual entropy (i.e., the amount of visual information in the image)). Different application-specific image metrics can be used, such as penalizing the occlusion of important structures or the overlap of vascular structures. Image quality metrics based on machine learning can be used, for example, trained by clinical experts on the ranking of views. The optimized view can then be used in the rendering, such as used by a photorealistic renderer (e.g., volume Monte Carlo path tracing). View optimization can be initialized using a laparoscope or other camera view.

[0068] Rendering can enhance laparoscopic camera views. View synthesis from neural fields can produce images for augmented reality headsets or multi-view autostereoscopic displays.

[0069] In other methods, the 3D representation includes or is a deformation field. The deformation field is applied to a volume dataset of another modality (e.g., preoperative CT or MR). The rendering is performed from the other volume dataset as deformed. Figure 5 This method is involved.

[0070] Combinations of implementations may be provided. For example, the same or different neural fields are trained to output multiple types of information, such as a 3D representation of ultrasound intensity, a 3D representation of segmentation, and / or a 3D deformation field. One or more images are rendered using any combination of the output information.

[0071] In action 130, one or more images are displayed on a display device. The image is presented to a viewer on a screen or printout. Different images may be presented, such as a multiplanar reconstruction showing three orthogonal views and a volume rendering in four quadrants. A segmented image may be displayed. The image may be a deformation field to inform the physician of changes due to the scan. The image may be or include a rendering or imaging from a preoperative scan or planning. Different renderings, renderings from different views, 2D images with or without 3D images, 3D images with or without 2D images, images from different types of renderings, stereoscopic views, augmented reality views, and / or other imaging may be provided.

[0072] Figure 5 A flow chart of one implementation of a method for surgical guidance using composite ultrasound imaging by a medical system is shown. A deformed preoperative dataset is rendered based on ultrasound captured in real time during the surgical procedure, which allows for use during surgery. The real-time deformation is derived from the modeling of the ultrasound acquisition, and then the deformation is applied to the preoperative dataset. For example, a 128x128x128 volume composite (see Figure 1 ) provides deformations in 3D which are then applied to the preoperative dataset for rendering or as part of rendering. Ultrasound composite volume updates can be provided at 10 Hz during ultrasound acquisition, with asynchronous visualization system updates provided at 60 Hz (e.g., during user interaction).

[0073] Figure 5 The method is performed by a medical system such as Figure 8 A medical scanner, ultrasound system or medical system or another medical system. For example, an ultrasound scanner uses a probe (transducer) to acquire ultrasound data. A sensor and / or image processor tracks the probe or field of view. The image processor uses the model to generate a 3D deformation field. The image processor and / or graphics processing unit uses the 3D deformation field to render an image in real time using the scan. A display displays the image. Other devices may perform Figure 1 Any action in the action.

[0074] The method is performed in the order shown (e.g., from top to bottom or numerical order), but other orders may be used. For example, the generation in act 520 may occur interleaved with or between repetitions of acts 500 and 510. Acts 500 and 510 may be performed simultaneously.

[0075] Additional, different or fewer actions may be provided. For example, actions 530 and / or 540 are not performed (such as in the case where the deformation field is used for measurement). As another example, actions for configuring another use of scanning, imaging by ultrasound and / or output are performed.

[0076] In act 500, the ultrasound system scans tissue of the patient. Figure 1 The same or a different scan as used in act 100. The scan produces a 2D representation.

[0077] In act 510, the pressure from the probe to the tissue is tracked. The pressure may be assumed to be one-dimensional, such as directed along a normal from the surface of the transducer or probe. Alternatively, the pressure is tracked in two or three dimensions.

[0078] In one implementation, the position of the probe or field of view is used as the pressure. The difference in position from placement against the tissue to when the tissue deforms due to compression after placement represents the pressure. In other implementations, the actual pressure is measured, such as using a strain gauge sensor. Any probe tracking in which the measurement value is related to pressure can be used, such as tracking position or pressure sensing.

[0079] The image processor uses the model to generate a 3D deformation field in act 520. The model uses the pressure and the 2D representation to generate the 3D deformation field.

[0080] Any model may be used. For example, a physical or biomechanical model is used to determine the deformation. Using fitting or optimization, the 2D representation and pressure are used to determine the deformation of the tissue at different sample locations (eg, voxels).

[0081] In one implementation, the deformation in 2D is determined. An image processing or machine learning model determines the deformation in 2D. The 2D deformation is then determined in 3D by combining the ultrasound-based 2D representation. In another implementation, the image processor generates an approximate 3D deformation field from the simplified 2D deformation by compounding the vector field in parallel with compounding of the ultrasound imaging. Both the 3D vector field and the 3D ultrasound representation are registered to the preoperative plan using computational methods, such as machine learning-based methods and physical tracking methods.

[0082] In another implementation, a neural network (such as a neural field) generates a three-dimensional deformation field. For example and as described above for Figure 1 As discussed, a neural field is trained to generate a 3D deformation field from a 2D representation and pressure or tracking input. The same or different neural fields can be compounded to form a 3D representation for ultrasound intensity and / or segmentation.

[0083] Figure 6 An example is shown. Image 600 shows a volume in a regular grid reflected by parallel lines. Image 610 shows a volume and grid rendered with a deformation field applied. Lines are displaced corresponding to nodes or intersections being displaced. Each node or voxel has a displacement representing the deformation.

[0084] The 3D deformation field is generated in real time with the scan (e.g., less than 0.1, 0.5, 1, or 3 seconds from the completion of the image scan of the 2D representation to the generation of the 3D deformation field). For real-time imaging, the 2D representation (e.g., 2D ultrasound image), the 3D representation (e.g., ultrasound volume data set of intensity), and the 3D deformation field are streamed to the renderer or graphics processing unit via a streaming interface. Standard interfaces (such as OpenIGTLink) or custom optimized interfaces for low-latency applications can be used. The interface can further implement compression, such as video codecs, differentiable compression schemes, wavelet-based compression, or machine learning-based methods such as neural volumes. The interface is configured to synchronize data for rendering for real-time visualization. Computing components (e.g., image processors and graphics processing units) can be coupled or linked to obtain higher performance at the expense of deployment flexibility. For example, a graphics processing unit-based ultrasound composite implementation can use back projection directly into the GPU memory of the rendering service.

[0085] In act 530, an image processor and / or graphics processing unit renders a preoperative image (an image based on a preoperative data set) in real time using the 3D deformation field with the scan of act 500. The primary target is surgical ultrasound and / or preoperative CT for surgical guidance or other data deformed to match ultrasound data or current tissue state. The one or more images rendered from the preoperative data reflect the current deformation as detected from the ultrasound data. Alternatively, the ultrasound data (e.g., 3D representation) is deformed to a state corresponding to the preoperative scan (i.e., to offset transducer pressure) by applying the inverse of the calculated 3D deformation field to fuse with the planning (preoperative) data.

[0086] In order to deform the preoperative data set as part of the rendering, the 3D deformation is used as an index texture in the rendering pipeline of the graphics processing unit. The 3D deformation volume is used as an index texture to displace the voxels of the target volume (e.g., CT data set or preoperative data) when sampling. The deformation volume is updated in real time based on the streaming data from the 3D composite, so the sampling is updated. The 3D deformation field is a separate volume in the rendering pipeline used for indexing to render based on the preoperative data set. The rendered image obtained from the preoperative data set includes the deformation.

[0087] In another implementation, the graphics processing unit calculates data sampling locations based on splines generated from the 3D deformation field. A parametric description of the deformation field, such as a spline or other curve, is calculated. Parameter representations of various orders (e.g., splines) are evaluated during rendering to calculate data sampling locations. As an optimization, a 3D deformation volume can be calculated from the parameter representation in a pre-processing step, such as in a compute shader of the graphics processing unit or pipeline. The deformation volume is then used during rendering to select sampling locations in the pre-operative data set to render an image from the pre-operative data set.

[0088] As another implementation, a graphics processing unit deforms a preoperative volume representation in a compute shader based on a deformation field. A preoperative volume representation is rendered from the deformed preoperative volume representation. In a preprocessing (pre-rendering) action, the preoperative volume is deformed in a compute shader. The deformed volume is then sampled during rendering without any additional computation and can be streamed to pipeline components that do not support deformed volumes, such as image processing or segmentation. This compute shader execution occurs each time deformation data is streamed in from the composite.

[0089] Other rendering-based applications of 3D deformation can be used to render the preoperative data set. Alternatively, the image processor deforms the preoperative data set (e.g., volume) before providing it to the graphics processing unit. The 3D deformation field is applied to the preoperative data set. The graphics processing unit then renders based on the deformed data set.

[0090] In action 540, the display displays the rendered pre-operative image. The displayed image includes at least a portion reflecting the deformation caused by the ultrasound probe. By rendering in real time with the ultrasound scan, the rendered pre-operative image at least partially reflects the changes or variations in the position of the tissue caused by the probe pressure.

[0091] The image may be an image rendered from the preoperative data set alone. Alternatively, the preoperative image is displayed together with the ultrasound image. Images from different modalities may be synchronized. Since modeled deformations are used together with rendering using deformations, real-time imaging of the preoperative data set with deformations to account for probe pressure is provided. The ultrasound image is generated in real-time, allowing the preoperative image to be synchronized with the ultrasound image. The current state of tissue deformation is reflected in both types of images to guide the surgeon in an easy comparison.

[0092] Various combinations of images can be displayed, including deformed and undeformed images, and images including preoperative and ultrasound. For example, surgical ultrasound and preoperative planning data are both visualized together for surgical guidance. In one implementation, one volume image is overlaid on top of another. The ultrasound image is rendered in a different classification in the same scene as the preoperative (e.g., CT) volume, or is composited with different opacities or coloring. This enables a quick comparison between the two at a glance.

[0093] In another implementation, the reference (ultrasound) volume is rendered as a contour on an image rendered from a fully colored deformed volume (deformed preoperative volume). The datasets can be visualized together with the preoperative (e.g., CT) image rendered in full colorization and the ultrasound image rendered as a contour to provide additional context without flooding the image with redundant and low detail information.

[0094] As another implementation, the magnitude of the displacement from the 3D deformation is used as a metric for coloring, giving a heat map of the amount of change at each voxel in the volume. Deformations can be applied to 3D volume, MPR, and / or segmentation mesh renderings. Changes in the mesh from the deformation can be animated to highlight changes in the structure in the region of interest. For example, liver vasculature may change over time, and the deformation data can be used to highlight these changes.

[0095] The image can be of a segmented object. The segmentation can be overlaid on an image of a region to highlight a specific structure or displayed alone.

[0096] Figure 7 An example display is shown. Image 700 is an ultrasound multi-planar reformatted (MPR) image of an ultrasound composite volume from a 3D volume rendering deeply embedded in the ultrasound composite volume. Image 710 contains a 3D volume rendering with an embedded MPR from a preoperative CT of the same patient as image 700, with the ultrasound field of view indicated by the embedded box indicator at the white arrow. The preoperative volume is deformed as part of the rendering to match the deformation of the tissue reflected in image 710. Image 720 is a rendering of a larger portion of the preoperative volume of the patient with or without the deformation. The field of view of ultrasound image 700 is visualized as the embedded box indicator at the arrow.

[0097] In another implementation, a multi-planar reconstruction is displayed. A deformation is applied to render the planar reconstruction from the pre-operative data. No deformation is applied to render from a 3D representation (eg, volume) of the pre-operative dataset.

[0098] Figure 8 An embodiment of a medical system for ultrasound compounding and / or deformation is shown. The medical system includes an image processor 800, a memory 810, a display 820, an ultrasound system 830, and / or a medical scanner 840. The display 820, the image processor 800, and / or the memory 810 may be part of a medical scanner 840 (e.g., a CT scanner), an ultrasound system 830, a computer, a server, a workstation, or another system for image processing of scanned medical images from a patient. A workstation or computer without a medical scanner 840 and / or an ultrasound system 830 may be used as a medical system.

[0099] Additional, different, or fewer components may be provided. For example, a computer network is included for remote image generation of locally captured image data. The machine learning model 802 is applied as a standalone application on a workstation or local device, or as a service deployed on a network (cloud) architecture. In another example, such as where a preoperative data set was previously acquired and stored in the memory 810, a medical scanner 840 is not provided.

[0100] The ultrasound system 830 is a diagnostic medical ultrasound scanner. The ultrasound system 830 includes a probe 202 for scanning a field of view 836. An optional sensor 834 is provided on the probe 202 or for sensing the probe 202. The ultrasound system 830 may include a transmit beamformer, a receive beamformer, a B-mode detector, a color flow or Doppler detector, filters, a scan converter, and / or other components.

[0101] The probe 202 is an ultrasound transducer having an array of transducer elements (such as a one-dimensional array). The probe 202 is sized, shaped, and made of material for hand-held or robotic use outside the patient, such as for pressing against the patient's skin for scanning. Alternatively, the probe 202 is sized, shaped, and made of material for use in an orifice (e.g., transesophageal), for insertion into a patient (e.g., an intraoperative probe or intracardiac catheter), or for intraoperative use. Other probes 202 may be used.

[0102] The ultrasound system 830 uses the probe 202 to scan the field of view 836. The scan provides a 2D representation. The scan is repeated to obtain multiple 2D representations. As the probe 202 shifts or moves, the field of view 836 shifts or moves to scan other planes in the patient's tissue.

[0103] The sensor 834 is mounted to the probe 202 and / or is mounted to sense the probe 202. The sensor 834 can be a camera, an electromagnetic sensor, an accelerometer, a strain gauge, a pressure sensor, and / or other sensors for detecting the probe posture. In another approach, the medical scanner 840 is a sensor that detects the probe 202 or a reference on the probe 202 when the probe 202 is in the patient or adjacent to the patient. In an alternative, the image processor 800 sequentially compares (e.g., associated with different shifts) 2D representations to determine changes in position or posture. The sensor 834 is an ultrasound probe tracker that is configured to track the ultrasound probe 202 during acquisition of ultrasound data (e.g., during a 2D scan of a plane).

[0104] Image processor 800 or another processor may operate with the sensor for tracking. Any image processing (eg, fiducial identification and position determination) is performed by the processor for tracking.

[0105] In one implementation, the image processor 800 or another processor performs segmentation or other processing on the ultrasound data.For example, the ultrasound data is converted to a 2D field of each signed distance and object in a 2D representation.

[0106] The medical scanner 840 is a diagnostic or therapeutic scanner (e.g., a CT, X-ray, or MR scanner). The scanner 840 operates according to one or more settings to scan a patient 842 lying on a bed or table 844. The settings control the scan, including transmission, reception, reconstruction, and image processing. A scan protocol is followed to generate data representing the patient 842. The patient 842 is imaged by the scanner 840 using the settings. The scanner 840 generates a preoperative or other planning data set, such as a volume representation of the patient 842 before being scanned by the ultrasound system 830.

[0107] Ultrasound data (e.g., a 2D representation such as an ultrasound image), segmentation (e.g., a 2D signed distance field), preoperative data, a composite 3D representation, a 3D deformation field, a machine learning model 802, values ​​of learning parameters, probe tracking data, rendered images, and / or other information are stored in a non-transitory computer-readable memory such as a memory 810. The memory 810 is configured to store a machine learning neural network (e.g., model 802) formed as a neural field. The neural field is trained using joint optimization of probe pose and parameters of the neural field. Other models 802 may be stored, such as for different domains (e.g., liver vs. kidney imaging).

[0108] The memory 810 is an external storage device, RAM, ROM, database and / or local memory (e.g., a solid state drive or hard disk drive). The same or different non-transitory computer readable media can be used for instructions and other data. The memory 810 can be implemented using a database management system (DBMS) and resides on a memory such as a hard disk, RAM or removable media. Alternatively, the memory 810 is located inside the processor 800 (e.g., a cache).

[0109] Instructions for implementing the training or application processes, methods, and / or techniques discussed herein are provided on a non-transitory computer-readable storage medium or memory, such as a cache, buffer, RAM, removable media, hard drive, or other computer-readable storage medium (e.g., memory 810). Non-transitory computer-readable storage media include various types of volatile and non-volatile storage media. The functions, actions, or tasks illustrated in the figures or described herein are performed in response to one or more instruction sets stored in or on a computer-readable storage medium. The function, action, or task is independent of a particular type of instruction set, storage medium, processor, or processing strategy, and may be performed by software, hardware, integrated circuits, firmware, microcode, and the like operating alone or in combination.

[0110] In one embodiment, instructions are stored on a removable media device for reading by a local or remote system. In other embodiments, instructions are stored in a remote location for delivery over a computer network. In yet other embodiments, instructions are stored in a given computer, CPU, GPU, or system memory. Because some of the constituent system components and method steps depicted in the accompanying drawings can be implemented in software, the actual connections between system components (or process steps) can be different depending on the way the present embodiment is programmed.

[0111] The image processor 800 is a controller, a control processor, a general processor, a microprocessor, a tensor processor, a digital signal processor, a three-dimensional data processor, a graphics processing unit, an application-specific integrated circuit, a field programmable gate array, an artificial intelligence processor, a digital circuit, an analog circuit, a combination thereof, or another device now known or later developed for processing image data. The image processor 800 is a single device, a plurality of devices, or a network of devices. For more than one device, a division of processing in parallel or sequentially can be used. Different devices that make up the image processor 800 can perform different functions. In one implementation, the image processor 800 is a control processor or other processor of a medical scanner 840 or an ultrasound system 830. The image processor 800 operates according to stored instructions, hardware, and / or firmware, and is configured by stored instructions, hardware, and / or firmware to perform various actions described herein.

[0112] The image processor 800 may include an interface 804. The interface 804 is configured to receive tracking data, ultrasound data, and / or preoperative data at different rates and / or force synchronization. The interface 804 may be a standard or custom interface for receiving data and / or transmitting (e.g., outputting) data (such as images) to a display 820 for visualization.

[0113] The image processor 800 includes shading calculation (shader 802 ), rasterization (rasterizer), and / or parallel processors forming a rendering pipeline. Alternatively, other hardware components or only a processor are provided.

[0114] The image processor 800 or another remote processor is configured to train a machine learning architecture, such as a model 802. For example, a neural field is trained using training data. As another example, a neural field is trained using joint optimization to compound from probe tracking and images of a patient. The machine learning model 802 is trained to generate a 3D representation, such as a compound or 3D deformation field.

[0115] Alternatively or additionally, the image processor 800 is configured to apply (one or more) machine learning models 802. The machine learning model 802 may include a machine learning network for 2D segmentation.

[0116] The image processor 800 is configured to use a neural field (model 802) to compound ultrasound data into a volume representation. The ultrasound data as compounded has refined position information from an ultrasound probe tracker or sensor 834. The compounding forms one or more 3D representations based on ultrasound data (2D or 3D) from the ultrasound system 830 and tracking from the sensor 834. The 3D representation is ultrasound intensity, segmented objects in 3D, and / or a 3D deformation field. For example, the model 802 for compounding ultrasound data also generates a 3D deformation field. As another example, the image processor 800 is configured to generate a 3D signed distance field or another segmentation of an object in response to inputting a 2D segmentation or 2D ultrasound data to the model 802. In another example, a different model (computational and / or machine learning model 802) generates a 3D deformation field based on an ultrasound scan.

[0117] The image processor 800 is configured to generate outputs, such as images showing segmentation. The image processor 800 uses the shader 802 or another (one or more) rendering component to render one or more images. The image can be rendered according to the 3D composite ultrasound data. Alternatively or additionally, the image is rendered according to the preoperative data. Rendering images according to the preoperative data can reflect the organization deformed as based on the rendering using the 3D deformation field. Real-time imaging can be provided. As another alternative or supplement, the image processor 800 is configured to generate a visualization of the segmentation (contour). Render the image according to the 3D segmentation.

[0118] The display 820 is a CRT, LCD, projector, plasma, printer, tablet computer, smart phone, or another display device now known or later developed for displaying output, such as one or more rendered images from the volume representation. A sequence of rendered images can be displayed during surgery to guide the surgeon. For example, preoperative images can be rendered in real time with ultrasound scans based on a sequence of deformation fields generated from the ultrasound scans (e.g., from the model 802 or another model 802 for compounding). As another example, ultrasound and / or preoperative dataset images are rendered as 3D segmentations using a signed distance field generated by the compound model 802.

[0119] Fig. 96 is a block diagram illustrating another example implementation of a system for compounding and / or rendering visualization of at least one image using 3D deformation using modeling (e.g., neural fields). An ultrasound B-mode 2D representation and probe pose 900 are provided to a processor for detection 910 at low latency and provided to a compounding server 920 at low latency. Detection 910 generates a segmented or labeled 2D representation 912. The labeled representation 912 can be further processed, such as filtering or separation (e.g., distinguishing different objects) 914. The compounding server 920 generates a 3D representation (e.g., an ultrasound composite or composite 3D representation 922, a labeled or segmented 3D representation 926, and / or a deformation field 3D representation 928) using a neural field, multiple neural fields, and / or other models. The server 920 can deliver or also generate a 2D ultrasound representation. Various representations are generated with the same or different delays. Visualization 630 is a graphics processing unit or other image generator for generating an image based on any one, combination, or all of the representations 922, 924, 926, and / or 928. Image generation may require different amounts of processing.

[0120] Fig.10 An embodiment of an artificial neural network 1000 according to one or more embodiments is shown. Alternative terms for "artificial neural network" are "neural network", "artificial neural network" or "neural network". The artificial neural network 1000 can be used in part, for example, in one or more machine learning-based networks for compounding or generating 3D deformations. The artificial neural network 1000 can be arranged or designed as a neural field for ultrasound 3D compounding and / or generating deformation fields.

[0121] The artificial neural network 1000 includes nodes 1020-1032 and edges 1040-1042, wherein each edge 1040-1042 is a directed connection from a first node 1020-1032 to a second node 1020-1032. Typically, the first node 1020-1032 and the second node 1020-1032 are different nodes 1020-1032, and it is also possible that the first node 1020-1032 and the second node 1020-1032 are the same. For example, in Fig.10 , edge 1040 is a directed connection from node 1020 to node 1023, and edge 1041 is a directed connection from node 1021 to node 1023. Edge 1040-1042 from a first node 1020-1032 to a second node 1020-1032 is also represented as an "ingoing edge" of the second node 1020-1032 and an "outgoing edge" of the first node 1020-1032.

[0122] In this embodiment, the nodes 1020-1032 of the artificial neural network 1000 may be arranged in layers 1050-1053, wherein the layers may include an inherent order introduced by the edges 1040-1042 between the nodes 1020-1032. In particular, the edges 1040-1042 may only exist between adjacent layers of nodes. Figure 5 In the embodiment shown in , there is an input layer 1050 including only nodes 1020-1022 without incoming edges, an output layer 1053 including only nodes 1031 and 1032 without outgoing edges, and hidden layers 1051, 1052 located between the input layer 1050 and the output layer 1053. Generally, the number of hidden layers 1051, 1052 can be selected arbitrarily. The number of nodes 1020-1022 in the input layer 1050 is generally related to the number of input values ​​of the neural network 1000, and the number of nodes 1031-1032 in the output layer 1053 is generally related to the number of output values ​​of the neural network 1000. For the neural field, the number of layers and nodes corresponds to the coordinate system. Each node can be connected to a neighboring node in any direction in a bidirectional edge. Input can be to all nodes, and output can be to all nodes.

[0123] In one approach, the network architecture is an MLP with a small number of fully connected layers, where the input to the network is the 3D position (mapped to a higher dimension using a position encoding), and the output of the network is a vector of parameters used by the volume renderer (attenuation, reflection, scattering at the 3D position).

[0124] A (real) number may be assigned as a value to each node 1020-1032 of the neural network 1000. Here, x (n) i represents the value of the i-th node 1020-1032 of the n-th layer 1050-1053. The values ​​of the nodes 1020-1022 of the input layer 1050 are equivalent to the input values ​​of the neural network 1000, and the values ​​of the nodes 1031-1032 of the output layer 1053 are equivalent to the output values ​​of the neural network 1000. In addition, each edge 1040-1042 may include a weight that is a real number, in particular, the weight is a real number in the interval [-1, 1] or in the interval [0, 1]. Here, w (m,n) i,j represents the weight of the edge between the i-th node 1020-1032 of the m-th layer 1050-1053 and the j-th node 1020-1032 of the n-th layer 1050-1053. In addition, for the weight w (n,n+1) i,j , define the abbreviation w (n) i,j .

[0125] Specifically, in order to calculate the output value of the neural network 1000, the input value is propagated through the neural network. Specifically, the values ​​of the nodes 1020-1032 of the (n+1)th layer 1050-1053 can be calculated based on the values ​​of the nodes 1020-1032 of the nth layer 1050-1053 by the following formula:

[0126]

[0127] Where function f is a transfer function (another term is "activation function"). Known transfer functions are step functions, sigmoid functions (e.g., logistic function, generalized logistic function, hyperbolic tangent, inverse tangent function, error function, smoothed step function), or rectified functions. Transfer functions are mainly used for normalization purposes.

[0128] In particular, values ​​are propagated through the neural network 1000 layer by layer or through neighboring nodes, where the value of the input layer 1050 is given by the input of the neural network 1000, where the value of the first hidden layer 1051 can be calculated based on the value of the input layer 1050 of the neural network 1000, where the value of the second hidden layer 1052 can be calculated based on the value of the first hidden layer 1051, and so on.

[0129] To set the edge value w (m,n) i,j , training data must be used to train the neural network 1000. In particular, the training data includes training input data and training output data (denoted as t i ). For the training step, the neural network 1000 is applied to the training input data to generate the calculated output data. In particular, the training data and the calculated output data include a plurality of values, the number of which is equal to the number of nodes of the output layer.

[0130] In particular, the comparison between the calculated output data and the training data is used to recursively adjust the weights within the neural network 1000 (back propagation algorithm). In particular, the weights are changed according to the following formula:

[0131]

[0132] where γ is the learning rate, and if the (n+1)th layer is not the output layer, then the number δ (n) j Can be based on δ (n+1) j is recursively calculated as:

[0133]

[0134] If the (n+1)th layer is the output layer 1053, then

[0135] where f' is the first derivative of the activation function, and y (n+1) j is the comparative training value of the j-th node of the output layer 1053.

[0136] The following is a list of non-limiting illustrative embodiments disclosed herein.Illustrative embodiments of one set or type (eg, method or system) may be provided within or combined with sets of other types of illustrative embodiments.

[0137] Illustrative embodiment 1: A method for surgical guidance using composite ultrasound imaging through an ultrasound system, the method comprising: scanning a patient's tissue through an ultrasound system, the scanning producing a two-dimensional (2D) representation; tracking the position of the 2D representation; composite the 2D representation by inputting the position and the 2D representation into a neural field, training the neural field using joint optimization of the posture and parameters of the neural field, and compositely providing a three-dimensional (3D) representation of the patient; rendering a first image based on the 3D representation; and displaying the first image.

[0138] Illustrative embodiment 2: The method of illustrative embodiment 1, wherein tracking includes tracking with a camera, probe detection, or electromagnetic sensing.

[0139] Illustrative Embodiment 3: The method of any of Illustrative Embodiments 1-2, wherein scanning comprises scanning while moving the transducer probe relative to the patient.

[0140] Illustrative Embodiment 4: The method of any of Illustrative Embodiments 1-3, wherein the 2D representation comprises an image of ultrasound intensity, wherein compounding comprises providing the 3D representation as voxels representing intensity distribution in three dimensions, and wherein rendering comprises rendering the first image as ultrasound intensity as a function of position in two dimensions.

[0141] Illustrative embodiment 5: The method of any of illustrative embodiments 1-4, wherein the 2D representation includes a 2D segmentation of the object, wherein the compounding includes providing a signed distance field representing a 3D segmentation of the object as a 3D representation through the neural field, and wherein rendering the first image includes rendering the first image based on the 3D segmentation of the object.

[0142] Illustrative Embodiment 6: The method of any of Illustrative Embodiments 1-5, wherein compounding comprises compounding by a neural field trained with a loss based on a comparison of a rendering of the output with one of the 2D representations.

[0143] Illustrative Embodiment 7: The method of any of Illustrative Embodiments 1-6, wherein compounding comprises compounding via a neural field, the neural field comprising a coordinate-based neural network.

[0144] Illustrative Embodiment 8: The method of Illustrative Embodiment 7, wherein the compounding comprises compounding wherein the coordinate-based neural network comprises sinusoidal, multi-resolution hashing, or another position encoding.

[0145] Illustrative Embodiment 9: The method of any of Illustrative Embodiments 1-8, wherein rendering comprises directly sampling the neural field.

[0146] Illustrative Embodiment 10: The method of any of Illustrative Embodiments 1-9, wherein rendering includes view optimization using differentiable rendering.

[0147] Illustrative Embodiment 11: The method of any of Illustrative Embodiments 1-10, wherein compounding comprises providing the 3D representation as a deformation field based on the scanning as real-time ultrasound, and wherein rendering comprises rendering the pre-operative image based on the deformation field.

[0148] Illustrative Embodiment 12: The method of Illustrative Embodiment 11, wherein compounding comprises compounding using a neural field trained using differentiable deformable volume rendering.

[0149] Illustrative embodiment 13: A medical system for ultrasound compounding, the medical system comprising: an ultrasound probe tracker configured to track an ultrasound probe during acquisition of ultrasound data; a memory configured to store a machine learning neural network formed as a neural field; an image processor configured to compound the ultrasound data into a volume representation using training of the neural field as a joint optimization of the probe pose and parameters of the neural field, such that the compounded ultrasound data has refined position information from the ultrasound probe tracker; and a display configured to display an image from the volume representation.

[0150] Illustrative Embodiment 14: The system of Illustrative Embodiment 13, wherein the volumetric representation comprises a deformation field, and wherein the image comprises a pre-operative image rendered using the deformation field.

[0151] Illustrative Embodiment 15: The system of any of Illustrative Embodiments 13-14, wherein the ultrasound data comprises a two-dimensional field of signed distances of the object, wherein the volume representation comprises the signed distance field of the object, and wherein the image comprises a three-dimensional segmentation rendered using the signed distance field.

[0152] Illustrative embodiment 16: A method for surgical guidance using composite ultrasound imaging through an ultrasound system, the method comprising: scanning a patient's tissue through an ultrasound system, the scanning producing a two-dimensional (2D) representation; tracking pressure corresponding to the 2D representation; generating a three-dimensional (3D) deformation field based on the pressure and the 2D representation through a model; using the 3D deformation field to render a preoperative image in real time with the scan; and displaying the preoperative image.

[0153] Illustrative Embodiment 17: The method of Illustrative Embodiment 16, wherein generating comprises generating by a model comprising a neural field.

[0154] Illustrative Embodiment 18: The method of any of Illustrative Embodiments 16-17, wherein rendering includes using the deformation field as an index texture.

[0155] Illustrative Embodiment 19: The method of any of Illustrative Embodiments 16-18, wherein rendering includes calculating data sampling positions according to a spline generated from the deformation field.

[0156] Illustrative Embodiment 20: The method of any of Illustrative Embodiments 16-19, wherein rendering comprises deforming the pre-operative volume representation based on a deformation field in a compute shader, and rendering according to the deformed pre-operative volume representation.

[0157] Illustrative Embodiment 21: The method of any of Illustrative Embodiments 16-20, wherein displaying includes displaying the pre-operative image synchronized with the ultrasound image from the scan.

[0158] The various improvements described herein may be used together or separately. Although illustrative embodiments of the present invention have been described herein with reference to the accompanying drawings, it is to be understood that the present invention is not limited to those precise embodiments and that various other changes and modifications may be made therein by those skilled in the art without departing from the scope or spirit of the present invention.

Claims

1. A method for surgical guidance using composite ultrasound imaging by an ultrasound system, the method comprising: The patient's tissue is scanned by an ultrasound system, and the scan produces a two-dimensional (2D) representation; Tracking the position of the 2D representation; compounding the 2D representation by inputting the position and the 2D representation into a neural field, training the neural field using joint optimization of the pose and parameters of the neural field, and compounding to provide a three-dimensional (3D) representation of the patient; rendering the first image according to the 3D representation; and The first image is displayed.

2. The method of claim 1, wherein tracking comprises tracking using a camera, probe detection, or electromagnetic sensing. The method of claim 1 , wherein scanning comprises scanning while moving the transducer probe relative to the patient.

4. The method of claim 1 , wherein the 2D representation comprises an image of ultrasound intensity, wherein compounding comprises providing the 3D representation as voxels representing intensity distribution in three dimensions, and wherein rendering comprises rendering the first image as ultrasound intensity as a function of position in two dimensions.

5. The method of claim 1 , wherein the 2D representation comprises a 2D segmentation of the object, wherein compounding comprises providing a signed distance field representing a 3D segmentation of the object as a 3D representation via the neural field, and wherein rendering the first image comprises rendering the first image based on the 3D segmentation of the object.

6. The method of claim 1, wherein compounding comprises compounding via a neural field trained with a loss based on a comparison of a rendering of the output with one of the 2D representations.

7. The method of claim 1, wherein compounding comprises compounding via a neural field, the neural field comprising a coordinate-based neural network.

8. The method of claim 7, wherein compounding comprises compounding wherein the coordinate-based neural network comprises sinusoidal or multi-resolution hashed position encoding.

9. The method of claim 1, wherein rendering comprises directly sampling the neural field.

10. The method of claim 1, wherein rendering comprises view optimization using differentiable rendering.

11. The method of claim 1, wherein compounding comprises providing a 3D representation as a deformation field based on the scanning as real-time ultrasound, and wherein rendering comprises rendering the pre-operative image based on the deformation field.

12. The method of claim 11, wherein compounding comprises compounding using a neural field trained using differentiable deformable volume rendering.

13. A medical system for ultrasound compounding, the medical system comprising: an ultrasound probe tracker configured to track the ultrasound probe during acquisition of ultrasound data; a memory configured to store a machine learning neural network formed as a neural field; an image processor configured to compound the ultrasound data into a volume representation using the training of the neural field as a joint optimization of the probe pose and the parameters of the neural field, such that the compounded ultrasound data has refined position information from an ultrasound probe tracker; and A display is configured to display an image from the volumetric representation.

14. The system of claim 13, wherein the volume representation comprises a deformation field, and wherein the image comprises a pre-operative image rendered using the deformation field.

15. The system of claim 13, wherein the ultrasound data comprises a two-dimensional field of signed distances of objects, wherein the volume representation comprises the signed distance field of objects, and wherein the image comprises a three-dimensional segmentation rendered using the signed distance field.

16. A method for surgical guidance using composite ultrasound imaging via an ultrasound system, the method comprising: The patient's tissue is scanned by an ultrasound system, and the scan produces a two-dimensional (2D) representation; Tracking the pressure corresponding to the 2D representation; Generate a three-dimensional (3D) deformation field from the pressure and 2D representation through the model; Using the 3D deformation field to render pre-operative images in real time with the scan; and Preoperative images are shown.

17. The method of claim 16, wherein generating comprises generating by a model comprising a neural field. The method of claim 16 , wherein rendering comprises using the deformation field as an indexed texture.

19. The method of claim 16, wherein rendering comprises computing data sampling locations from a spline generated from the deformation field.

20. The method of claim 16, wherein rendering comprises deforming the pre-operative volume representation based on a deformation field in a compute shader and rendering based on the deformed pre-operative volume representation.

21. The method of claim 16, wherein displaying comprises displaying the pre-operative image synchronized with the ultrasound image from the scan.