Ultrasound neuromodulation guided by artificial intelligence
Through the DSLVM trained on the target anatomical structure database, combined with an ultrasound probe and a neuromodulation beamformer, automatic targeting and neuromodulation control of medical ultrasound systems are realized, solving the problem of dependence on human operators in the prior art, and improving the accuracy and efficiency of the system.
Patent Information
- Application Number
- CN202510173435.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-25
- Filing Date
- 2025-02-17
- Publication Date
- 2025-08-15
AI Technical Summary
In existing medical ultrasound systems, targeted and neuromodulatory control mainly relies on human operators, and there is a lack of automated and efficient AI-assisted methods.
The domain-specific large vision model (DSLVM) trained on the target anatomical database is adopted, combined with an ultrasound probe and a neuromodulation beamformer, to achieve automatic targeting and neuromodulation control of real-time ultrasound data.
Efficient and automated targeting and neuromodulation control on the target anatomy are achieved, reducing dependence on human operators, and improving the accuracy and efficiency of the system.
Smart Images

Figure CN120477815A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63,554,004, filed on February 15, 2024, entitled “ULTRASOUND NEUROMODULATION GUIDED BY ARTIFICIAL INTELLIGENCE,” which is incorporated herein by reference in its entirety. Technical Field
[0003] The present invention relates to a medical ultrasound system that employs a Domain Specific Large Scale Vision Model (DSLVM) trained on a database of target anatomical structures, and in particular to a system that is adaptable to use volumetric or slice imaging. Background Art
[0004] In medical ultrasound systems, the key functions of targeting and neuromodulation control have previously been performed by human operators with the assistance of computer vision algorithms. In the present invention, new methods for performing these functions are proposed. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1 The training of a device-specific large-scale vision model (DSLVM) using data from an MRI database for target anatomy is demonstrated.
[0006] Figure 2 Instantiates DSLVM operating in inference mode.
[0007] Figure 3 We demonstrate an ultrasound system centered around a large vision model that performs inference using a volume-scanning ultrasound probe.
[0008] Figure 4 The example includes a slice geometry of an azimuth detecting element having an elevation angle greater than an azimuth angle.
[0009] Figure 5 We illustrate an alternative ultrasound system centered around a large vision model that uses a slice probe for inference. DETAILED DESCRIPTION
[0010] Two medical ultrasound systems are described, each centered around an artificial intelligence (AI) system that performs key functions such as targeting and neuromodulation control. Large-scale vision models (LVMs), a form of AI, evolved from large-scale language models (LLMs) and are oriented toward images rather than text. Early examples of LVMs include VGGNet, GoogleNet, and ResNet. In this application, domain-specific large-scale vision models (DSLVMs) are targeted at medical ultrasound systems.
[0011] The two ultrasound systems differ in the type of probe and data acquisition system used. The first system to be described (see Figure 3 ) uses a 2D or matrix probe that can acquire volumetric data. Here, the AI's encoded knowledge of a large database of MRI scans correlates the real-time volumetric data being acquired by the probe with the estimated MRI. Figure 3 The system does not require a pre-operative MRI scan.
[0012] The second system (see Figure 5 ) uses a 1D or other probe that produces slice data by electronically scanning in one dimension. It uses AI to correlate the real-time slice data being acquired by the probe with preoperative structural MRI scans.
[0013] Three architectures have emerged for DSLVMs: attention-based, convolutional, and multilayer perceptron architectures. Prior to this, transformer architectures have emerged as the building blocks for DSLVMs. Examples include the Vision Transformer (ViT, Google Brain), the Swin Transformer (Microsoft), and the VideoMAE (Tencent). These DSLVMs are pre-trained using a wide range of image datasets, enabling them to capture image content and extract semantic information. Meta's Segment Anything Model (SAM) serves as a generalizable image segmentation model.
[0014] Second, convolutional neural networks (CNNs) have historically been recognized for their superior computational efficiency compared to multilayer perceptrons (MLPs). CNNs have outperformed traditional algorithms in nearly every computer vision task, including object detection, segmentation, demosaicing, super-resolution, and deblurring. CNN architectures typically consist of alternating convolutional and pooling layers, followed by several fully connected layers.
[0015] Third, Google Brain developed the MLP-mixer architecture, another new paradigm for computer vision that uses neither the Transformer attention mechanism nor convolution operations.
[0016] This paper discloses an ultrasound treatment system centered around a DSLVM, in which neuromodulation guidance is based on real-time ultrasound combined with anatomical knowledge encapsulated in the DSLVM. The target anatomical structure can be a structure of the subject's brain or any other anatomical structure. The DSLVM can be constructed from any of the described architectures (transformer, CNN, or MLP mixer) or other methods.
[0017] Figure 1A training system 10 for a DSLVM 14 is illustrated, which employs a volumetric anatomical database 11 to illustrate the extent of anatomical structures in a clinical region of interest. Predicted target positions 17 corresponding to the trained DSLVM can be obtained using transfer learning from similar models trained using MRI and computed tomography (CT), such as V-Net (Reference 4). Alternatively, it can be trained from scratch using MRI, CT, or other volumetric scans. In either case, a number of ultrasound training data sets are created from individual MRI scans, each with a different transducer type, such as a slice probe with a 1D element row or a volume probe with a 2D element matrix. Further ultrasound synthetic training data is generated from the MRI scans at different probe geometries and center frequencies. The simulation used for model training includes many probe positions on the scalp.
[0018] The DSLVM is trained to understand the relationship between ultrasound slice or volume data and a volumetric map of the target anatomical structure. The synthetic ultrasound data indicated in the box "Ultrasound Volume or Slice Data" 13d can be reliably obtained from MRI or other volumetric scanning modalities in several steps. For example, this can be accomplished by first creating a synthetic CT image using established deep learning methods (Reference 7), and then estimating the acoustic properties of the target anatomical structure from the synthetic CT 13a (Reference 8). Alternatively, density maps and sound velocity maps 13a can be calculated from a database of CT scans. Other routes from volumetric medical imaging modalities to acoustic property datasets are also possible. Using this data, models such as pseudospectral simulation 13c can reliably model ultrasound propagation and scattering and create simulated ultrasound returns to the probe 13b to train the model. The transducer type and position obtained from the ultrasound probe's specifications 13b are also input to the pseudospectral simulation 13c. Other types of ultrasound simulation algorithms, such as finite element or finite difference models, can also be used. Labels or visual cues 15 may be applied to the volumetric anatomical database 11, but this may not be necessary as the DSLVM is able to identify relevant structures without labels.The output of the ultrasound simulation 13c is fed to the DSLVM in slice or volume form 13d.
[0019] Figure 2The system 20 is illustrated, including a DSLVM 21 operating in inference mode, as in production. An ultrasound probe 22 generates real-time ultrasound data 23 in various modes, which may include tissue harmonic imaging (THI), power Doppler (PD, vascular), strain (tissue stiffness), or a dual-transducer mode that transmits using one aperture and receives using a second aperture. The real-time ultrasound data 23 can be enhanced with optical or LIDAR data related to the subject's head geometry from a camera that sweeps around a target anatomical structure (not shown). The operator or clinician specifies the name of the structure to be treated 24, and the DSLVM outputs the target coordinates 25 to a real-time ultrasound controller 26, which drives a neuromodulation beamformer 27 after calculation of the beamforming coefficients. Calculating beamforming coefficients that match the ultrasound field to the coordinates used to bound the target structure(s) is a well-studied area in therapeutic ultrasound (Ref. 9).
[0020] DSLVM 21 also creates validation outputs 28. These outputs show operator aspects of the treatment, such as real-time anatomical illustrations highlighting the treated volume(s). DSLVM is trained to produce MRI-type anatomical estimates in the coordinate system of the ultrasound probe from ultrasound data acquired at any position or angle.
[0021] Using References Figure 1 and Figure 2 Background is described, and further details are provided in each of the two embodiments of the present disclosure: (i) a volume ultrasound approach and (ii) a slice ultrasound approach. Figure 3 The volumetric ultrasound method to be described is superior; it employs a probe comprising a 2D matrix of transducer elements and an application-specific integrated circuit (ASIC) to process the high information rate from the matrix array (Ref. 10). Figure 4 and Figure 5 The slice ultrasound method to be described produces much less real-time anatomical data, which is a limitation. However, the hardware system can be manufactured using conventional, low-cost transducers, which in practice may be an advantage when manufacturing products. Figure 3 Compared with the implementation of , it has lower hardware technology risk, but is more dependent on the capabilities of DSLVM.
[0022] Figure 3The volume ultrasound embodiment 30 of the present disclosure is illustrated. It illustrates how a volume scanning matrix probe 31b and an EEG or MEG sensor 31a provide real-time input 31 to the DSLVM. The ultrasound volume data prompts a model 32, the output 33 of which can be used to program a delay and amplitude control beamformer to provide ultrasound neuromodulation to the target anatomical structure. An AI processor 32a within the domain-specific large-scale visual model 32 generates ultrasound-reachable target contour coordinate points 33 for creating beamformer coefficients 34, and slow-time analog control 35 derived from the EEG / MEG 31a for providing timing information to the neuromodulation hardware 36. In an embodiment, the AI processor 32a includes a network of GPU and CPU processors. The DSLVM 32 acts like the kernel of an operating system; its functionality is similar to the way a Linux kernel interacts with peripherals attached to its CPU. Here, devices similar to peripherals in a Linux system include:
[0023] • An electroencephalogram (EEG) 31a or magnetoencephalogram (MEG) sensor array that can provide information about the response to neuromodulation of the target anatomy.
[0024] • A matrix array of ultrasound transducers and beamforming electronics capable of volumetric acquisition 31b of anatomical and blood flow data.
[0025] Desired targeting data 38 specified by the clinician through either:
[0026] o text strings such as "posterior cingulate cortex and amygdala", and
[0027] o A set of points marked by the clinician on the pre-operative MRI scan (if available). These points define the treatment volume in the coordinate system of the MRI (which is different from the coordinate system of the neuromodulation transducer).
[0028] • An operator display 37 provides access to several data sources, thereby informing the clinician of the system's progress toward reliable targeting of neuromodulation. This display may include:
[0029] o Volumetric ultrasound data 37a being received from imaging / guidance hardware.
[0030] o A "meter" 37b used to indicate to the clinician the DSLVM's confidence in the usefulness of the current probe position and orientation for the neuromodulation task.
[0031] o Validation output 37c for providing the clinician with further data regarding the overall confidence of the DSLVM's interpretation of the target anatomy revealed by real-time ultrasound.
[0032] Neuromodulation hardware 36, which includes the probe, drive electronics that generate appropriate neuromodulation drive voltages to be applied to the various detection elements, and a beamformer that converts a description of the volume to be treated into parameters that control the amplitude and time delay of the voltages applied to the detection elements to deliver the treatment specified by the DSLVM.
[0033] • Slow time neuromodulation control 35, which is informed by EEG or MEG information supplied to the LVM. These data can define when neuromodulation is applied with respect to activity in the target anatomy.
[0034] The behavior of the neuromodulation hardware 36 is controlled in two ways. Steering, focusing, and other control over the spatial extent of the beam is accomplished through selection of beamforming coefficients 34. Temporal control of neuromodulation to synchronize with aspects of physiological activity is achieved through separate slow-time neuromodulation controls 35 output from the DSLVM 32.
[0035] In the embodiments of the present disclosure, Figure 3 A use case of the volumetric ultrasound method includes the following steps, where the target anatomical structure is the cranial region of the subject:
[0036] Designated disposal area: This can be:
[0037] o Part of the prescription by the prescribing physician,
[0038] o Health apps that can be accessed without a prescription.
[0039] Place the probe on the subject's skull.
[0040] DSLVM converts volumetric ultrasound data into anatomical structures and marks treatment areas.
[0041] • If the target anatomy probe proves to be in a position that is not conducive to treatment, the operator display (in a clinical setting) or the subject's phone (during home use) can show how to reposition the probe.
[0042] When the probe is positioned appropriately for treatment, treatment begins.
[0043] In the embodiments of the present disclosure, Figure 4 A 1.75D probe 40 is illustrated including an array of detector elements 41 having elements with a height dimension 42 that is greater than an azimuth dimension 43. In an exemplary application, the array 40 is used to steer and focus ultrasound waves in a transcranial ultrasound (TUS) procedure. Four azimuth detector elements 44 are shown as a column of the array 41. The 1.75D configuration conserves the total number of detector elements by only considering steering in the azimuth dimension. Figure 5The present invention provides a system that utilizes a reduced number of elements to achieve the efficacy of the accompanying tFUS procedure. The smaller element count allows standard cables to connect the probe array 41 to the signal processor. The probe can be constructed using conventional, low-cost manufacturing techniques using bulk PZTs. The element size is insufficient to support steering in elevation, but the element size is sufficient to correct for aberrations in elevation as well as in azimuth. The data created using the probe array 41 can be referred to as "slice data," which has been corrected for cranial aberrations.
[0044] In the embodiments of the present disclosure, Figure 5 An example slice ultrasound method 50 is illustrated, in which the target anatomical structure is the cranial region of the subject. A 1.75D probe 40 is used as a real-time input 51, connected to slice ultrasound hardware 51a, and then connected to a DSLVM 52. The DSLVM can also receive EEG data 51b (or magnetoencephalography data, i.e., MEG data) and probe orientation data (e.g., generated by a MEMs gyroscope 51c). Fixed input 53 includes a preoperative MRI 53a (which is volumetric data) and a targeting command 53b (specified as text or a set of points marked on the MRI by the operator). The AI processor 52a within the DSLVM 52 outputs the next probe position / orientation 54a, data from an "indicator" 54b showing the degree of completion of the acquisition, a second indicator 54c showing the confidence level of the treatment coordinates, and a real-time head model 54d with a treatment outline to the operator display 54. In an embodiment, the AI processor 52a includes a network of GPU and CPU processors. The head model can be considered as a rotation and translation of the preoperative MRI data. The real-time head model 54d is overlaid with treatment contour data from coordinate points 56. Beamformer coefficients define the steering and focusing of the neuromodulation hardware 59 and are generated from the coordinate points 56. Slow-time analog control 58, for example to coordinate the timing of the neuromodulation beam with physiological events sensed by EEG or MEG 51b, can also control the neuromodulation hardware 59. The coordinate points can include fiducial structures or blood vessel locations within the target contour.
[0045] In the embodiments of the present disclosure, Figure 5 The use case of the slice ultrasound method depicted in (where the target anatomical structure is the cranial region of the subject) includes the following steps:
[0046] DSLVM reads preoperative MRI volumes.
[0047] • The operator defines the structures to be processed as text or by annotating the MRI volume.
[0048] The operator begins scanning the subject's head with the ultrasound probe.
[0049] The operator moves the probe over the subject's head to acquire slices for DSLVM. This is necessary because Figure 3 In contrast to the situation in the system of , the probe 40 is not acquiring volume data.
[0050] • The operator views the MRI 54d translated into the probe's coordinate system on the display 54, where the DSLVM generates the MRI to real-time probe coordinate transformation.
[0051] o The operator monitors the degree of completion of data acquisition so far by observing the indicator 54b while moving the probe over the subject's head.
[0052] o The DSLVM 52 provides instructions 54a informing the operator of the best way to move the probe to complete data acquisition.
[0053] • Once the indicator 54b indicates that enough data has been acquired, the probe position indicator 54a shows the ideal position for treatment.
[0054] The operator secures the probe in the indicated position and checks the indicator gauge 54c to determine whether the acquired data and probe position provide confidence in the treatment setup.
[0055] • If the clinical operator approves the data for display, they begin processing.
[0056] exist Figure 3 and Figure 5 Some differences are evident, reflecting the increased complexity of achieving full guidance control with a guidance array capable of producing only electronically scanned slice data. Specifically:
[0057] • Additional real-time input may be provided to the DSLVM via a MEMS gyroscope 51c which reports the orientation of the probe. Similarly, a MEMS accelerometer (not shown) may also be mounted on the probe to report its movement.
[0058] · Figure 5 Key to the operation of the device is a pre-operative MRI 53a of the subject's head. This provides anatomical information of each subject to correlate with the real-time ultrasound slice data.
[0059] The operator display includes a real-time head model with a treatment outline 54d. DSLVM creates the treatment volume outline based on text prompts or a set of points marked on the MRI by the clinician. DSLVM also determines how to rotate and translate the MRI to appear in the coordinate system of the neuromodulation transducer. This operation is simplified if the neuromodulation probe is identical to the guide probe or is rigidly attached to it.
[0060] • As will be clear from the previously described methods, the slice transducer may need to image the head from multiple viewpoints before the DSLVM has sufficient confidence that the probe position and orientation can be moved to a suitable plane for targeted neuromodulation.
[0061] • Another part of the display 54a shows the head outline and the position of the probe thereon, as well as arrows recommending translation and rotation for the new acquisition, which effectively provides a proper idea of the volumetric anatomy of the brain for successful targeting.
[0062] • Another display element 54b shows the DSLVM view of the completion of the acquisition.For slice acquisition, several views need to be acquired before the DSLVM can have confidence in mapping the pre-operative MRI to real-time ultrasound.
[0063] • Once enough slices are obtained for reliable targeting, the DSLVM specifies to the operator the optimal position of the probe for neuromodulation treatment. Again, arrows help show the operator how to move from the initial position to the optimal position and orientation.
Claims
1. A system for training large-scale visual models, comprising: Technical regulations for ultrasound probes; a database of volumetric anatomical scans; a cue string defining the anatomical target; at least one computing device; as well as an application executed by the at least one computing device, the application, when executed, causing the at least one computing device to at least: training a large-scale vision model, wherein input comprises the cue string and data obtained by simulating a field produced by a prescribed ultrasound probe and one of the volumetric anatomical scans; as well as The trained outputs produce weights that generate a set of target locations.
2. The system according to claim 1, wherein: The ultrasound probe defined comprises a two-dimensional array of elements.
3. The system according to claim 1, wherein: The ultrasound probe defined comprises a one-dimensional array of elements.
4. The system according to claim 1, wherein: The target anatomical structures are structures within the human brain.
5. The system according to claim 1, wherein: The large-scale vision model is trained on a database of cranial data.
6. A system comprising: Ultrasound probe; a neuromodulation beamformer coupled to the ultrasound probe; at least one computing device; as well as an application executed by the at least one computing device, the application, when executed, causing the at least one computing device to at least: performing effect inference using an AI processor within a large visual model (LVM) based on a database of target anatomy and data acquired by at least one of the ultrasound probe and the neuromodulation beamformer; and Coordinates of the target anatomical structure are generated.
7. The system according to claim 6, wherein: The ultrasound probe includes a two-dimensional array of elements.
8. The system according to claim 6, wherein: The ultrasound probe includes a one-dimensional element array.
9. The system of claim 6, further comprising a prompt to indicate the target anatomical structure.
10. The system of claim 6, further comprising a volume scan of the head acquired prior to surgery.
11. The system of claim 6, further comprising beamforming electronics for neuromodulation.
12. The system according to claim 6, wherein: Each generator comprises a rectangular element having a height greater than its aspect.
13. The system according to claim 6, wherein: The large visual model identifies the best match from a database of MRI volumes.
14. The system according to claim 6, wherein: The project is displayed to the clinician to confirm correct targeting.
15. The system according to claim 6, wherein: The large visual model is guided by ultrasound measured coordinates of at least one of a reference structure and a blood vessel location within the target outline.
16. The system according to claim 6, wherein: The user interface also includes a first indicator showing a degree of completion of acquisition of anatomical information and a second indicator showing a targeting confidence level.
17. The system according to claim 6, wherein: The application generates a user interface including at least one of a grayscale B-mode, a tissue harmonic imaging mode (THI mode), a color flow mapping mode, a power Doppler mode, and a strain imaging mode.
18. The system according to claim 6, wherein: The user interface includes additional data provided using a body scan produced by a camera or LiDAR device.
19. The system according to claim 6, wherein: The target anatomical structure is a structure within the human brain.
20. The system of claim 6, wherein: The timing of neuromodulatory emissions is controlled by the slow time modulation control provided by the LVM.