Natural language interface for imaging system
The machine learning system for natural language control of imaging devices addresses the complexity of user interfaces in surgical assistance systems by allowing users to control devices using spoken commands, enhancing accessibility and efficiency.
Patent Information
- Application Number
- PCT/US2024/057263
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-12
AI Technical Summary
Complex user interfaces in surgical assistance systems, such as robotic C-arm X-ray imaging devices, hinder ready use due to poor design and variability across devices, making high-level functions inaccessible to users.
A machine learning system for natural language control of imaging devices, which includes an imaging device interface, a language model interface, an audio interface, and a human model, allowing users to control imaging devices using spoken natural language commands.
The system enables users to efficiently control imaging devices by converting natural language commands into machine-readable instructions, reducing the complexity of user interfaces and making high-level functions more accessible.
Smart Images

Figure US2024057263_12062025_PF_FP_ABST
Abstract
Description
NATURAL LANGUAGE INTERFACE FOR IMAGING SYSTEMRelated Application
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 606,685, filed December 6, 2023, and entitled, “Natural Language Interface for Imaging System.”Field
[0002] This disclosure relates generally to imaging systems, such as medical imaging systems.Background
[0003] As surgical assistance systems become more capable, their user interfaces have likewise grown more complex. Poor interface design or interface variability across devices and manufacturers hinders ready use. High-level functions may exist, but if the actual steps to use them are not obvious, then their availability is merely theoretical, or “present-at-hand.” A hammer, by contrast, is “ready-to-hand,” in that its function and use are both available at a glance.
[0004] Robotic C-arm X-ray imaging devices are capable of precisely orienting themselves to achieve desired views and can provide guidance to many minimally invasive procedures in orthopedics, interventional radiology, and angiology, for example. However, the user interfaces for such systems are highly complex, e.g., often requiring manipulation of independent axes via joysticks. Moreover, higher-level functions, such as task-aware imaging, may reduce radiation or operating time, butoften hide inside device-specific menus, remaining hidden to all but dedicated power users of specific systems.Summary
[0005] According to various embodiments, a machine learning system for natural language control of an imaging device is presented. The machine learning system includes: an imaging device interface communicatively coupled to the imaging device; a language model interface communicatively coupled to a language model and to the imaging device interface; an audio interface communicatively coupled to the language model interface; and a human model communicatively coupled to the imaging device and to the language model interface, wherein: the audio interface is configured to receive an audio input from a user, transcribe the audio input to an audio input transcription, and provide the audio input transcription to the language model interface, the language model interface is configured to provide the audio input transcription to the language model, receive a response from the language model, and provide the response to the human model, the human model is configured to maintain a representation of a current position of an individual, receive the response from the language model interface, form an action request based on the response and the current position of the individual, and send the action request to imaging device interface, and the imaging device interface is configured to send the action request to the imaging device, whereby the imaging device acts on the action request.
[0006] Various optional features of the above system embodiments include the following. The system may include an image analysis module communicatively coupled to the imaging device interface and to the human model, wherein: the image analysis module is configured to receive images from the imaging device interface andgenerate image analysis data, and the human model is configured to update the representation of the current position of the individual based on the image analysis data received from the image analysis module. The image analysis module may be configured to apply automated landmark detection to the images received from the imaging device interface. The automated landmark detection may include fitting a statistical shape model based on the images received from the imaging device. The individual may include a patient, wherein the imaging device comprises a medical imaging device, and wherein the human model comprises a patient model. The imaging device may include one of: an X-ray machine, an ultrasound device, a Magnetic Resonance Imaging (MRI) machine, or a Computed Tomography (CT) machine. The audio interface may be further configured to receive a second audio input from the user, wherein the second audio input transcription represents a positionagnostic action, transcribe the second audio input to a second audio input transcription, and provide the second audio input transcription to the language model interface, the language model interface may be further configured to provide a second audio input transcription to the language model, receive a second action request from the language model, and provide the second action request to the imaging device interface, and the imaging device interface may be further configured to send the second action request to the imaging device, whereby the imaging device acts on the second action request. The second action request may include a shot action request. The language model interface may be further configured to provide the audio input transcription to the language model together with an instruction set, wherein the instruction set comprises natural language instructions. The instruction set may include instructions that the language model may request clarification. The imaging device interface may be further configured to send state information to the languagemodel interface, and the language model interface may be further configured to provide the state information to the language model, whereby the language model is provided with context information regarding the individual. The action request may include at least one of: a pose action request or a view action request. The language model may include a large language model.
[0007] According to various embodiments, a machine learning method of providing natural language control of an imaging device is presented. The method is implemented by an imaging device interface comprising an electronic processor and communicatively coupled to the imaging device, a language model interface comprising an electronic processor and communicatively coupled to a language model and to the imaging device interface, an audio interface comprising an electronic processor and communicatively coupled to the language model interface, and a human model comprising an electronic processor and communicatively coupled to the imaging device and to the language model interface. The method includes: receiving, by the audio interface, an audio input from a user; transcribing, by the audio interface, the audio input to an audio input transcription; providing, by the audio interface, the audio input transcription to the language model interface; providing, by the language model interface, the audio input transcription to the language model; receiving, by the language model interface, a response from the language model; providing, by the language model interface, the response to the human model; maintaining, by the human model, a representation of a current position of an individual; receiving, by the human model, the response from the language model interface; forming, by the human model, an action request based on the response and the current position of the individual; sending, by the human model, the action request to imaging deviceinterface; and sending, by the imaging device interface, the action request to the imaging device, whereby the imaging device acts on the action request.
[0008] Various optional features of the above method embodiments include the following. The method may be implemented using an image analysis module communicatively coupled to the imaging device interface and to the human model, and the method may include: receiving, by the image analysis module, images from the imaging device interface and generate image analysis data; and updating, by the human model, the representation of the current position of the individual based on the image analysis data received from the image analysis module. The method may include applying, by the image analysis module, automated landmark detection to the images received from the imaging device interface. The automated landmark detection may include fitting a statistical shape model based on the images received from the imaging device. The individual may include a patient, wherein the imaging device comprises a medical imaging device, and wherein the human model comprises a patient model. The imaging device may include one of: an X-ray machine, an ultrasound device, a Magnetic Resonance Imaging (MRI) machine, or a Computed Tomography (CT) machine. The method may include receiving, by the audio interface, a second audio input from the user, wherein the second audio input transcription represents a position-agnostic action; transcribing, by the audio interface, the second audio input to a second audio input transcription; providing, by the audio interface, the second audio input transcription to the language model interface; providing, by the language model interface, a second audio input transcription to the language model; receiving, by the language model interface, a second action request from the language model; providing, by the language model interface, the second action request to the imaging device interface; and sending, by the imaging device interface, the secondaction request to the imaging device, whereby the imaging device acts on the second action request. The second action request may include a shot action request. The method may include providing, by the language model interface, the audio input transcription to the language model together with an instruction set, wherein the instruction set comprises natural language instructions. The instruction set may include instructions that the language model may request clarification. The method may include sending, by the imaging device interface, state information to the language model interface; and providing, by the language model interface, the state information to the language model, whereby the language model is provided with context information regarding the individual. The action request may include at least one of: a pose action request or a view action request. The language model may be a large language model.
[0009] Combinations, (including multiple dependent combinations) of the above-described elements and those within the specification have been contemplated by the inventors and may be made, except where otherwise indicated or where contradictory.Brief Description of the Drawings
[0010] Various features of the examples can be more fully appreciated, as the same become better understood with reference to the following detailed description of the examples when considered in connection with the accompanying figures, in which:
[0011] Fig. 1 provides a high-level overview of use of a system for natural language control of an imaging device, according to various embodiments;
[0012] Fig. 2 is a block and data flow schematic diagram of a system for natural language control of an imaging device, according to various embodiments;
[0013] Fig. 3 presents non-limiting example messages and responses in a system for natural language control of an imaging device, according to various embodiments;
[0014] Fig. 4 illustrates a Brainlab Loop-X X-ray system as used in a system for natural language control of an imaging device, according to various embodiments;
[0015] Fig. 5 illustrates automated landmark detection to determine the patient pose with respect to a statistical shape model, according to various embodiments; and
[0016] Fig. 6 illustrates a study that implemented a system for natural language control of an imaging device, according to various embodiments.Description of the Examples
[0017] Reference will now be made in detail to example implementations, illustrated in the accompanying drawings. Wherever convenient, the same reference numbers will be used throughout the drawings to refer to the same or like parts. In the following description, reference is made to the accompanying drawings that form a part thereof, and in which is shown by way of illustration specific exemplary examples in which the invention may be practiced. These examples are described in sufficient detail to enable those skilled in the art to practice the invention and it is to be understood that other examples may be utilized and that changes may be made without departing from the scope of the invention. The following description is, therefore, merely exemplary.
[0018] Some embodiments provide a voice interface for controlling a medical imaging system, such as a robotic X-ray imaging system. According to some embodiments, the voice interface provides a convenient, flexible way for a user, such as a surgeon, to manipulate a medical imaging system based on desired outcomes,rather than device-specific processes. Some embodiments allow a user to provide natural language commands to express their desired outcome, rather than remembering the steps necessary to achieve it, allowing for direct access to, for example, task-aware, patient-specific C-arm X-ray functionality. Some embodiments provide a fully integrated system, which uses a large language model (LLM) to convert natural language spoken commands into machine-readable instructions, allowing low- level commands, like “tilt back a bit” to increase the angular tilt, or patient-specific directions, like “go to the obturator oblique view of the right ramus,” based on automated image analysis. Some embodiments make complex surgical assistance systems ready-to-hand, which may encourage adoption and use of potentially time- and risk-reducing capabilities.
[0019] A non-limiting example embodiment is presented herein in reference to a study, as shown and described in reference to Fig. 6, which implemented a fully integrated interface for controlling a robotic X-ray system using natural language spoken commands, with additional support for patient-specific imaging in X-ray guided pelvic surgery. Based on observations of X-ray guided surgery, the example embodiment incorporated a limited set of episodes indicating to a LLM how to adjust a robotic X-ray system given certain commands. The example embodiment provided support for patient-specific imaging, which chose the optimal pose to achieve a desired view based on automated image analysis. When provided with these episodes and a short set of instructions for standardizing communication, the example embodiment was able to generate machine-readable commands for positioning a Brainlab Loop-X device, a fully robotic X-ray system with six degrees of freedom. The example embodiment utilized a transcription model to convert spoken commands into text prompts, which were relayed to the LLM. The resulting commands were either relayeddirectly to the Loop-X, in the case of a low-level movement, or referred to an image analysis engine for determining the patient-specific view.
[0020] The inventors evaluated the example embodiment of the study using 212 prompts provided by an attending physician. The evaluation included a real-time study in which the attending physician placed orthopedic hardware along desired trajectories through an anthropomorphic phantom, interacting solely with an X-ray system via voice. The evaluation revealed that the system performed satisfactory actions 97.17% of the time.
[0021] These and other features and advantages are shown and described herein in reference to the accompanying figures.
[0022] Fig. 1 provides a high-level overview of use of a system for natural language control of an imaging device, according to various embodiments. As shown in Fig. 1 , the system allows a user to control independent axes of a robotic X-ray system, e.g., a robotic C-arm X-ray such as the Brainlab Loop-X, using spoken natural language. Patient-specific views, such as the obturator oblique, are supported via automated image analysis. As shown in Fig. 1 at 102, a user such as a physician vocalizes a task-aware patient-specific view command in natural language, such as “go to the obturator oblique view of the right ramus.” At 104, the user vocalizes, “now take a shot.” At 106, the user vocalizes, “looks good.”
[0023] Fig. 2 is a block and communication flow schematic diagram of a system 200 for natural language control of an imaging device, according to various embodiments. As shown in Fig. 1 , the system 200 is used by a user such as a surgeon user 220. Shown in Fig. 2 are an X-ray device 202, a device server 204, a device client 206, an image streamer 230, an image analysis module 240, a patient model 250, an LLM server 260, an LLM client 262, a speaker 264, and a transcriber 266.Any, or any combination, of the aforementioned elements may be included in one or more computing devices. For example, according to some embodiments, the aforementioned elements may be included in the X-ray device 202. According to some embodiments, the aforementioned elements, excepting the X-ray device 202, may be included in a separate computing device. According to some embodiments, the aforementioned elements, excepting the X-ray device 202, the device server 204, and the LLM server 260, may be included in a separate computing device.
[0024] The X-ray device 202 may be any type of partially- or fully-robotic controlled X-ray device. For example, the X-ray device 202 may be a robotic C-arm X-ray such as the Brainlab Loop-X. A non-limiting specific example of a suitable X- ray device is shown and described herein in reference to Fig. 4. More generally, embodiments may control any medical imaging device, such as, by way of non-limiting example: a magnetic resonance imaging (MRI) device, a computed tomography (CT) imaging device, or an ultrasound device.
[0025] According to some embodiments, the X-ray device 202 may utilize a client-server architecture to facilitate communication, such as command and control communication, between the X-ray device 202 and other elements of the system 200. The client-server architecture may utilize any client-server protocol, e.g., TCP / IP, Remote Function Call Protocol (RFCP), HTTP, File Transfer Protocol (FTP), Point Tunneling Protocol (PPTP), or Telnet. Embodiments may include a device server 204, which may be present in, or external to, the X-ray device 202. The device server 204 may communicate with a device client 206. The device client 206 is a non-limiting example of an interface for the X-ray device 202. According to various embodiments, the device client may be present in, or external to, the X-ray device 202 and / or a separate computer, e.g., a computer that also includes one or more of: the LLM server260, the LLM client 262, the speaker 264, the transcriber 266, the patient model 250, and / or the image analysis module 240.
[0026] The system 200 is also shown in Fig. 2 as including an image streamer 230. The image streamer 230 may provide a buffer for obtaining and temporarily storing images obtained from the X-ray device 202. The image streamer 230 may implement any appropriate communication protocol for streaming digital images.
[0027] The LLM server 260 may be any large language model implementation. The LLM server 260 may be hosted remote from other elements of the system 200, or may be implemented in a computer that also includes the LLM client 262, for example. The LLM server 260 inputs natural language text, which it embeds into a tokenized large language model. It converts these latent-space embeddings into machine- readable instructions, which it outputs to the LLM client 262 for consumption by other elements of the system 200 as described in detail herein.
[0028] The LLM client 262 may be implemented to communicate with the LLM server 260 using any appropriate protocol, e.g., TCP / IP, RFCP, HTTP, FTP, PPTP, or Telnet. The LLM client 262 is a non-limiting example of an interface for the LLM server 260. According to some embodiments, the LLM client 262 may be implemented in the same computer as other elements of the system 200, such as the patient model 250. According to some embodiments, the LLM client 262 formats communications received from the transcriber 266 and the device client 206 into standard-form communications, as shown and described herein in reference to Fig. 3.
[0029] The patient model 250 maintains a representation of the current position of the patient being imaged by the X-ray device 202. The patient model 250 may be implemented using a statistical shape model (SSM), which may be specific to a particular anatomical feature, such as a pelvis. The patient model 250 may base itsrepresentation of the current position of the patient at least partially on an analysis of a current image of the patient received from the image analysis module 240. The patient model 250 also receives machine-readable instructions from the LLM client 262 and forms an action request based on the instructions and its representation of the current position of the patient. The patient model 250 then sends the action request to X-ray device client 206. A detailed description of a suitable patient model 250 appears below in reference to Fig. 5.
[0030] The image analysis module 240 identifies anatomical landmarks on one or more images generated by the X-ray device 202 and acquired from the image streamer 230. According to some embodiments, the image analysis module 240 may identify one or more recent X-ray images from which it identifies the anatomical landmarks. The image analysis module 240 provides the identified landmarks to the patient module 250, which registers them to its internal representation of the patient. A detailed description of a suitable image analysis module 240 appears below in reference to Fig. 5.
[0031] The system 200 also includes a speaker 264 and a transcriber 266 that facilitate audio communication between the surgeon user 220 and the LLM client 262. The speaker 264 may include a standard electromagnetic speaker coupled to a speech synthesizer. The transcriber 266 may include a speech-to-text module, which may be implemented as an executing transcription program. According to some embodiments, the transcriber 266 may be implemented using artificial intelligence. The transcriber 266 is a non-limiting example of an audio interface; according to some embodiments, such an audio interface may also include the speaker 264.
[0032] An example use case of the system 200 is shown and described presently in reference to both Fig. 2 and Fig. 3.
[0033] Fig. 3 presents non-limiting example messages 310 and responses 320 in a system for natural language control of an imaging device, according to various embodiments. In Fig. 3, communications to the LLM server 260 are referred to as messages 310, and communications from the LLM server 260 are referred to as responses. As shown in Fig. 3, and by way of non-limiting example, communications include single-line strings delimited by semicolons, with the first element indicating the topic or mes sage_type.
[0034] During usage of the system 200, a surgeon user 220 may provide a command in the form of an audio input to the transcriber 266. The transcriber 266 transcribes the speech to corresponding text, which it passes to the LLM client 262. The device client 206 may also pass the current patient position from the patient model 250 and the current X-ray system pose from the device client 206 to the LLM client 262 to provide additional context. The LLM client 262 formats the transcribed audio input from the transcriber 266, and for some inputs also the current patient position from the patient model 250 and / or the current X-ray system state from the device client 206, into a command 312, which it passes to the LLM server 260. The LLM server 260 can provide a response, which may be an action or a question 326, the latter in the case in which the LLM server 260 is unsure how to proceed.
[0035] Actions can include an axis movement (pose 322), patient-specific view (view 324), image acquisition (shot), or no-op action (none). Patient-agnostic actions, such as image acquisition and no-op actions, may be passed from the LLM server to the X-ray device 202 by way of the LLM client 262, the device client 206, and the device server 204. Actions that depend on the patient (e.g., the current pose of the patient) are passed from the LLM server 260 to the X-ray device 202 by way of theLLM client 262, the patient model 250, the device client 206, and the device server204. According to some embodiments, e.g., for safety reasons, the system 200 may await confirmation from the user 220, e.g., in the form of activation of a physical button before the X-ray device takes a shot according to the passed action. Once the user provides such confirmation, the system 200 may pass a conf irmation 316 from the X-ray device 202 to the LLM server 260 by way of the device server 204, the device client 206, and the LLM client 262.
[0036] Questions are passed from the LLM server 260 to the user 220 by way of the LLM client 262, which may reformat the question into natural language from its original communication format, and the speaker 264. In the case of a question, the user 220 may respond with a clari f ication 314, which the system 200 may store for future reference.
[0037] As shown in Fig. 3, communications may be concise in order to minimize token usage, a common bottleneck in existing LLMs.
[0038] According to some embodiments, one or more (e.g., every) message sent to the LLM server 260 is accompanied by a set of full instructions. A non-limiting example of such instructions appears at the end of this written description under the heading “Example LLM Instructions.” According to various embodiments, example interactions are appended after the full text of the instructions. (For the study shown and described herein in reference to Fig. 6, thirty-five such example interactions were included.) Each example interaction includes the initial command, followed by an anticipated response from the LLM, as well as a final summary of the interaction. By way of non-limiting illustrative example, the following interaction details the “push in” command, which should move the C-arm X-ray laterally toward the surgeon. (The patient_s ide is provided at start-time here, but could be inferred from imaging or external camera sources.)
[0039] command; True ; supine ; patient_right ; 180 . 0 ; 0 . 0 ; 0 . 0 ; 0 . 0 ;0 . 0 ; 0 . 0 ; Push in .
[0040] action; pose ; 180 . 0 ; 0 . 0 ; 0 . 0 ; 10 . 0 ; 0 . 0 ; 0 . 0
[0041] summary; " Push in" means moving along X toward the surgeon . Since the surgeon is on patient right , head_f irst=True , and patient_pose=supine , you should increase X by 10 .
[0042] Fig. 4 illustrates a Brainlab Loop-X X-ray system 400 as used in a system for natural language control of an imaging device, according to various embodiments. Note that the Brainlab Loop-X X-ray system 400 is presented by way of non-limiting example. In particular, the Brainlab Loop-X X-ray system 400 was used for an embodiment as part of the study show as shown and described herein in reference to Fig. 6. Note further that although various embodiments as presented herein (e.g., in reference to the Example LLM Instructions) control the degrees of freedom present in this system 400, embodiments may be implemented for other medical imaging systems due to the brevity of the instructions, which are less than 250 lines including examples.
[0043] As shown in Fig. 4, the Loop-X X-ray system 400 includes six independent axes. Unlike conventional C-arm systems, the Loop-X is an O-arm-like device with independently moving X-ray source and detector arms (source_angle and detector_angle). It moves on the floor in any direction, in a coordinate system specified by the lateral (x) and longitudinal (y) position. It can also rotate freely (yaw). Finally, the gantry tilts from -30° to 60° about x. Communication with the system 400 may utilize the RFCP, which allows for queries and commands based on these axes. The RFCP also provides functionality to achieve precise views specified by anorientation vector in coordinate frame of the ring fring, with which to align the principle ray r, and acquire navigated 2D images semi-automatically.
[0044] Fig. 5 illustrates automated landmark detection to determine the patient pose with respect to a statistical shape model (SSM) according to various embodiments. By way of non-limiting example, Fig. 5 illustrates techniques that were implemented for the study shown and described herein in reference to Fig. 6. To determine patient-specific views, the example embodiment utilized automated landmark detection to register the anatomy to an SSM of the pelvis, in which standard views were defined relative to the anterior pelvic plane (APP) coordinate frame. First, the embodiment identified a suitable set of recent acquisitions from which to triangulate points. Because X-ray guided surgery often includes multiple images acquired from the same standard view, the principle ray directions of all acquired images were clustered using the angular distance metric. Choosing the most recent image from each cluster, then, obtained landmarks {uz, I e IR3|Z e £ ,is the set of landmarks present in image / . The embodiment used the available pipeline from Killeen, et al., Pelphix: Surgical Phase Recognition from X-ray Images in Percutaneous Pelvic Fixation, arXiv (2023), available at doi.org / 10.48550 / arXiv.2304.09285 (hereinafter, “Killeen”) to detect anatomical landmarks automatically. Finally, for each landmark present in at least two images, the embodiment minimized the reprojection error to estimate the 3D position xf, I e IR3according to the following formula: x(= argmax(||P(x - u||2, where P(is the projection matrix of image / relative to an optical marker body fixed to the patient, and ~ indicates the homogeneous point. The transformation from the marker body coordinates to / ring was provided by the RFCP. Once anatomical landmarks had been obtained in 3D, the embodiment fit an SSM of the pelvis to the triangulated points.This ensured that the anterior superior iliac spine (ASIS) landmarks could be estimated, from which the APP frame was determined. The principle ray for each standard view, including the AP, lateral, inlet, outlet, and obturator oblique views, were defined according to the conventions of Killeen. Fig. 5 shows the image projections, detected landmarks, and standard views for three acquisitions sampled randomly, such as might be acquired in the course of fluoro-hunting.
[0045] Fig. 6 illustrates a study 600 that implemented a system for natural language control of an imaging device, according to various embodiments. The study 600 utilized a phantom, and an attending orthopedic surgeon was able to emulate X-ray guided percutaneous fixation of the sacroiliac joint by interfacing with a Loop-X X-ray system solely through voice commands.
[0046] The inventors evaluated the embodiment used in the study 600 in three ways. First, the inventors evaluated the language interface by determining whether the embodiment chose satisfactory actions based on 212 prompts commonly occurring in pelvic trauma surgery. Second, for actions that requested a patient specific view, the inventors evaluated the embodiment’s ability to attain this view based on 1 , 2, and 3 random prior images, using a rating system of 1 (wrong view) to 5 (no adjustment needed), as rated by an attending physician. Finally, the inventors evaluated the embodiment of the study 600 with spoken commands carried out by a Brainlab Loop-X X-ray device.
[0047] To generate prompts in the test set, the inventors began with the base prompts listed in Table 1 , below. These were paraphrased from conversations with an attending orthopedic surgeon and professional C-arm technologists. Additional variations on the base prompts were then generated by asking a LLM for more ways to say the same phrase, being sure to inject alternate phrases and synonyms. Forexample, the simplest command to achieve more inlet tilt was “more inlet,” based on which variations were introduced such as “provide a downward tilt away from the patient’s head,” and “adjust the C-Arm for a lower view, away from the head.” The resulting list was then culled of 100 commands to ensure they still aligned with sensible outputs for controlling the Loop-X system, resulting in 53 prompts not seen during the example episodes. For each prompt, the embodiment supplied with the instructions was tested on four randomly sampled Loop-X poses as the input pose. In general, the embodiment performed quite well, choosing the correct response for (206 / 212) 97.17% commands. In general, failures resulted from extreme Loop-X pose inputs dissimilar to those supplied in the example episode.Table 1 : Actions Given Responses
[0048] The real images obtained for patient-specific views achieved during these tests were also evaluated on a radiopaque anthropomorphic phantom of the pelvis, in terms of their alignment with the standard view as evaluated on a 10-point scale by an attending orthopedic surgeon (see Table 2). The AP, lateral, and inlet views were successfully or very nearly achieved based on voice commands, with an average rating of 8.8, while the outlet and obturator oblique views demonstrated room for improvement. This was in part due to the fact that the Loop-X was limited to a 30° tilt in the outlet direction, although the outlet view is typically closer to 40° away from AP.Fig. 2: Physician Rating for Patient-Specific Views
[0049] The study also included an overall evaluation, in which the embodiment was used to emulate percutaneous fracture fixation in the pelvis. An attending orthopedic surgeon interfaced with the embodiment solely via voice. A pair of Shokz OpenRun wireless headphones streamed audio to a 2019 Macbook Pro laptop, which transcribed commands in real time using the Whisper ASR model. A Linux server with an RTX 3090 GPU was responsible for automated landmark detection andtriangulation, as well as communication with the Loop-X X-ray system via the RFCP. Images acquired on the Loop-X X-ray system were relayed in DICOM format, including navigation information, to the server directly. While interfacing with the Loop-X, the surgeon aimed to align a surgical pointer with the S2 corridor, a narrow bony corridor often used for stabilize sacroiliac fractures. Alternating between inlet and outlet views, the surgeon positioned the pointer, which was held in place with a passive positioning arm to allow for trajectory verification. In total, 33 X-ray images were acquired during the study, and successful placement along the S2 corridor was evaluated with a CT scan by the Loop-X.
[0050] As shown in Fig. 6, the study included moving from an inlet view 602 to a patient-specific AP view 604. Note that because the Loop-X X-ray system source and detector move independently, it was possible to take non-isocentric images that were post-processed to correct for skew, accounting for the angled detector. In the study, the embodiment requested clarifications for commands it did not understand, which could be avoided in further embodiments by specifying the amount of movement desired for fine-grained adjustments. The LLM server of the embodiment also responded with question messages in the case of a transcription error, such as confusing “inlet” for “in that.” Once the user provided a clarification, the LLM server was able to compensate for similar errors.
[0051] Embodiments have been disclosed herein that provide many advantages. For example, embodiments may facilitate patient-specific imaging, which can reduce the number of images acquired, e.g., in the process of fluoro-hunting, thereby reducing the radiation exposure for patients and staff as well as the time under general anesthesia. The natural language interface of some embodiments provides not only a seamless way to access such functions, emulating the existingcommunication strategies that exist between surgeons and operating room support staff, but also can accelerate prototyping and testing of high-level functions by avoiding the need for an intuitive graphical user interface, which can be difficult and costly to develop in addition to the challenges associated with potential new capabilities.
[0052] Many variations of the embodiments explicitly disclosed herein are possible. For example, although in some embodiments disclosed herein the full instruction set, with examples, is included in every exchange with the LLM server, other embodiments may utilize a fine-tuned LLM server, which can support exchanges that persist throughout the surgery, with life-long learning that accommodates individual surgeon preferences. Additionally, using a fine-tuned LLM server can allow for more sophisticated protocols with longer instruction sets, improving the robustness of low-level control by relegating pose changes to an underlying function, rather than relying on the LLM server.
[0053] Certain examples can be performed using a computer program or set of programs. The computer programs can exist in a variety of forms both active and inactive. For example, the computer programs can exist as software program(s) comprised of program instructions in source code, object code, executable code or other formats; firmware program(s), or hardware description language (HDL) files. Any of the above can be embodied on a transitory or non-transitory computer readable medium, which include storage devices and signals, in compressed or uncompressed form. Exemplary computer readable storage devices include conventional computer system RAM (random access memory), ROM (read-only memory), EPROM (erasable, programmable ROM), EEPROM (electrically erasable, programmable ROM), flash memory, and magnetic or optical disks or tapes.
[0054] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented using computer readable program instructions that are executed by an electronic processor.
[0055] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the electronic processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.
[0056] In embodiments, the computer readable program instructions may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++, or the like, andprocedural programming languages, such as the C programming language or similar programming languages. The computer readable program instructions may execute entirely on a user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
[0057] As used herein, the terms “A or B” and “A and / or B” are intended to encompass A, B, or {A and B}. Further, the terms “A, B, or C” and “A, B, and / or C” are intended to encompass single items, pairs of items, or all items, that is, all of: A, B, C, {A and B}, {A and C}, {B and C}, and {A and B and C}. The term “or” as used herein means “and / or.”
[0058] As used herein, language such as “at least one of X, Y, and Z,” “at least one of X, Y, or Z,” “at least one or more of X, Y, and Z,” “at least one or more of X, Y, or Z,” “at least one or more of X, Y, and / or Z,” or “at least one of X, Y, and / or Z,” is intended to be inclusive of both a single item (e.g., just X, or just Y, or just Z) and multiple items (e.g., {X and Y}, {X and Z}, {Y and Z}, or {X, Y, and Z}). The phrase “at least one of” and similar phrases are not intended to convey a requirement that each possible item must be present, although each possible item may be present.
[0059] The techniques presented and claimed herein are referenced and applied to material objects and concrete examples of a practical nature that demonstrably improve the present technical field and, as such, are not abstract, intangible or purely theoretical. Further, if any claims appended to the end of this specification contain one or more elements designated as “means for [perform]ing [a function]...” or “step for [performing [a function]...”, it is intended that such elements are to be interpreted under 35 U.S.C. § 112(f). However, for any claims containingelements designated in any other manner, it is intended that such elements are not to be interpreted under 35 U.S.C. § 112(f).
[0060] While the invention has been described with reference to the exemplary examples thereof, those skilled in the art will be able to make various modifications to the described examples without departing from the true spirit and scope. The terms and descriptions used herein are set forth by way of illustration only and are not meant as limitations. In particular, although the method has been described by examples, the steps of the method can be performed in a different order than illustrated or simultaneously. Those skilled in the art will recognize that these and other variations are possible within the spirit and scope as defined in the following claims and their equivalents.
[0061] Example LLM Instructions
[0062] Please only receive and respond to messages in the following single-line format with no additional text . Do not explain your reasoning except in a summary message when requested .
[0063] Messages are a single line with the following structure :
[0064]
[0065] message_type ; . . .
[0066]
[0067] where message_type= { command | question | action | clari fication ! accepted ! s ummary }
[0068] The ' command' message type has this structure :
[0069]
[0070] command; patient_head_f irst ; patient_pose ; surgeon_side; source_angle ; detector_angle ; tilt ; x; y ; yaw; user_input
[0071]
[0072] where patient_head_f irst= { True, False } , patient_pose= { supine | prone } , and surgeon_side= { left | right } . source_angle, detector_angle, tilt, x, y, and yaw are all floats giving the loopx pose. user_input is a string.
[0073] You have two options for a response. Please always reply with one of the following:
[0074] 1. If you are confused, send a 'question' message:
[0075]
[0076] question; question_text
[0077]
[0078] where question_text is a string.
[0079] 2. If you are ready, reply with an 'action' message:
[0080]
[0081] action; action_type ; . . .
[0082]
[0083] where action_type= {pose | shot | view | none }
[0084]
[0085] - A 'pose' action moves the loopx to the given pose:
[0086]
[0087] action; pose ; source_angle ; detector_angle ; tilt ; x; y ; yaw
[0088]
[0089] - A 'shot' action is has no extra parameters and just takes a shot.
[0090]
[0091] action; shot
[0092]
[0093] - A 'view' action requests a specific view:
[0094]
[0095] action; view; view_name ; anatomy
[0096]
[0097] where view_name= {ap | lateral | inlet | outlet | oblique_lef t | oblique_right | teardrop} and anatomy= { hip_lef t | hip_right | f emur_lef t | f emur_right | sacrum | vert ebra_L5 | pelvis | sl_left| sl_right | si | s2 | ramus_left | ramus_right | teardrop_lef t | teardrop_ right }
[0098] - 'none' performs no action.
[0099] The user will then send one of the following:
[0100] If a question was sent, the user will send a'clarification' message:
[0101]
[0102] clarification; cl ar if ication_text
[0103]
[0104] where clarif ication_text is a string with the explanation. You can then reply with an 'action' message as defined above.
[0105] Otherwise, the user might send a 'finished' message:
[0106]
[0107] finished; accepted
[0108]
[0109] where accepted= { True | False } . You should respond with a ' summary' message giving your reasoning
[0110]
[0111] reasoning; reasoning_text
[0112]
[0113] where reasoning_text is a string with your reasoning. Only send a 'summary' message when the user sends a 'finished' message. Otherwise, no explanation is needed.
Claims
What is claimed is:1 . A machine learning system for natural language control of an imaging device, the system comprising: an imaging device interface communicatively coupled to the imaging device; a language model interface communicatively coupled to a language model and to the imaging device interface; an audio interface communicatively coupled to the language model interface; and a human model communicatively coupled to the imaging device and to the language model interface, wherein: the audio interface is configured to receive an audio input from a user, transcribe the audio input to an audio input transcription, and provide the audio input transcription to the language model interface, the language model interface is configured to provide the audio input transcription to the language model, receive a response from the language model, and provide the response to the human model, the human model is configured to maintain a representation of a current position of an individual, receive the response from the language model interface, form an action request based on the response and the current position of the individual, and send the action request to imaging device interface, and the imaging device interface is configured to send the action request to the imaging device, whereby the imaging device acts on the action request.
2. The system of claim 1 , further comprising an image analysis module communicatively coupled to the imaging device interface and to the human model, wherein: the image analysis module is configured to receive images from the imaging device interface and generate image analysis data, and the human model is configured to update the representation of the current position of the individual based on the image analysis data received from the image analysis module.
3. The system of claim 2, wherein the image analysis module is configured to apply automated landmark detection to the images received from the imaging device interface.
4. The system of claim 3, wherein the automated landmark detection comprises fitting a statistical shape model based on the images received from the imaging device.
5. The system of claim 1 , wherein the individual comprises a patient, wherein the imaging device comprises a medical imaging device, and wherein the human model comprises a patient model.
6. The system of claim 5, wherein the imaging device comprises one of: an X-ray machine, an ultrasound device, a Magnetic Resonance Imaging (MRI) machine, or a Computed Tomography (CT) machine.
7. The system of claim 1 , wherein: the audio interface is further configured to receive a second audio input from the user, wherein the second audio input transcription represents a position-agnostic action, transcribe the second audio input to a second audio input transcription, and provide the second audio input transcription to the language model interface, the language model interface is further configured to provide a second audio input transcription to the language model, receive a second action request from the language model, and provide the second action request to the imaging device interface, and the imaging device interface is further configured to send the second action request to the imaging device, whereby the imaging device acts on the second action request.
8. The system of claim 7, wherein the second action request comprises a shot action request.
9. The system of claim 1 , wherein the language model interface is further configured to provide the audio input transcription to the language model together with an instruction set, wherein the instruction set comprises natural language instructions.
10. The system of claim 9, wherein the instruction set comprises instructions that the language model may request clarification.11 . The system of claim 1 , wherein:the imaging device interface is further configured to send state information to the language model interface, and the language model interface is further configured to provide the state information to the language model, whereby the language model is provided with context information regarding the individual.
12. The system of claim 1 , wherein the action request comprises at least one of: a pose action request or a view action request.
13. The system of claim 1 , wherein the language model is a large language model.
14. A machine learning method of providing natural language control of an imaging device, the method implemented by an imaging device interface comprising an electronic processor and communicatively coupled to the imaging device, a language model interface comprising an electronic processor and communicatively coupled to a language model and to the imaging device interface, an audio interface comprising an electronic processor and communicatively coupled to the language model interface, and a human model comprising an electronic processor and communicatively coupled to the imaging device and to the language model interface, the method comprising: receiving, by the audio interface, an audio input from a user; transcribing, by the audio interface, the audio input to an audio input transcription;providing, by the audio interface, the audio input transcription to the language model interface; providing, by the language model interface, the audio input transcription to the language model; receiving, by the language model interface, a response from the language model; providing, by the language model interface, the response to the human model; maintaining, by the human model, a representation of a current position of an individual; receiving, by the human model, the response from the language model interface; forming, by the human model, an action request based on the response and the current position of the individual; sending, by the human model, the action request to imaging device interface; and sending, by the imaging device interface, the action request to the imaging device, whereby the imaging device acts on the action request.
15. The method of claim 14, further implemented using an image analysis module communicatively coupled to the imaging device interface and to the human model, the method further comprising: receiving, by the image analysis module, images from the imaging device interface and generate image analysis data; andupdating, by the human model, the representation of the current position of the individual based on the image analysis data received from the image analysis module.
16. The method of claim 15, further comprising applying, by the image analysis module, automated landmark detection to the images received from the imaging device interface.
17. The method of claim 16, wherein the automated landmark detection comprises fitting a statistical shape model based on the images received from the imaging device.
18. The method of claim 14, wherein the individual comprises a patient, wherein the imaging device comprises a medical imaging device, and wherein the human model comprises a patient model.
19. The method of claim 18, wherein the imaging device comprises one of: an X-ray machine, an ultrasound device, a Magnetic Resonance Imaging (MRI) machine, or a Computed Tomography (CT) machine.
20. The method of claim 14, further comprising: receiving, by the audio interface, a second audio input from the user, wherein the second audio input transcription represents a position-agnostic action; transcribing, by the audio interface, the second audio input to a second audio input transcription;providing, by the audio interface, the second audio input transcription to the language model interface; providing, by the language model interface, a second audio input transcription to the language model; receiving, by the language model interface, a second action request from the language model; providing, by the language model interface, the second action request to the imaging device interface; and sending, by the imaging device interface, the second action request to the imaging device, whereby the imaging device acts on the second action request.21 . The method of claim 20, wherein the second action request comprises a shot action request.
22. The method of claim 14, further comprising providing, by the language model interface, the audio input transcription to the language model together with an instruction set, wherein the instruction set comprises natural language instructions.
23. The method of claim 22, wherein the instruction set comprises instructions that the language model may request clarification.
24. The method of claim 14, further comprising: sending, by the imaging device interface, state information to the language model interface; andproviding, by the language model interface, the state information to the language model, whereby the language model is provided with context information regarding the individual.
25. The method of claim 14, wherein the action request comprises at least one of: a pose action request or a view action request.
26. The method of claim 14, wherein the language model is a large language model.
Citation Information
Patent Citations
Complex image data analysis using artificial intelligence and machine learning algorithms
US20210264212A1
Radiology report editing method and system
US20230096939A1
Speech control of a medical apparatus
US20230105362A1
Cited By
Patient positioning for medical image acquisition
DE102025105347A1
Apparatus for and method of controlling visualization of anatomical image data
US12726698B1
Technique for medical imaging control based on a request message
US20260088156A1