Voice-controlled surgical system
The voice-activated surgical command system addresses noise and command recognition issues by using phased arrays for efficient, hands-free surgical device control, reducing personnel and contamination risks.
Patent Information
- Application Number
- JP2023547087
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-05
- Filing Date
- 2022-02-01
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2042-02-01
AI Technical Summary
Current voice-controlled surgical devices face challenges in recognizing and prioritizing voice commands from multiple surgical staff, are affected by background noise, and require predefined command syntax, leading to inefficiencies and increased personnel and contamination risks.
A voice-activated surgical command system using phased microphone and loudspeaker arrays for active noise reduction, echo cancellation, and sound source directionality, enabling conversational control of surgical devices with natural language interaction and user-specific command mapping.
Facilitates hands-free operation of surgical systems, reduces personnel requirements, and minimizes contamination risks by accurately identifying and prioritizing voice commands from different staff members, enhancing surgical efficiency and safety.
Smart Images

Figure 0007808611000001 
Figure 0007808611000002 
Figure 0007808611000003
Abstract
Description
[Technical Field]
[0001] Priority claims This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 146,126, entitled "VOICE-CONTROLLED SURGICAL SYSTEM," filed February 5, 2021, inventors Steven T. Charles and Paul R. Hallen, which is incorporated by reference in its entirety as if fully and completely set forth herein.
[0002] The present disclosure relates generally to surgical devices and systems, and more particularly to voice-activated control systems for surgical devices and systems. [Background technology]
[0003] Many surgical procedures, including ophthalmic procedures such as refractive cataract surgery and vitreoretinal surgery, are extremely challenging and require multiple surgical staff members to coordinate various equipment in the operating room, such as lighting, the operating table, microscopes, display devices, and surgical tools and / or consoles. The presence of multiple surgical staff allows surgeons to continue their work without having to shut down and change settings on desired equipment. However, simultaneous and seamless operation of separate devices or systems by multiple surgical staff is a significant challenge during surgical procedures, particularly ophthalmic procedures. Furthermore, the additional personnel add to the cost of the procedure and additional strain on operating room resources, such as floor space, while also increasing the risk of contamination within the operating room.
[0004] In recent years, voice-activated applications have been utilized to alleviate some of the challenges of complex surgical procedures, enabling voice control of surgical devices without the need for physical interaction by surgical staff, thereby reducing the number of personnel required for a surgical procedure. Nevertheless, current voice-controlled surgical devices and systems have several limitations, such as an inability to recognize, distinguish, and prioritize personnel providing voice commands to control the devices. Additionally, in certain instances, the number and spatial placement of microphones in noisy operating rooms is suboptimal, leading to non-detection or misinterpretation of voice commands from surgical staff due to issues with background noise and sound intelligibility. Furthermore, in certain instances, surgical staff must learn a predetermined command input syntax to effectively execute desired device functions, rather than using spoken or natural language commands. Summary of the Invention [Problem to be solved by the invention]
[0005] Therefore, there is a need in the art for an improved voice-controlled surgical system. [Means for solving the problem]
[0006] The present disclosure relates to surgical devices and systems, and more particularly to voice-activated control systems for surgical devices and systems.
[0007] According to certain embodiments, a surgical command system is provided. The surgical command system includes a processor, one or more microphones configured to convert sound waves within the surgical environment into one or more audio input signals relayed to the processor, one or more loudspeakers configured to generate sound waves within the surgical environment based on one or more audio output signals received directly or indirectly from the processor, and a memory in data communication with the processor, the memory including executable instructions. The processor is configured to execute instructions to cause the surgical command system to receive, directly or indirectly, one or more audio input signals from the one or more microphones, identify one or more speech commands in the one or more audio input signals, map at least one of the one or more speech commands to a user within the surgical environment, and identify one or more actions associated with at least one of the one or more speech commands. The processor is further configured to indicate the one or more actions to a surgical device, causing the surgical device to perform the one or more actions, generate one or more audio output signals based on the one or more actions, and generate an outgoing speech response to the one or more loudspeakers based on the one or more audio output signals.
[0008] So that the above-described features of the present disclosure can be understood in detail, a more particular description of the present disclosure briefly summarized above can be had by reference to embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only exemplary embodiments and should not be considered as limiting the scope thereof, as other equally effective embodiments may be recognized. [Brief explanation of the drawings]
[0009] [Figure 1] 1 illustrates a surgical setting having a voice-controlled surgical command system according to certain embodiments of the present disclosure. [Figure 2]1 shows a schematic diagram of a voice-controlled surgical command system according to certain embodiments of the present disclosure; [Figure 3] 3 illustrates exemplary components of the surgical command system of FIG. 2 in accordance with certain embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0010] To facilitate understanding, the same reference numerals are used wherever possible to indicate identical elements that are common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further description.
[0011] In the following description, details are set forth as examples to facilitate understanding of the disclosed subject matter. However, it should be apparent to those skilled in the art that the disclosed implementations are examples and do not encompass all possible implementations. Therefore, it should be understood that reference to the described examples is not intended to limit the scope of the present disclosure. Any changes and further modifications to the described devices, apparatuses, and methods, and any further applications of the principles of the present disclosure, are fully contemplated as would normally occur to one skilled in the art to which the present disclosure pertains. In particular, it is fully contemplated that features, components, and / or steps described with respect to one implementation can be combined with features, components, and / or steps described with respect to other implementations of the present disclosure.
[0012] Embodiments of the present disclosure generally relate to a voice-controlled system for controlling devices and systems in a surgical setting, such as an ophthalmic surgical setting. In certain aspects, the voice-controlled system includes one or more phased microphone arrays (phased microphone array refers to multiple microphones arranged in a phased array) and one or more phase-loud speaker arrays (phase-loud speaker array refers to multiple loudspeakers arranged in a phased array). In certain aspects, the one or more phased microphone arrays are distributed throughout the surgical setting to identify and receive voice commands from surgical staff. Additionally, the one or more phase-loud speaker arrays may be distributed throughout the surgical setting to output voice responses and other audible signals to the surgical staff. In certain aspects, the one or more phased microphone arrays are coordinated and synchronized to perform active noise reduction, echo cancellation, and sound source directionality determination. In certain aspects, the voice-controlled system performs tasks such as activating surgical devices primarily through conversational or natural language interaction with surgical staff via one or more phased microphone and loudspeaker arrays. In certain aspects, the voice-controlled system is configured to decipher, learn, recognize, and prioritize verbal commands from different surgical staff, and further provides for pre-programming of device settings and surgical staff presets.
[0013] As used herein, the term "surgical setting" can refer to any environment in which a surgical procedure is performed. For example, the term "surgical setting" can refer to an operating room with one or more surgeons and surgical staff involved in the surgical setting.
[0014] As used herein, the term "surgical system" can refer to any surgical system, console, or device for performing a surgical procedure. For example, the term "surgical system" can refer to a surgical tool or system, such as a phacoemulsification console, a laser system, an imaging system, an intraocular lens (IOL) alignment system, a biometer, an optical coherence tomography (OCT) machine, or a vitrectomy console.
[0015] The devices and systems described herein are generally described with reference to ophthalmic surgical settings, but may be implemented in other settings and contexts, such as other surgical settings, without departing from the scope of this application.
[0016] As used herein, the term "about" may refer to a ±10% variation from the nominal value. It is understood that such a variation may be included in any value provided herein.
[0017] 1 illustrates an example of a surgical setting 100, such as an ophthalmic surgical setting, with a surgeon 150, one or more additional surgical staff, and a patient 112 having a voice-controlled surgical command system 102 in accordance with certain embodiments of the present disclosure. While one surgeon 150 is shown in FIG. 1, multiple surgeons and / or surgical staff may use the surgical command system 102. For example, in certain embodiments, surgical assistants and / or circulating nurses (i.e., circulators) may also be present in the surgical setting 100 and utilize the surgical command system 102.
[0018] As shown, the surgical command system 102 includes a surgical command controller 104 that directly or indirectly communicates with one or more surgical systems, consoles, and / or devices within the surgical setting 100 (e.g., integrated into an internet-enabled surgical suite), such as a surgical table 120, a surgical console 122, a heads-up display 124, and a microscope system 126. Examples of suitable surgical systems that may be included in a surgical suite include surgical consoles for performing vitreoretinal surgery, cataract surgery, corneal transplants, glaucoma surgery, LASIK (laser-assisted keratomileusis) surgery, refractive lens exchange, trabeculectomy, and refractive surgery, among other consoles, imaging devices, laser devices, diagnostic devices, and accessories identifiable by one of ordinary skill in the art.
[0019] In certain embodiments, the surgical command controller 104 is a stand-alone device or module (including a processor and memory) that communicates wirelessly or wired to one or more surgical systems physically located within the surgical setting 100. However, in certain other embodiments, the surgical command controller 104 includes one or more processors and / or memory integrated into one or more of the surgical systems physically located within the surgical setting 100. For example, the surgical command controller 104 may be integrated into the surgical console 122, heads-up display 124, and / or microscope 126, as shown by the virtual element 104 in FIG. 1 . In certain aspects, the surgical command controller 104 refers to a set of software instructions configured to be executed by a processor associated with at least one of the surgical systems physically located within the surgical setting 100. In certain aspects, the operation of the surgical command controller 104 may be performed in part by a processor associated with the surgical command controller 104 and in part by a processor residing in a private or public cloud.
[0020] The surgical command system 102 further includes one or more microphones 106 arranged in a phased array 136 and one or more loudspeakers 108 arranged in a phased array 138. The microphones 106 and loudspeakers 108 communicate wirelessly or via wires to the surgical command controller 104, thereby enabling the controller 104 to receive voice commands provided by the surgeon 150 and other surgical staff and to generate directional audible responses to the voice commands. In certain embodiments, the microphones 106 and / or loudspeakers 108 are distributed within the surgical setting 100 in close proximity to the desired users (e.g., the surgeon 150, surgical assistants, and / or circulating nurses) who will receive the voice commands. However, in certain embodiments, the microphones 106 and / or loudspeakers 108 are widely distributed within the surgical setting 100 to provide wider coverage of the surgical setting 100.
[0021] Similar to the surgical command controller 104, the microphone 106 and / or loudspeaker 108 may be a stand-alone device or may be physically integrated into one or more other surgical systems within the surgical setting 100. For example, the microphone 106 and / or loudspeaker 108 may be physically integrated into various components of the surgical console 122, head-up display 124, and / or microscope 126, as shown by virtual elements 106 and 108 in FIG. 1. In certain embodiments, one or more phased microphone arrays 136 orient (i.e., face) toward one or more users within the surgical setting 100 for directional listening, thus enabling focusing of desired sounds (i.e., voice commands) and suppression of undesired sounds during capture and identification of voice commands in noisy environments. For example, in certain embodiments, the microscope 126 may include at least one set of “forward-facing” microphones 106 for listening to the surgeon 150 and at least another set of “peripheral” microphones 106 oriented 90° or 180° relative to the forward-facing microphones 106 for listening to surgical staff members (e.g., during a surgical procedure, a surgical assistant may be positioned 90° or 180° from the surgeon 150 relative to the microscope 126). In further embodiments, one or more sets of microphones 106 may be physically integrated into a surgical face mask 130, headset, cap, lavalier, visor, glasses, or other accessory worn by the surgeon 150 or other surgical staff during a surgical procedure. In such embodiments, the microphones 106 may be disposable microphones or combination reusable and disposable microphones configured for use with replaceable disposable filters.
[0022] The microphones 106 and / or loudspeakers 108 may include any suitable microphones and / or loudspeakers arranged in a phased array to perform beamforming (e.g., beamsteering) to facilitate directional listening, sound source localization and speech recognition, directional audio output (e.g., text-to-speech output), and improved audio signal quality. For example, the microphones 106 may perform receive beamforming, or receive-side beamforming, while the loudspeakers 108 may perform transmit beamforming, or transmit-side beamforming. In certain embodiments, the beamforming microphones 106 arranged in a phased array 136 may enable the surgical command controller 104 to continuously detect and localize the position of a desired sound source (e.g., a user providing a voice command) from among many sources, and further capture and amplify sound waves emitted by the source while reducing or ignoring background noise, reverberation, and feedback, improving signal-to-noise ratio and speech recognition accuracy. For example, during a typical ophthalmic surgery, the surgeon 150 sits near a microscope or other surgical device either to the side or above the patient's head, while a surgical assistant is near the side of the patient's head and a circulating nurse moves between several different locations within the surgical setting 100. In such an example, the microphone 106 can detect that a given voice command originates from one of the above-mentioned locations, thereby enabling the surgical command system 102 to identify the source of the voice command as either the surgeon 150, the surgical assistant, or the circulating nurse.
[0023] Additionally, the loudspeaker 108, in synchronization with the microphone 106, enables the surgical command controller 104 to issue directional audible responses, outputs, signals, and alerts to desired users within the surgical setting 100. For example, upon detecting a voice command and the location of either the surgeon 150, surgical assistant, or circulating nurse issuing the voice command, the surgical command controller 104 may optionally transmit a directional text-to-speech response toward the user, which in certain embodiments may be preceded by the user's name or other identifier. The directional response is output from the loudspeaker 108 and, in some examples, is directed via beamforming toward the user issuing the voice command, thereby increasing the likelihood that the user issuing the command will be able to hear the response more clearly compared to, for example, others in the room. The directional response is also beneficial to the user issuing the command because the response may appear to originate from the system or device intended by the user issuing the command. Thus, the directionality of responses and other outputs promotes improved hearing by surgical staff within the surgical setting 100 and further reduces the likelihood of eliciting a response from or disturbing the patient 112.
[0024] Further, to improve the quality of the audio received by the microphone 106, the microphone 106, the loudspeaker 108, and the surgical command controller 104 are synchronized to perform active noise reduction ("ANR") to eliminate or reduce continuous background noise, such as noise caused by heating, ventilation, and cooling (HVAC) systems and surgical systems and / or computer cooling fans. Additionally, the microphone 106, the loudspeaker 108, and the surgical command controller 104 are configured to synchronize and perform reverberation suppression or echo cancellation to more clearly capture voice commands provided by the surgeon 150 and / or surgical staff. In certain embodiments, the microphones 106 in the surgical setting 100 are tuned for high-frequency and / or low-frequency sounds and may be utilized independently or in combination with other microphones 106. Similarly, the loudspeakers 108 in the surgical setting 100 may be utilized independently or in combination with other loudspeakers 108.
[0025] In certain embodiments, microphones 106 and / or loudspeakers 108 distributed within the surgical setting 100 may be utilized by the surgeon 150 and / or other surgical staff to listen to music and to make or receive phone calls. In such embodiments, the microphones 106 and / or loudspeakers 108 may be connected to a user's (e.g., the surgeon's) mobile device via a Bluetooth® connection. When a user makes or receives a phone call, the surgical command system 102 may prioritize the call and automatically mute the music playing through the loudspeakers 108.
[0026] In a further embodiment, the surgical command system 102 includes one or more microphones 106 and / or loudspeakers 108 pointed toward and positioned near the patient 112, shown reclining on the operating table 120 in FIG. 1. For example, the one or more microphones 106 and / or loudspeakers 108 (arranged as individual devices or in a phased array) may be positioned under or integrated into the operating table 120, the patient's headrest, and / or the drape 114 that covers the patient 112 during the surgical procedure. The microphones 106 and / or loudspeakers 108 may provide a communication channel between the patient 112 and surgical staff members inside or outside the surgical setting 100, such as the surgeon 150, surgical assistants, and / or circulating nurses, to improve intelligibility during communications therebetween and to accommodate hearing loss or removed patient hearing aids. For example, the microphone 106 and / or loudspeaker 108 may enable the surgeon 150 to speak more clearly to the patient 112 to provide instructions and information to the patient 112 or to calm the patient 112 during a surgical procedure. In certain embodiments, the loudspeaker 108 may be utilized to provide soothing music to the patient 112, further reducing the chance that the patient 112 will hear conversations between surgical staff members. In such embodiments, music for the patient may be provided on a separate channel from music provided to the surgeon 150 and / or other surgical staff, thus allowing the patient's music to continue while making or receiving phone calls via other microphones 106 and / or loudspeakers 108 within the surgical setting 100.
[0027] While the microphones 106 and loudspeakers 108 are generally described above as being arranged in a beamforming phased array, individual directional microphones 106 and loudspeakers 108 placed near the surgeon 150, surgical assistants, circulating nurses, and / or patient 112 (e.g., under the drape 114) are also within the scope of this disclosure.
[0028] As discussed in further detail below with reference to Figures 2 and 3, the surgical command system 102 interfaces (e.g., wirelessly or wired) with one or more devices and / or systems within the surgical setting 100 (e.g., surgical console 122, head-up display 124, and microscope 126) to perform one or more actions to operate the devices and / or systems based on voice commands provided by a surgeon 150 or other surgical staff within the surgical setting 100, which voice commands are received by one or more phased microphone arrays 136 distributed within the surgical setting. In certain embodiments, the surgical command system 102 can control various devices and / or systems within the surgical setting 100 (e.g., start, stop, change operating parameters, etc.), adjust device settings, navigate a display dashboard (graphical user interface (GUI) / human-machine interface (HMI)), control diagnostic devices, and provide alerts and / or advice to the surgeon 150 and other surgical staff via the phased array of loudspeakers 108. With speech interaction, the surgical command system 102 essentially enables "hands-free" surgical system control, thus improving surgical efficiency by reducing the amount of movement and / or "hands-on" device manipulation required for the surgeon 150 and other surgical staff whose hands are already occupied with other tasks.
[0029] 2 shows a schematic operational diagram 200 of a voice-controlled surgical command system 102, in accordance with certain embodiments of the present disclosure. As previously described, the surgical command system 102 may be located in a surgical setting 100, such as an operating room, with one or more users 210 (e.g., a surgeon 150 and / or surgical staff). During a surgical procedure, the user 210 issues voice commands 220. The voice commands, along with other sounds in the surgical setting 100 (e.g., ambient noise, speech from other users), are captured (e.g., picked up) by the phased microphone array 136 of the surgical command system 102. The voice commands 220 may be simple phrases, such as, but not limited to, "start," "stop," "increase," "decrease," or the voice commands may be complex phrases or sentences, such as, but not limited to, "increase 10%," "decrease 10%," "1 millimeter to the left," "1 millimeter to the right," or even more complex phrases and / or sentences. In certain embodiments, the voice commands 220 have a layered command architecture, where the user 210 selects a desired system, a desired tool (e.g., device), a desired tool mode, and / or a desired task to be performed by the tool while in, for example, a tool mode. Each layer of the layered command architecture (e.g., system, tool, tool mode, and task) may correspond to a set of instructions (e.g., software instructions) that cause the corresponding surgical system to perform one or more actions during surgery. In still further embodiments, the voice commands 220 are natural language voice commands 220 that are interpreted via a natural language processing (NLP) module in the surgical command controller 104, with or without pre-programming, as described in further detail below.
[0030] It should be noted that while the exemplary voice commands 220 described above are in English, the surgical command controller 104 may be configured to support any number of suitable languages, including but not limited to English, Mandarin, Hindi, Spanish, French, Arabic, Portuguese, Russian, etc.
[0031] As described above, the phased microphone array 136 receives sound waves of the voice command 220 and other sounds within the surgical setting 100 and converts the sound waves into one or more audio input signals 230, which are then relayed directly or indirectly to the surgical command controller 104 of the surgical command system 102. Upon receiving the audio input signals 230, the surgical command controller 104 identifies the voice command 220 in the audio input signals 230 via a speech recognition module, identifies the source of the voice command 220 via a user 210 location detection and / or user identification module, analyzes the voice command 220 via, for example, an NLP module, and maps the voice command 220 to the user 210 (e.g., a user profile) and a predefined (e.g., user-defined) rule set for the user 210. The predefined rule set determines what instructions are sent to the corresponding surgical system to perform the desired action indicated by the voice command 220. In certain embodiments, the predefined rule sets may include user-defined actions to be performed by specific voice commands, as well as user-preferred system settings, tool modes, tool sub-modes, task parameters, etc. Because multiple users may be present within the surgical setting 100 and may be using the surgical command system 102 simultaneously, the surgical command controller 104 is configured to identify (e.g., recognize) and distinguish between voice commands from each user within the surgical setting 100.
[0032] Voice identification is possible in part due to directional listening with one or more beamforming and phased microphone arrays 136 distributed within the surgical setting 100 to facilitate sound source localization (i.e., location detection), and suppression of unwanted operating room noise, such as speech by surgical staff other than the surgeon 150, in certain circumstances. Additionally, the surgical command controller 104 further includes a user identification module, described in more detail below with reference to FIG. 3, which works in conjunction with the NLP module and may be pre-programmed prior to a surgical procedure using voice recognition. Pre-programming of the user identification module may include a series of short natural language conversations initiated between the surgical command system 102 and one or more users within the surgical setting 100, with the surgical command controller 104 learning the speech patterns for each of the one or more users. Subsequently during the surgical procedure, the user identification module may perform a speech recognition algorithm on the voice commands 220 picked up by the phased microphone array 136 to identify the source (e.g., user) 210 of the voice commands based on the learned speech patterns.
[0033] The ability to identify the source of each voice command allows the surgical command controller 104 to store and associate predetermined sets of commands and / or rules with each user, where each set of commands and / or rules may correspond to a different set of instructions to be performed by the corresponding device. For example, after the surgical command controller 104 analyzes and identifies a voice command 220 issued by a particular user 210, the surgical command system controller 104 can map the voice command 220 to a predetermined rule set for the user 210, which may include preset system settings, tool modes, tool submodes, task parameters, etc., for each system and / or device in the surgical setting 100. In certain embodiments, the predetermined rule set for the user 210 includes an association between simple phrases issued by the user 210 and a complex set of predetermined instructions. For example, a simple phrase such as "invert display" can cause the heads-up display image to be inverted along with a particular color preset preferred by the user 210. A predetermined set of rules for each user may be pre-programmed and stored within the surgical command controller 104, in cooperation with the voice recognition pre-programming sequence described above, which learns the speech patterns for each user. For example, during cooperation of the voice recognition and user pre-programming sequence, the surgical command system 102 may first request each user to state the name and device the user wishes to target, followed by a request for the desired system and / or device mode and numerical parameters associated with the corresponding voice command from the user.
[0034] Identifying the source of each voice command further enables the surgical command controller 104 to rank the voice commands and prioritize voice commands from particular users over other users. This may be particularly beneficial when multiple users issue voice commands simultaneously or within a short time frame. Thus, the surgical command controller 104 may store a predetermined hierarchy (e.g., user profiles) of users who can receive voice commands, and the predetermined hierarchy may be utilized to prioritize voice commands from particular users over other users. Alternatively, the surgical command controller 104 may suppress certain voice commands from users that are determined not to have priority.
[0035] The surgical command controller 104, in cooperation with a speech recognition implementation, analyzes the voice command 220 to determine its content and the intent of the user 210. In certain embodiments, the analysis includes matching the voice command 220 with one or more commands preprogrammed by the user 210. However, because the surgical command system 102 also supports natural language-type interaction with the user 210, more complex analysis can be performed by an NLP module in the surgical command controller 104 to process and decipher (i.e., understand) complex natural language uttered by the user 210. Thus, the surgical command system 102 is easy to use without requiring preprogramming or extensive syntax training by the user 210.
[0036] After analyzing and mapping the voice command 220 to the user 210, the surgical command controller 104 identifies one or more instructions associated with the voice command 220 and the user 210 based on a predetermined set of rules for the user 210. The instructions generally cause one or more actions to be performed or initiated by one or more surgical systems in the surgical setting 100, such as the operating table 120, surgical console 122, heads-up display 124, and microscope 126 shown in FIG. 1. For example, in certain embodiments, the actions include system and / or device mode selection, system and / or device parameter selection, system and / or device activation and deactivation, actions to control the operation of surgical tools, data transfer initiation, data recall, data entry, patient profile selection, surgical parameter selection or modification, video and photo control functions (e.g., record, pause, stop, snapshot), display control functions (e.g., image inversion, color preset selection), surgical note dictation control functions (e.g., record, stop, save, delete), telephone control functions (e.g., answering or making a call), and other procedural functions. In the case of ophthalmic surgery, actions may also include intraocular pressure control via injection system control, laser (eg, retinal laser) parameter selection or modification, and the like.
[0037] 2 , the surgical command controller 104 may optionally generate one or more audio output signals 240 based on the identification or non-identification of an action associated with the voice command 220, which are relayed directly or indirectly to the phased loudspeaker array 138. The audio output signals 240 are converted by the plurality of loudspeakers 108 into sound waves that form an audible response 250 that confirms whether the action was identified or not by the surgical command controller 104. For example, in certain embodiments, the loudspeakers 108 may generate the response 250, which may include a summary of the identified action (e.g., in simple or complex terms), a simple repetition of the voice command 220, or a request for further information, clarification, or verification from the user 210, to which the user 210 may respond with another voice command. In certain embodiments, the response 250 may include a request from the user 210 to confirm or verify that the identified action is correct. In still further embodiments, the response 250 may inform the user 210 that the voice command 220 was not clearly received and / or the action was not identified by the surgical command controller 104. As previously described, the loudspeaker 108 may include any suitable beamforming loudspeaker positioned and configured to direct the response 250 to the desired user within the surgical setting 100. Thus, the sound waves (i.e., sound wavefront) of the response 250 may be directed directly toward the user 210, thereby improving the user 210 who issued the voice command's ability to hear the response 250 and confirm, modify, clarify, or further complement the voice command 220. In certain embodiments, the response 250 may also be beamformed to direct the response so that it appears as if it originated from the surgical system to which the preceding voice command 220 was directed.
[0038] In certain aspects, following providing optional response 250, surgical command controller 104 generates instructions 260 corresponding to the identified task and based on a set of predetermined commands and / or rules for user 210. The instructions 260 are provided directly or indirectly to one or more desired surgical systems 270 associated with the identified task for execution. Thus, the instructions 260 indicate the task identified by surgical command controller 104 to the appropriate surgical systems 270 and cause the surgical systems 270 to perform the identified action, thereby fulfilling the purpose of voice command 220. The instructions 270 may be provided to separate surgical systems 270 or to surgical systems 270 integrated within a single console 272, as shown in FIG. 2 . In certain embodiments, a single set of instructions 260 may cause multiple surgical systems 270 to perform one or more actions simultaneously or sequentially.
[0039] In certain embodiments, the surgical command system 102 further includes a feedback mechanism that facilitates communication between one or more surgical systems in the surgical setting 100 and the user 210 during or after an identified action. For example, in certain embodiments, the surgical command controller 104 may generate an audio output signal 280 that is converted by a plurality of loudspeakers 108 into an audible response 290 that may convey to the user 210 warning alerts, alarms, progress or status indicators, and any other information related to the performance of an action by a surgical system in the surgical setting 100.
[0040] By including the components and systems disclosed herein, surgical settings can be reliably controlled, at least in part, by a voice-activated application, thus providing hands-free operation of the surgical system during the performance of a surgical procedure. The disclosed embodiments allow the surgeon to control various functions of the surgical system without having to stop the surgical procedure to do so. Furthermore, by reducing the amount of physical interaction required to operate the surgical device, the number of personnel required for the surgical procedure can be reduced, while also reducing the potential risk of bacterial and viral infections (e.g., contamination) resulting from personnel contact with the surgical system.
[0041] Embodiments of the present disclosure advantageously provide voice control of a surgical setting through the overall use of a system that utilizes a processor and memory within the controller of the surgical system, as shown in Figure 3, which illustrates exemplary components of the surgical command system of Figures 1-2, according to certain embodiments.
[0042] FIG. 3 shows an exemplary diagram illustrating how the various components of the surgical command system 102 of FIGS. 1-2 communicate and operate together. As shown, the surgical command system 102 includes, but is not limited to, a surgical command controller 104, a microphone 106, and a loudspeaker 108. The surgical command controller 104 includes an interconnect 310, a network interface 312 for connection to a data communication network 350, and at least one I / O device interface 314 that allows various I / O devices (e.g., the microphone 106, the loudspeaker 108, and the surgical system 270) to be connected to the surgical command controller 104. The surgical command controller 104 further includes a central processing unit (CPU) 316, a memory 318, and storage 320. The CPU 316 may retrieve and store application data residing in the memory 318. Interconnect 310 transmits programming instructions and application data between CPU 316, network interface 312, I / O device interface 314, memory 318, and storage 320, etc. CPU 316 may represent a single CPU, multiple CPUs, a single CPU with multiple processing cores, etc. Additionally, memory 318 represents random access memory.
[0043] Storage 320 may be a disk drive. While storage 320 is shown as a single unit, it may be a combination of fixed or removable storage devices, such as a fixed disk drive, a removable memory card or optical storage, a network attached storage (NAS), or a storage area network (SAN). Additionally, storage 320 may contain trained voice models 332 of users within a surgical setting, including user presets 334. User presets 334 include separate rule sets associated with each user within a surgical setting, which are applied by surgical command system 102 to generate instructions to be performed by corresponding systems and / or devices in response to voice commands by the users.
[0044] The memory 318 includes a command module 322 containing instructions that, when executed by the processor, perform operations for controlling the surgical command system 102 as described in the embodiments herein. For example, according to the embodiments described herein, the memory 318 includes a speech recognition module 324 containing executable instructions for recognizing (i.e., identifying) words, such as voice commands, in an audio input signal received from the microphone 106. In addition, the memory 318 includes a user identification module 326 having a voice model trainer 330. The user identification module contains executable instructions for pre-programming voice commands, learning user speech patterns, and mapping speech identified by the speech recognition module 324 to a corresponding user. The memory 318 further includes a natural language processing (NLP) module 328 containing executable instructions for analyzing and interpreting natural language voice commands (e.g., matching natural language to tasks). Additionally, the memory 318 includes a response module 332 that includes executable instructions for generating an audio output signal based on information received from the speech recognition module 324 to enable two-way communication between the surgical command system 102 and a user.
[0045] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single elements. By way of example, "at least one of a, b, or c" is intended to cover a, b, c, ab, ac, bc, and abc, as well as any combination of multiples of the same elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other permutation of a, b, and c).
[0046] The foregoing description is provided to enable any person skilled in the art to practice the various embodiments described herein. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Accordingly, the claims are not intended to be limited to the embodiments shown herein, but are to be accorded the full scope consistent with the language of the claims.
[0047] In the claims, reference to an element in the singular is not intended to mean "one and only one," but rather "one or more," unless specifically stated otherwise. The term "some" refers to one or more, unless specifically stated otherwise. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later become known to those skilled in the art are expressly incorporated by reference herein and are intended to be encompassed by the claims. Furthermore, nothing disclosed herein is intended as a dedication to the public, regardless of whether such disclosure is expressly recited in the claims. No element of a claim shall be construed under the provisions of 35 U.S.C. 112(f) unless the element is expressly recited using the phrase "means for," or, in the case of a method claim, the element is recited using the phrase "step for." As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects. The present disclosure also includes the following inventions. The first aspect is 1. A surgical command system comprising: one or more microphones coupled to the processor and configured to convert sound waves within the surgical environment into one or more audio input signals; one or more loudspeakers coupled to the processor and configured to generate sound waves within the surgical environment based on one or more audio output signals received directly or indirectly from the processor; Memory containing executable instructions and the processor is in data communication with the memory and executes the executable instructions to provide the surgical command system with: receiving, directly or indirectly, the one or more audio input signals from the one or more microphones; identifying one or more speech commands in the one or more audio input signals; mapping at least one of the one or more speech commands to a user within the surgical environment; identifying one or more actions associated with the at least one of the one or more speech commands; indicating the one or more actions to a surgical device and causing the surgical device to perform the one or more actions; generating the one or more audio output signals based on the one or more actions; generating a transmitted speech response to the one or more loudspeakers based on the one or more audio output signals; and A surgical command system configured to cause The second aspect is The surgical command system of a first aspect, further configured to determine a location of a source of the one or more speech commands in the one or more audio input signals. The third aspect is In a second aspect, the location of the source is utilized to map the one or more speech commands to the user, the surgical command system. The fourth aspect is The surgical command system of a first aspect, wherein the processor is further adapted to actively reduce continuous ambient noise in the one or more audio input signals from the one or more microphones. The fifth aspect is The surgical command system of a first aspect, wherein the processor is further configured to perform echo cancellation on the one or more audio input signals from the one or more microphones. The sixth aspect is In a first aspect, the one or more actions include one or more of a surgical device selection, a mode selection, and a task selection. A seventh aspect is The surgical command system of a first aspect, wherein the one or more actions are identified based at least in part on a user profile associated with the user and accessible to the processor. The eighth aspect is A surgical command system in a seventh aspect, wherein the user profile includes a mapping of different actions to different speech commands, the mapping including a mapping between one or more of the different actions and one or more of the different speech commands. A ninth aspect is wherein the processor being configured to map the at least one of the one or more speech commands to the user further includes the processor being configured to map each of the one or more speech commands to a corresponding user in a group of users that includes the user; A surgical command system in a first aspect, wherein the processor is further configured to prioritize the at least one of the one or more speech commands and / or the one or more actions based on a predetermined hierarchy of the group of users, the predetermined hierarchy indicating a higher rank for the user compared to other users associated with the one or more speech commands. A tenth aspect is The surgical command system in a first aspect, wherein the transmitted speech response includes a notification of the one or more actions indicated on the surgical device. An eleventh aspect is A surgical command system according to a tenth aspect, wherein the transmitted speech response is relayed from the one or more loudspeakers in a direction of the user mapped to the identified one or more speech commands. A twelfth aspect is The surgical command system of a first aspect, wherein the one or more loudspeakers include a plurality of loudspeakers arranged in a phased loudspeaker array. A thirteenth aspect is A surgical command system according to a twelfth aspect, wherein the phase loudspeaker array is arranged to perform transmit beamforming for transmitting acoustic waves within the surgical environment. A fourteenth aspect is The surgical command system of a first aspect, wherein the one or more microphones include a plurality of microphones arranged in a phased microphone array. A fifteenth aspect is A surgical command system in a fourteenth aspect, wherein the phased microphone array is arranged to perform receive beamforming for receiving acoustic waves within the surgical environment. A sixteenth aspect is In a first aspect, the surgical command system is one in which the one or more microphones are located in a surgical mask worn by the user in the surgical environment.
Claims
1. 1. A surgical command system comprising: one or more microphones coupled to the processor and configured to convert sound waves within the surgical environment into one or more audio input signals; one or more loudspeakers coupled to the processor and configured to generate sound waves within the surgical environment based on one or more audio output signals received directly or indirectly from the processor; Memory containing executable instructions and the processor is in data communication with the memory and executes the executable instructions to provide the surgical command system with: receiving, directly or indirectly, the one or more audio input signals from the one or more microphones; identifying one or more speech commands in the one or more audio input signals; mapping at least one of the one or more speech commands to a user within the surgical environment; identifying one or more actions associated with the at least one of the one or more speech commands; indicating the one or more actions to a surgical device and causing the surgical device to perform the one or more actions; generating the one or more audio output signals based on the one or more actions; generating a transmitted speech response to the one or more loudspeakers based on the one or more audio output signals; configured to cause wherein the processor being configured to map the at least one of the one or more speech commands to the user further includes the processor being configured to map each of the one or more speech commands to a corresponding user in a group of users that includes the user; The processor is further configured to prioritize the at least one of the one or more speech commands and / or the one or more actions based on a predetermined hierarchy of the group of users, the predetermined hierarchy indicating a higher rank for the user compared to other users associated with the one or more speech commands.
2. The surgical command system of claim 1 , further configured to determine a location of a source of the one or more speech commands in the one or more audio input signals.
3. The surgical command system of claim 2 , wherein the location of the source is utilized to map the one or more speech commands to the user.
4. The surgical command system of claim 1 , wherein the processor is further adapted to actively reduce continuous ambient noise in the one or more audio input signals from the one or more microphones.
5. The surgical command system of claim 1 , wherein the processor is further adapted to perform echo cancellation on the one or more audio input signals from the one or more microphones.
6. The surgical command system of claim 1 , wherein the one or more actions include one or more of a surgical device selection, a mode selection, and a task selection.
7. The surgical command system of claim 1 , wherein the one or more actions are identified based at least in part on a user profile associated with the user and accessible to the processor.
8. 8. The surgical command system of claim 7, wherein the user profile includes a mapping of different actions to different speech commands, the mapping including a mapping between one or more of the different actions and one or more of the different speech commands.
9. The surgical command system of claim 1 , wherein the delivered speech response includes a notification of the one or more actions indicated on the surgical device.
10. The surgical command system of claim 9, wherein the outgoing speech response is relayed from the one or more loudspeakers in a direction of the user mapped to the identified one or more speech commands.
11. The surgical command system of claim 1 , wherein the one or more loudspeakers comprise a plurality of loudspeakers arranged in a phased loudspeaker array.
12. The surgical command system of claim 11 , wherein the phase-loud speaker array is arranged to perform transmit beamforming for transmitting acoustic waves within the surgical environment.
13. The surgical command system of claim 1 , wherein the one or more microphones include a plurality of microphones arranged in a phased microphone array.
14. The surgical command system of claim 13 , wherein the phased microphone array is positioned to perform receive beamforming for receiving acoustic waves within the surgical environment.
15. The surgical command system of claim 1 , wherein the one or more microphones are located in a surgical mask worn by the user in the surgical environment.
Citation Information
Patent Citations
Universal distributed surgery room control system
JP2001104336A
Steam pop detection
JP2018149294A
Surgical system with voice control
US20180168755A1
Workflow assistant for image guided procedures
US20190090954A1
Dynamic muting audio transducer control for wearable personal communication nodes
US20190105114A1