CONTROL OF SURGICAL VISUALIZATION SYSTEMS USING MULTI-MODAL USER EXPRESSION
The method enhances surgical visualization system control using multimodal user inputs with reduced latency and temporal relationships, addressing inaccuracies in existing systems.
Patent Information
- Application Number
- DE102024118626
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2026-01-08
- Estimated Expiration
- 2044-07-01
AI Technical Summary
Existing surgical visualization systems face inaccuracies due to noise in eye-tracking data and delays in speech recognition, compromising the reliability and accuracy of control.
A method for controlling surgical visualization systems using multimodal user inputs, including verbal commands and eye movements, with latency reduced to less than 0.5 seconds, and incorporating temporal relationships between user utterances to enhance precision.
Improves the reliability and accuracy of surgical visualization system control by minimizing latency and incorporating temporal relationships to reduce unintended actions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL AREA
[0001] Several examples of the disclosure relate to aspects of controlling a surgical visualization system, for example, a visualization system with an operating microscope. Several examples specifically concern the processing of multimodal user inputs occurring over different time periods, such as verbal commands and eye movements, for controlling a surgical visualization system. BACKGROUND
[0002] In surgical environments, precise and reliable control of surgical visualization systems, such as optical operating microscopes, is essential. Some systems use eye-tracking and voice commands for control. However, noise in the eye-tracking data and delays in speech recognition often lead to inaccurate control. These inaccuracies can compromise the reliability and accuracy of the surgical display systems.
[0003] Systems for controlling devices with gesture and voice control are known from the publications DE 10 2010 063 392 A1, DE 10 2015 117 824 A1, DE 10 2017 011 498 A1, DE 10 2021 006 023 B3 and US 2023 / 0 077 245 A1. SUMMARY
[0004] Therefore, there is a need for improved techniques for controlling surgical visualization systems that overcome or mitigate at least some of the aforementioned limitations and disadvantages.
[0005] This problem is solved by the features of the independent patent claims. The dependent patent claims define further advantageous embodiments.
[0006] The solution according to the invention is described below with respect to both the claimed methods for controlling surgical visualization systems and the corresponding control devices and surgical visualization systems. Furthermore, corresponding computer programs and electronically readable storage media are provided. It should be understood that features, advantages, and alternative embodiments can be assigned to the other categories, and vice versa. For example, the control devices or surgical visualization systems can be improved by features used in the described methods for controlling surgical visualization systems, and vice versa.
[0007] A computer-implemented method for controlling a surgical visualization system is provided. A surgical visualization system includes, for example, a visualization system used in surgical environments or in a surgical application to present visual information to a user.
[0008] A surgical visualization system can be used in conjunction with endoscopic, microscopic, or any other imaging or examination procedures. Such imaging systems are used, for example, in surgical settings to display visual information related to a surgical area and / or procedure to a user, for instance, in real time. Specifically, the visual information can be provided within the field of view of the surgical visualization system.
[0009] The visual information may include, for example, an image, multiple images and / or a video stream, which may be based on 2D or 3D image data, and may further include a representation of a target area (region of interest, ROI) within an operational area.
[0010] Visual information, such as the operating area and, if applicable, a surgical instrument, can be displayed to the user by one or more of the following devices: a purely optical system (e.g., comprising an eyepiece) for optical displays; a display on one or more electronic display devices (e.g., a screen or monitor) based on captured image data; a display in a virtual reality system (e.g., a virtual display in a virtual environment); a display by a 3D display device that visually represents depth information for the user; or a display in an augmented reality system that projects virtual elements into real space. These various display devices make it possible to present and access the visual information for the user.
[0011] In general, a surgical visualization system can therefore comprise a combination of hardware and software components that capture imaging data from examination devices such as endoscopic cameras, microscopes, or other imaging equipment and visually display this image data (in real time) for the user. These systems support the user by providing clear and detailed visual information, particularly of subsets of the imaging data. For example, they can display additional visual information about the actual surgical field, such as virtual instructions or navigation aids, to assist the user. Other examples include text, descriptions, icons, overlays, etc. These can be displayed, for example, optically or virtually.
[0012] The procedure for controlling the surgical visualization system includes the following steps: In one step, an initial user input is obtained. This obtaining can involve capture or detection by one or more initial sensors. The initial user input can be obtained in the form of raw sensor data or processed sensor data. It is conceivable that the initial user input could be transmitted electronically, for example, based on an external sensor measurement.
[0013] The first user utterance can be received in real time. The latency between the first user utterance and the receipt of corresponding data that indexes the user utterance can be, for example, less than 0.5 seconds, optionally less than 0.05 seconds, and further optionally less than 0.005 seconds.
[0014] The first user utterance can include an action or instruction from a user, or more generally, information regarding a user who is using the surgical visualization system to display visual content. The first user utterance can be captured to control the surgical visualization system.
[0015] The first user utterance may include an intentional, or possibly an unintentional, action that extends over a time interval of a period, in particular over the entire duration of the time interval.
[0016] In some examples, the first user utterance may include one or more of the following: a verbal utterance by the user, a gesture of a body part of the user, a touch gesture at a touch interface, a brain-computer interface (BCI) signal, or a multimodal combination of the above.
[0017] The first user utterance can serve to express a user intention or purpose. For example, the first user utterance can serve to determine what kind of control action the user expects from the surgical visualization system.
[0018] The first user utterance corresponds to a first user utterance type, in other words, a first (user utterance) modality. A user utterance type can refer to a category or type of user action that the user performs.
[0019] Examples of initial user utterances that can express user intent include verbal utterances from which, through processing, voice commands such as "Enlarge this area," "Focus on this structure," "Zoom in here," or "Show this as an overview view" are derived. These verbal utterances can indicate which display change the user desires. Gestures of at least one part of the user's body, for example, to define a section of the image, can also express user intent. Another example would be gestures performed using an instrument or tool, such as a surgical tool. For instance, a pointer could be used to define the initial user utterance.
[0020] Gestures involving at least one part of the user's body can take various forms and involve a multitude of body parts. Hand gestures are a common example, where the user moves or positions their hands or fingers in a specific way to indicate an intention or action. For instance, a particular body part or one or more instruments might be moved into a specific area or zone and then remain there. Hand or finger movements according to specific shapes or patterns are also conceivable. Winking patterns or gestures on a touchscreen are likewise possible.
[0021] The first user utterance can serve to determine an initial user input for the surgical visualization system, specifying an action, particularly a control action, of the surgical visualization system. This initial user utterance can be processed to determine, for example, a desired control action for the surgical visualization system. The processing of a user utterance to determine a user input can be rule-based or machine learning-based, for instance.
[0022] Based on the initial user utterance, an initial user input can be determined. This user input might include, for example, a user intention. The user input could, for instance, include a control command for a desired action within the surgical visualization system.
[0023] The first user utterance spans a time interval. A time interval refers to a specific period of time during which the user utterance occurs. The first user utterance therefore takes place throughout the entire time interval; it extends over the entire time interval, i.e., from the beginning to the end of the time interval. Accordingly, the surgical visualization system detects the first user utterance via sensors throughout the entire time interval, i.e., from the beginning to the end of the time interval.
[0024] Time information that characterizes the time interval, such as a start time, an end time, and / or a midpoint, can also be recorded and stored. For example, the start time indicates the beginning of the time interval, the end time indicates its end, and the midpoint indicates its middle. The start time, end time, and midpoint are exemplary reference points for the time interval. Furthermore, the length of the time interval can also be recorded and stored as time information.
[0025] The time interval lies within a larger period, i.e., it is contained within the period. The period encompasses the time interval. The period can be the same length as the time interval or longer.
[0026] In a further step, a variety of different second user utterances are received, for example, captured or detected, which are assigned to at least one second user utterance type. It is possible that the second user utterances are transmitted electronically by at least one second sensor.
[0027] The terms "first user utterance" and "multiple second user utterances" are not intended to imply any order or hierarchy between the first and second user utterances. For example, it would be conceivable that the multiple second user utterances are received before the first user utterance.
[0028] Every second user utterance can be received in real time. The latency between each second user utterance and the receipt of the corresponding data that indexes this second utterance can be, for example, less than 0.5 seconds, optionally less than 0.05 seconds, and further optionally less than 0.005 seconds.
[0029] Each of the numerous second user utterances is associated with a specific piece of time information. This time information can be indicative of a point in time or a duration during which the second user utterance occurred. Such time information is also preserved. In principle, the time information can be preserved implicitly or explicitly. For example, timestamps could be obtained. However, it would also be conceivable to deduce the specific times at which a particular second user utterance occurred based on a sampling rate and the sequence of the second user utterances.
[0030] The timing information can, in particular, indicate how each second user utterance is positioned in relation to the time interval over which the first user utterance extends. Based on the timing information and the time interval, a temporal relationship between each of the second user utterances and the first user utterance can therefore be determined.
[0031] The second user utterances are therefore temporally related to the first user utterance, as can be indicated by the temporal relation. For example, there may be an intentional or content-related connection between the first and second user utterances, which is represented by the temporal relation. For example, the first and second user utterances may be in a shared action or interaction context of a user-system interaction, which is represented by the temporal relation. For example, the first and second user utterances may be related to a shared user intention.
[0032] The procedure can involve determining the temporal relationship between each of the multitude of second user utterances and the first user utterance, which spans the time interval. This temporal relationship is determined based on the time information available for each of the second user utterances.
[0033] The second user utterances of the multitude of second user utterances are cached. Specifically, it is possible to cache the second user utterances until it is possible to determine, based on the time information for the second user utterances, temporal relationships between each of the multitude of second user utterances and the time interval spanned by the first user utterance. For example, it is conceivable that one or more second user utterances are received before the end of the time interval; in such cases, it may not yet be possible, at least in some examples, to definitively determine their temporal relationship to the time interval (which is not yet over or whose end will occur at an unknown future time).Accordingly, it is then possible to buffer the second user utterances until it is finally possible to determine temporal relationships between each of the multitude of second user utterances and the time interval. Furthermore, it is possible to buffer the second user utterances for a maximum or exactly a predefined or determinable buffer duration.
[0034] In some examples, the second user utterances and their corresponding time information can be associated and temporarily stored in a buffer data structure (buffer) for a specific buffer duration. This buffer data structure can, for example, be a FIFO (First-In-First-Out) storage system. The size of the buffer, or the buffer duration for which data is retained, can be predefined or determined, for example, based on the type of the first user utterance. The buffer duration can be implemented as a sliding time window. In particular, the buffer duration can be chosen to be larger than a permissible maximum value for the time interval of the first user utterance.
[0035] Next, aspects related to the second user utterances are described. The at least one second user utterance type differs from the user utterance type of the first user utterance.
[0036] The second user utterances are distributed across the (larger) time period. In other words, the multitude of second user utterances is distributed across the time period. This means that the second user utterances occur at different times within the (larger) time period and are appropriately recorded. The time information associated with the second user utterances indicates how the second user utterances are distributed across the time period.
[0037] The second user utterances comprise different second user utterances of each of at least one second user utterance type.
[0038] Second user utterances can be captured by at least one second sensor or multiple second sensors, for example, using sensor fusion, to control the surgical visualization system. The second sensor can be different from the first sensor. However, it is also conceivable that the second sensor could encompass the first sensor, allowing the first sensor to also be used to capture the second user utterances. The second user utterances can also include a processed signal / sensor data from one or more sensors.
[0039] The second user utterances can encompass multiple different actions or states of a user, i.e., different information regarding the user, occurring at different times throughout the period. The second user utterances occur multiple times throughout the period, i.e., at different times. The multitude of second user utterances comprises different (i.e., varying) second user utterances of the same second user utterance type to control the surgical visualization system.
[0040] For example, several of the second user utterances may fall within the time interval covered by the first user utterance. These second user utterances, which fall within the time interval covered by the first user utterance, may vary from one another.
[0041] In some examples, the second user utterances can serve to specify a target point, or target region, in the image data for the user intent of the first user utterance.
[0042] For example, the second user utterances can include the user pointing to a target region. The target region can be located within the field of view of the surgical visualization system. The field of view can be a displayed or the actual field of view of the surgical visualization system. For example, the user might point to a representation of the target region in an image captured by the surgical visualization system. However, it would also be possible for the user to point directly at the target region, for example, with a finger or a surgical instrument. In the latter case, the pointing can be observed in an image captured by the surgical visualization system.
[0043] Secondary user expressions can include both intentional and unintentional actions or states of the user, such as gazes, postures or body orientations, positions and / or orientations of at least one part of the user's body (e.g., within a coordinate system of the surgical visualization system, or particularly within a field of view of the surgical visualization system), touches at a mechanical interface or a touch interface, or other information about the user briefly captured by sensors. These secondary user expressions serve to determine secondary user inputs for the surgical visualization system, which it then uses to trigger or parameterize a control action determined by the primary user expression.
[0044] For example, it would be conceivable for the second user utterance to indicate the user's gaze direction. A user's gaze direction can be determined, for instance, by pupil tracking (also known as eye tracking). It would be conceivable to determine the gaze direction with a specific sampling rate, chosen high enough to detect even rapid changes in gaze direction. For example, the sampling rate for determining the gaze direction could be at least 100 Hz, or optionally at least 200 Hz. This would allow even rapid, jerky eye movements, such as those made by moving the eyes between two points to realign the field of vision, to be captured. Such movements typically occur within 20 to 40 ms due to physiological factors.Accordingly, a large number of second user utterances will be obtained if a person's gaze direction is determined at a sufficiently high sampling rate. It should be understood that it is not necessary for every second user utterance to specify a different user input, such as a different eye direction. In other words, different second user utterances can be the same or at least comparable to the multitude of second user utterances (e.g., if the eyes gaze in one direction for a longer period).
[0045] In preferred examples, the first user utterance may comprise a verbal utterance by the user, and the second user utterances may comprise at least a first group of second user utterances encompassing the user's gaze direction, and at least a second group of second user utterances encompassing the position and / or orientation of a surgical instrument operated by the user. Second user inputs can be determined with greater accuracy from such a multi-modal combination of second user utterances of different second user utterance types.
[0046] Second user utterances can be captured at specific times or with a predetermined sampling rate, so that different second user utterances exist, or at least can exist, at these times within the time period. The sampling rate can enable the determination of timing information for each second user utterance.
[0047] As described above, different second user utterances (containing different time information) can be different from each other or identical. In any case, a chronological sequence of second user utterances of the same second user utterance type, captured at different times, can be obtained in this way.
[0048] The second user utterance can provide additional parameters or context information for control based on the first user utterance.
[0049] The second user utterances can, for example, specify target locations, target areas, directions, or selection areas in connection with a control action for the display of the surgical visualization system.
[0050] For example, the second user utterances may include a direction of gaze or the user's gaze position, which can easily be converted into one another, for example in relation to a representation of the surgical visualization system or another reference coordinate system (perhaps defined by a represented or actual field of view of the surgical visualization system).
[0051] The position and / or orientation of the user's head or part of the head can also be captured as a second user utterance of a further second user utterance type. Similarly, the position and / or orientation of a body part, e.g., at least one of the user's fingers, can be captured as a second user utterance.
[0052] Another example of a second user utterance can include the position and / or orientation of an instrument at a specific time. For example, the position and / or orientation of the surgical instrument can be determined relative to the field of view of the surgical visualization system or relative to the surgical area. Furthermore, the position of a surgical instrument, or at least a part of it (e.g., the tip), can be determined, for example, by sensors in the machine coordinate system (i.e., in 3D world coordinates) or another reference coordinate system. Such a position of a surgical instrument can also be determined, for example, from image data of an imaging system of the surgical visualization system that depicts at least that part of the surgical instrument.
[0053] Second user utterances can be shorter relative to the first user utterance. Second user utterances can be captured more frequently relative to the first user utterance. For example, the time interval between two consecutive second user utterances can be less than 30%, or less than 20%, or less than 10%, or less than 5%, or less than 1% of the time interval covered by the first user utterance. Capturing a (optionally each) second user utterance can require less time compared to capturing the first user utterance. For example, less than 50%, or 30%, or 20%, or 10%, or 5%, or 1% of the time required to capture the first user utterance. Second user utterances can be recorded at a higher temporal density than the first user utterance. Second user utterances can be less complex than the first user utterance.Second user utterances can be recorded at a higher temporal density than the first. They can also be captured at shorter intervals than the duration of the first utterance. These utterances complement the user intent determined based on the first utterance, providing details for precise control of the visualization system. Second user utterances can also be used to determine the initial user input. For example, they can serve as continuous feedback mechanisms, enabling the system to respond to dynamic changes in user behavior in response to the determination and execution of the control action.
[0054] In some examples, each of the second user utterances has its own temporal extent, defined by a specific duration. Each of these durations is shorter than the first-mentioned time interval over which the first user utterance extends. Thus, each of the second user utterances, considered individually, has a shorter duration than the first user utterance. The respective durations of the second user utterances can be represented by the time information or timestamps associated with each of them.
[0055] Preferably, the length of each of the time intervals may be less than 50% of the length of the first-mentioned time interval, particularly preferably less than 10%.
[0056] Second user utterances can be used to provide spatial and / or temporal parameters for a control action. Corresponding second user inputs can be defined for one or more second user utterances. For example, it would be conceivable to define a corresponding second user input for every second user utterance in a multitude of second user utterances, and then, for instance, temporarily store it.
[0057] In some examples, every second user input can be determined by second user utterances of more than one, more than two, or more than three different second user utterance types.
[0058] For example, both a gaze direction and an instrument position can be processed together (as second user inputs of different user utterance types) to determine a target position of the user (user input).
[0059] The second user utterances and / or user inputs can supplement or characterize the first user input. The second user utterances and / or user inputs can parameterize an action, in particular a control command or a control action, of the surgical visualization system.
[0060] For example, at least one second user input can include temporal control parameters for the control, such as controlling the timing of a control action of the surgical visualization system, for example, defining the start or end of the control action, or parameterizing and / or triggering individual phases of the control action.
[0061] The procedure further includes prioritizing at least one second user utterance from the multitude of second user utterances based on temporal relationships. Subsequently, the surgical visualization system can be controlled based on the first user utterance and the prioritized second utterance. In particular, it would be possible for the surgical visualization system to be controlled based on at least one second user input, which is determined based on the prioritized second utterance.
[0062] Prioritizing at least one second user utterance can involve selecting that utterance. This means that only the selected second utterance is considered when controlling the surgical visualization system; unselected second utterances are ignored. Alternatively, prioritizing a second user utterance could involve giving greater weight to that utterance compared to unprioritized utterances. For example, user input could be defined for different second utterances; such input could, for instance, indicate a specific position.A weighted average of such positions could then be determined.
[0063] The temporal relationships of all second user utterances can therefore be used to perform the control, whereby either at least one of the second user utterances is used for control based on its temporal relationship, and / or a different control parameter is determined based on the temporal relationship of the second user utterances, and / or several (e.g. all) second user utterances are weighted differently for control, and / or certain second user utterances are excluded from control based on their temporal relationship.
[0064] By prioritizing at least one second user utterance based on its temporal relationship, it is possible to prevent an unintended second user utterance from being considered for controlling the surgical visualization system during a relatively long time interval over which the first user utterance extends. This avoids unintended control actions. In detail, several techniques described herein are based on the understanding that with relatively long time intervals over which the first user utterance extends—typically the case with voice input, which can last several seconds—some users may have difficulty consistently maintaining uniform second user utterances occurring on a shorter timescale.An example would be the combination of speech input (first user utterance) with gaze direction (second user utterances). During the speech input, which might last, for example, 3 to 5 seconds, the user may be tempted to briefly change their gaze direction. The gaze direction can be captured at a sampling rate of 100 Hz, so that the brief change in gaze direction is reflected in different second user utterances during the time interval. By considering the temporal relationship between the different second user utterances and the first user utterance, it is possible, for example, to specifically identify second user utterances that occur at the beginning, shortly before the end, or at the end of the first user utterance. Other second user utterances, such as those that occur at the beginning, shortly before the end, or at the end of the first user utterance, can be included.Events occurring at the beginning or in the middle of a time interval may not be considered or may only be given less weight. Similarly, prioritization can be applied to other temporal relationships.
[0065] A temporal relationship can represent a temporal connection between the first user utterance and the respective second user utterance. This temporal relationship can include information indicating the temporal sequence or arrangement of the user utterances. In some examples, the temporal relationship specifies the time interval between each second user utterance and the first user utterance, or, in particular, a reference point (e.g., the endpoint or midpoint) of the time interval over which the first user utterance spans.
[0066] A temporal relationship can be defined, for example, as the time interval between the (first-mentioned) time interval over which the first user utterance extends, or a first (reference) time point assigned to this time interval and / or the first user utterance, and a second time point or second time interval assigned to at least one second user utterance. This time interval can be positive, for example, if the second user utterance occurs after the first, or negative if the second user utterance occurs before the first. Therefore, the times of the respective second user utterances, or second user inputs, can be temporally related to the time interval of the first user utterance.
[0067] The temporal relationship can also describe whether the second user utterance was captured during the time interval, or at least partially during the time interval. In this case, the second user utterance occurs completely or partially within the time interval in which the first user utterance also occurs. The temporal relationship can thus indicate, for example, whether the first and at least one second user utterance occur at least partially concurrently.
[0068] For example, the temporal relationship can also be characterized by a predetermined time interval or threshold. This can specify the maximum or minimum time interval that can occur between the first and second user utterances for the second utterance to be considered relevant for control.
[0069] It is then possible to perform a check for every second user utterance. This allows verification of whether the respective temporal relationship assigned to the corresponding second user utterance fulfills one or more predefined test criteria. The selection of at least one second user utterance to subsequently be considered in the control of the surgical visualization system can then be based on the results of these checks, which are performed for all second user utterances. In other words, for example, those second user utterances can be selected for which the check yields a positive result.
[0070] Examples of such a check are described below. For instance, it could be specified that only second user utterances occurring before the end of the first user utterance are considered. Or, only those second user utterances could be used that occur at most 1 second, or at most a certain decimal multiple (e.g., 1.5 times or twice the length of the time interval spanning the first user utterance), before the start of the first user utterance.
[0071] By buffering the second user utterances along with their associated time information, flexibility is maintained to retrospectively determine and verify the temporal relationships between the various second user utterances after the first user utterance has finished—that is, after the time interval has ended. This also allows for the evaluation of temporal relationships with respect to one or more test criteria that relate to the end of the time interval—or, more generally, to a reference point that is only established once the time interval has ended. Thus, reference points determined by the end of the time interval can be used to define a specific time range. It can then be verified whether a point in time associated with a particular second user utterance falls within this time range.
[0072] Another advantage of caching the second user utterance along with the time information is the ability to further process the first and / or second utterances, for example, to determine user input for one or more of them. Speech recognition, for instance, might require a certain amount of time. Such further processing can take time due to limited computing resources. Without caching, this latency can lead to inaccuracies in determining the temporal relationships; caching prevents such inaccuracies.
[0073] For example, second user utterances could be prioritized or specifically selected if they fall within a defined time span starting from the end of the time interval. Time spans starting from the midpoint of the time interval could also be considered. For instance, second user utterances within the last 20 ms of the time interval, or within a time span extending 50 ms from the midpoint of the time interval back to the beginning of the time interval, could be prioritized or specifically selected. Furthermore, the one or more test criteria could include a check to see if a time point associated with a corresponding second user utterance, and indexed by the time information, lies within a specific time range before the end of the time interval covered by the first user utterance.
[0074] It can also be checked whether a time point associated with a corresponding second user utterance and indexed by the time information lies at most at the end of the time interval over which the first user utterance extends.
[0075] Determining the first user input based on the first user utterance typically takes some time, for example, due to limited computing resources. The point at which the first user input is determined may therefore be after the end of the time interval covered by the first user utterance. Using this test criterion, for example, those second user utterances that occur at or after the end of the time interval covered by the first user utterance can be disregarded.
[0076] In general, by considering temporal relationships and temporal criteria—that is, by determining temporal thresholds or valid time windows or ranges relative to the time interval—the search space for relevant second user utterances can be narrowed down. This allows for a more targeted selection of second user utterances that are relevantly temporally related to the first user utterance.
[0077] In some variations, additional test criteria could also be checked. For example, the number of second user comments that should occur before, during, or after the time interval of the first user comment. It could be required, for instance, that a predetermined percentage, such as at least 50%, or a predetermined number of second user comments fall within the time interval. Furthermore, it could be required that a minimum time interval ("sampling rate") between second user comments that fall within the time interval of the first user comment is maintained.Instead of the time interval of the first user utterance, an equivalent time interval can also be used as a reference, which may be shifted in a predefined way compared to the time interval of the first user utterance, in particular, at least partially, optionally completely, before the end of the time interval of the first user utterance.
[0078] For example, a weighting (as a form of prioritization) of the second user utterances could be carried out depending on their temporal position relative to the first user utterance. That is, the weighting of the second user utterances could depend on the temporal relationship of each utterance to the first. For example, those second user utterances that are closer to the beginning, middle, or end of the first user utterance could be weighted more heavily. In particular, the weighting of the second user utterances could depend on whether the respective second user utterance is closer to the beginning, middle, or end of the time interval of the first user utterance.
[0079] For example, a selection can be made of the second user utterance that is closest in time to the beginning, middle, or end of the first user utterance.
[0080] For example, a selection of related sequences of second user utterances could be made that exhibit specific temporal patterns relative to the first user utterance. For instance, a search could be conducted for a sequence that begins shortly before and ends shortly after the first user utterance.
[0081] By considering temporal relationships, particularly by specifying suitable (temporal) test criteria, the system can prioritize those second user utterances that are in a relevant temporal relationship to the first user utterance and are therefore likely to be relevant for control. The precise selection of the temporal criteria can depend on factors such as the type of the first user utterance, the application context, or user preferences. This can, for example, prevent an unintentional second user utterance from being considered for controlling the surgical visualization system during a relatively long time interval over which the first user utterance extends. Unintended control actions are thus avoided.
[0082] By considering and verifying the temporal relationships between the first and subsequent user utterances when controlling the surgical visualization system, the temporal sequence of user utterances can be incorporated into the control process. This improves the interaction between user and system, as the temporal context of the user utterances is taken into account.
[0083] Alternatively or in addition to the temporal relationship between the first user utterance and the second user utterances, temporal relationships between the second user utterances themselves can also be taken into account when controlling the surgical visualization system.
[0084] If multiple second user utterances are recorded, they are not only temporally related to the first user utterance, but also to each other. The time intervals between successive second user utterances can vary and can also be used for control purposes.
[0085] For example, the order of second user comments within a given time period can be taken into account. If no second user comments occur within a valid time window, this can also be a parameter for the control mechanism.
[0086] Even more complex temporal patterns across the entire captured sequence of second user utterances can be recognized and processed in relation to the time interval.
[0087] In some examples, the temporal relationship between all recorded second user utterances can be taken into account for control purposes, for example to select at least one second user utterance and / or to exclude at least one second user utterance.
[0088] Determining user input using captured user utterances can be done in different ways. One possibility is a rule-based approach, where predefined rules are used to recognize certain characteristics or patterns in the user utterance and derive the corresponding user input from them.
[0089] For example, certain keywords or gestures could be permanently linked to specific inputs. Other conceivable approaches include pattern recognition, statistical models, knowledge-based systems, or multimodal fusion, where information from different modalities of an initial user utterance (e.g., speech and gestures) is combined to improve input recognition.
[0090] Another way to determine user input is through the use of machine learning methods, particularly neural networks, i.e., using machine learning models. Here, the system is trained on training data to learn the correlation between user utterances and inputs. After training, the system can then generate appropriate inputs for new user utterances. Neural networks are capable of recognizing even complex user inputs in the data, for example, based on speech recognition.
[0091] Controlling the surgical visualization system can involve triggering an action within the system. The type of action can be specified by the first user input. At least one second user input can trigger and / or parameterize the action.
[0092] While the first user input specifies the type of action (e.g., "Zoom"), the second user input can initiate or trigger the execution of the action or specify a context for that action. Triggering the action increases security against user errors ("two-factor authentication"). Parameterization allows for more precise implementation of the action.
[0093] Typically, the information content of the first user input is greater than that of the second. This is because the first user input typically has to select the type of action from a relatively large candidate space: there are often many actions to be performed. In contrast, the second user input only has to confirm the execution of the specific action or set it within a limited parameter space, and such an input often has less information content (for example, clicking a button or saying "OK!" might suffice). This discrepancy in information content is one possible reason why only a single first user input occurs during the same period, while many second user inputs occur.
[0094] The multiple user inputs can define coordinates in a continuous space, where the coordinates exhibit a change over time. The process can further include applying a filter, particularly a low-pass filter or a Kalman filter, to the coordinates, thereby smoothing the change.
[0095] In some examples, second user inputs may include different types of second user utterances, such as the user's gaze direction or position on the display, or the position of the instrument tip. Such second user utterances provide information about the user's current area of interest, which can be characterized by a target position (POI) or region of interest (ROI) within the surgical field; however, they may contain noise and inaccuracies.
[0096] To obtain a more robust estimate of POl, a low-pass filter can be used in some examples.
[0097] To obtain a more robust estimate of the POI, a Kalman filter can be used, combining the information from both signals, i.e., both types of second user utterances. The Kalman filter can estimate the most probable POI, taking into account the uncertainties of the individual measurements and the dynamics of the POI over time. For this purpose, a state model can be defined that describes the position and possibly also the velocity of the POI in space. The two types of second user utterances are modeled as observations with corresponding uncertainties. Combining the data from two different types of second user utterances can provide a more accurate and stable estimate of the POI than would be possible with either signal alone.
[0098] It can be advantageous to consider at least three, or at least four, or more different second user utterance types.
[0099] A filter parameter of the filter can depend on a type of content displayed on a display of the surgical visualization system, and / or a phase of a surgical workflow, and / or the first user utterance type, and / or the second user utterance type, and / or a type of first user input determined based on the first user utterance.
[0100] A surgical visualization system is configured to execute any procedure or combination of procedures as disclosed herein. The procedure can, for example, be executed by a control unit, i.e., a shared or dedicated computing device of the surgical visualization system.
[0101] A control unit of a surgical visualization system is configured to execute any procedure or combination of procedures according to the present disclosure.
[0102] The control unit comprises a processor, memory, and an interface for receiving and providing sensor signals and / or user inputs, wherein the memory unit comprises instructions executable by the computing unit which, when executed by the computing unit, cause it to perform the steps of any method or any combination of methods according to the present disclosure.
[0103] A computer program, or computer program product, comprises instructions which, when executed by a processor, cause the processor to perform the steps of any method or combination of methods as disclosed herein.
[0104] The described techniques can control a surgical visualization system to process and display image data based on user input, thus providing imaging support for a surgical procedure. This includes, for example, targeted navigation within an image dataset, such as a 2D or 3D image of a person, to display a target area. However, these techniques do not themselves include or require any steps of a surgical procedure. The described techniques are based on capturing and processing user input to determine control commands for modifying the display of a surgical visualization system. These interactions between a user and the surgical visualization system can also occur, for example, before or after a surgical procedure.
[0105] Although the features described in the summary above and the detailed description below are described in the context of specific examples, it should be understood that the features can not only be used in the respective combinations, but can also be used in isolation or in any combination, and features from different examples of medical imaging systems and coil arrangements can be combined and thus correlate with each other, unless expressly stated otherwise.
[0106] The above summary is therefore intended only to provide a brief overview of some features of certain embodiments and implementations and should not be understood as a limitation. Other embodiments may include features beyond those described above. BRIEF DESCRIPTION OF THE FIGURES
[0107] The invention is explained in more detail below with reference to preferred embodiments and the accompanying drawings, where identical reference numerals denote identical or similar elements. The figures are schematic representations of various embodiments of the invention, and the elements depicted in the figures are not necessarily shown to scale. Rather, the various elements shown in the figures are represented in such a way that their function and general purpose are understandable to a person skilled in the art. Fig. Figure 1 schematically illustrates a surgical visualization system, according to various examples. Fig. Figure 2 schematically illustrates a temporal progression of speech command recognition and a change in gaze position over several phases, according to various examples. Fig. Figure 3 is a flowchart of an example procedure.
[0108] The properties, features and advantages of this invention described above, as well as the manner in which they are achieved, will be explained more clearly and understandably in connection with the following description of the exemplary embodiments, which are related to the drawings.
[0109] It should be noted that the description of the exemplary embodiments is not to be understood in a limiting sense. The scope of the invention is not intended to be limited by the exemplary embodiments described below or by the figures, which serve only for illustration. DETAILED DESCRIPTION
[0110] The present invention is explained in more detail below with reference to preferred embodiments and the drawings. Connections and couplings between functional units and elements shown in the figures can also be implemented as indirect connections or couplings. A connection or coupling can be implemented as a wired or wireless connection. Functional units can be implemented as hardware, software, or a combination of hardware and software.
[0111] Various techniques for a surgical visualization system are described. However, it should be understood that the described techniques are applicable to any system in which control or manipulation is based on multimodal user inputs of varying durations and frequencies. The techniques are therefore not limited to the surgical context but can be used in a wide variety of human-machine interfaces, for example, in interactions with computers, robots, vehicles, or other technical systems. The method according to the invention provides a general framework for processing and fusing user inputs with different temporal properties.
[0112] Fig. Figure 1 schematically illustrates a surgical visualization system 10, according to various examples.
[0113] As in Fig. As can be seen in Figure 1, the surgical visualization system 10 comprises a control unit 11, which controls the individual components of the system, in particular a representation of an operating area 13 on a display 2 of the system. The display in Fig. 1 is an external display, however it would also be possible to display the image in optical eyepieces, as an image on a 3D monitor, as an image in digital eyepieces, or as an image in AR glasses for user 1.
[0114] A patient or an object to be examined can be positioned on an operating table 9, encompassing an operating area which is captured by the imaging system 6 and displayed by the visualization system 10 on a screen 2. In this example, the imaging system 6 includes an operating microscope (OPMI) for magnifying the operating area 12.
[0115] The control of the surgical visualization system 10 is based on multimodal user inputs 3.5 different user utterance types with different durations and at different times.
[0116] User 1 of the surgical visualization system 10 operates a surgical instrument 4. The surgical instrument 4 is positioned within an operating area 12. The operating area 12 is displayed by the surgical visualization system 10 on the display 2 as an imaged operating area 13 for user 1. The operating area 12 is located within the field of view of the surgical visualization system 10. Furthermore, the surgical instrument 4 is also imaged by the surgical visualization system 10 and displayed on the display 2 as an imaged surgical instrument 14 for user 1. The imaged system can, for example, include a camera 6.
[0117] To control the surgical visualization system, the user interacts with System 10 via various modalities.
[0118] One of these modalities is the verbal user utterance 5, which represents an initial user utterance and serves as a voice control for the system. The verbal user utterance is captured via a microphone 8 and defines a desired control action, such as "Focus on this area." The verbal utterance 5 extends over a time interval within a longer (reference) period. From the verbal utterance, the control device uses speech recognition to determine a voice command that specifies the desired control action for the surgical visualization system 10.
[0119] Examples of voice commands could be “Hey, go there” – for example, the camera should perform a linear translation to center the POI in the field of view, “Look” – the camera should tilt around its optical center point and center the POI in the field of view, or “Autofocus there” – the autofocus algorithm should focus on pixels near the POI.
[0120] At least partially in parallel with the recording of the linguistic utterance 5 and the determination of the speech command or control action, a large number of second user utterances of different second user utterance types are recorded.
[0121] A first type of user expression is the gaze direction of user 3 in relation to the representation of the operating area relative to the display 2. The gaze direction is captured and tracked by an eye-tracking system comprising a camera 7. It would also be possible to capture and track the orientation, in particular the forward direction of the user's head, using camera 7.
[0122] From the viewing direction 3, as the second user input, a viewing position 15 on the display 2 can be determined by the known spatial setup of the surgical visualization system 10. Furthermore, based on the imaging parameters of the system, a target point 16 (or point of interest, POI) can be determined from this as user input, and from this, in turn, a target region (region of interest, ROI) can be determined with respect to the surgical area 12 or with respect to the displayed image data of the surgical area 12. Thus, the viewing position 15 on the display 2 corresponds to a POI in the field of view of the OPMI 16, and this correspondence can be readily calculated by a person skilled in the art.
[0123] Another second type of user expression is the position of a surgical instrument 4 in relation to the surgical field 12, which is determined via the imaging system 6b that also images the surgical field. The point of interest (POI) 16 in the surgical field 12 can also be determined from tracking the position and / or orientation of the surgical instrument 4.
[0124] The viewing direction and instrument position data are repeatedly recorded over time and temporarily stored along with associated timestamps (as an example of time information). These are used to track user 1's POI 16 in operating area 12.
[0125] Both the direction of gaze and the positions of the surgical instrument 4 can thus be captured and processed as second user utterances in order to determine a target point (POI) or target area (ROI) of the user 1 in the operating area 12.
[0126] The surgical visualization system 10 and the control unit 11 are configured to perform any procedure or any combination of procedures according to the present disclosure.
[0127] In this example, the buffered POls 16 from both input modalities are filtered and merged, resulting in a robust and smooth estimate of the user's POI 16. The POls 16 can exist (as data points) in a time series and / or be stored or buffered.
[0128] The merged POls are stored by the control unit 11 with respective time information in a buffer data structure and thus represent sequences of data points of the respective second user utterances.
[0129] Based on the specified voice command and the buffered POI data, a control command is generated for the surgical visualization system. This command is used to adjust the visualization according to the user's intention, for example, by focusing on the specified target area. The result is displayed to the surgeon on the screen.
[0130] The control can include, for example, control of the OPMI 6, such as XY movement by linear translation of the OPMI 6, rotation of the OPMI camera, digital cropping of the image region, or autofocus setting at a specific position in the operating area 12.
[0131] However, the eye-tracking signal can be noisy, for example due to subconscious eye movements by the user, and can easily be distorted, for example if the user is distracted and looking somewhere other than the intended gaze position. In this case, using the raw, unfiltered gaze positions as a point of interest (POI) is unreliable.
[0132] Furthermore, gaze position changes over time. When eye tracking is used in conjunction with voice command recognition, there is a duration for the spoken sentence and a delay in speech recognition. During this time, the user's gaze position may already have changed.
[0133] To address these challenges, gaze and instrument position data are repeatedly acquired over a period of time and stored in a buffer data structure along with timing information. By filtering, merging, and incorporating the timing information of the buffered POls from both input modalities, a more robust and smoother estimate of the user's actual point of interest within the operating area can be achieved.
[0134] The control process takes into account a temporal correlation between the individual gaze and instrument data points and the voice command. The time interval of the voice input is included to select only the POls relevant to the voice command from the time series of buffered POls for control purposes. This improves the reliability and control, as demonstrated by... Fig. 2 is described in more detail.
[0135] Fig. Figure 2 schematically illustrates a temporal progression of voice control 20 and changes in raw gaze positions 31-34 and filtered gaze positions 41-44 over several phases of voice command recognition, according to various examples.
[0136] The top line in Fig. Figure 2 illustrates the temporal progression of a voice control system, divided into different phases with regard to the linguistic utterance 5 (first user utterance). Fig. 1.
[0137] At time point 21, the user's linguistic utterance 5, for example, a spoken sentence or sentence fragment, begins. At time point 22, the linguistic utterance ends. Thus, the time interval between times 21 and 22 corresponds to the time interval 25, over which the linguistic utterance extends and during which linguistic utterance 5 is recorded. Time interval 25 lies within a larger period 26, which can include times before and / or after time interval 26.
[0138] Between time points 22 and 23, the captured speech utterance 5 is processed (speech recognition) to determine a voice command or user intention (corresponding to an initial user input). At time point 23, the speech recognition is complete and the voice command has been recognized.
[0139] Between time 23 and time 24, a desired control action of the surgical visualization system 10 can be determined based on the voice command; for example, parameterization of the control action can be carried out.
[0140] Based on the recognized voice command, a control action of the surgical visualization system 10 is initiated and subsequently executed with a time delay from time 24, for example the control action is activated or triggered at time 24.
[0141] In Fig. In the middle and bottom rows, changes in the user's gaze positions (31-34 and 41-44, respectively) during the voice control phases are shown (second user utterance of a second user utterance type, "gaze direction"). These second user utterances are shown here during time interval 25; however, they may also occur, at least partially, before and / or after time interval 25 within period 26.
[0142] The middle row of the Fig. Figure 2 shows changes in the raw, i.e., unprocessed or unfiltered, gaze positions 31-34 of the user, associated with the respective time points 51-54 at which the respective unfiltered gaze positions were captured. Each gaze position 31-34 represents a measurement at a specific time point. In general, temporal information can be associated with each of the gaze positions, which may, for example, represent a specific time or duration with which the respective gaze positions are associated.
[0143] The bottom line of the Fig. Figure 2 shows changes in the user's filtered gaze positions 41-44, associated with the respective time points 51-54, where the filtered gaze positions 44-44 are determined from the unfiltered gaze positions 31-34 by filtering. In the example of the Fig. 2. The times of the filtered gaze positions correspond to those of the unfiltered gaze positions 32-34; however, other times for the filtered gaze positions 41-44 could also be determined from the times 51-54 by a corresponding filtering process. Each gaze position 32-34 and 42-44 is thus associated with time information indicating when the respective gaze positions were recorded.
[0144] The eye-tracking system 7 captures the user's gaze position. Each point represents a gaze measurement at a specific time. It can be seen that these measurements are distributed across the display and also vary during the voice command.
[0145] As in Fig. As can be seen in Figure 2, the noise and variability are reduced in the filtered gaze positions 41-44 by applying a filter. The filtered gaze positions are more strongly focused on relevant areas of the display 2.
[0146] The present disclosure is based on the finding that gaze positions 33, 34, 43, and 44, which occur after recognition of the speech command 23, are less relevant for the execution of the control action than gaze positions 31, 41, and possibly also gaze positions 32 and 42, due to their temporal proximity to the speech utterance. There is a time delay between the end 22 of the speech utterance 5 and the activation of the command 24 by the control unit (10). During this delay, the gaze positions continue to change. Since the gaze positions change continuously, especially even after the end of the speech command, a temporal correlation is crucial for determining the relevant gaze position.
[0147] For example, the user might look at the target position and begin to speak the command (e.g., "Move to this position" or "Automatically focus the view there"), but look away before the entire sentence is finished. Another example would be if the surgeon continues to look at the target position until the entire sentence is finished, but the speech recognition algorithm takes a second to recognize the command and activate the movement, at which point the surgeon has already looked away. In such situations, using the last specified gaze position as a point of interest (POI) is unreliable when the voice command is activated.
[0148] Simply using the last recorded gaze position when activating the control action is therefore problematic if the user is already looking at a different area of the display at that time. Instead, selecting the gaze position to use based on the temporal context of the voice command and the gaze positions is advantageous. Alternatively, the different gaze positions can be prioritized in another way – for example, a relative weighting.
[0149] For example, gaze positions can be stored together with timestamps in a buffer data structure. When the voice command is activated, the relevant gaze position 3 can then be selected based on the temporal relationship between the buffered gaze positions 3 and the time interval of the voice command 5, even if this does not coincide with the end time of the voice command 5. This combination of modalities, taking their temporal relationships into account, achieves robust and reliable control of the surgical visualization system.
[0150] From a filtering perspective, a low-pass filter can be applied to the individual measurements of the second user utterances to smooth the target position. Non-limiting examples of such filtering include a simple PT1 low-pass filter, sensor fusion or a more general combination of different measurements to determine the POI (eye tracking, head direction, position of the surgical instrument in the image), a Kalman filter, dynamic filtering based on the phases of a voice command, or filtering of the POI that depends on the image content and the surgical phase.
[0151] In some examples, a low-pass filter can be applied to smooth the gaze positions, or even the target positions (POIs). This filter can be implemented in various ways to smooth the temporal evolution of the viewpoints, i.e., the viewpoint trajectory or POI trajectory.
[0152] In some examples, a simple PT1 low-pass filter can be used. This first-order filter can perform a weighted averaging of positions over time, with the influence of older measurements decaying exponentially. By choosing a suitable time constant, the smoothing effect can be adapted to the dynamics of eye movements.
[0153] In some examples, sensor fusion can be applied, combining second user utterances from two or different second user utterance types to determine the target positions (POIs). This can involve, for example, using data from eye tracking, head direction, and the position of the surgical instrument in the image. By combining this information, a more robust POI estimate can be achieved, one that is less susceptible to interference or inaccuracies in individual modalities.
[0154] In some examples, a Kalman filter can be used to smooth the gaze or target positions. The Kalman filter is a recursive algorithm that can estimate the state of a system from noisy measurements. In this case, the state can include the position and velocity of the point of interest (POI).
[0155] In some examples, dynamic filtering of the target position can be performed based on the phases of voice control. For instance, gaze positions can be weighted more heavily while a sentence is being spoken, as gaze position can be particularly relevant for interpreting the user's intent during this phase. Conversely, during the speech recognition delay after the sentence has finished, gaze positions can be weighted less heavily, as the user may already be looking at a different area of the screen.
[0156] In some examples, the filtering of target positions can also depend on the currently displayed image content and the surgical phase. For example, in phases where particularly fine structures are visible in the image, the filter time constants can be reduced to allow for a higher resolution of the viewing positions.
[0157] In phases with coarser structures or in overview views, longer time constants can be used to achieve greater smoothing. The type of surgical procedure and the associated typical eye movement patterns can also be taken into account when selecting the filter parameters.
[0158] Therefore, filtering the viewing positions or target positions to determine a smoothed time series of POls can improve the reliability and accuracy of the control.
[0159] To increase the reliability and user-friendliness of POl estimation through eye movements and thus improve the user experience, buffering (caching) is performed for the raw eye positions 31-34 and / or the filtered eye positions 42-44.
[0160] The trajectory of gaze positions over time can be buffered so that the system can retrieve gaze positions from earlier phases of voice control when the voice command is activated.
[0161] For this purpose, the specified unfiltered gaze positions 31-34, or the filtered gaze positions 41-44, are temporarily stored in a buffer data structure (e.g., a first-in-first-out buffer data structure) for a specific time period (e.g., 10 seconds). The system stores the timestamps associated with the buffered gaze positions in the buffer data structure. The system also receives time information that characterizes the time interval, such as the start and end times of the interval. If the spoken utterance is recognized as a valid voice command, the system selects one or more gaze positions relevant to the voice command from the buffered gaze directions and uses these to control the surgical visualization system.
[0162] In the example of the Fig. 2 can, for example, capture the gaze directions 31 or 41 that were captured at the same time as the first user utterance, and / or the gaze directions 31, 32 or 41, 42 that were captured within a valid time window 27 around the beginning or end of the time interval 25, and uses these gaze positions for control.
[0163] There are various ways in which the buffer data structure can be implemented. For example, the buffer data structure can store the raw or filtered eye-tracking positions, gaze positions, head-forward direction, or the resulting target positions (POIs) within the image dataset or the surgical field. Buffering can, for instance, store data points at the system frequency, half the system frequency, fixed time intervals, dynamic sampling rate, or fixed distance intervals. Furthermore, the user or an AI algorithm can configure the system to use the buffered element that best matches the given time interval. This can be achieved, for example, by comparing the timestamps with the start or end of the speech utterance, an intermediate point in time, or a fixed time window relative to the speech command's time interval.
[0164] Fig. Figure 3 is a flowchart of an exemplary procedure for controlling a surgical visualization system.
[0165] The procedure begins in step S10.
[0166] In step S20, the first user utterance of a first user utterance type is received. The first user utterance spans a time interval within a period.
[0167] In step S30, a multitude of second user utterances of at least one distinct second user utterance type is retrieved. These second user utterances are distributed across the time period and encompass a variety of different second user utterance types. Step S30 also retrieves time information for each of these second user utterances.
[0168] The data obtained in step S30 is temporarily stored, for example in a buffer storage in FIFO structure.
[0169] In step S35, temporal relationships between each of the multiple second user utterances and the first user utterance are determined. This can be done, for example, based on a comparison between timestamps and the time interval in which the first user utterance occurs.
[0170] In step S37, at least one of the second user utterances is prioritized based on their temporal relationships. This prioritization could, for example, involve selecting the at least one second user utterance and discarding the one or more unselected second user utterances. It could also include assigning comparatively higher weights to the at least one second user utterance.
[0171] In step S40, the surgical visualization system is controlled based on the first user utterance, the prioritized at least one second user utterance, and a temporal relationship between the first user utterance and the at least one second user utterance.
[0172] The procedure ends in step S50. REFERENCE MARK LIST 1 user 2 ads 3. Second user comment - direction of gaze 4 surgical instruments 5. First user utterance - verbal utterance 6 Imaging system 7 Eye-Tracking System 8 microphones 9 Examination table 10 surgical visualization systems 11 Control unit 12 Operating area 13. Operating area shown 14 illustrated surgical instruments 15 Viewing position 16 Destination (Point of Interest, POI) 20 phases of voice control 21 Time: Start of the first user comment 22 Time: End of the first user comment 23 Time: Voice command recognized 24 Time: Start of the tax action 25 Time interval 26 period 27 valid time slots 31-33 unfiltered viewing positions 41-44 filtered viewing positions 51-54 Time points associated with gaze positions S10 to S50 process steps
Claims
[1] Computer-implemented method for controlling a surgical visualization system (10), comprising: - Receiving a first user utterance (5) of a first user utterance type, wherein the first user utterance (5) extends over a time interval (25) within a period (26); - Receiving and caching a plurality of second user utterances (3) of at least one second user utterance type and a respective associated time information for each of the plurality of second user utterances (3), wherein the second user utterances (3) are distributed within the time period (26) and comprise different second user utterances (3) of each of the at least one second user utterance type; - based on the time information and the time interval (25), determining temporal relations between each of the multitude of second user utterances (3) and the first user utterance (5), - Prioritizing at least one second user utterance from the multitude of second user utterances (3) based on the temporal relations, and - Control of the surgical visualization system (10) based on the first user utterance (5) and the prioritized at least one second user utterance. [2] Computer-implemented method according to claim 1, further comprising: - for every second user utterance of the multitude of second user utterances (3): Perform a check to see if the respective temporal relation meets one or more predefined test criteria, whereby the at least one second user utterance is prioritized depending on a result of the checks to see if the temporal relations meet the one or more test criteria. [3] Computer-implemented method according to claim 2, wherein the one or more test criteria comprise a verification of whether a time point associated with a corresponding second user utterance (3) and indexed by the time information lies within a certain time range before the end of the time interval (25). [4] Computer-implemented method according to any one of the preceding claims, further comprising: - Determining an initial user input based on the initial user utterance (5); and - Determine at least one second user input for the prioritized at least one second user utterance, wherein the surgical visualization system (10) is controlled based on the first user input and at least one second user input. [5] Computer-implemented method according to claim 4, where controlling the surgical visualization system (10) includes triggering an action of the surgical visualization system, where the type of action is specified by the first user input (5), and where at least one second user input (3) triggers and / or parameterizes the action. [6] Computer-implemented method according to any one of the preceding claims, wherein the multitude of second user utterances (3) defines coordinates in a continuous space, wherein the coordinates exhibit a development during the period (26), the procedure further includes: - Applying a filter to the coordinates, especially a low-pass filter or a Kalman filter, which smooths the development. [7] Computer-implemented method according to claim 6, wherein a filter parameter of the filter depends on a type of content displayed on a display screen (2) of the surgical visualization system (10), a phase of a surgical workflow and / or the first user utterance type. [8] Computer-implemented method according to one of the preceding claims, wherein several of the plurality of second user utterances (3) are different from each other and lie within the time interval (25). [9] Computer-implemented method according to one of the preceding claims, wherein each of the second user utterances (3) extends over a respective time period, wherein each of the time periods is shorter than the time interval (25) over which the first user utterance (5) extends, wherein preferably the length of each of the time periods is less than 50% of the length of the time interval (25) over which the first user utterance (5) extends, and particularly preferably less than 10%. [10] Computer-implemented method according to any one of the preceding claims, wherein the first user utterance (5) comprises one of the following: - a linguistic utterance by the user; - a gesture of a body part of the user; - a touch gesture at a touch interface; - a brain-computer interface signal; or - a multimodal combination of the above. [11] Computer-implemented method according to one of the preceding claims, wherein the second user utterances (3) comprise a user indicating a target region in a field of view of the surgical visualization system (10). [12] Computer-implemented method according to any one of the preceding claims, wherein the second user utterances (3) comprise one of the following: - a position and / or orientation of at least one body part of the user; - a user's direction of view (3); and - a position and / or orientation of a surgical instrument (4) operated by the user (1), or - a multimodal combination of the above. [13] Computer-implemented method according to claim 4, wherein the first user utterance (5) comprises a verbal utterance of the user (1), wherein the second user utterances (3) comprise at least a first group of second user utterances comprising gaze directions (3) of the user, and wherein the second user utterances comprise at least a second group of second user utterances comprising a position and / or orientation of a surgical instrument (4) operated by the user (1), wherein the at least one second user input is determined using at least one second user utterance from each of the first and second groups of second user utterances. [14] Control unit for a surgical visualization system (10) configured to perform the method according to any one of claims 1-13. [15] Surgical visualization system (10) comprising a control unit according to claim 14.
Citation Information
Patent Citations
Microscope with sensor screen, associated control unit and operating procedure
DE102010063392A1
Device and method for setting a focus distance in an optical observation device
DE102015117824A1
Method for operating an assistance system and an assistance system for a motor vehicle
DE102017011498A1
Procedures for operating a speech dialogue system and speech dialogue system
DE102021006023B3
Artificial intelligence apparatus
US20230077245A1