Controlling a surgical visualization system using multi-modal user representations

The method addresses inaccuracies in surgical visualization systems by employing multimodal user expressions for real-time processing of verbal commands and gestures, ensuring precise and reliable control.

JP2026015227AActive Publication Date: 2026-01-29CARL ZEISS MEDITEC AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025109471
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-01
Filing Date
2025-06-27
Publication Date
2026-01-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing surgical visualization systems face inaccuracies due to noise in eye tracking data and delays in voice recognition, compromising the reliability and accuracy of control processes.

Method used

A method for controlling surgical visualization systems using multimodal user expressions, including real-time acquisition and processing of verbal commands, gestures, and gaze direction, with latency less than 0.5 seconds, to enhance control precision and accuracy.

Benefits of technology

The method provides precise and reliable control of surgical visualization systems by integrating multiple user inputs, reducing latency and enhancing the accuracy of control actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015227000001_ABST
    Figure 2026015227000001_ABST
Patent Text Reader

Abstract

To provide control of a surgical visualization system using multi-modal user utterances.SOLUTION: A computer-implemented method of controlling a surgical visualization system is provided. A first user utterance of a first user utterance type is received, and the first user utterance lasts for a time interval within a defined time period. A plurality of second user utterances of at least one different second user utterance type are received, which are captured in a distributed manner within the time period and which vary within the time period. The surgical visualization system is controlled using the first user utterance and at least one second user utterance that is prioritized based on a temporal relationship of the second user utterance to the time interval.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Various examples of the present disclosure relate to aspects of control of a surgical visualization system, such as a visualization system including a surgical microscope. The various examples particularly relate to multi-modal user interaction techniques, which may occur over different durations, for controlling a surgical visualization system. expression , for example, in the processing of verbal commands and, on the other hand, eye movements. [Background technology]

[0002] In surgical environments, accurate and reliable control of surgical visualization systems, such as optical surgical microscopes, is important. Some systems use, for example, eye tracking and voice commands for control. However, noise in the eye tracking data and delays in voice recognition often lead to an inaccurate control process. These inaccuracies can compromise the reliability and accuracy of the surgical viewing system. Summary of the Invention [Problem to be solved by the invention]

[0003] Therefore, there is a need for improved techniques for controlling surgical visualization systems that overcome or mitigate at least some of the limitations and drawbacks discussed above. [Means for solving the problem]

[0004] This object is achieved by the features of the independent patent claims. The dependent patent claims define further advantageous embodiments.

[0005] The inventive solution is described below with respect to both a claimed method for controlling a surgical visualization system and a corresponding controller and surgical visualization system. Additionally, a corresponding computer program and electronically readable storage medium are also provided. It should be understood that features, advantages, and alternative exemplary embodiments may be assigned to other respective categories, and vice versa. For example, a controller or surgical visualization system may be enhanced with features used as part of the described method for controlling a surgical visualization system, and vice versa.

[0006] A computer-implemented method for controlling a surgical visualization system is provided, including, for example, a visualization system used to provide visual information to a user in a surgical environment or application.

[0007] Surgical visualization systems may be used in conjunction with endoscopic, microscopic, or any other imaging or inspection method. Such imaging systems are used, for example, in a surgical environment to provide a user with visual information, e.g., in real time, regarding a surgical field and / or a surgical sequence. In particular, this visual information may be provided regarding the field of view of the surgical visualization system.

[0008] The visual information may include, for example, an image, multiple images and / or a video stream, which may be based on 2D or 3D image data, and may further include a representation of a target region (region of interest, ROI) within the surgical field.

[0009] Visual information, such as the surgical field and possibly surgical instruments, is displayed to the user, for example, using one or more of the following devices: a purely optical system (including, for example, an eyepiece) for optical display, a display on one or more electronic display devices (e.g., displays or screens) based on acquired image data, a display in a virtual reality system (e.g., a virtual display in a virtual environment), a display provided by a 3D display device that visually represents depth information for the user, or a display in an augmented reality system that projects virtual elements into real space. These various display devices allow visual information to be displayed to the user and for the user to access it.

[0010] In general, a surgical visualization system may therefore include a combination of hardware and software components that acquire image data from inspection equipment, such as an endoscopic camera, microscope, or other imaging device, and visually display (in real time) this image data to a user. These systems assist the user by providing clear and detailed visual information, particularly portions of the image data. For example, they may display additional visual information about the actual surgical field, such as virtual instructions or navigational aids to assist the user. Other examples include, for example, text, descriptions, icons, overlapping images, etc. These may be overlaid, for example, optically or virtually.

[0011] The method for controlling a surgical visualization system includes the following steps: In one step, a first user expression is acquired. The acquisition may include capturing or detecting using one or more first sensors. expression may be obtained in the form of raw data from a sensor or processable sensor data. expression It is also envisaged that the signal is transmitted electronically, for example based on measurements of an external sensor.

[0012] First User expression can be obtained in real time. expression and the user expressionThe latency between the acquisition of the corresponding data indicative of the time may be, for example, less than 0.5 seconds, optionally less than 0.05 seconds, also optionally less than 0.005 seconds.

[0013] First User expression may include actions or instructions for a user or, in general, information about a user displaying visual content using a surgical visualization system. expression can be captured to control the surgical visualization system.

[0014] First User expression may include an intentional, but possibly unintentional, action that continues for a time interval of a period, particularly the entire duration of that time interval.

[0015] In some instances, the first user expression may include one or more of the following: the language spoken by the user expression , gestures made by the user's body parts, touch gestures on a touch interface, brain / computer interface (BCI) signals, or a multimodal combination of these.

[0016] First User expression By this, the user can express his / her intention or intent. For example, expression may serve to specify what type of control actions the user expects from the surgical visualization system.

[0017] First User expression is the first user expression Type, in other words, the first (user expression ) modality. User expression The type may refer to a category or type of user action performed by a user.

[0018] The first user who can express the user's intention expression Examples of languages expressionfrom which processing can be used to identify voice commands such as "magnify this area," "focus on this structure," "zoom here," or "show this in overview." expression may indicate what display changes the user wants to make. For example, a gesture made by at least one body part of the user to define an image section may also express the user's intention. As a further example, it would be possible to use, for example, a command or a gesture made by a tool, for example a surgical tool. For example, a pointer command may be used to indicate to a first user expression It becomes possible to define

[0019] A gesture performed by at least one body part of a user may take various forms and may involve multiple body parts. A hand gesture is a common example in which a user moves or positions their hand or fingers in a particular way to indicate an intention or action. For example, a particular body part or one or more implements may be moved to a particular area or a particular zone and remain there. For example, hand or finger movements in a particular shape or pattern may be envisioned. Blinking patterns or gestures on a touchscreen may also be envisioned.

[0020] First User expression may serve to identify for the surgical visualization system a first user input that is used to specify the operation, and in particular the control operation, of the surgical visualization system. expression may be processed therefrom to identify desired control actions of, for example, a surgical visualization system. expression The processing may be, for example, rule-based or based on machine learning.

[0021] The first user input is expression The user input may include, for example, a user's intent. The user input may include, for example, a control command for a desired control action for the surgical visualization system.

[0022] First User expression The time interval lasts for a certain period of time. expression The first user expression is therefore performed over the entire time interval, which continues throughout the entire time interval, i.e., from the start to the end of the time interval. expression is therefore captured by the surgical visualization system using sensors over the entire time interval, i.e., from the beginning to the end of the time interval.

[0023] Time information characterizing the time interval, such as the start time, end time, and / or midpoint of the time interval, may additionally be acquired and stored. For example, the start time indicates the start of the time interval, the end time indicates the end of the time interval, and the midpoint indicates the middle of the time interval. The start time, end time, and midpoint are exemplary reference times for the time interval. Furthermore, the length of the time interval may also be acquired and stored as time information.

[0024] A time interval is within or contained within a longer period. A period contains a time interval. A period may have the same length as the time interval or may be longer.

[0025] In another step, a plurality of different second users expression are acquired, e.g., captured or detected, and these are transmitted to at least one second user. expression Assigned to type: Secondary user expression can be electronically transmitted from the at least one second sensor.

[0026] "First User expression " and "Multiple Second Users expression " instruction to the first and second users expression No order or hierarchy between them is intended to be implied. For example, multiple secondary user interfaces may be used in conjunction with a primary user interface. expression It is also expected that the information will be acquired before the deadline.

[0027] Second User expression Each of the second users may be obtained in real time. expression and this second user expression The latency between the acquisition of the corresponding data indicative of the time may be, for example, less than 0.5 seconds, optionally less than 0.05 seconds, also optionally less than 0.005 seconds.

[0028] Multiple Secondary Users expression Each of the time information is associated with a respective time information. expression Such time information may also be acquired. In principle, the time information may be acquired implicitly or explicitly. For example, a timestamp may be acquired. However, if a specific second user expression Each time this occurs, the second user expression It is also envisaged that the estimation is based on the sampling rate and sequence of .

[0029] The time information is particularly expression But the first user expression The second user can then see how the expression Each of the first user expression The temporal relationship between may be specified based on the time information and the time interval.

[0030] Therefore, the second user expression is the first user expression The first and second users are in a time correlation with each other, which may be indicated by a time relationship. For example, expression There may be an intentional or content-based correlation between, which is represented by a temporal relationship. For example, expression may relate to a common action or interaction of a user / system interaction expressed by a temporal relationship. For example, expression can jointly correlate with the user's intent.

[0031] The method includes: expression and the first user for each of the time intervals expression In this case, the temporal relationship may include determining a temporal relationship between the second user expression are identified based on the time information available for each of the

[0032] Multiple Secondary Users expression The second user expression is stored in the buffer. expression is the second user expression Based on the time information about the second user expression and the first user expression The time intervals may be stored in a buffer until it is possible to determine the time relationship between the time intervals. For example, one or more second users may expression It is also envisioned that the second user may obtain the time interval before the end of the time interval, and therefore, in at least some instances, it may not be possible to definitively determine the temporal relationship for that time interval (which may not yet end or which will end at an unknown future point). expression Finally, multiple second users expression The time relationships between each of the time periods and the time intervals can be determined. expression can be stored in the buffer for up to or exactly a predefined or specifiable buffer duration.

[0033] In some instances, the second user expression The time information and the time information may be correlated with each other and stored in a buffer data structure (buffer) for a specific buffer duration. The buffer data structure may include, for example, a FIFO (first-in, first-out) memory. The size of the buffer or the buffer duration for which the data is held may in this case be predetermined or may be specified by, for example, a first user. expression The primary user of expressionThe buffer duration may be specified based on the type of the buffer. The buffer duration may be in the form of a sliding time window. The buffer duration may be specified based on the type of the buffer. expression may be chosen to be greater than the maximum allowable time interval.

[0034] Then, the second user expression Aspects relating to the second user will be described. expression At least one second user of expression Type the first user expression The primary user of expression Different from the type.

[0035] Second User expression is distributed over a (longer) period. In other words, multiple second users expression is distributed over the period. Therefore, this is expression occurs in a distributed manner at different times within a (longer) period and can be conveniently captured. expression The time information associated with the second user expression is distributed over that period.

[0036] Second User expression At least one second user expression Type each one different second user expression Includes.

[0037] Second User expression may be captured by at least one second sensor or multiple second sensors to control the surgical visualization system, for example, using sensor fusion. The at least one second sensor may be different from the first sensor. However, it is also envisioned that the at least one second sensor includes the at least one first sensor, whereby the at least one first sensor is also captured by the second user. expression can be used to capture the second user expression may also include processed signals / processed sensor data from one or more sensors.

[0038] Second User expression may include several different actions or statements of the user, i.e., different items of information about the user, made in a scattered manner at different times during the period. expression occurs at several points in time, i.e., different points in time during the period. expression Allows the same second user to control the surgical visualization system expression A second user of a different type (i.e., changing) expression Includes.

[0039] For example, multiple second users expression is the first user expression These second users may be included in the following time interval. expression is the first user expression are within the following time intervals and may vary between them.

[0040] In some instances, the second user expression is the first user expression The image data may serve to specify a target point or region in the image data that is relevant to the user's intent.

[0041] For example, the second user expression The surgical visualization system may include having a user point to a target area. The target area may be located within a field of view of the surgical visualization system. The field of view may be a displayed or actual field of view of the surgical visualization system. For example, the user may point to a representation of the target area in an image captured by the surgical visualization system. However, the user could also point to the target area directly, for example with a finger or a surgical instrument or simply an instrument. In the latter case, the pointing may be observed in the image captured by the surgical visualization system.

[0042] Second User expressionmay include both intentional and unintentional actions or states of the user, such as gaze, posture or body orientation, the position and / or orientation of at least one body part of the user (e.g., within the coordinate system of the surgical visualization system or specifically within the field of view of the surgical visualization system), touching a mechanical interface or a touch interface, or other information of the user that is obtained instantaneously using sensors. expression is responsible for identifying a second user input for the surgical visualization system and, based thereon, expression Trigger or parameterize the control action specified by

[0043] For example, the second user expression is assumed to indicate the user's gaze direction. The user's gaze direction can be determined, for example, by pupil tracking (also called eye tracking). Therefore, it is assumed that the gaze direction is determined at a certain sampling rate, which is selected to be high enough to be able to identify rapid changes in gaze direction. For example, the sampling rate used to determine the gaze direction is assumed to be in the range of 100 Hz or higher, or optionally 200 Hz or higher. This can also capture more rapid saccadic eye movements, in which the eyes move between two points to readjust the field of view. For example, such movements are physiologically induced and typically occur within 20 to 40 ms. Therefore, if the gaze direction of a person is determined at a high sampling rate corresponding to each case, it is possible to detect the gaze direction of a large number of second users. expression This is obtained by the second user expression It should be understood that this means that it is not necessary to specify different user inputs, i.e., different gaze directions of the eyes, for example. expression A different second user in expression may be the same or at least comparable (eg, if the eye gazes in one direction for a longer period of time).

[0044] In a preferred example, the first user expressionis the language used by the user expression and a second user expression is the second user's gaze direction, including the user's gaze direction. expression and a second user, which may include at least one first group of a surgical instrument operated by the user, and which may include a position and / or orientation of the surgical instrument. expression The second group may include at least one different second user. expression Type of second user expression From such a multimodal combination, the second user input can be identified with greater accuracy.

[0045] Second User expression may be captured at specific times or at specific sampling rates, thereby allowing different second users to capture different second users at these times in the period. expression Depending on the sampling rate, the second user expression It may be possible to determine time information for each of the

[0046] As mentioned above, a different second user expression (with different time information) can be different from each other or the same. In either case, it allows the same second user to capture different time points. expression Type of second user expression It is possible to obtain a temporal sequence of

[0047] Second User expression is the first user expression This may provide additional parameters or context information for control based on the

[0048] Second User expression may specify, for example, a target location, target region, direction, or selection region that relates to a control action for a display of a surgical visualization system.

[0049] For example, the second user expressionmay include the user's line of sight or gaze position, which may be easily transformed to and from, for example, with respect to the surgical visualization system's display or other reference frame (e.g., defined by the displayed or actual field of view of the surgical visualization system).

[0050] The position and / or orientation of the user's head or part of the head may also be transmitted to a further second user. expression Type of second user expression Similarly, for example, the position and / or orientation of a body part of a user, at least one finger, may be captured as a second user's expression can be captured as:

[0051] Second User expression Further examples of the position and / or orientation of the surgical instrument at a particular time may include the position and / or orientation of the instrument. For example, the position and / or orientation of the surgical instrument may be determined relative to the field of view or surgical field of the surgical visualization system. Furthermore, the position of the surgical instrument or at least a portion of the surgical instrument, e.g., the tip, may be determined in a machine coordinate system, i.e., 3D world coordinates, or other reference coordinate system, e.g., using a sensor. Such a position of the surgical instrument may also be determined from image data from, for example, an imaging system of the surgical visualization system, which image at least a portion of the surgical instrument.

[0052] Second User expression is the first user expression The second user may have a shorter time expression is the first user expression So, for example, two consecutive second users expression The time difference between the first user expression may be less than 30%, or less than 20%, or less than 10%, or less than 5%, or less than 1% of the time interval that the second user expression (optionally each of them) in each case is expression For example, the first user expressionless than 50%, or 30%, or 20%, or 10%, or 5%, or 1% of the time required to capture the second user. expression is the first user expression It can be recorded at a higher temporal density. expression is the first user expression It may not be as complicated as expression is the first user expression It can be recorded at a higher temporal density. expression The time difference between the first user expression These may be shorter than the duration of the first user. expression The second user complements the user's intent, which is identified based on the second user's intent, and provides details for precise control of the visualization system. expression may also be used to identify a first user input. For example, they may serve as a continuous feedback mechanism that allows the system to accommodate dynamic changes in user behavior in response to the identification and execution of control actions.

[0053] In some instances, the second user expression Each of these has its own time range defined by a respective duration, where each of these durations is expression Therefore, when considered separately, the second user expression The duration of the first user expression Shorter than the second user expression The duration range of each of the expression The time information or timestamp associated with the timestamp may be represented by time information or timestamps associated with the timestamp.

[0054] Preferably, the length of each of the durations may be less than 50%, particularly preferably less than 10% of the length of the first-mentioned time interval.

[0055] Second User expressionmay serve to provide spatial and / or temporal parameters for the control operation. The associated second user input may be one or more second user inputs. expression For example, the associated second user inputs may be identified for multiple second users. expression Second user of expression are identified for each of the sigma-based ...

[0056] In some examples, each of the second user inputs may be associated with two or more, three or more, or four or more different second user inputs. expression Type of second user expression can be identified from

[0057] For example, both gaze direction and instrument position may be used together (for different users) to identify the user's target position (user input). expression The user may then use the 'select' command to select the desired user.

[0058] Second User expression and / or the user input may complement or characterize the second user input. expression And / or user input may parameterize the operation of the surgical visualization system, particularly control commands or control actions.

[0059] For example, the at least one second user input may include a temporal control parameter for controlling, for example, the timing of a sequence of control operations of the surgical visualization system, for example, defining the start or end of a control operation, or parameterizing and / or triggering individual phases of a control operation.

[0060] This method involves multiple second users expression At least one second user in expression The surgical visualization system then prioritizes the first user based on their temporal relationship. expression and at least one preferred second user. expressionIn particular, the surgical visualization system may be controlled based on at least one preferred second user. expression The control may be based on at least one second user input determined based on the at least one second user input.

[0061] At least one second user expression Such a preference is given to at least one second user expression This may include selecting the at least one second user. expression Only the second user is considered and not selected when controlling the surgical visualization system. expression This means that the surgical visualization system is not taken into account when controlling the surgical visualization system. However, at least one second user expression This prioritization allows the corresponding second user to control the surgical visualization system. expression Secondary users who are not preferred expression For example, user input may include various secondary user inputs. expression For example, such user input may indicate a particular location, and a weighted average of such locations may then be determined.

[0062] Therefore, all secondary users expression The temporal relationship between the two can be used to exercise control, expression At least one of the parameters is used for control based on the temporal relationship, and / or a different control parameter is used for control based on the temporal relationship of the second user. expression and / or multiple (e.g., all) second users based on a temporal relationship between expression are weighted differently for control and / or specific second expression are excluded for control based on their temporal relationships.

[0063] At least one user expression are prioritized based on their temporal relationships, so the first user expressionDuring a relatively long time period during which expression The first user's actions can be prevented from being taken into account for control of the surgical visualization system, thereby avoiding unintentional control actions. expression For relatively long time intervals where , ... expression One example is, for example, voice input (first user expression ) and gaze direction (second user expression For example, during speech input, which may last 3-5 seconds, the user may only briefly change their gaze direction. The gaze direction may be captured at a sampling rate of 100 Hz, and a brief change in gaze direction may itself indicate a different second user during that time interval. expression Various secondary users expression and the first user expression Since the time relationship between the first user and the second user is taken into account, for example, expression A second user who is located shortly before or at the start or end of expression For example, other second user inputs at the beginning or middle of the time interval can be taken into account. expression may not be considered or may be considered only with less weight. Similarly, other temporal relationships may be prioritized.

[0064] The temporal relationship is in each case the first user expression and each second user expression The temporal relationship may represent a temporal relationship between the user expression In some examples, the temporal relationship may include information indicating the temporal sequence or arrangement of the respective second users. expression But the first user expression or the first user expressionindicates the time difference captured relative to a reference time (e.g., an end point or midpoint) of the duration of the time interval.

[0065] The temporal relationship may be, for example, expression or this time interval and / or the first user expression a first (reference) time allocated to at least one second user; expression This time difference may be defined as a time difference between the second time or second time interval assigned to the second user. expression is the first expression If done later, the second user expression is the first expression Therefore, for example, each second user expression Or the time of the second user input is the time of the first user input. expression The time interval may have a temporal relationship with the time interval.

[0066] The temporal relationship is expression It may also describe whether the second user was captured during the time interval or at least partially captured during the time interval. expression is, in whole or in part, the first user expression The temporal relationship can thus be, for example, expression may indicate whether or not the

[0067] For example, the temporal relationship may be characterized by a predetermined time difference or threshold value, which may be taken into account for the control. expression First and second users of expression It may indicate the maximum or minimum size of the time gap allowed between

[0068] Therefore, the second user expression Therefore, it is possible to check each of the following. expressionIt is possible to check whether each of the temporal relationships identified corresponding to the at least one second user, which will subsequently be taken into account when controlling the surgical visualization system, satisfies one or more predetermined check criteria. expression Therefore, all second users expression In other words, for example, the second user whose check result is positive may be selected based on the results of these checks. expression It is possible to select:

[0069] Examples of these checks are described below. For example, the first user expression The second user who made expression Alternatively, it may be specified that only the first user expression At most 1 second from the start of expression a second user who is a certain fractional multiple, e.g., 1.5 or 2 times the length of the time interval that expression It would also be possible to use only

[0070] Second User expression are stored in the buffer together with the corresponding associated time information, thereby expression After the time interval has elapsed, various second users expression This allows flexibility in retrospectively identifying and checking the time relationship with respect to one or more check criteria for the end of a time interval, i.e., a time relationship with respect to a reference time that typically first determines that the time interval has ended. Therefore, it is possible to define a specific time range using a reference time identified by the end of a time interval. Therefore, it is possible to determine the time range of a specific second user. expression It is possible to check whether the time associated with is within this time range or not.

[0071] Second User expression A further advantage of storing the time information in the buffer is that the first and / or second user expression, for example, by the first and / or second user expression The user input may be further processed to identify one or more of the user inputs. For example, speech recognition may require a certain amount of time. Such further processing may require a certain amount of time due to limited computational resources. This latency would otherwise lead to distortions in identifying temporal relationships, which can be avoided by buffering.

[0072] For example, a second user within a particular time interval defined as beginning at the end of the time interval. expression It would be possible to prioritize or specifically select a second user who continues within the last 20 ms of a time interval or within a time interval that lasts 50 ms from the middle of the time interval towards the start of the time interval. expression Furthermore, one or more check criteria may be used to prioritize or specifically select the corresponding second user. expression The time indicated by the time information is associated with the first user. expression is indicated by time information that falls within a specific time range before the end of the time interval that lasts.

[0073] Corresponding second user expression and the time indicated by the time information is at most expression It is also possible to check whether is within a certain time range before the end of the continuing time interval.

[0074] Typically, the first user input is expression The time at which the first user input is identified may therefore be determined based on the first user input. expression The test criteria may be used to determine whether the first user expression The second user, starting at or after the end of the time interval during which expressioncan be ignored.

[0075] Generally, by considering the time relationship and the time check criteria, i.e., by specifying a valid time window or time range in terms of a time threshold or time interval, the second user concerned can be identified. expression This allows the first user to limit the search space. expression A second user with a good correlation with expression can be more targeted.

[0076] In some variations, further check criteria may also be checked. For example, the first user expression A second user assumed to be before, during or after the time interval expression For example, a predetermined percentage, for example, at least 50% or a predetermined number of second users expression There may be a requirement that the first user expression The second user exists within the time interval expression There may also be a requirement that the time difference (sampling rate) between the first user and the second user must not be less than expression Instead of the time interval, for example, in each case, for example, the first user expression It is also possible to use as a reference a time interval of equal length which can be shifted in a predetermined manner with respect to the time interval of the first user, in which case said time interval is in particular at least partially, optionally completely expression before the end of the time interval.

[0077] For example, the second user expression is the first user expression In other words, the second user expression The weighting of each second user expression and the first user expression For example, the first user expressionA second user who is closer to the beginning of, in the middle of, or at the end of expression can be weighted more heavily. expression The weighting of each second user expression is the primary user expression This may depend on whether the time interval is closer to the beginning, the middle, or the end.

[0078] For example, the first user expression The second user who is closest in time to the start, middle, or end of expression You can select:

[0079] For example, the first user expression A second user having a specific time pattern regarding expression For example, the first user may select a consecutive sequence of expression It would be possible to search for a sequence that starts slightly before and ends slightly after.

[0080] By taking into account the time relationship and specifying particularly appropriate (temporal) check criteria, the system expression A second user who has a suitable time correlation with the expression The exact choice of time criteria can be made in this case by the first user. expression This may depend on factors such as the type of application, the content of the application, or the user's preferences. expression During a relatively long time period during which expression is not taken into account for the control of the surgical visualization system, thereby avoiding unintentional control actions.

[0081] When controlling a surgical visualization system, the primary user expression and the second user expression By considering the time relationship between expressionThis allows for the time sequence of events to be incorporated into the control, which improves the interaction between the user and the system. expression This is because the temporal context of the event is taken into account.

[0082] First User expression and the second user expression Alternatively or additionally, the second user may control the surgical visualization system. expression The time relationships between the two can also be taken into account.

[0083] Multiple Secondary Users expression Once captured, these are expression Not only is there a relationship with the second user, but there is also a mutual relationship. expression The time difference or interval can be varied and used for control purposes as well.

[0084] For example, the second user in the period expression The second user can take into account the sequence expression If the time window is not valid, this can be a parameter for control as well.

[0085] Second user captured expression It is also possible to recognize more complex temporal patterns across the sequence and process these in terms of time intervals.

[0086] In some instances, all captured second users expression The temporal relationship between, for example, at least one second user expression and / or at least one second user expression can be considered in the control to exclude

[0087] Captured Users expression Identifying the user input using can be done in various ways. One possibility is a rule-based scheme, where the user expressionPredefined rules are used to recognize specific features or patterns in the data and derive corresponding user inputs from them. For example, specific keywords or gestures can be permanently linked to specific inputs. Further possible approaches include, among others, pattern recognition, statistical models, knowledge-based systems, or multimodal fusion, in which case the first user expression Information from different modalities (e.g., speech and gesture) is combined to improve input recognition.

[0088] Another possibility for identifying the user input is the use of machine learning methods, in particular neural networks, i.e. machine learning models. In this case, the system determines the user's input based on training data. expression After training, the system is then trained to learn the mapping between the user inputs. expression Neural networks can also recognize complex user inputs in the data, for example based on speech recognition.

[0089] Controlling the surgical visualization system may include triggering an operation of the surgical visualization system. The type of operation may be specified by a first user input. At least one second user input may trigger and / or parameterize the operation.

[0090] Thus, a first user input specifies the type of action (e.g., "zoom"), while a second user input may initiate or trigger the execution of this action or specify the content of this action. Triggering provides increased safety with respect to erroneous actions ("two-element control"). Parameterization allows for more precise execution of actions.

[0091] Typically, the information content in a first user input is greater than the information content in a second user input. This is because the first user input typically must select a type of action from a relatively large candidate space; that is, there are often a large number of available actions to perform. On the other hand, the second user input may simply confirm the execution of a particular action or set it within a limited parameter space, and the information content of such an input is less (e.g., only clicking a button or saying "OK!"). This difference in information content can be attributed to the fact that one first user input has to select a type of action from a relatively large candidate space; that is, there are often a large number of available actions to perform. expression only during that period, whereas many second users expression This is why it is done.

[0092] The plurality of second user inputs may define coordinates in continuous space, the coordinates having variations over time. The method may further comprise applying a filter, in particular a low pass filter or a Kalman filter, to the coordinates to smooth the variations.

[0093] In some examples, the second user input may be a different second user input. expression This may include a type of second user input, such as the user's gaze direction or the user's gaze position on the display. expression provides information about the user's current region of interest, which may be characterized by the user's target position (POI) or region of interest (ROI) in the surgical field, but the information may contain noise and inaccuracies.

[0094] To obtain a more robust estimation of the POI, a low pass filter may be used in some instances.

[0095] To obtain a more robust estimation of the POI, a Kalman filter can also be used, which takes into account both signals, i.e. both types of second users. expressionThe Kalman filter is able to estimate the most likely POI, taking into account the uncertainty of the individual measurements and the dynamics of the POI over time. For this purpose, a state model can be defined that describes the position and possibly also the velocity of the POI in space. Two types of second users expression are modeled as observations with corresponding uncertainties. expression Combining data from these types can provide a more accurate and more stable estimate of the POI than would be possible with either signal alone.

[0096] At least three or at least four or more different second users expression It may be advantageous to consider the type.

[0097] The filter parameters of the filter may be based on the type of content to be displayed on the display of the surgical visualization system, and / or the phase of the surgical workflow, and / or the first user expression Type and / or Secondary User expression Type and / or primary user expression The first user input type may depend on the first user input type identified based on

[0098] The surgical visualization system may be configured to perform any method or any combination of methods according to the present disclosure, which may be performed, for example, by a control unit, i.e., a shared or dedicated computing device, of the surgical visualization system.

[0099] The control unit of the surgical visualization system is configured to perform any method or any combination of methods according to the present disclosure.

[0100] The control unit comprises a processor, memory and a controller for controlling the sensor signals and / or the user. expressionand an interface for receiving and providing the instructions, wherein the memory unit contains instructions executable by a computing unit that, when executed by the computing unit, cause the computing unit to perform the steps of any method or any combination of methods according to the present disclosure.

[0101] The computer program or computer program product comprises instructions which, when executed by a processor, cause the processor to perform the steps of any method or any combination of methods according to the present disclosure.

[0102] The described technology expression The surgical visualization system can be controlled to process and display image data based on the image data, thus providing image assistance for surgery, such as targeted navigation through an image dataset, e.g., a 2D or 3D image dataset of a person, to display a target area, but without including or requiring the steps of the surgery itself. expression Based on the capture and processing of the surgical visualization system, control commands are identified for changing the display of the surgical visualization system, and such interaction between the user and the surgical visualization system may occur before the start of a surgical procedure or even after the surgical procedure has concluded.

[0103] Although the features described in the summary above and in the detailed description below are described in relation to particular examples, it is to be understood that these features may not only be used in each combination, but may also be used individually or in any desired combination, and features from different examples may be combined with each other and therefore may be related to each other, unless specifically stated to the contrary.

[0104] As such, the above summary provides merely a brief overview of some features of some example embodiments and examples and should not be taken as limiting. Other embodiments may include features in addition to those described above.

[0105] The present invention will now be described in more detail based on preferred exemplary embodiments with reference to the accompanying drawings, in which like reference numerals designate the same or similar elements. The drawings are schematic illustrations of various embodiments of the present invention, and the elements shown in the drawings are not necessarily drawn to scale. Rather, the various elements shown in the drawings are drawn in such a way that their function or general purpose will be apparent to those skilled in the art. [Brief explanation of the drawings]

[0106] [Figure 1] 1A-1D are schematic illustrations of various exemplary surgical visualization systems; [Figure 2] 10A-10C show schematic illustrations of voice command recognition and temporal profiles of gaze position changes over several phases according to various examples. [Figure 3] 1 is a flowchart of an exemplary method. DETAILED DESCRIPTION OF THE INVENTION

[0107] The foregoing characteristics, features and advantages of the present invention and the manner in which they are achieved will become more apparent and will be more clearly understood from the following description of exemplary embodiments set forth in more detail in conjunction with the drawings.

[0108] It should be noted here that the description of the exemplary embodiments should not be understood in a limiting sense, and the scope of the present invention is not intended to be limited by the following exemplary embodiments or the drawings, which serve merely as examples.

[0109] The present invention will be described in more detail below based on preferred embodiments with reference to the drawings. The connections and couplings between the functional units and elements shown in the drawings can also be implemented as indirect connections or couplings. The connections or couplings can be implemented wired or wirelessly. The functional units can be implemented as hardware, software, or a combination of hardware and software.

[0110] Various techniques are described for surgical visualization systems. However, it should be understood that the described techniques are applicable to any system in which control or operation is performed based on multimodal user inputs with different durations and frequencies. These techniques are therefore not limited to surgery, but may also be used in a variety of human-machine interfaces, such as when interacting with computers, robots, vehicles, or other technological systems. The method according to the present invention provides a general framework for processing and fusing user inputs with different temporal characteristics.

[0111] FIG. 1 illustrates a schematic diagram of a surgical visualization system 10, according to various examples.

[0112] As can be seen in Figure 1, surgical visualization system 10 includes a controller 11 that controls the individual components of the system, in particular the display of a surgical field 13 on the system's display 2. The display in Figure 1 is an external display, but it would also be possible to display images to user 1 in an optical eyepiece, as images on a 3D monitor, in a digital eyepiece, or in AR glasses.

[0113] A patient or an examination object may be placed on the operating table 9, said object including the surgical field, captured by the imaging system 6 and displayed on the display 2 by the visualization system 10. In this example, the imaging system 6 includes an operating microscope (OPMI) for magnified imaging of the surgical field 12.

[0114] The surgical visualization system 10 may be used by different users with different durations. expression It is controlled based on multimodal user input at different times.

[0115] A user 1 of surgical visualization system 10 manipulates a surgical instrument 4. The surgical instrument 4 is positioned within a surgical field 12. The surgical field 12 is displayed by surgical visualization system 10 to user 1 on display 2 as a surgical field 13 imaged by an imaging system. The surgical field 12 is positioned within the field of view of surgical visualization system 10. Furthermore, the surgical instrument 4 is also imaged by surgical visualization system 10 and displayed to user 1 on display 2 as an imaged surgical instrument 14. The imaging system may include, for example, a camera 6.

[0116] To control the surgical visualization system, a user interacts with the system 10 using a variety of modalities.

[0117] One of these modalities is the language user expression 5, which in this case is the first user expression Represents the user's voice and is used for voice control of the system. expression is captured using microphone 8 and defines the desired control action, e.g. "focus on this area". expression 5 continues for the duration of a time interval within a longer (reference) period. expression From there, processing performed by the controller using voice recognition is used to identify voice commands indicating desired control actions for the surgical visualization system 10.

[0118] Examples of voice commands could be "go there," in which case the camera would perform, for example, a linear translational movement to center the POI in the field of view, or "look," in which case the camera would tilt around its optical midpoint to center the POI in the field of view, or "autofocus there," in which case the autofocus algorithm would focus on pixels near the POI.

[0119] language expression 5 and at least partially in parallel with the identification of the voice command or control action, a different second user expression Type multiple secondary users expression is captured.

[0120] First Second User expression The type is the gaze direction of the user 3 with respect to the representation of the surgical field on the display 2. The gaze direction is captured and tracked using an eye tracking system including a camera 7. The camera 7 can also be used to capture and track the head orientation of the user 1, in particular the forward direction.

[0121] Second User expression From the gaze direction 3 as the gaze direction 3, the known spatial setup of the surgical visualization system 10 can be used to identify a gaze position 15 on the display 2, from which, based on the imaging parameters of the system as user input, a target point 16 (or point of interest, POI) or hence a target region (region of interest, ROI) can be identified for the surgical field 12 or the displayed image data of the surgical field 12. The gaze position 15 on the display 2 therefore corresponds to a POI in the field of view of the OPMI 16, and this correspondence can be easily calculated by one skilled in the art.

[0122] Further secondary users expression The type is the position of the surgical instrument 4 relative to the surgical field 12, which is determined using an imaging system 6b that also images the surgical field. The POI 16 of the surgical field 12 can also be determined from tracking the position and / or orientation of the surgical instrument 4.

[0123] Gaze direction data and instrument position data are repeatedly acquired over a period of time and stored in a buffer with associated timestamps (as an example of time information) and are used to track the POI 16 of the user 1 within the surgical field 12.

[0124] Therefore, both the line of sight and the position of the surgical instrument 4 are controlled by the second user. expression and processed to identify a target point of interest (POI) or a target region of interest (ROI) of the user 1 within the surgical field 12.

[0125] The surgical visualization system 10 and controller 11 are designed to perform any method or any combination of methods according to the present disclosure.

[0126] In this example, the buffered POIs of both input modalities are filtered and merged, resulting in a robust and smooth estimation of the user's POI 16. The POI 16 may be present in a time series (as data points) and / or saved or buffered.

[0127] The merged POIs are stored in a buffer data structure by the control unit 11 together with their respective time information, and thus are available to each second user. expression represents a sequence of data points.

[0128] Control commands for the surgical visualization system are generated based on the identified voice commands and the buffered POI data. These commands are used to adapt the visualization to the user's intent, for example, by focusing on a specified target area. The results are presented to the surgeon on the display.

[0129] Control may include, for example, controlling the OPMI 6, such as XY movement by linearly translating the OPMI 6, rotating the OPMI camera, digitally cropping the image area, or setting autofocus at a specific position in the surgical field 12.

[0130] However, eye tracking signals can be noisy, e.g., due to the user's unconscious gaze wandering, and can be slightly distorted if, for example, the user is distracted and looks somewhere other than the intended gaze position. In this case, using raw, unfiltered gaze positions as POIs is unreliable.

[0131] Furthermore, gaze position changes over time. When eye tracking is used in conjunction with voice command recognition, there is a delay in the duration of the spoken sentence and in the voice recognition. During this period, the user's gaze position may already be changing.

[0132] To address these challenges, gaze position data and instrument position data are repeatedly acquired over a period of time and stored in a buffer data structure along with time information. By filtering, merging, and taking into account the time information of the POIs stored in the buffers of both input modalities, a more robust and smoother estimation of the user's actual points of interest within the surgical field can be achieved.

[0133] The control process takes into account the temporal correlation between individual gaze and instrument data points and voice commands. The time interval of the voice input is now incorporated to select, from the buffered time series of POIs, only those POIs relevant to the voice command for control. This improves reliability and control, as will be explained in more detail with reference to FIG. 2.

[0134] FIG. 2 shows schematically the temporal profile of the voice control process 20 and the evolution of raw and filtered gaze positions 31-34 and 41-44 over the phases of voice command recognition, according to various examples.

[0135] The top row of Figure 2 shows the languages ​​in Figure 1. expression 5 (First user expression ) shows the temporal profile of the voice control process 20 divided into different phases.

[0136] Time 21, language used by the user expression 5, e.g., represents the beginning of a spoken sentence or sentence fragment. expression The time interval between times 21 and 22 is therefore expression corresponds to a time interval of 25, during which the language expression5 is captured. Time interval 25 is within a longer period 26 that may include time before and / or after time interval 25.

[0137] Language captured between times 22 and 23 expression 5 is processed (speech recognition) and a voice command or user intent (corresponding to the first user input) is identified therefrom. At time 23, the voice recognition is completed and the voice command has been recognized.

[0138] Between time 23 and time 24, desired control actions of surgical visualization system 10 may be specified based on the voice commands, eg, the control structure may be parameterized.

[0139] Based on the recognized voice command, a control action of the surgical visualization system 10 is initiated and then executed based on the voice command, with a time delay beginning at time 24, e.g., the control action is activated or triggered at time 24.

[0140] In the center and bottom rows of FIG. 2, the user's gaze positions 31 to 34 and 41 to 44 (second user) expression Second user of type "Gaze direction" expression ) during the voice control phase. expression are shown here during time interval 25, they may also occur at least partially before and / or after time interval 25 within time period 26.

[0141] 2 shows the user's raw, unprocessed, or unfiltered gaze positions 31-34 in association with respective times 51-54 at which each unfiltered gaze position was acquired. Each gaze position 31-34 represents a measurement at a particular time. In general, time information may be associated with each of the gaze positions, which may represent, for example, a respective time or a respective duration with which each gaze position is associated.

[0142] The bottom row of Fig. 2 shows the evolution of the user's filtered gaze positions 41-44 associated with times 51-54, respectively, where the filtered gaze positions 41-44 are determined by filtering from the unfiltered gaze positions 31-34. In the example of Fig. 2, the times of the filtered gaze positions correspond to those of the unfiltered gaze positions 32-34, but other times for the filtered gaze positions 41-44 may also be determined from the times 51-54 through corresponding filtering operations. Each gaze position 32-34 and 42-44 is therefore associated with time information indicating when the respective gaze position was captured.

[0143] In this case, the gaze position is captured by the eye tracking system 7. Each point represents a gaze measurement at a particular time. It can be seen that these measurements are distributed across the display and also change during voice commands.

[0144] 2, the filtered gaze positions 41-44 have reduced noise and variability due to the application of the filter. The filtered gaze positions are concentrated by the relevant region of the display 2.

[0145] The present disclosure is directed to the fact that the gaze positions 33, 34, 43, 44 obtained after recognition of the voice command 23 are more linguistically relevant than the gaze positions 31, 41 and possibly the gaze positions 32 and 42. expression This is based on the finding that language is less relevant for the execution of control actions because it is temporally closer to the expression There is a time delay between the end 22 of the voice command and the activation 24 of the command by the control unit 10. During this delay, the gaze position continues to change. Since the gaze position changes continuously, especially after the end of the voice command, temporal correlation is very important to identify the relevant gaze position.

[0146] For example, this can occur when the user is looking at a target location and expressionThis may be the case when a practitioner starts to say a command (e.g., "move to that location" or "autofocus view there") but looks away before the entire sentence is finished. Another example is when a practitioner continues to look at a target location until the entire sentence is finished, but the voice recognition algorithm takes a second to recognize the command and activate the movement, by which time the practitioner has already looked away. In such a situation, using the last identified gaze location as the POI when a voice command is activated is unreliable.

[0147] Thus, simply using the last captured gaze position when activating a control action is problematic if the user is already looking at other areas of the display 2 at this point in time 24. Instead, it is advantageous to select a gaze position taking into account the temporal content of the voice command and the gaze position. Different forms of prioritization of different gaze positions with respect to one another, for example relative weighting, can also be implemented.

[0148] For example, gaze positions may be stored in a buffer data structure along with a timestamp. When a voice command is activated, the associated gaze position 3 may therefore be selected based on the temporal relationship between the gaze positions 3 stored in the buffer and the time interval of the voice command 5, which may not coincide with the end time of the voice command 5. Combining modalities in this way, taking into account their temporal correlation, allows for robust and reliable control of the surgical visualization system.

[0149] From the viewpoint of filtering, for example, a low-pass filter can be used as a second user expression may be applied to individual measurements to smooth the target location. Non-limiting examples of such filtering include, for example, a simple PT1 low pass filter, sensor fusion or general combination of various measurements (eye tracking, head direction, position of surgical tool in the image) to identify the POI, Kalman filter, dynamic filtering based on the phase of the voice command or filtering of the POI, which depends on the content of the image and the surgical phase.

[0150] In some examples, a low-pass filter may be applied to smooth the gaze position or point of interest (POI). This filter may be implemented in various ways to smooth the temporal evolution of the gaze point, i.e., the gaze point trajectory or POI trajectory.

[0151] In some instances, a simple PT1 low-pass filter can be used. This first-order filter can perform a weighted average of the position over time, and the influence of older measurements can decay exponentially. By choosing an appropriate time constant, the smoothing effect can be tailored to the dynamics of eye movement.

[0152] In some instances, sensor fusion may be applied, where two or different second users expression Type of second user expression are combined to identify the point of interest (POI). For this purpose, for example, data from eye tracking, head direction, and the position of the surgical tool in the image can be used. By combining this information, a more robust estimation of the POI can be achieved, which may be less affected by interference or uncertainties of the individual modalities.

[0153] In some examples, the gaze or target position can be smoothed using a Kalman filter, which is a regression algorithm that can estimate the state of a system from noisy measurements, where the state can include the position and velocity of the POI.

[0154] In some examples, target locations may be dynamically filtered based on the phase of voice control, where, for example, gaze locations may be weighted more heavily when a sentence is currently being spoken, since gaze locations may be particularly relevant to interpreting the user's intent during this phase, whereas gaze locations may be weighted less heavily during the speech recognition delay after the sentence ends, since the user may already be looking at a different image region during this phase.

[0155] In some examples, the target location may also be filtered depending on the content of the currently displayed image and the phase of the surgery. For example, in phases where particularly fine structures are visible in the image, the filter time constant may be reduced to allow for a higher resolution of the gaze position. On the other hand, in phases or overviews of coarser structures, a longer time constant may be used to achieve higher smoothing. The type of surgical intervention and the typical gaze movement patterns associated therewith may also be taken into consideration when selecting the filter parameters.

[0156] Therefore, by filtering the gaze or target positions to identify a smoothed time series of POIs, the reliability and accuracy of the control may be improved.

[0157] To increase the reliability and operability of POI estimation using gaze movements and thus improve the user experience, buffering is performed on the raw gaze positions 31-34 and / or the filtered gaze positions 42-44.

[0158] The trajectory of gaze position over time may be stored in a buffer so that the system can recall gaze positions from earlier phases of a voice command when the voice command is activated.

[0159] For this purpose, the identified unfiltered gaze positions 31-34 or filtered gaze positions 41-44 are stored in a buffer data structure (e.g., a first-in, first-out buffer data structure) for a specific period of time (e.g., 10 seconds). The system may store timestamps associated with the gaze positions stored in the buffer in the buffer data structure. The system also obtains time information characterizing the time interval, e.g., the start and end times of the time interval. expression is recognized as a valid voice command, the system selects one or more gaze positions from the buffered gaze directions that are closely related to the voice command and uses them to control the surgical visualization system.

[0160] In the example of FIG. 2, for example, the first user expression The simultaneously captured gaze directions 31 and 41 and / or gaze directions 31, 32 and 41, 42 within a valid time window 27 near the beginning or end of the time interval 25 may be used for control here.

[0161] There are various possibilities how to implement the buffer data structure. For example, the buffer data structure may store raw or filtered eye tracking positions, or gaze positions, or head forward movement directions, or points of interest (POIs) identified therefrom in the image dataset or within the surgical field. The buffer storage may, for example, store data points at the system frequency or half the system frequency, or at a fixed time variable, or at a dynamic sampling rate or a fixed sampling variable. Furthermore, the user or an AI algorithm may configure the system to use such buffer storage elements that best correspond to the time interval. For example, a timestamp may be stored in the audio expression has begun or voice expression or the time between them, or by comparing with a fixed time window with respect to the time interval between the voice commands.

[0162] FIG. 3 is a flowchart of one exemplary method for controlling a surgical visualization system.

[0163] The method begins at step S10.

[0164] In step S20, the first user expression First user of type expression The first user expression continues for a certain time interval within a certain period.

[0165] In step S30, at least one second different user expression Type multiple secondary users expression The second user expression are distributed in a distributed manner within the period and are used by different second users. expression In step S30, a plurality of second users expression Time information is also available for each of the

[0166] The data obtained in step S30 is stored in a buffer memory having a FIFO structure, for example.

[0167] In step S35, a plurality of second users expression Each of the first user expression A temporal relationship between the timestamp and the first user is identified. expression may be performed based on a comparison with the time interval in which the

[0168] In step S37, the second user expression At least one of the second users is therefore prioritized based on a temporal relationship. For example, such prioritization may be expression and one or more second users who were not selected. expression Such prioritization includes discarding the at least one second user. expression This may also include setting a relatively high weight for

[0169] In step S40, the surgical visualization system expression , at least one preferred secondary user expression and the first user expression and at least one second user expression The control is based on the temporal relationship between

[0170] The method ends in step S50. [Explanation of symbols]

[0171] 1 user 2. Display 3 Secondary users expression -View direction 4 Surgical instruments 5. Primary User expression -language expression 6 Imaging System 7. Eye Tracking System 8 microphones 9 Examination table 10 Surgical Visualization System 11 Control device 12 Surgical field 13 Displayed surgical field 14 Displayed surgical instruments 15 Gaze position 16 target points (points of interest, POI) 20 Voice Control Phase 21 hours: First user expression Start of 22 hours: First user expression End of 23 hours: Recognize voice commands 24 hours: Start of control action 25 time intervals 26 period 27 Valid Time Windows 31-33 Unfiltered gaze position 41-44 Filtered gaze position 51-54 Time associated with gaze position S10~S50 Method steps

Claims

1. 1. A computer-implemented method for controlling a surgical visualization system (10), comprising: - acquiring a first user utterance (5) of a first user utterance type, said first user utterance (5) lasting for a time interval (25) within a period (26); - acquiring and storing in a buffer a plurality of second user utterances (3) of at least one second user utterance type and respective associated time information for each of said plurality of second user utterances (3), said second user utterances (3) being distributed within said time period (26) and comprising different second user utterances (3) for each of said at least one second user utterance type; - determining a temporal relationship between each of said plurality of second user utterances (3) and said first user utterance (5) based on said time information and said time interval (25); prioritizing at least one second user utterance from said plurality of second user utterances (3) based on said temporal relationship; - controlling the surgical visualization system (10) based on the first user utterance (5) and the prioritized at least one second user utterance; 10. A computer-implemented method comprising:

2. - performing a check for each second user utterance of said plurality of second user utterances (3) to identify whether said respective temporal relations satisfy one or more predetermined check criteria; and wherein the at least one second user utterance is prioritized depending on a result of the check to determine whether the temporal relationship satisfies the one or more check criteria.

3. 3. The computer-implemented method of claim 2, wherein the one or more check criteria include a check to determine whether a time associated with a corresponding second user utterance (3) and indicated by the time information is within a specific time range before the end of the time interval (25).

4. - identifying a first user input based on said first user utterance (5); - identifying at least one second user input for said preferred at least one second user utterance; 4. The computer-implemented method of claim 1, further comprising: controlling the surgical visualization system based on the first user input and the at least one second user input.

5. controlling the surgical visualization system (10) includes triggering operation of the surgical visualization system; the type of action is specified by the first user input (5); The computer-implemented method of claim 4 , wherein the at least one second user input (3) triggers and / or parameterizes the action.

6. the plurality of second user utterances (3) define coordinates in a continuous space, the coordinates having a change during the period (26); The method comprises: applying a filter, in particular a low-pass filter or a Kalman filter, to said coordinates, thereby smoothing said variations; The computer-implemented method of any one of claims 1 to 5, further comprising:

7. 7. The computer-implemented method of claim 6, wherein filter parameters of the filter depend on a type of content displayed on a display screen (2) of the surgical visualization system (10), a phase of a surgical workflow, and / or the first user utterance type.

8. The computer-implemented method of any one of claims 1 to 7, wherein some of the plurality of second user utterances (3) are different from each other and within the time interval (25).

9. 9. The computer-implemented method according to claim 1, wherein each of the second user utterances (3) lasts for a respective duration, each of the durations being shorter than the time interval (25) during which the first user utterance (5) lasts, preferably the length of each of the durations being less than 50%, particularly preferably less than 10%, of the length of the time interval (25) during which the first user utterance (5) lasts.

10. The first user utterance (5) is - a verbal utterance made by the user, - gestures made by body parts of said user, - Touch gestures on touch interfaces, - Brain / computer interface signals, or - Multimodal combinations of the above The computer-implemented method of any one of claims 1 to 9, comprising one of:

11. 11. The computer-implemented method of claim 1, wherein the second user utterance (3) includes directing the user to a target area within a field of view of the surgical visualization system (10).

12. The second user utterance (3) is the position and / or orientation of at least one body part of said user, - the user's gaze direction (3), and the position and / or orientation of the surgical instrument (4) operated by said user (1), or - Multimodal combinations of the above The computer-implemented method of any one of claims 1 to 11, comprising one of:

13. The first user utterance (5) comprises a linguistic utterance made by the user (1), the second user utterance (3) comprises at least one first group of second user utterances comprising a gaze direction (3) of the user, and the second user utterance comprises at least one second group of second user utterances comprising a position and / or orientation of a surgical instrument (4) operated by the user (1); 5. The computer-implemented method of claim 4, wherein the at least one second user input is identified using at least a second user utterance from the first and second groups of second user utterances in each case.

14. A control unit for a surgical visualization system (10) configured to perform the method of any one of claims 1 to 13.

15. A surgical visualization system (10) including a control unit according to claim 14.