Augmented reality display processing
The integrated display processing system addresses inefficiencies in medical suites by prioritizing information and adapting display configurations based on user gaze and network conditions, enhancing operational efficiency and sterility in medical environments.
Patent Information
- Application Number
- PCT/US2025/023303
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-05
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-09
AI Technical Summary
Conventional medical interventional and surgical suites face challenges in efficiently managing and coordinating information from multiple medical systems, often resulting in replicated displays, inefficient interaction, and sterility constraints that hinder effective use.
An integrated display processing system that utilizes a head-mounted display (HMD) to manage multiple video inputs, prioritize regions of interest, adjust image and audio quality, and adapt display configurations based on user gaze direction and network bandwidth, ensuring efficient information presentation while maintaining sterility.
Enhances the coordination and efficiency of medical information display by prioritizing critical information, optimizing resource usage, and maintaining sterility, thereby improving operational effectiveness in medical environments.
Smart Images

Figure US2025023303_09102025_PF_FP_ABST
Abstract
Description
[0001] AUGMENTED REALITY DISPLAY PROCESSING
[0002] CROSS REFENCE TO RELATED APPLICATION
[0003]
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 575,390, filed on April 5, 2024, which is incorporated herein by reference in its entirety for all purposes.
[0004] TECHNICAL FIELD
[0005]
[0002] This disclosure generally relates to an integrated display processing system for an augmented reality environment.
[0006] BACKGROUND
[0007]
[0003] In conventional medical interventional and surgical suites, there are often considerations where one or more medical systems are operated or installed in various locations in or near a sterile surgical field while other medical systems are stationed in a separate non-sterile area. Each medical system may have one or more medical displays and input devices to provide control by an operator. The coordinated management of information from the multiple medical systems and coordinated control of multiple medical systems by one or more operators presents obstacles for efficient and effective use. For example, the same information may be replicated on multiple displays in multiple locations, without the ability to interact effectively with the information being displayed.
[0008] BRIEF DESCRIPTION OF THE FIGURES
[0009]
[0004] The disclosed embodiments have advantages and features which will be more readily apparent from the detailed description, the appended claims, and the accompanying figures (or drawings). A brief introduction of the figures is below.
[0010]
[0005] Figure 1 illustrates an example system environment for a display processing system according to various embodiments.
[0011]
[0006] Figure 2 illustrates displayed information from a medical system according to an embodiment.
[0012]
[0007] Figure 3 illustrates a display configuration with information from multiple medical systems according to an embodiment.
[0013]
[0008] Figure 4 illustrates an adjustment of displays according to an embodiment.
[0014]
[0009] Figure 5 illustrates gaze-based user interaction with a virtual display according to an embodiment.
[0010] Figure 6 illustrates a process for prioritizing inputs to the display processing system according to various embodiments.
[0015] SUMMARY
[0016] [Oil] In an embodiment, a method includes receiving a first video input from a first input source. The method further includes determining a region of interest of the first video input. The method further includes receiving a second video input from a second input source. The method further includes determining an arrangement of the region of interest of the first video input and the second video input in an augmented reality display. The method further includes providing the augmented reality display for presentation by a head-mounted display (HMD). The method further includes determining a gaze direction using sensor data from the HMD. The method further includes updating the augmented reality display to track the gaze direction, wherein content from the region of interest of the first video input remains at a fixed position relative to content from the second video input in the arrangement.
[0017]
[0012] In an embodiment, the method further includes determining that the first video input is higher priority than the second video input; and adjusting image quality of the first video input to be greater than image quality of the second video input.
[0018]
[0013] In an embodiment, the method further includes determining a network bandwidth of the HMD, wherein adjusting image quality of the first video input to be greater than image quality of the second video input is in response to determining that the network bandwidth is less than a threshold bandwidth.
[0019]
[0014] In an embodiment, the method further includes providing first audio associated with the first video input and second audio associated with the second video input for presentation by the HMD; and responsive to determining that the first video input is higher priority than the second video input: amplifying the first audio, and suppressing the second audio.
[0020]
[0015] In an embodiment, the determining that the first video input is higher priority than the second video input is responsive to determining that the gaze direction is directed to the content from the region of interest of the first video input.
[0021]
[0016] In an embodiment, determining the region of interest of the first video input is based on one or more image processing techniques.
[0022]
[0017] In an embodiment, determining the region of interest of the first video input is based on one or more preset locations.
[0018] In an embodiment, the method further includes determining an additional region of interest of the first video input that partially overlaps with the region of interest of the first video input.
[0023]
[0019] In an embodiment, the method further includes determining a region of interest of the second video input, wherein when updating the augmented reality display to track the gaze direction, the content from the region of interest of the first video input remains at a fixed position relative to content from the region of interest of the second video input in the arrangement.
[0024]
[0020] In another embodiment, a system includes a head-mounted display (HMD) and a non-transitory computer-readable storage medium storing instructions, the instructions when executed by one or more processors cause the one or more processors to: receive a first video input from a first input source; determine a region of interest of the first video input; receive a second video input from a second input source; determine an arrangement of the region of interest of the first video input and the second video input in an augmented reality display; provide the augmented reality display for presentation by the HMD; determine a gaze direction using sensor data from the HMD; and update the augmented reality display to track the gaze direction, wherein content from the region of interest of the first video input remains at a fixed position relative to content from the second video input in the arrangement.
[0025]
[0021] In various embodiments, a non-transitory computer-readable storage medium stores instructions that when executed by one or more processors cause the one or more processors to perform steps of any of the methods described herein.
[0026] DETAILED DESCRIPTION
[0027]
[0022] A display processing system 100 provides graphics from one or more medical systems for presentation on one or more combinations of physical or virtual displays within the operating environment. Displays may be configured to render one or more display configurations. One or more displays may be a physical display (e.g., LCD or OLED display) within the sterile or non-sterile environment. One or more displays may be a virtual display within the virtual environment of a head-mounted display (HMD). Each display may be communicatively connected by a wired or wireless connection to a computer controller or processor. In various embodiments, one or more medical systems may provide a video output (e.g., VGA, DVI, HDMI, etc.) to the system, and receive user input (e.g., keyboard, mouse, USB, etc.) from the display processing system 100. Medical systems may also provide additional bi-directional communication interface (e.g., network, KVM over IP, USB, etc.) connections to the system. A display processing system 100 may receive user input from medical systems (e.g., over ethernet, USB, Bluetooth, etc.), connected physical displays (e.g., touchscreen, mouse, keyboard, microphone, speaker, etc.), or virtual displays (e.g., gaze, voice, hand, head motion, etc.). The display processing system 100 may record to non-transitory data storage any number of inputs and outputs for synchronous playback, review, or analysis.
[0028]
[0023] In an embodiment, a method responsive to an input, configures the virtual or physical displays to a specific display configuration to support the operating physician. In some embodiments, the display configuration parameters may be modified responsive to user input. The display configuration consists of one or more displays of medical information in a specific arrangement of relative resolution, region of interest, orientation, position, or scale. In one embodiment, the input is direct input from a user interface (UI) to the display processing system 100 through a physical display. In another embodiment, the input is from a UI element with the virtual environment of the HMD, responsive to input modalities supported by the HMD (e.g. head direction, eye-direction, hand position or pose, voice, etc.). In another embodiment, the input is computed through inspection of the content or change in displayed content of one more display of the medical systems. In another embodiment, the input is received from a medical system through the bidirectional communication interface. In one embodiment, the display configuration is a fixed preset. A user can input presets to the display processing system 100 before the start of a medical procedure or during the medical procedure. Presets define one or more configurations for displayed information during the medical procedure. For example, a preset indicates what types of information are displayed (e.g., electrocardiogram, ultrasound, augmented reality models, photos, videos, or other sensor measurements) and how many different screens are displayed. A user can input this preset based on what types of information are more relevant for a certain procedure and opt not to display other types of information to reduce clutter in the display.
[0029]
[0024] In various embodiments, presets include timing information. For example, based on a timing preset, the display processing system 100 provides a particular display for a predetermined duration of time. The duration of time can be set as the full procedure. This enables the display processing system 100 to create a “virtual boom” that has a similar use to a physical boom display. Information displayed in a “virtual boom” moves together as to simulate how information in physical boom display would move together if the display (e.g., a display monitor) were relocated. A user can move their head to look around a “virtual boom” similar to how they would move to look around a physical boom display. Alternatively, if the duration of time is set to one or more specific portions of the procedure, the display processing system 100 will not provide the particular display outside of the one or more specific portions. A user can take advantage of timing presets to display information only when it is relevant for a procedure, which reduces clutter in the display when that same information is no longer needed. This can free up space in the display area for other information.
[0030]
[0025] In various embodiments, presets include sizing and 2D or 3D location information. For example, the display processing system 100 provides a virtual display at a predetermined virtual screen size or level of zoom. The display processing system 100 can also provide the virtual display at a fixed location (e.g., to avoid obstructing the view of another object or displayed information) or a location relative to a reference point in an augmented display environment of the user. Location presets can be based on a location of the user in a room. The reference point can be another screen in a virtual display, e.g., a first screen is preset to be displayed adjacent to a second screen.
[0031]
[0026] In some embodiments, the display processing system 100 determines presets to recommend to a user. For example, the display processing system 100 can determine the presets based on video resolution, text size, feature size, or historical patterns of usage and viewing of content by a certain user or a group of users.
[0032]
[0027] In another embodiment, the display processing system 100 responsive to an input, determines the medical system display of interest and adjusts the position, orientation, scale, or region of interest of the display algorithmically. In some embodiments, the display processing system 100 replicates a display configuration across one or more displays or configures a distinct configuration for one or more displays responsive to the input. In some embodiments, the display processing system 100 annotates one or more displays with a cursor (e.g., gaze, touch, mouse, etc.) representation for each user interacting with the one or more displays.
[0033]
[0028] In one embodiment, a method comprises one or more displays connected by a wired or wireless network. Responsive to the number of connected displays and display configuration, the display processing system 100 or a computer controller or processor of one or more displays applies a priority to one or more displays. The display processing system 100 may have a preset total display capacity based on system specification that may be affected by current operating conditions within the medical environment (e.g., ablation generators, x-rays, fluoroscopy, electromagnetic instrument tracking systems, etc.). Responsive to the measured transmission latency and effective image quality, the display processing system 100 may calculate total display capacity available. During periods of exceeded display capacity, the display processing system 100 may adjust the available capacity to displays based on priority. In one embodiment, the priority is assigned responsive to the connection type of the medical system to the display processing system 100. In another embodiment, the priority is assigned responsive to the display processing system 100 conducting an automated analysis of the content of the medical information to be transmitted. In these embodiments, priority may be also responsive to the current display configuration of the displays. The allocated capacity of one or more sources may be adjusted by modulating encoding parameters (e.g., bitrate, encoding quality, frame rate, resolution, latency, group of pictures, etc.). In some embodiments, the allocated capacity of one or more sources may also be adjusted by assigning a quality-of-service tag to the data during transmission. The allocated capacity of one or more sources may also be adjusted by negotiating a different network transmission channel.
[0034]
[0029] In some use cases, an operating physician must maintain sterility within the sterile field. During one period of a procedure, the operating physician may need access to specific information from one or more specific medical systems. During another period, the operating physician may need access to only a subset of the information available from one or more medical systems. During other periods, the physician may require access to information from all available medical systems for situational awareness, with a specific focus on a subset of the available data. In some environments the operating physician may require both hands on medical instruments, e.g., catheters, guidewires, etc. In other environments, the operating physician may have free use of their hands but must maintain sterility. The requirement of sterility may limit the location, number, and type of medical systems with which the operating physician can interact. The sterility requirement may also limit the interaction of the operating physician with any support personnel operating the various medical systems. The sterility and interaction limitations, as well as the physical limitations of the environment may also limit the physical configuration of the screens relative to each other, relative to the patient, and relative to the operating physician limiting efficient or ergonomic access to medical information during the procedure.
[0035]
[0030] In conventional systems, the number of displays of medical systems may exceed the capacity of the information distribution system resulting in undesired events. During some operating conditions during a case, an event may result in specific display of data from one or more medical systems may be delayed. In other operating conditions, an event may be display of data from one or more medical systems may be subject to image quality artifacts during transmission. The operating physician may prioritize significance of one or more medical information systems or regions of interest of medical information systems during these conditions. The operating physician may also prioritize significance by physically moving the display of highest priority in the center of their field of view, perpendicular to their gaze direction to reduce visual artifacts due to the operating environment (e.g., glare, foreshortening, occlusion, etc.). In some embodiments, the display processing system 100 adapts to user specific attributes (e.g., adjusting a virtual screen’s depth and position such that the virtual screen falls into the region of sharper focus in the user’s prescription multifocal vision correction glasses).
[0036] I. SYSTEM OVERVIEW
[0037]
[0031] Figure (FIG.) 1 illustrates an example system environment for a display processing system 100 according to various embodiments. The system environment shown in FIG. 1 includes the display processing system 100, one or more medical systems 110 communicatively coupled to the display processing system 100, and one or more client devices 120 which may consist of head-mounted displays (HMDs) 120a or 2-dimensional (2D) display based devices 120b (TV, smartphone, Tablet, boom display, laptop, etc.), and may be connected via network 130, e.g., the Internet, WIFI, BLUETOOTH, or another type of wired or wireless network connection. The display processing system 100 comprises one or more medical system interfaces, network interfaces, databases 105, non-transitory storage, memory, system clocks, and computer processors, among other components. The display processing system 100 may operate in an augmented reality system.
[0038]
[0032] In various embodiments, the display processing system 100 processes sensor data or user input information from medical systems 110 or client devices 120 for communication or display. In some embodiments, the display processing system 100 may use user input information from HMDs 120a (e.g., head vector, eye-vector, voice, hand gesture, etc.) to communicate a shared cursor representation of user attention or intent. In some embodiments, the display processing system 100 may use user input from 2D displays 120b (e.g., touchscreen, keyboard, mouse, etc.) to communicate a shared cursor representation of user attention or intent. In some embodiments, the display processing system 100 communicates user intent with medical systems 110 using application programming interfaces (e.g., API, TCP / IP, USB, etc.), or user input emulation (e.g., HID, IP over KVM, etc.). In various embodiments, the display processing system 100 stores any number of displays or user inputs on non-transitory storage for playback, review, or analysis. These synchronized recordings may be replayed through the display processing system 100, to allow visualization or analysis of synchronized video, user input, display configuration, etc. from the perspective of any client device 120 (e.g., 2D display, HMD, etc.). User input may be further processed for playback user input analysis (e.g., histogram, heatmap) to render additional annotations for user attention for any number of users of the display processing system 100. User input may be assisted through analysis of the displayed content to facilitate interaction in coarse input modalities. Movement of the user input cursor may be modified responsive to display content or display configuration. For example, if user interface elements are detected present within a display of a client device 120, the display cursor may snap to user interface elements to facilitate control using an HMD 120a input modality. In another embodiment, the display cursor may change appearance or behavior responsive to detected content or display configuration. For example, if electrogram information is detected within a display, the cursor may snap to specific regions of interest detected in the underlying display (e.g., detected peaks, derivative changes, regions of high contrast, etc.).
[0033] The client device 120 comprises one or more displays, computer processors, system clocks, and connections communicatively coupled with the display processing system 100. In some embodiments, a client device 120 comprises a 2D display (e.g., LCD panel) connected via a video cable (e.g., DVI or HDMI) to the system clock and computer processors of display processing system 100. In some embodiments, a client device 120 comprises one or more displays, computer processors, and system clocks are communicatively coupled with the display processing system 100 via a wired or wireless network connection. In some embodiments, these client device 120 display functions may be provided by an HMD 120a. In some embodiments, these client device 120 display functions may be provided by a laptop, desktop, all-in-one or tablet computing device with a 2D display 120b. In some embodiments, a client device 120 may contain any number of additional sensors (e.g., cameras, accelerometers, microphones, etc.) to enable user input to the display processing system 100, such as sensed head vector, eye vector, voice, and hand gesture. The client device 120 system clock, computer processors and connection support the communications, monitoring, decoding, and display functions of the client device 120.
[0039] II. REGION OF INTEREST SELECTION
[0040]
[0034] A region of interest of a display of a medical system 110 is a modifiable sub image of the entire display, up to and including the entire region of the display. In some embodiments, the region of interest is a fixed selection performed by the user by specifying a horizontal offset, vertical offset, width, and height. In other embodiments, the region of interest is automatically extracted by the display processing system 100 using image processing techniques. In some embodiments, the region of interest is extracted using user interface detection techniques by extracting user interface components from the visible display. User interface detection may be accomplished using image processing techniques such as performing edge detection and extraction, connected component identification, and automatic optical character recognition. In other embodiments the region of interest is extracted by identifying content based on appearance. In various embodiments, the display processing system 100 achieves content identification using image processing techniques, such as template matching, feature extraction (e.g., SIFT, SURF, ORB, etc.). Extracted features may be used directly, or as inputs into pre-trained machine learning networks (e.g., Convolutional Neural Networks) in addition to raw image information.
[0041]
[0035] In various embodiments, any number of regions of interest may be generated from a single display of a medical system 110. Regions of Interest may be overlapping or nonoverlapping. In some embodiments the regions-of-interest may dynamically move and resize over the course of the medical procedure (e.g., to follow a desktop window of a medical system as the user moves and / or resizes it on the medical system). Displaying a region of interest is advantageous when network bandwidth is constrained because data for the entire display frame (from image or video input) does not need to be transmitted to the HMD 120a. In addition, displaying a region of interest instead of an entire display frame requires fewer resources by the HMD 120a, which allows the HMD 120a hardware to be more compact in size and thus lighter in weight.
[0042]
[0036] Figure 2 illustrates displayed information from a medical system 110 according to an embodiment. For example, the displayed information is real-time or previously processed data collected by the medical system 110 during a medical procedure. The displayed information takes up a portion of the full display area available. The display processing system 100 can determine that the displayed information is a region of interest, or the user can manually select this region of interest.
[0043] III. DISPLAY CONFIGURATION
[0044]
[0037] Figure 3 illustrates a display configuration with information from multiple medical systems 110 according to an embodiment. In various embodiments, a display configuration comprises one or more medical system 110 display configurations and associated user input controls. Each medical system 110 display configuration comprises of any number of scale, rotation, or region of interest values. In some embodiments, responsive to user input, a previously stored display configuration may be loaded. In some embodiments, the display configuration may be loaded responsive to the content identification of any number of displays. In various embodiments, the display processing system 100 achieves content identification using image processing and classification techniques (e.g., template matching, CNN, R-CNN, OCR, etc.). Controls available for each display within a display configuration may be enabled or disabled responsive to display configuration, display client type, or identified content within the display. In various embodiments, the display processing system 100 updates the display configuration responsive to a client device 120 position or pose to reduce display foreshortening. In various embodiments, the display processing system 100 modifies the display configuration (e.g., position, pose, scale, transparency, brightness contrast, color balance, black-level, etc.) responsive to user input during operation.
[0045]
[0038] In some embodiments, the display processing system 100 automatically adjusts virtual screens presented to a user. For example, the display processing system 100 determines a depth of virtual screens to match a focal distance of the user’s HMD 120a. When there are multiple virtual screens displayed simultaneously, the display processing system 100 can also ensure that the virtual screens do not overlap each other. In another embodiment, the display processing system 100 determines to allow a certain amount of overlap between the virtual screens to reduce the distance that the user’s gaze direction needs to travel to look at the different virtual screens. The display processing system 100 can prioritize areas of a virtual screen that require more frequent monitoring and configure the display such that the prioritized areas are not obstructed. In some embodiments, the display processing system 100 determines areas that require more frequent monitoring by tracking the user’s eye movements or gaze direction. Alternatively, the display processing system 100 can determine high priority areas based on a user preset. The display processing system 100 can determine other areas of a virtual screen that are lower priority and configure the display to obstruct the lower priority areas if needed due to space constraints. The display processing system 100 can update areas of virtual screens determined to be higher or lower priority over the course of a procedure because the importance of certain information may change during the procedure.
[0046]
[0039] Figure 4 illustrates an adjustment of displays according to an embodiment. In some embodiments, the display processing system 100 updates virtual screens to continuously face toward the user. In the example illustrated in Fig. 4, the display processing system 100 updates the virtual screens shown in the top panel such that the virtual screens are angled towards the user as shown in the bottom panel. The display processing system 100 can also move the virtual screens away responsive to a user input. For example, if a user looks toward a physical object (instead of an augmented reality or virtual object), the display processing system 100 relocates virtual screens out of the user’s line of sight to avoid obstructing the physical object. The display processing system 100 can re-adjust the location of the virtual screens once the user looks away from the physical object.
[0047]
[0040] In some embodiments, the display processing system 100 uses eye tracking convergence depth to determine when a user is looking behind a virtual screen, e.g., at a real- world object or another augmented reality object. Responsive to this determination, the display processing system 100 adjusts the transparency of the virtual screen displayed to the user. By doing so, the virtual screen is de-emphasized in the display because the user is not focusing on the virtual screen while looking behind the virtual screen. The display processing system 100 can also use eye tracking convergence depth to determine when a user is looking at a virtual screen. In this situation, the display processing system 100 determines that the user is focusing on the virtual screen. Responsive to this determination, the display processing system 100 adjusts the display of the virtual screen to make it appear more salient, e.g., increasing the brightness, opaqueness, or resolution of the display.
[0048] IV. DISPLAY RENDERING
[0049]
[0041] Figure 5 illustrates gaze-based user interaction with a virtual display according to an embodiment. The client device 120 may be render displays from distinct perspectives responsive to the type of display or user input. In some embodiments, the client device 120 may render the displays within a 3D environment according to the relative position of a display of the client device 120 and the position, pose, or scale of displays within the display configuration. In some embodiments, the display processing system 100 modifies a display configuration responsive to relative user or display position to improve readability by constraining distance or perspective through rotation of the display (e.g., spherical billboarding, cylindrical billboarding) or warping of the display (e.g., curved display, flexible display, or folded display). In some embodiments, the client device 120 may render the display configuration from the perspective of another client device 120. For example, a 2D display 120b may render the display configuration from the perspective of an HMD 120a within the 3D environment. In some embodiments, the 2D display 120b may render the display configuration with displays in a flattened, normalized 2D environment. In various embodiments the cursors of any number of display clients may be rendered superimposed on the rendered display configuration in the location of user interaction. In the example shown in Fig. 5, a cursor is displayed based on the user’s gaze direction.
[0050]
[0042] Figure 6 illustrates a process for prioritizing inputs to the display processing system according to various embodiments. In various embodiments, the display processing system 100 optimizes display configurations based on one or more factors using a prioritized stream processor. For example, to reduce the computational resources required to render a display in real time, the display processing system 100 displays only one or more subsets of a virtual screen corresponding to regions of interest, instead of displaying an entire frame of the virtual screen. The regions of interest can be based on a preset, e.g., manual cropping by a user to set the coordinates or size. The display processing system 100 can also automatically determine or update a region of interest by tracking when a virtual screen is relocated or resized. In addition, the display processing system 100 can use computer vision techniques to determine whether specific virtual screens are required for a medical procedure (e.g., ablation parameter graph, catheter force widget, electrograms, live ultrasound scan) and set those virtual screens as regions of interest within a broader display. Optimizing display based on regions of interests is advantageous when network bandwidth is constrained. In some embodiments, the display processing system 100 processes full frames from a video input, not limited to regions of interest that were displayed in real time and stores the processed frames for playback at a later time. Since the playback does not need to occur in real time, the display processing system 100 can complete the processing asynchronously, e.g., when there is more bandwidth available.
[0051]
[0043] In various embodiments, the display processing system 100 adapts to low network bandwidth situations using a network monitor by prioritizing certain video inputs or client devices 120 over others that are lower priority. The display processing system 100 can determine that a video input is higher priority by determining that the user is actively looking at information from the video input. Lower priority devices may include HMDs 120a that are not currently worn by a user or devices of observers in a medical procedure. In some embodiments, the display processing system 100 encodes multiple fidelity, sample rate, or resolution versions of each input video stream, which saves bandwidth to client devices 120 that are lower priority. In some embodiments, the display processing system 100 concatenates multiple input video frames or regions of interest into an aggregate frame to reduce transmission and rendering overhead, compared to separately rendering each input video frame or region of interest.
[0052]
[0044] In various embodiments, the display processing system 100 adjusts a black level of a virtual screen (e.g., of a HMD 120a) to make augmented reality content in the virtual screen appear more readable against a real-word background. The display processing system 100 monitors a virtual screen’s displayed range of brightness and colors and can use this information to adjust the virtual screen to appear more readable. The display processing system 100 can determine that an amount of black level increase may be constant or vary over time and position. The entire frame may have the same or spatially varying levels of modification to improve readability, e.g., increasing the color contrast in a certain region of interest within a virtual screen to draw attention to that region.
[0053]
[0045] In some embodiments, the HMD 120a or other client devices 120 may include a camera. The display processing system 100 may use the camera video from the HMD 120a or client device 120 to estimate the legibility of the virtual screens against their real -world background, and alter contrast, edges, or other image corrections. Adjusting brightness and / or colors can increase or optimize legibility. For example, if a virtual screen is displayed in front of a black-painted wall with a white wall clock, the portion of the virtual screen presented in front of the wall clock may have lower legibility than the portion presented in front of the black wall. The display processing system 100 may compensate by making the portion of the virtual screen in front of the wall clock brighter.
[0054] V. PERFORMANCE FEEDBACK
[0055]
[0046] The display processing system 100 receives display information from medical system 110 connections and provides information for display of graphics or objects by a client device 120. In some embodiments, the processor of the display processing system 100 uses received display information to compare decoded video (e.g., actual image quality from a decoder and Tenderer) to determine effective image performance metrics including image quality, image latency, and encoding efficiency, e.g., using an image quality monitor. The display processing system 100 determines image quality using one or more metrics including, for example, image resolution, dynamic range, frame rate, latency, and compression metrics (e.g., Mean-squared error, Peak signal-to-noise ratio, structural similarity, among other metrics). The display processing system 100 may also receive effective image performance metrics from client devices 120 to determine effective image performance of the display processing system 100. The display processing system 100 may also receive effective network performance metrics from client devices 120 to determine effective network performance of the display processing system 100. The display processing system 100 uses effective image performance metrics, effective system image performance metrics, effective system network performance, display priorities, and display configuration to determine overall performance of the display processing system 100. Responsive to overall performance, the display processing system 100 may modulate priority and encoding rate (e.g., using encoder and network assignment) of each display to improve overall performance of the system in achieving target system performance metrics.
[0047] In some embodiments, a computer processor of client device 120 communicatively coupled with the display processing system 100 determines a common reference time for measuring display system metrics. The common reference time may be used to calculate time related metrics including jitter or latency of network transmission. The common reference time may also be used to calculate time related to overall latency of image encoding to image display in the client device 120. These metrics may be communicated to the display processing system 100 as feedback for monitoring overall system performance. VI. AUDIO RENDERING
[0056]
[0048] In various embodiments involving interventional or operating rooms, medical systems 110 communicate information using audio signals. Medical systems 110 may emit sound cues to communicate various events or states to the users (e.g., energy is being delivered to the patient, x-ray radiation is present in the room, etc.). A medical system 110 can present the visual information in a different position than where the audio speaker is presenting sound cues. Multiple medical systems 110 can simultaneously emit sounds cues. Medical systems 110 may communicate voice signals between other participants within the interventional or operating rooms (e.g., intercom systems, etc.). Medical systems 110 may communicate voice signals between participants within the interventional or operating rooms and other physically distinct locations (e.g., Voice over Internet Protocol (VOIP), etc.).
[0057]
[0049] In some embodiments, the display processing system 100 may receive audio signals from the medical systems 110 and render the received audio in the HMD 120a (using 3D spatialized audio processing) such that they appear to emanate from the same position as visual information of the medical system 110 (e.g., be it a virtual screen or physical display). This can make it more salient to the user which particular medical system 110 is emitting which audio signals. In some embodiments, when the display processing system 100 detects that the user is looking at one medical system 110 or a corresponding physical or virtual display, the display processing system 100 may amplify (e.g., increasing the volume) audio signals from that medical system 110 while partially suppressing audio coming from other medical systems 110 (e.g., lowering the volume). In some embodiments, when the display processing system 100 detects that the user is not looking at the source of an audio signal, a visual cue may be displayed to direct the user’s attention to the source of the spatialized audio signal in 3D space.
[0058] VI. ALTERNATIVE CONSIDERATIONS
[0059]
[0050] The foregoing description of the embodiments of the invention has been presented for the purpose of illustration; it is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Persons skilled in the relevant art can appreciate that many modifications and variations are possible in light of the above disclosure.
[0060]
[0051] Some portions of this description describe the embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Furthermore, it has also proven convenient at times, to refer to these arrangements of operations as modules, without loss of generality. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combinations thereof.
[0061]
[0052] Any of the steps, operations, or processes described herein may be performed or implemented with one or more hardware or software modules, alone or in combination with other devices. In one embodiment, a software module is implemented with a computer program product including a computer-readable non-transitory medium containing computer program code, which can be executed by a computer processor for performing any or all of the steps, operations, or processes described.
[0062]
[0053] Embodiments of the invention may also relate to a product that is produced by a computing process described herein. Such a product may include information resulting from a computing process, where the information is stored on a non-transitory, tangible computer readable storage medium and may include any embodiment of a computer program product or other data combination described herein.
[0063]
[0054] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.
Claims
What is claimed is:
1. A method comprising: receiving a first video input from a first input source; determining a region of interest of the first video input; receiving a second video input from a second input source; determining an arrangement of the region of interest of the first video input and the second video input in an augmented reality display; providing the augmented reality display for presentation by a head-mounted display (HMD); determining a gaze direction using sensor data from the HMD; and updating the augmented reality display to track the gaze direction, wherein content from the region of interest of the first video input remains at a fixed position relative to content from the second video input in the arrangement.
2. The method of claim 1, further comprising: determining that the first video input is higher priority than the second video input; and adjusting image quality of the first video input to be greater than image quality of the second video input.
3. The method of claim 2, further comprising: determining a network bandwidth of the HMD, wherein adjusting image quality of the first video input to be greater than image quality of the second video input is in response to determining that the network bandwidth is less than a threshold bandwidth.
4. The method of claim 2, further comprising: providing first audio associated with the first video input and second audio associated with the second video input for presentation by the HMD; and responsive to determining that the first video input is higher priority than the second video input: amplifying the first audio, and suppressing the second audio.
5. The method of claim 2, wherein determining that the first video input is higher priority than the second video input is responsive to determining that the gaze direction is directed to the content from the region of interest of the first video input.
6. The method of claim 1, wherein determining the region of interest of the first video input is based on one or more image processing techniques.
7. The method of claim 1, wherein determining the region of interest of the first video input is based on one or more preset locations.
8. The method of claim 1, further comprising: determining an additional region of interest of the first video input that partially overlaps with the region of interest of the first video input.
9. The method of claim 1, further comprising: determining a region of interest of the second video input, wherein when updating the augmented reality display to track the gaze direction, the content from the region of interest of the first video input remains at a fixed position relative to content from the region of interest of the second video input in the arrangement.
10. A non-transitory computer-readable storage medium storing instructions, the instructions when executed by one or more processors cause the one or more processors to: receive a first video input from a first input source; determine a region of interest of the first video input; receive a second video input from a second input source; determine an arrangement of the region of interest of the first video input and the second video input in an augmented reality display; provide the augmented reality display for presentation by a head-mounted display (HMD); determine a gaze direction using sensor data from the HMD; and update the augmented reality display to track the gaze direction, wherein content from the region of interest of the first video input remains at a fixed position relative to content from the second video input in the arrangement.
11. The non-transitory computer-readable storage medium of claim 10, storing further instructions that when executed by the one or more processors cause the one or more processors to: determine that the first video input is higher priority than the second video input; and adjust image quality of the first video input to be greater than image quality of the second video input.
12. The non-transitory computer-readable storage medium of claim 11, storing further instructions that when executed by the one or more processors cause the one or more processors to: determine a network bandwidth of the HMD, wherein adjusting image quality of the first video input to be greater than image quality of the second video input is in response to determining that the network bandwidth is less than a threshold bandwidth.
13. The non-transitory computer-readable storage medium of claim 11, storing further instructions that when executed by the one or more processors cause the one or more processors to: provide first audio associated with the first video input and second audio associated with the second video input for presentation by the HMD; and responsive to determining that the first video input is higher priority than the second video input: amplify the first audio, and suppress the second audio.
14. The non-transitory computer-readable storage medium of claim 11, wherein determining that the first video input is higher priority than the second video input is responsive to determining that the gaze direction is directed to the content from the region of interest of the first video input.
15. The non-transitory computer-readable storage medium of claim 10, wherein determining the region of interest of the first video input is based on one or more image processing techniques.
16. The non-transitory computer-readable storage medium of claim 10, wherein determining the region of interest of the first video input is based on one or more preset locations.
17. The non-transitory computer-readable storage medium of claim 10, storing further instructions that when executed by the one or more processors cause the one or more processors to: determine an additional region of interest of the first video input that partially overlaps with the region of interest of the first video input.
18. The non-transitory computer-readable storage medium of claim 10, storing further instructions that when executed by the one or more processors cause the one or more processors to: determine a region of interest of the second video input, wherein when updating the augmented reality display to track the gaze direction, the content from the region of interest of the first video input remains at a fixed position relative to content from the region of interest of the second video input in the arrangement.
19. A system comprising: a head-mounted display (HMD); and a non-transitory computer-readable storage medium storing instructions, the instructions when executed by one or more processors cause the one or more processors to: receive a first video input from a first input source; determine a region of interest of the first video input; receive a second video input from a second input source; determine an arrangement of the region of interest of the first video input and the second video input in an augmented reality display; provide the augmented reality display for presentation by the HMD; determine a gaze direction using sensor data from the HMD; and update the augmented reality display to track the gaze direction, wherein content from the region of interest of the first video input remains at afixed position relative to content from the second video input in the arrangement.
Citation Information
Patent Citations
World-locked display quality feedback
US20150317832A1
Devices, Methods, and Graphical User Interfaces for Displaying Objects in 3D Context
US20200357184A1
Systems and methods for pinning content items to locations in an augmented reality display based on user preferences
US20240071001A1