A method and system for capturing images of scenes, such as medical scenes, and tracking objects within those scenes.

The camera array system with integrated depth sensors and trackers addresses the challenge of accurate and low-latency object tracking in mediated reality systems by synthesizing a virtual field of view and processing a region of interest, improving surgical precision.

JP2026091852APending Publication Date: 2026-06-04PROPRIO INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PROPRIO INC
Filing Date
2026-02-27
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing mediated reality systems face challenges in accurately tracking objects within a scene while maintaining low system latency, and slight misalignments between multiple cameras can cause undesirable distortion in reconstructed images.

Method used

A camera array system with integrated depth sensors, high-resolution RGB cameras, and infrared trackers mounted on a common frame, which processes image data to synthesize a virtual field of view and track objects using a combination of optical and image-based methods, allowing for high-precision object tracking with low latency.

Benefits of technology

The system achieves high-precision, low-latency object tracking by processing only a region of interest in the image, enhancing accuracy and reducing computational complexity, thereby improving surgical procedures with enhanced visual assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026091852000001_ABST
    Figure 2026091852000001_ABST
Patent Text Reader

Abstract

A method and system for capturing images of scenes, such as medical scenes, and tracking objects within those scenes. To provide. [Solution] The camera array includes a support structure having a center and a depth sensor mounted on the support structure close to the center. The camera array may further include a plurality of cameras mounted on the support structure radially outward from the depth sensor and a plurality of trackers mounted on the support structure radially outward from the cameras. The cameras are configured to capture image data of the scene, and the trackers are configured to capture position data of tools in the scene. The image data and position data can be processed to generate a virtual view of the scene, including a graphic representation of the tools at the determined positions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] 〔Cross-Reference to Related Applications〕 This application claims the benefit of U.S. Patent Application No. 15 / 930,305, filed on May 12, 2020, entitled "METHODS AND SYSTEMS FOR IMAGING A SCENE, SUCH AS A MEDICAL SCENE, AND TRACKING OBJECTS WITHIN THE SCENE", which is hereby incorporated by reference in its entirety.

[0002] This technology generally relates to camera arrays, and more specifically to camera arrays for (i) generating a virtual field of view of a scene for a mediated reality observer and (ii) tracking objects within the scene.

Background Art

[0003] In a mediated reality system, an image processing system adds, subtracts, and / or modifies visual information representing an environment. In surgical applications, the mediated reality system enables a surgeon to view a surgical site from a desired field of view using context information that assists the surgeon in performing surgical tasks more efficiently and accurately. Such context information can include the position of objects such as surgical tools within the scene. However, it is difficult to accurately track objects while maintaining low system latency. Additionally, such mediated reality systems rely on multiple camera angles to reconstruct an image of the environment. However, even a slight relative movement and / or misalignment between multiple cameras can result in undesirable distortion in the reconstructed image.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

[0005] Many aspects of this disclosure can be better understood by referring to the following drawings. The components in the drawings are not necessarily to scale. Rather, the focus is on clearly illustrating the principles of this disclosure. [Brief explanation of the drawing]

[0006] [Figure 1] This is a schematic diagram of an imaging system according to an embodiment of this technology. [Figure 2] This is a perspective view of a surgical environment employing an imaging system for surgical application according to an embodiment of this technology. [Figure 3A] This is a side view of the camera array and movable arm of an imaging system according to an embodiment of this technology. [Figure 3B] This is an isometric view of the camera array and movable arm of the imaging system according to an embodiment of this technology. [Figure 4A] This is an isometric view of a camera array according to an embodiment of this technology. [Figure 4B] This is a bottom view of a camera array according to an embodiment of this technology. [Figure 4C] This is a side view of a camera array according to an embodiment of this technology. [Figure 5A] This is an isometric view of the camera array with the housing removed, according to an embodiment of this technology. [Figure 5B] This is a bottom view of the camera array with the housing removed, according to an embodiment of this technology. [Figure 5C] This is a side view of the camera array with the housing removed, according to an embodiment of this technology. [Figure 6] This is a front view of an imaging system in a surgical environment during surgical application, according to an embodiment of this technology. [Figure 7] This is a flowchart illustrating a process or method for tracking the tip of a tool using an imaging system, according to an embodiment of this technology. [Figure 8] This is an isometric view of a co-calibration target used in the calibration of an imaging system according to an embodiment of this technology. [Figure 9A] This is a partial schematic side view of a tool illustrating various steps of the method shown in Figure 7, according to an embodiment of this technology. [Figure 9B] This is a partial schematic side view of a tool illustrating various steps of the method shown in Figure 7, according to an embodiment of this technology. [Modes for carrying out the invention]

[0007] Aspects of this technology relate to mediated reality imaging systems generally used in surgical procedures and the like. In some embodiments described below, the imaging system includes, for example, a camera array having (i) a depth sensor, (ii) a plurality of cameras, and (iii) a plurality of trackers. The depth sensor, cameras, and trackers can each be mounted on a common frame and arranged within a housing. In some embodiments, the depth sensor is mounted near the center of the frame. The cameras can be mounted on the frame radially outward from the depth sensor and configured to capture image data of the scene. In some embodiments, the cameras are high-resolution RGB cameras. The trackers can be mounted on the frame radially outward from the cameras and configured to capture positional data of one or more objects in the scene, such as surgical tools. In some embodiments, the trackers are infrared imagers configured to image and track reflective markers attached to objects in the scene. Thus, in one aspect of this technology, the camera array may include a camera system and an optical tracking system integrated on a common frame.

[0008] The imaging system may further include a processing unit communicatively coupled to the camera array. The processing unit may be configured to synthesize a virtual image corresponding to a virtual field of view of the scene based on image data from at least a subset of the cameras. The processing unit may further determine the positions of objects in the scene based on position data from at least a subset of trackers. In some embodiments, the imaging system may further include a display device configured to display a graphic representation of the object at the determined position in the virtual image.

[0009] In some embodiments, the imaging system is configured to track the tool tip in a scene using data from both a tracker and a camera. For example, the imaging system can estimate the three-dimensional (3D) position of the tool tip based on positional data from the tracker. The imaging system can then project the estimated 3D position onto a two-dimensional (2D) image from the camera and define a region of interest (ROI) in each image based on the projected position of the tool tip. The imaging system can then process the image data within the ROI of each image to determine the position of the tool tip within the ROI. Finally, the determined position of the tool tip within the ROI of the image can be triangulated (or otherwise mapped into 3D space) to determine an updated, high-precision position of the tool tip.

[0010] In one aspect of this technology, since the camera has a higher resolution than the tracker, the tool tip position determined from camera data can be more accurate than the position determined from the tracker alone. In another aspect of this technology, tracking can be performed at a high frame rate and low latency because only the ROI in the image from the camera needs to be processed, rather than the entire image, because the ROI is initialized using a 3D estimate of the tool tip position from the tracker. Without using an ROI, the processing requirements for the image from the camera become very large, making low-latency processing difficult or impossible.

[0011] In this specification, specific details of some embodiments of the present technology will be described with reference to FIGS. 1 to 9B. However, the present technology can also be implemented without some of these specific details. In some cases, well-known structures and techniques often related to camera arrays, light field cameras, image reconstruction, object tracking, etc. are not shown in detail so as not to obscure the present technology. The terms used in the following description are intended to be interpreted in the broadest reasonable form, even if they are used in the detailed description of some specific embodiments of the present invention. Although some terms are emphasized below, terms intended to be interpreted in any limited form are expressly and specifically defined as such in this detailed description section.

[0012] The accompanying drawings show embodiments of the present technology, which are not intended to limit the scope of the present technology. The sizes of the various elements shown are not necessarily to scale, and these various elements may be arbitrarily enlarged for better readability. In the figures, details of components may be abstracted and details such as the position of components and specific exact connections between such components may be excluded when they are not necessary for a complete understanding of the method of implementation and use of the present technology. Many of the details, dimensions, angles, and other features shown in the figures are merely illustrative of specific embodiments of the present disclosure. Therefore, other embodiments can have other details, dimensions, angles, and features without departing from the spirit or scope of the present technology.

[0013] The headings shown in this specification are merely for convenience and should not be construed as limiting the subject matter disclosed.

[0014] I. Selective Embodiments of the Imaging System Figure 1 is a schematic diagram of an imaging system 100 ("System 100") according to an embodiment of the present technology. In some embodiments, System 100 may be a synthetic augmented reality system, a mediated reality imaging system, and / or a computational imaging system. In the illustrated embodiment, System 100 includes one or more display devices 104, one or more input controllers 106, and a processing unit 102 operably / communicatively coupled to a camera array 110. In other embodiments, System 100 may include further, fewer, or different components. In some embodiments, System 100 may include some features that are substantially the same as or identical to those of an imaging system disclosed in U.S. Patent Application No. 16 / 586,375, filed September 27, 2019, entitled "CAMERA ARRAY FOR A MEDIATED-REALITY SYSTEM," which is incorporated herein by reference in its entirety.

[0015] In the illustrated embodiment, the camera array 110 includes a plurality of cameras 112 (individually identified as cameras 112a - 112n) configured to each capture an image of the scene 108 from a different field of view. The camera array 110 further includes a plurality of dedicated object trackers 114 (individually identified as trackers 114a - 114n) configured to capture position data of one or more objects, such as a tool 101 having a tip 103 (e.g., a surgical tool), and track the movement and / or orientation of the object(s) through / within the scene 108. In some embodiments, the cameras 112 and trackers 114 are arranged in fixed positions and orientations (e.g., poses) relative to each other. For example, the cameras 112 and trackers 114 can be structurally fixed by an attachment structure (e.g., a frame) at predetermined fixed positions and orientations (as will be described in more detail below, for example, with reference to FIGS. 3A - 5C). In some embodiments, the cameras 112 can be arranged such that adjacent cameras 112 share overlapping views of the scene 108. Similarly, the trackers 114 can also be arranged such that adjacent trackers 114 share overlapping views of the scene 108. Thus, all or a subset of the cameras 112 and trackers 114 can have different external parameters, such as position and orientation.

[0016] In some embodiments, cameras 112 in the camera array 110 are synchronized to capture images of scene 108 substantially simultaneously (e.g., within a threshold time error). In some embodiments, all or a subset of cameras 112 may be light-field / prenooptic / RGB cameras configured to capture bright-field information (e.g., information about the intensity of light rays in scene 108 and information about the direction in which the light rays are moving in space) arising from scene 108. Thus, in some embodiments, the images captured by cameras 112 can encode depth information representing the surface shape of scene 108. In some embodiments, cameras 112 are substantially identical. In other embodiments, cameras 112 may include multiple cameras of different types. For example, different subsets of cameras 112 may have different intrinsic parameters such as focal length, sensor type, and optical components. Cameras 112 may have charge-coupled device (CCD) and / or complementary metal-oxide-semiconductor (CMOS) image sensors and associated optical systems. Such optical systems can include a variety of configurations, including larger macro lenses, microlens arrays, prisms, and / or negative lenses, as well as individual image sensors with or without lenses.

[0017] In some embodiments, the tracker 114 is an imaging device such as an infrared (IR) camera, each configured to capture images of the scene 108 from a different field of view compared to other trackers 114. Thus, the trackers 114 and cameras 112 may have different spectral sensitivities (e.g., infrared wavelengths versus visible wavelengths). In some embodiments, the tracker 114 is configured to capture image data of multiple optical markers in the scene 108 (e.g., a reference marker, a marker ball, etc.), such as a marker 105 coupled to a tool 101.

[0018] In the illustrated embodiment, the camera array 110 further includes a depth sensor 116. In some embodiments, the depth sensor 116 includes (i) one or more projectors 118 configured to project a structured light pattern onto / into the scene 108, and (ii) one or more cameras 119 (e.g., a pair of cameras 119) configured to detect the structured light projected onto the scene 108 by the projectors 118 to estimate the depth of surfaces in the scene 108. The projectors 118 and cameras 119 can operate at the same wavelength, and in some embodiments they can operate at different wavelengths than the tracker 114 and / or camera 112. In other embodiments, the depth sensor 116 and / or cameras 119 can be separate components not incorporated into an integrated depth sensor. In yet another embodiment, the depth sensor 116 may include other types of dedicated depth sensing hardware, such as a LiDAR detector, for estimating the surface shape of the scene 108. In other embodiments, the camera array 110 may omit the projector 118 and / or the depth sensor 116.

[0019] In the illustrated embodiment, the processing unit 102 includes an image processing unit 107 (e.g., an image processor, image processing module, image processing unit, etc.) and a tracking processing unit 109 (e.g., a tracking processor, tracking processing module, tracking processing unit, etc.). The image processing unit 107 is configured to (i) receive images (e.g., brightfield images, brightfield image data, etc.) captured by the camera 112 of the camera array 110, and (ii) process these images to synthesize an output image corresponding to a selected virtual camera field of view. In the illustrated embodiment, the output image corresponds to an approximation of an image of scene 108 captured by a camera positioned at any position and orientation corresponding to the virtual camera field of view. In some embodiments, the image processing unit 107 is further configured to receive depth information and / or calibration data from a depth sensor 116 and synthesize an output image based on the image, depth information and / or calibration data. Specifically, the depth information and calibration data can be used / combined with the image from the camera 112 to synthesize an output image as a 3D (or stereoscopic 2D) rendering of scene 108 as seen from the virtual camera field of view. In some embodiments, the image processing apparatus 107 may synthesize an output image using any of the methods disclosed in U.S. Patent Application No. 16 / 457,780, filed June 28, 2019, entitled “Synthesizing an Image from a Virtual Perspective Using Pixels from a Physical Imager Array Weighted Based on Depth Error Sensitivity,” which is incorporated herein by reference in its entirety.

[0020] The image processing device 107 can synthesize an output image from images captured by a subset (e.g., two or more) of the cameras 112 in the camera array 110, and does not necessarily need to utilize images from all cameras 112. For example, the processing device 102 can select a stereoscopic pair of images from two of the cameras 112 that are positioned and oriented to most closely match the virtual camera field of view for a given virtual camera field of view. In some embodiments, the image processing device 107 (and / or depth sensor 116) is configured to estimate the depth of each surface point in the scene 108 relative to a common origin and generate a point cloud and / or 3D mesh representing the surface shape of the scene 108. For example, in some embodiments, the camera 119 of the depth sensor 116 can estimate the depth information of the scene 108 by detecting structured light projected onto the scene 108 by the projector 118. In some embodiments, the image processing device 107 can estimate depth from multiview image data from camera 112 using techniques such as light field correspondence, stereo block matching, photometric symmetry, correspondence, defocus, block matching, texture-assisted block matching, and structured light, whether or not it utilizes information collected by the depth sensor 116. In other embodiments, depth can be acquired by a special set of cameras 112 that perform the methods described above at different wavelengths.

[0021] In some embodiments, the tracking processor 109 can process positional data captured by the tracker 114 to track an object (e.g., a tool 101) in the vicinity of the scene 108. For example, the tracking processor 109 can determine the position of a marker 105 in a 2D image captured by two or more trackers 114 and calculate the 3D position of the marker 105 via triangulation of the 2D positional data. Specifically, in some embodiments, the tracker 114 includes dedicated processing hardware for determining positional data, such as the centroid of the marker 105 in the captured image, from the captured image. The tracker 114 can then transmit the positional data to the tracking processor 109 so that the 3D position of the marker 105 can be determined. In other embodiments, the tracking processor 109 can receive raw image data from the tracker 114. For example, in a surgical application, the object being tracked may include surgical instruments, the hand or arm of a doctor or assistant, and / or another object to which the marker 105 is attached. In some embodiments, the processing unit 102 can recognize that the tracked object is away from the scene 108 and can apply visual effects to distinguish the tracked object, such as highlighting the object, labeling the object, or applying transparency to the object.

[0022] In some embodiments, the functions of the processing unit 102, the image processing unit 107, and / or the tracking processing unit 109 can be implemented by substantially two or more physical devices. For example, in some embodiments, a synchronization controller (not shown) controls the image displayed by the projector 118 and sends a synchronization signal to the camera 112 to ensure synchronization between the camera 112 and the projector 118, enabling high-speed, multi-frame, multi-camera structured optical scanning. Such a synchronization controller may also act as a parameter server storing hardware-specific settings such as structured optical scanning parameters, camera settings, and camera calibration data specific to the camera configuration of the camera array 110. The synchronization controller may be implemented in a separate physical device from the display controller that controls the display device 104, or these devices may be integrated with each other.

[0023] The processing unit 102 may include a processor and a non-temporary computer-readable storage medium for storing instructions that, when executed by the processor, perform functions of the processing unit 102 as described herein. While not required, aspects and embodiments of the present technology may be described in the general context of computer-executable instructions, such as routines executed by general-purpose computers, such as servers or personal computers. Those skilled in the art will understand that the present technology can be implemented in conjunction with other computer system configurations, including internet appliances, handheld devices, wearable computers, cellular or mobile phones, multiprocessor systems, microprocessor-based or programmable appliances, set-top boxes, network PCs, minicomputers, and mainframe computers. The present technology can be embodied in a dedicated computer or data processor specifically programmed, configured, or constructed to execute one or more of the computer-executable instructions described in detail below. In practice, the term “computer” (and similar terms) as used herein means any of the above-mentioned devices, as well as any network-communicating device, including data processors or consumer electronic products such as game consoles, cameras, or other electronic devices having components such as processors and network communication circuits.

[0024] This technology can also be implemented in a distributed computing environment in which tasks or modules are executed by remote processing units linked through a communication network such as a local area network ("LAN"), a wide area network ("WAN"), or the Internet. In a distributed computing environment, program modules or subroutines can be located on both local and remote memory storage devices. The embodiments of this technology described below can be stored or distributed on computer-readable media, including magnetically and optically readable and removable computer disks stored on chips (e.g., EEPROM or flash memory chips). Alternatively, embodiments of this technology can be electronically distributed over the Internet or other networks (including wireless networks). Those skilled in the art will recognize that some parts of this technology may reside on a server computer and corresponding parts on client computers. Data structures and data transmission specific to embodiments of this technology are also included in the scope of this technology.

[0025] The virtual camera field of view can be controlled by an input controller 106 that provides control inputs corresponding to the position and orientation of the virtual camera field of view. The output image corresponding to the virtual camera field of view is output to a display device 104. The display device 104 is configured to receive the output image (e.g., a composite 3D rendering of scene 108) and display the output image for viewing by one or more observers. The processing unit 102 processes the input received from the input controller 106 and the images captured from the camera array 110 to generate the output image corresponding to the virtual field of view in substantially real time (e.g., at least the same speed as the frame rate of the camera array 110) that is perceived by the observer on the display device 104. The display device 104 can also display a graphic representation of any tracked object (e.g., tool 101) in scene 108 on / within the image of the virtual field of view.

[0026] The display device 104 may include, for example, a head-mounted display device, a monitor, a computer display, and / or another display device. In some embodiments, the input controller 106 and the display device 104 are integrated into the head-mounted display device, and the input controller 106 includes a motion sensor for detecting the position and orientation of the head-mounted display device. A virtual camera field of view can then be derived that corresponds to the position and orientation of the head-mounted display device 104 at a calculated depth (e.g., calculated by the depth sensor 116) within the same reference frame, so that the virtual field of view corresponds to the field of view seen by an observer wearing the head-mounted display device 104. Thus, in such embodiments, the head-mounted display device 104 can provide a real-time rendering of the scene 108 as seen by an observer not wearing the head-mounted display device 104. Alternatively, the input controller 106 may include a user-controlled control device (e.g., a mouse, pointing device, handheld controller, gesture recognition controller, etc.) that allows the observer to manually control the virtual field of view displayed by the display device 104.

[0027] Figure 2 is a perspective view of a surgical environment employing System 100 for surgical application according to an embodiment of the present technology. In the illustrated embodiment, a camera array 110 is positioned on a scene 108 (e.g., a surgical site) and supported / positioned via a movable arm 222 operably coupled to a workstation 224. In some embodiments, the arm 222 can be manually moved to position the camera array 110, and in other embodiments, the arm 222 can be robotically controlled in response to an input controller 106 (Figure 1) and / or another controller. In the illustrated embodiment, the display device 104 is a head-mounted display device (e.g., a virtual reality headset, an augmented reality headset, etc.). The workstation 224 may include a computer that controls various functions of the processing unit 102, the display device 104, the input controller 106, the camera array 110 and / or other components of System 100 shown in Figure 1. Thus, in some embodiments, the processing unit 102 and the input controller 106 are integrated into the workstation 224, respectively. In some embodiments, the workstation 224 includes a secondary display 226 that can display a user interface for performing various configuration functions, a mirror image of the display on the display device 104, and / or other useful visual images / instructions.

[0028] II. Selective Embodiments of Camera Arrays Figures 3A and 3B are side and isometric views, respectively, of the camera array 110 and arm 222 of Figures 1 and 2 according to an embodiment of the present technology. Referring to both Figures 3A and 3B, in the illustrated embodiment, the camera array 110 is movably coupled to a base 330 via a plurality of rotatable joints 332 (each individually identified as the first to fifth joints 332a to 332e) and extensions 334 (each individually identified as the first extension 334a and the second extension 334b). The base 330 can be securely mounted to a movable dolly / cart or the like at a desired location, such as in an operating room (e.g., the floor of the operating room or other rigid part). The joints 332 allow the camera array 110 to articulate and / or rotate relative to the scene 108 so that the camera 112 and tracker 114 (Figure 1) can be positioned to capture data from different parts / volumes of the scene 108. Referring to Figure 3B, for example, the first joint 332a allows the camera array 110 to rotate around axis A1, the second joint 332b allows the camera array 110 to rotate around axis A2, and so on. The joints 332 can be controlled manually (e.g., by a surgeon, operator, etc.) or robotically. In some embodiments, the arm 222 has more than three degrees of freedom so that it can be positioned in any selected orientation / position relative to the scene 108. In other embodiments, the arm 222 may include more or fewer joints 332 and / or extensions 334.

[0029] Figures 4A to 4C are isometric, bottom, and side views, respectively, of the camera array 110 according to an embodiment of the present technology. Referring to both Figures 4A to 4C, the camera array 110 includes a housing 440 (e.g., shell, casing, etc.) that surrounds the various components of the camera array 110. Figures 5A to 5C are isometric, bottom, and side views, respectively, of the camera array 110 with the housing 440 removed, according to an embodiment of the present technology.

[0030] Referring together to Figures 5A to 5C, the camera array 110 includes a support structure such as a frame 550 to which the cameras 112 (individually identified as the first to fourth cameras 112a to 112d), the trackers 114 (individually identified as the first to fourth trackers 114a to 114d), and the depth sensor 116 are coupled (e.g., mounted and securely attached). The frame 550 can be formed of metal, composite material, or other suitable strong and rigid material. The cameras 112, trackers 114, and depth sensor 116 can be coupled to the frame via bolts, brackets, adhesives, and / or other suitable fasteners. In some embodiments, the frame 550 is configured to function as a heat sink for the cameras 112, trackers 114, and / or other electronic components of the camera array 110, for example, to uniformly distribute heat around the camera array 110 and minimize thermally induced deflection / deformation.

[0031] In the illustrated embodiment, a depth sensor 116, including a projector 118 and a pair of cameras 119, is coupled to the central portion (e.g., radially inward) of the frame 550 and is generally aligned along the central axis AC of the frame 550 (Figure 5B). In one aspect of the technology, the placement of the depth sensor 116 in the center or near the center of the camera array 110 can help ensure that the scene 108 (Figure 1) is sufficiently illuminated by the projector 118 for depth estimation during operation.

[0032] The cameras 112 and trackers 114 can be distributed radially outward from the depth sensor 116 around the frame 550. In some embodiments, the trackers 114 are mounted on the frame radially outward from the camera 112. In the illustrated embodiment, the cameras 112 and trackers 114 are arranged symmetrically / evenly around the frame 550. For example, each of the cameras 112 and trackers 114 can be equally spaced from (i) the central axis AC and (ii) the longitudinal axis AL extending perpendicular to the central axis AC. In one aspect of the technology, this spacing can simplify the processing performed by the processing unit 102 (Figure 1) when synthesizing the output image corresponding to the virtual camera field of view of scene 108, as described in detail above. In another aspect of the technology, the arrangement of the cameras 112 can roughly maximize the parallax of the cameras 112, which can help facilitate depth estimation using image data from the cameras 112. In other embodiments, the camera array 110 may include more or fewer cameras 112 and / or trackers 114, and / or the cameras 112 and trackers 114 may be arranged in different ways around the frame 550.

[0033] In the illustrated embodiment, the camera 112 and tracker 114 are oriented / angled inward toward the central portion of the frame 550 (e.g., toward axes AC and AL). In other embodiments, the frame 550 can be configured (e.g., molded, angled) to orient the camera 112 and tracker 114 inward without requiring them to be angled relative to the frame 550. In some embodiments, the camera 112 can be approximately in focus on a first focal point in the scene 108, and the tracker 114 can also be approximately in focus on a second focal point in the scene 108, which may be different from or the same as the first focal point of the camera 112. In some embodiments, the field of view of each camera 112 can at least partially overlap with the field of view of one or more other cameras 112, and the field of view of each tracker 114 can at least partially overlap with the field of view of one or more other trackers 114. In some embodiments, the field of view of individual cameras 112 can be selected to vary the effective spatial resolution of the camera 112 (for example, through the selection of an accessory lens). For example, the field of view of the camera 112 can be narrowed to increase the effective spatial resolution of the camera 112 and the resulting accuracy of the system 100.

[0034] In the illustrated embodiment, the cameras 112 are identical and have, for example, the same focal length, depth of field, resolution, color characteristics, and other unique parameters. In other embodiments, some or all of the cameras 112 may be different. For example, the first and second cameras 112a, 112b (e.g., the first pair of cameras 112) may have different focal lengths or other characteristics than the third and fourth cameras 112c, 112d (e.g., the second pair of cameras 112). In some such embodiments, the system 100 can render / generate a stereoscopic view independently for each pair of cameras 112. In some embodiments, the cameras 112 may have a resolution of about 10 megapixels or more (e.g., 12 megapixels or more). In some embodiments, the cameras 112 may have relatively small lenses (e.g., about 50 millimeters) compared to typical high-resolution cameras.

[0035] Referring to Figures 4A to 5C, the housing 440 includes a bottom surface 442 having (i) a first opening 444 aligned with the camera 112, (ii) a second opening 446 aligned with the tracker 114, and (iii) a third opening 448 aligned with the depth sensor 116. In some embodiments, some or all of the openings 444, 446, and 448 can be covered with a transparent panel (e.g., glass or plastic panel) to prevent dust, dirt, etc. from entering the camera array 110. In some embodiments, the housing 440 is configured (e.g., molded) so that the transparent panel is positioned perpendicular to the angles of the camera 112, tracker 114, and depth sensor 116 to reduce distortion of the captured data caused by reflection, diffraction, scattering, etc. of light passing through the transparent panel across each of the openings 444, 446, and 448.

[0036] Referring again to Figures 5A and 5C, the camera array 110 may include integrated electrical components, communication components, and / or other components. For example, in the illustrated embodiment, the camera array 110 further includes a circuit board 554 (e.g., a printed circuit board) and an in / out (I / O) circuit box 556 coupled to the frame 550. The I / O circuit box 556 can be used to connect the camera 112, tracker 114, and / or depth sensor 116 to other components of the system 100, such as a processing unit 102, via one or more connectors 557 (Figure 5B).

[0037] Figure 6 is a front view of the system 100 in a surgical environment during a surgical application according to an embodiment of the present technology. In the illustrated embodiment, the patient 665 is positioned at least partially within a scene 108 below the camera array 110. The surgical application may be a procedure performed on a part of the patient's body of interest, such as a spinal procedure performed on the spine 667 of the patient 665. The spinal procedure may be, for example, a spinal fixation procedure. In other embodiments, the surgical application may target another part of the patient's body of interest.

[0038] Referring together to Figures 3A to 6, in some embodiments, the camera array 110 can be moved to a position above the patient 665 by articulating / moving one or more of the joints 332 and / or extension portions 334 of the arm 222. In some embodiments, the camera array 110 can be positioned such that the depth sensor 116 is roughly aligned with the patient's spine 677 (for example, the spine 677 is roughly aligned with the central axis AC of the camera array 110). In some embodiments, the camera array 110 can be positioned such that the depth sensor 116 is located at a distance D above the spine 677 of the patient 665, corresponding to the primary focal depth / plane of the depth sensor 116. In some embodiments, the focal depth D of the depth sensor is approximately 75 centimeters. In one aspect of the art, this positioning of the depth sensor 116 ensures accurate depth measurement, which facilitates accurate image reconstruction of the spine 667.

[0039] In the illustrated embodiment, each camera 112 has a field of view 664 of scene 108, and each tracker 114 has a field of view 666 of scene 108. In some embodiments, the fields of view 664 of the cameras 112 can overlap at least partially with each other to define an imaging volume together. Similarly, the fields of view 666 of the trackers 114 can overlap at least partially with each other (and / or with the fields of view 664 of the cameras 112) to define a tracking volume together. In some embodiments, the trackers 114 are positioned to maximize the overlap of the fields of view 666, and the tracking volume is defined as the volume over which all fields of view 666 overlap. In some embodiments, the tracking volume is larger than the imaging volume because (i) the field of view 666 of the trackers 114 is larger than the field of view 664 of the cameras 112, and / or (ii) the trackers 114 are positioned further radially outward along the camera array 110 (e.g., near the periphery of the camera array 110). For example, the field of view 666 of the tracker 114 can be approximately 82 × 70 degrees, and the field of view 664 of the camera 112 can be approximately 15 × 15 degrees. In some embodiments, the overlapping regions are tiled so that the fields of view 666 of the cameras 112 do not completely overlap, and there is a selective volume in which the imaging volume covered by all cameras 112 exists as a subset of the volume covered by the tracker 114. In some embodiments, each camera 112 has a focal axis 668, which converges at a point generally below the focal depth D of the depth sensor 116 (for example, about 5 centimeters below the focal depth D of the depth sensor 118). In one aspect of the art, the convergence / alignment of the focal axes 668 can generally maximize the disparity measurements between the cameras 112. In another aspect of this technology, the arrangement of cameras 112 centered on camera array 110 provides high-angle resolution of the patient's 665 spine 667, enabling the processing unit 102 to reconstruct a virtual image of the scene 108 including the spine 667.

[0040] III. Selective Embodiments of High-Precision Object Tracking Referring again to Figure 1, the system 100 is configured to track one or more objects, such as the tip 103 of a tool 101 in a scene 108, via (i) an optical-based tracking method using a tracker 114, and / or (ii) an image-based tracking method using a camera 112. For example, a processing unit 102 (e.g., a tracking processing unit 109) can process data from the tracker 114 to determine the position (e.g., position and orientation) of a marker 105 in a scene 108. Specifically, the processing unit 102 can triangulate the three-dimensional (3D) position of the marker 105 from images captured by multiple trackers 114. The processing unit 102 can then estimate the position of the tip 103 of the tool 101 based on a known (e.g., predetermined, calibrated, etc.) model of the tool 101, for example, by determining the centroids of various markers 105 and applying a known offset between the centroids and the tip 103 of the tool 101. In some embodiments, the tracker 114 operates at a wavelength (e.g., near-infrared) that allows for easy identification of the marker 105 in the image from the tracker 114, significantly simplifying the image processing required to identify the position of the marker 105.

[0041] However, in order to track a rigid body such as tool 101, at least three markers 105 must be attached so that the system 100 can track the center of gravity of various markers 105. In many cases, due to practical constraints, multiple markers 105 must be placed on the opposite side of the tip 103 of tool 101 (e.g., the working part of tool 101) without interfering with the user, while remaining visible when tool 101 is gripped by the user. Therefore, the known offset between the marker 105 and the tip 103 of tool 101 must be relatively large in order to maintain visibility of the marker 105, and any error in the determined position of the marker 105 will propagate along the length of tool 101.

[0042] In addition to or instead of this, the processing unit 102 can also process image data from the camera 112 (e.g., visible wavelength data) to determine the position of the tool 101. Such image-based processing can achieve relatively higher accuracy than optical-based methods using the tracker 114, but the frame rate is lower due to the complexity of the image processing. This is especially true for high-resolution images such as those captured by the camera 112. Specifically, the camera 112 is configured to capture high-frequency details on the surface of the scene 108, which function as feature points that represent characteristics of the object being tracked. However, there is a tendency for the computational requirements to increase further due to an excessive number of image features that must be filtered to reduce erroneous correspondences that degrade tracking accuracy.

[0043] In some embodiments, the system 100 is configured to track the tip 103 of the tool 101 with high accuracy and low latency by using tracking information from both the tracker 114 and the camera 112. For example, the system 100 can (i) process data from the tracker 114 to estimate the position of the tip 103, (ii) define a region of interest (ROI) in the image from the camera 112 based on this estimated position, and then (iii) process the ROI in the image to determine the position of the tip 103 with higher accuracy (e.g., sub-pixel accuracy) than the estimated position from the tracker 114. In one aspect of the technique, since the ROI contains only a small portion of the image data from the camera 112, the image processing of the ROI is computationally inexpensive and fast.

[0044] Specifically, Figure 7 is a flowchart of a process or method 770 according to an embodiment of the present technology for tracking the tip 103 of a tool 101 using tracking / location data captured by a tracker 114 and image data captured by a camera 112. Some features of method 770 will be described for illustrative purposes in the context of the embodiments shown in Figures 1 to 6, but those skilled in the art will readily understand that method 770 can also be carried out using other suitable systems and / or apparatus described herein. Similarly, although this specification refers to tracking a tool 101, method 770 can also be used to track all or part of other objects in a scene 108 including reflective markers (e.g., a surgeon's arm, further tools, etc.).

[0045] Method 770 includes, in block 771, internally and externally calibrating the system 100, and calibrating the parameters of the tool 101 to enable accurate tracking of the tool 101. In the illustrated embodiment, the calibration includes blocks 772-775. Method 770 includes, in blocks 772 and 773, calibrating the camera 112 and tracker 114 of the system 100, respectively. In some embodiments, the processing unit 102 performs a calibration process for the camera 112 and tracker 114 to detect the respective positions and orientations of the camera 112 / tracker 114 in 3D space with respect to a shared origin and / or the amount of overlap of their respective fields of view. For example, in some embodiments, the processing unit 102 may (i) process captured images including reference markers placed in the scene 108 from each of the cameras 112 / tracker 114, and (ii) perform optimizations across camera parameters and distortion coefficients to minimize reprojection errors of keypoints (e.g., points corresponding to reference markers). In some embodiments, the processing unit 102 can perform the calibration process by correlating feature points across different camera views. The features to be correlated may be, for example, reflective marker centroids from a binary image, or scale-invariant feature transforms (SIFT) features from a grayscale or color image. In some embodiments, the processing unit 102 can extract feature points from a ChArUco target and process them using an OpenCV camera calibration routine. In other embodiments, such calibration can be performed using a Halcon circle target or other custom targets with well-defined feature points having known locations. In some embodiments, further calibration refinement can be performed using bundle analysis and / or other suitable techniques.

[0046] Method 770 includes co-calibrating camera 112 and tracker 114 in block 774 so that tool 101 can be tracked within a common reference frame using data from both camera 112 and tracker 114. In some embodiments, camera 112 and tracker 114 can be co-calibrated based on imaging of a known target in scene 108. For example, Figure 8 is an isometric view of a co-calibrated target 880 according to an embodiment of the art. In some embodiments, the spectral sensitivities of camera 112 and tracker 114 do not overlap. For example, camera 112 may be a visible wavelength camera and tracker 114 may be an infrared imager. Thus, in the illustrated embodiment, target 880 is a multispectral target including (i) a pattern 882 visible to camera 112 and (ii) a plurality of retroreflective markers 884 visible to tracker 114. Pattern 882 and marker 884 share a common origin and coordinate frame so that camera 112 and tracker 114 can co-calibrate to measure their positions (e.g., of tool 101) in the common origin and coordinate frame. That is, the resulting external co-calibration of camera 112 and tracker 114 can be represented in the common reference frame or using the measured transformation between these reference origins. In the illustrated embodiment, pattern 882 is a printed black and white Halcon circle target pattern. In other embodiments, pattern 882 can be another black and white (or other high-contrast color combination) ArUco, ChArUco, or Halcon target pattern.

[0047] In other embodiments, the target 880 measured by camera 112 and tracker 114 does not need to be precisely aligned and can be determined separately using a hand-eye calibration technique. In yet another embodiment, the ink or material used to form two high-contrast regions of pattern 882 can exhibit similar absorption / reflection to the measurement wavelength used by both camera 112 and tracker 114. In some embodiments, blocks 772-774 can be combined into a single calibration step based on imaging of a target 880 configured (e.g., molded, sized, and precisely manufactured) to allow uniform sampling of calibration points over a desired tracking volume.

[0048] Method 770 includes calibrating the tool 101 (and / or any further object being tracked) in block 775 to determine the position of the tip 103 relative to the spindle of the tool 101 and various attached markers 105. In some embodiments, the calibration of the system 100 (block 771) may be performed only once, provided that the cameras 112 and trackers 114 remain spatially fixed (e.g., firmly fixed to the frame 550 of the camera array 110) and their optical properties do not change. However, the optical properties of the cameras 112 and trackers 114 may change slightly due to vibration and / or thermal cycling. In such cases, the system 100 can be recalibrated.

[0049] Blocks 776-779 show processing steps for determining the position of the tip 103 of tool 101 in scene 108 with high accuracy and low latency. Figures 9A and 9B are partial schematic side views of tool 101 showing various steps of method 770 of Figure 7 according to an embodiment of the present technology. Accordingly, some aspects of method 770 will be described in the context of Figures 9A and 9B.

[0050] Method 770 includes, in block 776, estimating the 3D position of the tip 103 of the tool 101 using the tracker 114. For example, the tracker 114 can process acquired image data to determine the centroid of the marker 105 in the image data. The processing unit 102 can (i) receive centroid information from the tracker 114, (ii) triangulate the centroid information to determine the 3D position of the marker 105, (iii) determine the principal axis of the tool 101 based on the calibration of the tool 101 (block 775), and then (iv) estimate the 3D position of the tip 103 based on the principal axis and the calibrated offset of the tip 103 relative to the marker 105. For example, as shown in Figure 9A, the system 100 estimates the position and orientation of the tool 101 based on the determined / measured position of the marker 105 (e.g., shown as a dashed line as the tool position 101' with respect to the orthogonal XYZ coordinate system) and models the tool 101 as having a principal axis AP. Next, system 100 estimates the position of tip 103 (indicated as tip position 103') based on a calibrated offset C from marker 105 along the main axis AP (e.g., from the centroid of marker 105). Data from at least two trackers 114 is required so that the position of marker 105 can be triangulated from the position data. In some embodiments, system 100 can estimate the position of tip 103 using data from each tracker 114. In other embodiments, the processing performed to estimate the 3D position of tip 103 can be divided differently between the tracker 114 and the processing unit 102. For example, the processing unit 102 can be configured to receive raw image data from the tracker and determine the centroid of the marker in the image data.

[0051] Method 770 includes, in block 777, defining a region of interest (ROI) in images from one or more cameras 112 based on the estimated tip position 103 determined in block 776. For example, as shown in Figure 9B, system 100 can define an ROI 986 around the estimated 3D tip position 103'. Specifically, the estimated 3D tip position 103' is used to initialize a 3D volume (e.g., a cube, sphere, cuboid, etc.) with determined limit dimensions (e.g., radius, area, diameter, etc.). The 3D volume is then mapped / projected onto 2D images from camera 112. In some embodiments, the limit dimensions can be fixed based on, for example, the known geometry of system 100 and the motion parameters of tool 101. As further shown in Figure 9B, the actual 3D position of the tip 103 of tool 101 may differ from the estimated tip position 103' due to measurement errors (e.g., propagating along the length of tool 101). In some embodiments, the dimensions and / or shape of ROI986 are selected such that the actual position of tip 103 always or almost always lies within ROI986. In other embodiments, as will be described in detail below with reference to block 778, the system 100 may initially define ROI986 to have a minimum size and iteratively increase the size of ROI986 until it is determined that the position of tip 103 lies within ROI986.

[0052] In some embodiments, ROI processing can be performed on data from only one camera 112, such as one camera 112 specifically positioned to capture images from tool 101. In other embodiments, ROI processing can be performed on two or more (e.g., all) cameras 112 in the camera array 110. That is, an ROI can be defined within one or more images from each camera 112.

[0053] Method 770 includes determining the location of the tip 103 of the tool 101 within the (single or multiple) ROI in block 778. In some embodiments, the processing unit 102 can determine the location of the tip 103 by identifying a set of feature points directly from the ROI image using the Scale Invariant Feature Transform (SIFT) method, the Speeded Up Robust Features (SURF) method, and / or the Oriented Fast and Rotated Brief (ORB) method. In other embodiments, the processing unit 102 can use a histogram to pinpoint the location of the tip 103 within the (single or multiple) ROI. In yet another embodiment, the processing unit 102 can determine the location of the tip of the tool 101 along the principal axis by (i) determining / identifying the principal axis of the tool 101 using, for example, the Hough transform or principal component analysis (PCA), and then (ii) searching for the tip 103 along the principal axis using, for example, a method using feature points or image gradients (e.g., a Sobel filter). In yet another embodiment, the processing unit 102 can utilize a gradient-based method to enable sub-pixel localization of the tip 103.

[0054] Finally, in block 779, the determined (single and multiple) positions of the tip 103 in the (single and multiple) ROI are used to determine an updated / refined 3D position of the tip 103 that is more accurate than the position estimated by the tracker 114 in block 776, for example. In some embodiments where system 100 processes image data from multiple cameras 112 (e.g., the ROI is determined within images from multiple cameras 112), the 3D position of the tip 103 can be directly triangulated based on the determined position of the tip 103 in the 2D images from the cameras 112. In other embodiments where system 100 processes image data from only one camera 112 (e.g., the ROI is determined within images from only one camera 112), the 3D position of the tip 103 can be determined by using the calibration of the camera 112 to project the position of the tip 103 in the 2D images from the camera 112 into a 3D line (block 772). In some embodiments, the system 100 then determines the position of the tip 103 as the closest point or intersection of (i) this 3D line determined by the tip position in the camera image and (ii) the spindle AP (Figure 9A) determined from the tracker 114 (block 776) and the calibration of the tool 101 (block 775).

[0055] In one aspect of this technology, the system 100 can determine the updated 3D position of the tip 103 with greater accuracy by processing image data from multiple cameras 112 rather than from only one camera 112. Specifically, the refined 3D position of the tip 103 is determined directly by triangulation using multiple cameras 112 and does not depend on data from the tracker 114 (e.g., the position and orientation of the main axis AP). In another aspect of this technology, the use of multiple cameras 112 increases the flexibility of the layout / arrangement of the camera array 110, as the relative orientation between the camera array 110 and the tool 101 is not constrained as long as the tool 101 is visible to at least two cameras 112.

[0056] In another aspect of this technology, camera 112 has a higher resolution than tracker 114, and although tracker 114 can cover a wider field of view than camera 112, it has a lower resolution, so the position of tip 103 determined from camera 112 (block 779) is more accurate than the position determined from tracker 114 (block 776). Furthermore, tracker 114 usually works well when marker 105 is imaged with enough pixels and the edge between marker 105 and the background of scene 108 is clear. On the other hand, camera 112 can have a very high resolution, such that it has a much higher effective spatial resolution while covering a similar field of view to tracker 114.

[0057] In some embodiments, after the system 100 has determined the 3D position of the tip 103, it may overlay a graphic representation of the tool 101 (for example, provided on a display device 104) onto a virtual rendering of the scene 108. Method 770 can then return to block 776 to update the 3D position of the tool 101 in real time or near real time with high accuracy. In one aspect of the technique, since the 3D estimate of the position of the tip 103 from the tracker 114 is used to initialize the ROI, it is only necessary to process the ROI in the image data from the camera 112, rather than the entire image, allowing for updates at high frame rates and low latency. Without using an ROI, the processing requirements for the image from the camera 112 would be very high, making low-latency processing difficult or impossible. Alternatively, the resolution of the camera 112 could be reduced to lower the processing requirements, but the resulting improvement in accuracy of the system would be negligible compared to the case of the tracker 114 alone.

[0058] IV. Further Examples The following examples illustrate multiple embodiments of the present technology. 1. A camera array, A support structure with a center, A depth sensor attached to the support structure close to the center, Multiple cameras are mounted on a support structure facing radially outward from the depth sensor and configured to capture image data of the scene, Multiple trackers are mounted on a support structure facing radially outward from the camera and configured to capture positional data of tools in the scene, A camera array equipped with a camera array.

[0059] 2. The camera array according to Example 1, wherein the multiple cameras include four cameras, and the multiple trackers include four trackers, and the cameras and trackers are arranged symmetrically around a support structure.

[0060] 3. The camera array according to Example 1 or Example 2, wherein the support structure includes a central axis and a longitudinal axis extending perpendicular to the central axis, the depth sensors are aligned along the central axis, and each camera and each tracker is equally spaced from the central axis and the longitudinal axis.

[0061] 4. The tracker is an infrared imager, as described in any one of Examples 1 to 3 of the camera array.

[0062] 5. The camera array according to any one of Examples 1 to 4, wherein the tracker is an imager having a different spectral sensitivity than the camera.

[0063] 6. A camera array according to any one of Examples 1 to 5, wherein the camera and the tracker each have fields of view, the camera's field of view at least partially overlapping to define the imaging volume, and the tracker's field of view at least partially overlapping to define a tracking volume larger than the imaging volume.

[0064] 7. A camera array according to any one of Examples 1 to 6, wherein each camera is angled radially inward toward the center of the support structure.

[0065] 8. The camera array according to Embodiment 7, wherein the depth sensor has a focal plane, each camera has a focal axis, and the focal axes of the cameras converge at a point below the focal plane of the depth sensor.

[0066] 9. A mediating reality system, A support structure with a center, A depth sensor attached to the support structure close to the center, Multiple cameras are mounted on a support structure facing radially outward from the depth sensor and configured to capture image data of the scene, Multiple trackers are mounted on a support structure facing radially outward from the camera and configured to capture positional data of tools in the scene, A camera array including, An input controller configured to control the position and orientation of the virtual field of view of the scene, The camera array and input controller are communicatively coupled, A virtual image corresponding to the virtual field of view is synthesized based on image data from at least two of the cameras. The tool's position is determined based on location data from at least two of the trackers. A processing apparatus configured as follows, A display device that is communicatively coupled to a processing unit and configured to display a graphic representation of a tool at a determined position in a virtual image, A mediating reality system that provides support.

[0067] 10. The mediating reality system according to Embodiment 9, wherein the processing device is further configured to determine the position of a tool based on image data from at least one of the cameras.

[0068] 11. The processing device is The initial 3D position of the tool is estimated based on positional data from at least two of the trackers. Based on the tool's initial 3D position, define the region of interest within the image data from at least one of the cameras. The image data in the region of interest is processed to determine the position of the tool in the region of interest. Determine the updated 3D position of the tool based on the determined position of the tool in the region of interest. A mediating reality system according to Example 10, configured to determine the position of a tool by means of...

[0069] 12. A method for imaging a subject within a scene, The camera array's depth sensor is aligned with the subject's area of ​​interest so that the area of ​​interest is positioned near the depth of focus of the depth sensor. This involves using multiple cameras in a camera array to capture image data of a scene that includes the subject's area of ​​interest, Using multiple trackers in a camera array to capture positional data of tools in the scene, It receives input regarding the selected position and orientation of the virtual field of view of the scene, Based on image data, a virtual image corresponding to the virtual field of view of the scene is synthesized, Determining the tool's position based on location data, In a display device, the graphical representation of a tool at a determined position within a virtual image is displayed, A method that includes this.

[0070] 13. The method according to Example 12, wherein the scene is a surgical scene and the subject's area of ​​interest is the subject's spine.

[0071] 14. The method according to Example 12 or Example 13, wherein each camera has a focal axis, and the focal axes of the cameras converge at a point below the depth of focus of the depth sensor.

[0072] 15. Determining the position and orientation of the tool is further based on the image data, according to the method of any one of Examples 12-14.

[0073] 16. Determining the position of the tool is This involves estimating the initial 3D position of the tool based on location data, and Based on the tool's initial 3D position, one or more regions of interest are defined within the image data, Processing image data in one or more regions of interest to determine the position of a tool in one or more regions of interest, Determining the updated 3D position of a tool based on the determined position of the tool in one or more regions of interest, The method according to Example 15, including the method described in Example 15.

[0074] 17. A method for determining the position of the tip of a tool in a scene, The tool's location data must be received from at least two trackers, This involves estimating the 3D position of the tool's tip based on positional data, Receiving images of the scene from one or more cameras, For images from one or more cameras, Based on the estimated 3D position of the tool's tip, define the region of interest within the image, Processing images in the region of interest to determine the position of the tool tip in the region of interest, Determining the updated 3D position of the tool tip based on the determined position of the tool in the region of interest of one or more images, A method that includes this.

[0075] The method according to Example 17, wherein receiving images from each of 18.1 or two or more cameras includes receiving images from corresponding cameras among multiple cameras, and determining the updated 3D position of the tool tip includes triangulating the updated 3D position based on the determined position of the tool tip in the region of interest in the images from multiple cameras.

[0076] 19. Estimating the 3D position of the tool tip is The main focus of the tool is to determine the core of the tool, Estimating the 3D position based on a known offset along the main axis, The method according to Example 17 or Example 18, including the method described in Example 18.

[0077] 20.1 Receiving images from each of two or more cameras includes receiving an image from one camera, and determining the updated 3D position of the tool's tip is Based on camera calibration, the position of the tool tip in the region of interest is projected onto a 3D line, Based on the intersection of the 3D line and the tool's main axis, the updated 3D position of the tool tip is determined, The method according to Example 19, including the method described in Example 19.

[0078] 21. A camera array, A support structure having a central region, A first camera is mounted on a support structure in the central region and configured to capture first image data of at least a portion of the scene, Multiple second cameras are mounted on a support structure, away from the central region and facing outward from the first camera, and are configured to capture second image data of at least a portion of the scene. A third camera is mounted on a support structure facing outward from the second camera, away from the central region, and is configured to capture third image data of at least a portion of the scene. A processing device that is communicatively coupled to the first camera, the second camera, and the third camera, The processing device comprises, and based on the first image data, the second image data and / or the third image data, Determine the depth of at least part of the scene, Determine the position of the tools in the scene. A camera array configured as follows.

[0079] 22. The camera array according to Example 21, wherein the first camera, the second camera, and the third camera each have at least one distinct intrinsic parameter.

[0080] 23. The camera array according to Example 22, wherein the second cameras are substantially identical to each other, and the third cameras are substantially identical to each other.

[0081] 24. A camera array according to any one of Examples 21 to 23, wherein the first camera, the second camera, and the third camera operate at different wavelengths.

[0082] 25. The camera array according to Example 24, wherein the second cameras are substantially identical to each other, and the third cameras are substantially identical to each other.

[0083] 26. The camera array according to any one of Examples 21 to 25, wherein the processing device is further configured to synthesize a virtual image corresponding to a virtual field of view of a scene based on a first image data, a second image data and / or a third image data.

[0084] 27. A camera array according to any one of Examples 21 to 26, wherein each of the second and third cameras is angled radially inward toward the central region of the support structure.

[0085] 28. A camera array according to any one of Examples 21 to 27, wherein the second camera and the third camera each have fields of view, the fields of view of the second camera at least partially overlap to define a first volume, and the fields of view of the third camera at least partially overlap to define a second volume that is larger than the first volume.

[0086] 29. A camera array according to any one of Examples 21 to 28, wherein the first camera has a focal plane, and each second camera has a focal axis, the focal axes of the second cameras converging at a point different from the focal plane of the first camera.

[0087] 30. A camera array according to any one of Examples 21 to 29, wherein the first camera has a focal plane, the second camera each has a focal axis, and the focal axes of the second cameras converge at a point below the focal plane of the first camera.

[0088] 31. A camera array, Support structure and, A depth sensor having a focal plane is attached to a support structure, Multiple cameras are mounted on a support structure and configured to capture image data of a scene, each having a focal axis that converges at a point different from the focal plane of the depth sensor. Multiple trackers, mounted on a support structure and configured to capture positional data of tools in the scene, A camera array equipped with a camera array.

[0089] 32. The camera array according to Embodiment 31, wherein the focal axis of the camera converges at a point below the focal plane of the depth sensor.

[0090] 33. The depth sensor is a camera array of Example 31 or Example 32, mounted in the central region of the support structure.

[0091] 34. The camera array according to Embodiment 33, wherein the cameras and trackers are arranged radially outward from the depth sensor, away from the central region.

[0092] 35. The camera array of Example 33 or Example 34, wherein the tracker is positioned radially outward from the camera, away from the central region.

[0093] 36. The camera array according to any one of Examples 31 to 35, wherein the tracker is positioned in close proximity to the periphery of the support structure.

[0094] 37. A camera array according to any one of Examples 31 to 36, wherein the camera and the tracker each have fields of view, the camera's fields of view at least partially overlap to define an imaging volume, and the tracker's fields of view at least partially overlap to define a tracking volume larger than the imaging volume, and the tracker is an imager having a different spectral sensitivity than the camera.

[0095] 38. A camera array according to any one of Examples 31 to 37, wherein the support structure includes a central axis and a longitudinal axis extending perpendicular to the central axis, the depth sensors are aligned along the central axis, and each camera and each tracker is spaced approximately equally apart from the central axis and the longitudinal axis.

[0096] 39. A method for determining the position of the tip of a tool in a scene, Receiving location data from the tool, This involves estimating the 3D position of the tool's tip based on positional data, Receiving two or more images of the scene, For each of two or more images, Based on the estimated 3D position of the tool's tip, define the region of interest within the image, Processing images in the region of interest to determine the position of the tool tip in the region of interest, Determining the updated 3D position of the tool tip based on the determined position of the tool tip in the region of interest in two or more images, A method that includes this.

[0097] The method according to Example 39, wherein receiving 40.2 or 3 or more images includes receiving images from corresponding cameras among multiple cameras, and determining the updated 3D position of the tool tip includes triangulating the updated 3D position based on the determined position of the tool tip in the region of interest in 2 or 3 or more images.

[0098] 41. A camera array, A support structure having a central region, Depth sensing is attached to the support structure at the center, Multiple cameras are mounted on a support structure, away from the central region and facing outward from the depth sensor, and are configured to capture image data of the scene. Multiple trackers are mounted on a support structure, away from the central area and facing outward from the camera, and are configured to capture positional data of tools in the scene. A camera array equipped with a camera array.

[0099] 42. The camera array according to Example 42, wherein the multiple cameras include four cameras and the multiple trackers include four trackers, and the cameras and trackers are arranged symmetrically around a support structure.

[0100] 43. The camera array according to Example 41 or Example 42, wherein the support structure includes a central axis and a longitudinal axis extending perpendicular to the central axis, the depth sensors are aligned along the central axis, and each camera and each tracker is equally spaced from the central axis and the longitudinal axis.

[0101] 44. The tracker is an infrared imager, as described in any one of Examples 41-43 of the camera array.

[0102] 45. The camera array according to any one of Examples 41 to 44, wherein the tracker is an imager having a different spectral sensitivity from the camera.

[0103] 46. ​​A camera array according to any one of Examples 41 to 45, wherein the camera and the tracker each have fields of view, the fields of view of the camera at least partially overlap to define an imaging volume, and the fields of view of the tracker at least partially overlap to define a tracking volume larger than the imaging volume.

[0104] 47. A camera array according to any one of Examples 41 to 46, wherein each camera is angled radially inward toward the central region of the support structure.

[0105] 48. A camera array, A support structure with a center, A depth sensor having a focal plane, mounted on a support structure close to the center, Multiple cameras are mounted on a support structure radially outward from the depth sensor and configured to capture image data of the scene, each having a focal axis that converges at a point below the focal plane of the depth sensor. Multiple trackers are mounted on a support structure facing radially outward from the camera and configured to capture positional data of tools in the scene, A camera array equipped with a camera array.

[0106] 49. A mediating reality system, A support structure with a center, A depth sensor attached to the support structure in the central region, Multiple cameras are mounted on a support structure, away from the central region and facing outward from the depth sensor, and are configured to capture image data of the scene. Multiple trackers are mounted on a support structure, away from the central region and radially outward from the camera, and are configured to capture positional data of tools in the scene. A camera array including, An input controller configured to control the position and orientation of the virtual field of view of the scene, The camera array and input controller are communicatively coupled, A virtual image corresponding to the virtual field of view is synthesized based on image data from at least two of the cameras. The tool's position is determined based on location data from at least two of the trackers. A processing apparatus configured as follows, A display device that is communicatively coupled to a processing unit and configured to display a graphic representation of a tool at a determined position in a virtual image, A mediating reality system equipped with these features.

[0107] 50. The mediating reality system according to Embodiment 49, wherein the processing device is further configured to determine the position of a tool based on image data from at least one of the cameras.

[0108] 51. The processing unit is The initial 3D position of the tool is estimated based on positional data from at least two of the trackers. Based on the tool's initial 3D position, define the region of interest within the image data from at least one of the cameras. The image data in the region of interest is processed to determine the position of the tool in the region of interest. Determine the updated 3D position of the tool based on the determined position of the tool in the region of interest. A mediating reality system according to Example 50, configured to determine the position of a tool by means of a tool.

[0109] 52. A method for imaging a subject within a scene, Align the depth sensor with the subject's area of ​​interest so that the area of ​​interest is positioned near the depth of focus of the depth sensor, This involves using multiple cameras to capture image data of scenes that include the subject's area of ​​interest, This involves capturing the positional data of tools in a scene using multiple trackers, wherein the depth sensor, camera, and trackers are commonly mounted on a support structure for the camera array. It receives input regarding the selected position and orientation of the virtual field of view of the scene, Based on image data, a virtual image corresponding to the virtual field of view of the scene is synthesized, Determining the tool's position based on location data, In a display device, the graphical representation of a tool at a determined position within a virtual image is displayed, A method that includes this.

[0110] 53. The method according to Example 52, wherein the scene is a surgical scene and the subject's area of ​​interest is the subject's spine.

[0111] 54. A method for imaging a subject within a scene, The camera array's depth sensor is aligned with the subject's area of ​​interest so that the area of ​​interest is positioned near the depth of focus of the depth sensor. The method involves using multiple cameras in a camera array, each having a focal axis that converges at a point below the depth of field of a depth sensor, to acquire image data of a scene containing the subject's area of ​​interest, Using multiple trackers in a camera array to capture positional data of tools in the scene, It receives input regarding the selected position and orientation of the virtual field of view of the scene, Based on image data, a virtual image corresponding to the virtual field of view of the scene is synthesized, Determining the tool's position based on location data, In a display device, the graphical representation of a tool at a determined position within a virtual image is displayed, A method that includes this.

[0112] V-knot The above detailed description of embodiments of the present technology is not exhaustive, nor does it limit the present invention to the exact forms disclosed above. While specific embodiments and examples of the present technology are described above for illustrative purposes, various equivalent modifications are possible within the scope of the present technology, as will be apparent to those skilled in the art. For example, the steps are shown in a given order, but in other embodiments, the steps may be performed in a different order. Further embodiments can also be provided by combining the various embodiments described herein.

[0113] From the above, it will be understood that while specific embodiments of the present technology have been described for illustrative purposes, well-known structures and functions have not been illustrated or described in detail in order to avoid unnecessarily obscuring the description of the embodiments of the present technology. Where contextually possible, singular or plural terms may include plural or singular terms, respectively.

[0114] Furthermore, unless the word “or” is explicitly limited to mean only one item without including other items when referring to a list of two or more items, the use of “or” in such a list should be interpreted as including (a) any one item in the list, (b) all items in the list, or (c) any combination of items in the list. Moreover, throughout, the term “comprising” is used to mean including at least the described (singular or plural) features, and does not exclude any many more of the same features and / or other features of further types. Also, while specific embodiments have been described in this specification for illustrative purposes, it will be understood that various modifications can be made without departing from the Art. Furthermore, while advantages related to some embodiments of the Art have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and it is not necessary for all embodiments to exhibit such advantages in order to be included in the scope of the Art. Accordingly, this disclosure and related art may include other embodiments not expressly illustrated or described herein. [Explanation of symbols]

[0115] 100 Imaging Systems 101 Tools 102 Processing Unit 103 Tool Tips 104 Display device 105 Marker 106 Input Controller 107 Image Processing Device 108 scenes 109 Tracking Processing Unit 110 Camera Array 112a~n Camera 114a~n Tracker 116 Depth Sensor 118 Projectors 119 (single / multiple) cameras

Claims

1. It is a camera array, A support structure with a center, A depth sensor attached to the support structure in close proximity to the center, Multiple cameras are attached to the support structure radially outward from the depth sensor and configured to capture image data of the scene, Multiple trackers are attached to the support structure radially outward from the camera and configured to capture positional data of tools within the scene, A camera array characterized by comprising the following features.

2. The plurality of cameras includes four cameras, the plurality of trackers includes four trackers, and the cameras and trackers are arranged symmetrically around the support structure. The camera array according to claim 1.

3. The support structure includes a central axis and a longitudinal axis extending perpendicular to the central axis, the depth sensors are aligned along the central axis, and each of the cameras and each of the trackers is spaced equally apart from the central axis and the longitudinal axis. The camera array according to claim 1.

4. The aforementioned tracker is an infrared imager. The camera array according to claim 1.

5. The tracker is an imager having a different spectral sensitivity than the camera. The camera array according to claim 1.

6. The camera and the tracker each have a field of view, the field of view of the camera at least partially overlapping to define an imaging volume, and the field of view of the tracker at least partially overlapping to define a tracking volume larger than the imaging volume. The camera array according to claim 1.

7. Each of the cameras is angled radially inward toward the center of the support structure. The camera array according to claim 1.

8. The depth sensor has a focal plane, and each camera has a focal axis, the focal axis of the camera converging at a point below the focal plane of the depth sensor. The camera array according to claim 7.

9. It is a mediating reality system, A support structure with a center, A depth sensor attached to the support structure in close proximity to the center, Multiple cameras are attached to the support structure radially outward from the depth sensor and configured to capture image data of the scene, Multiple trackers are attached to the support structure radially outward from the camera and configured to capture positional data of tools within the scene, A camera array including, An input controller configured to control the position and orientation of the virtual field of view in the aforementioned scene, The camera array and the input controller are communicated together, A virtual image corresponding to the virtual field of view is synthesized based on image data from at least two of the aforementioned cameras. The position of the tool is determined based on position data from at least two of the trackers. A processing apparatus configured as follows, A display device is communicatively coupled to the processing device and configured to display a graphic representation of the tool at the determined position in the virtual image, A mediating reality system characterized by comprising the following features.

10. The processing device is further configured to determine the position of the tool based on the image data from at least one of the cameras. The mediating reality system according to claim 9.

11. The processing unit is The initial three-dimensional (3D) position of the tool is estimated based on the position data from at least two of the trackers. Based on the initial 3D position of the tool, a region of interest is defined within the image data from at least one of the cameras. The image data in the region of interest is processed to determine the position of the tool in the region of interest. Based on the determined position of the tool in the region of interest, the updated 3D position of the tool is determined. The mediating reality system according to claim 10, configured to determine the position of the tool by doing so.

12. A method for imaging subjects within a scene, The depth sensor of the camera array is aligned with the subject's area of ​​interest so that the area of ​​interest is positioned near the depth of focus of the depth sensor. The camera array is used to capture image data of the scene, including the portion of interest of the subject, The camera array is used to capture positional data of tools in the scene, The system receives input regarding the selected position and orientation of the virtual field of view in the aforementioned scene, Based on the aforementioned image data, a virtual image corresponding to the virtual field of view of the scene is synthesized, The position of the tool is determined based on the aforementioned position data, In a display device, the graphic representation of the tool at the determined position within the virtual image is displayed, A method characterized by including the following.

13. The aforementioned scene is a surgical scene, and the subject's area of ​​interest is the subject's spine. The method according to claim 12.

14. Each of the cameras has a focal axis, and the focal axes of the cameras converge at a point below the focal depth of the depth sensor. The method according to claim 12.

15. Determining the position and orientation of the tool is based further on the image data, The method according to claim 12.

16. Determining the position of the tool is Based on the aforementioned position data, the initial three-dimensional (3D) position of the tool is estimated, Based on the initial 3D position of the tool, one or more regions of interest are defined within the image data. Processing the image data in the one or more regions of interest to determine the position of the tool in the one or more regions of interest, Determining the updated 3D position of the tool based on the determined position of the tool in the one or more regions of interest, The method according to claim 15, including the method described in claim 15.

17. A method for determining the position of the tip of a tool in a scene, The tool receives location data from at least two trackers, Based on the position data, estimate the three-dimensional (3D) position of the tip of the tool, Receiving images of the scene from each of one or more cameras, Regarding the images from each of the one or more cameras, A region of interest is defined within the image based on the estimated 3D position of the tip of the tool, Processing the image in the region of interest to determine the position of the tip of the tool in the region of interest, The updated 3D position of the tip of the tool is determined based on the determined position of the tool in the region of interest of the one or more images, A method characterized by including the following.

18. Receiving the image from each of the one or more cameras includes receiving the image from a corresponding camera among the multiple cameras, and determining the updated 3D position of the tip of the tool includes triangulating the updated 3D position based on the determined position of the tip of the tool in the region of interest in the images from the multiple cameras. The method according to claim 17.

19. Estimating the 3D position of the tip of the tool is To determine the main focus of the aforementioned tool, Estimating the 3D position based on a known offset along the principal axis, The method according to claim 17, including the method described in claim 17.

20. Receiving the image from each of the one or more cameras includes receiving the image from one camera, and determining the updated 3D position of the tip of the tool is Based on the calibration of the camera, the position of the tip of the tool in the region of interest is projected into the 3D line, Based on the intersection of the 3D line and the main axis of the tool, the updated 3D position of the tip of the tool is determined. The method according to claim 19, including the method described in claim 19.