System and method for masking an identified object

CN122842003APending Publication Date: 2026-09-29INTUITIVE SURGICAL OPERATIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010651.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-01-20
Filing Date
2021-01-18
Publication Date
2026-09-29

AI Technical Summary

Benefits of technology

[0009]示例性非暂态计算机可读介质存储指令,所述指令在被执行时引导计算设备的处理器在将合成元素应用于原始图像期间掩蔽所识别的对象。更具体地,所述指令指示所述处理器访问在场景的原始图像中描绘的所识别的对象的模型;将所述模型与所识别的对象相关联,以及生成呈现数据供呈现系统使用以呈现所述原始图像的增强版本,其中基于与所识别的对象相关联的所述模型来防止添加到所述原始图像的合成元素遮挡所识别的对象的至少一部分。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842003A_ABST
    Figure CN122842003A_ABST
Patent Text Reader

Abstract

The subject matter of this disclosure is systems and methods for masking identified objects. An example object masking system is configured to mask an identified object during application of a synthetic element to an original image. For example, the object masking system accesses a model of an identified object depicted in an original image of a scene. The object masking system associates the model with the identified object. The object masking system then generates presentation data for use by a presentation system to present an enhanced version of the original image in which a synthetic element added to the original image is prevented from occluding at least a portion of the identified object based on the model associated with the identified object. In this way, the synthetic element is made to appear as if it is located behind the identified object. Corresponding systems and methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application 2021800210243 (PCT / US2021 / 013826), entitled "System and method for masking identified objects," filed on January 18, 2021, with an international filing date of January 18, 2021, and entering the national phase on September 14, 2022.

[0002] Cross-reference to related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 963,249, filed January 20, 2020, the contents of which are incorporated herein by reference in their entirety. Background Technology

[0003] Image capture devices with various imaging modalities are used to capture images in a variety of scenarios and for a variety of use cases. These images include objects and scenery visible from different parts and locations. As an example, a camera can be used to capture photographs of people attending an event, cars driving along a road, the interior of a house for sale, etc. As another example, an endoscope or other medical imaging modalities can be used to capture endoscopic images of surgical sites, such as surgical spaces within a patient's body.

[0004] The images captured by these devices may ultimately be presented to viewers. Referring to the examples above, for instance, photos of people can be shared with friends and family, photos of cars can be used for print ads in magazines, and photos of the interior of a house can be included in a real estate posting, etc. In the surgical imaging example, endoscopic images can be presented to the surgical team through different types of display devices, thereby facilitating the surgical team's visualization of the surgical space while performing surgical procedures.

[0005] In any of these or various other examples, it may be desirable to enhance a captured image by adding composite elements (e.g., augmented reality overlays, etc.). For example, composite elements such as depictions of objects, images, information, and / or other enhancements by the capturing device about the rest of the image that were not actually captured (e.g., overlaid on other images). However, it may not always be desirable to present such composite elements in front of all other images depicted in a particular image. Summary of the Invention

[0006] The following description presents a simplified overview of one or more aspects of the systems and methods described herein. This invention is not an exclusive overview of all anticipated aspects and is neither intended to identify key or essential elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present one or more aspects of the systems and methods described herein as a prelude to the detailed description presented below.

[0007] An exemplary system includes: a memory storing instructions; and a processor communicatively coupled to the memory and configured to execute the instructions to mask identified objects during the application of compositing elements to an original image. More specifically, the exemplary system accesses a model of the identified objects depicted in an original image of a scene; associates the model with the identified objects; and generates rendering data for use by a rendering system to render an enhanced version of the original image, wherein compositing elements added to the original image are used to prevent at least a portion of the identified objects from being occluded based on the model associated with the identified objects.

[0008] An exemplary method for masking identified objects during the application of composite elements to an original image is performed by an object masking system. The method includes: accessing a model of the identified object depicted in an original image of a scene; associating the model with the identified object; and generating rendering data for use by a rendering system to render an enhanced version of the original image, wherein the model associated with the identified object is used to prevent composite elements added to the original image from occluding at least a portion of the identified object.

[0009] An exemplary non-transitory computer-readable medium storage instruction, when executed, directs a processor of a computing device to mask identified objects during the application of compositing elements to an original image. More specifically, the instruction instructs the processor to access a model of the identified object depicted in an original image of a scene; associate the model with the identified object; and generate rendering data for use by a rendering system to render an enhanced version of the original image, wherein the compositing elements added to the original image are used to prevent at least a portion of the identified object from being occluded based on the model associated with the identified object. Attached Figure Description

[0010] The accompanying drawings illustrate various embodiments and are part of the specification. The illustrated embodiments are merely examples and do not limit the scope of the invention. Throughout the drawings, the same or similar reference numerals designate the same or similar elements.

[0011] Figure 1 Exemplary images depicting various objects according to the principles described herein are shown.

[0012] Figure 2 This illustrates the principles described herein. Figure 1 An exemplary enhanced version of the image, wherein synthetic elements are applied to the image.

[0013] Figure 3This paper illustrates exemplary aspects of how depth data can be detected and used to mask objects included in an image when synthetic elements are applied to an image, based on the principles described herein.

[0014] Figure 4 An exemplary object masking system is shown, based on the principles described herein, for masking identified objects during the application of synthetic elements to an original image.

[0015] Figure 5 An exemplary configuration based on the principles described herein is shown, in which... Figure 4 The object masking system can be operated to mask identified objects during the application of composite elements to the original image.

[0016] Figure 6 An exemplary representation of a segmented image, based on the principles described herein, and configured for use by a rendering system, is shown.

[0017] Figure 7 This illustrates the principle described herein in the application of... Figure 6 The masking data is applied to an exemplary composite element after the composite element.

[0018] Figure 8 Exemplary aspects of how depth data (including depth data from a model of the identified object) can be detected and used to improve the masking of the identified object when applying synthetic elements to the original image, based on the principles described herein.

[0019] Figure 9 An exemplary computer-assisted surgical system based on the principles described herein is shown.

[0020] Figure 10 The use of the principles described herein is illustrated. Figure 9 An exemplary aspect of masking identified objects during the application of synthetic elements to an original image depicting a surgical site in a computer-aided surgical system.

[0021] Figure 11 An exemplary method for masking identified objects during the application of synthetic elements to an original image, based on the principles described herein, is shown.

[0022] Figure 12 An exemplary computing device based on the principles described herein is shown. Detailed Implementation

[0023] This paper describes systems and methods for masking identified objects during the application of compositing elements to an original image. When an original image is enhanced by one or more compositing elements (e.g., augmented reality overlays, etc.), conventional systems add the compositing elements to other content of the original image in a manner that places the compositing element in front of or on top of all other content in the image (i.e., in a graphics layer, the compositing element is visible while covering other layers of content behind or behind the overlay content layer). However, the systems and methods described herein help to address scenes where certain content of the image (e.g., one or more specific objects depicted in the image) is depicted as if in front of (or on top of) the augmenting material.

[0024] As used herein, augmented reality technologies, scenes, images, etc., will be understood to include original elements (e.g., images captured from a real-world scene) and augmentations (e.g., composite elements that are not actually present but make appear to be present) in a presentation of reality that blends original and augmented elements in any suitable manner. Therefore, it will be understood that the term "augmented reality" can refer to any type of augmented, blended, virtual, or other extended reality at any point on the virtual spectrum that may be used in a particular implementation, and is not limited to any particular definition of "augmented reality" as may be used in the art.

[0025] The systems and methods described herein improve augmented reality scenarios in which enhancements, such as three-dimensional (“3D”) anatomical models, are presented in augmented endoscopic images captured during a surgical procedure. In this context, members of the surgical team benefit from the augmented reality experience, where enhancements (e.g., 3D anatomical models) are depicted behind surgical instruments shown in the endoscopic view. For example, this layering can make enhancements appear more natural and less distracting or disorienting compared to enhancements superimposed in front of surgical instruments in an augmented reality image, and / or may otherwise facilitate the surgical procedure or be preferred by surgical team members viewing the endoscopic augmented reality image.

[0026] Therefore, the systems and methods described herein for masking identified objects during the application of synthetic elements to an original image may involve tracking the 3D position of a particular object (e.g., identified objects such as surgical instruments in the surgical procedure example above) by generating a dense, accurate depth map of the instrument surface based on a 3D model of the instrument that can be used in the system (e.g., a computer-aided design (“CAD”) model, a pre-scanned 3D representation, etc.). For example, the depth map in this example can be generated by using kinematic data and / or computer vision techniques that allow the system to track the position and orientation of surgical instruments in space, and by using camera calibration data (e.g., inherent parameters defined for an endoscope, etc.) to determine how to depict the surgical instruments in an image captured by the endoscope based on the position and orientation of the instruments in space. The depth map generated or acquired in this way can be used to efficiently and accurately create a mask that allows the object to appear in front of the overlay when the overlay is applied to the image (e.g., overlaid on an image or otherwise integrated with an image), rather than the overlay appearing to be in front of the object as in conventional designs.

[0027] An exemplary object masking system according to embodiments described herein may include a memory storing instructions; and a processor communicatively coupled to the memory and configured to execute the instructions to perform various operations to mask identified objects during the application of compositing elements to an original image. More specifically, the object masking system may access a model of the identified object depicted in an original image of the scene; associate the model with the identified object; and generate rendering data for use by a rendering system to render an enhanced version of the original image, wherein compositing elements added to the original image are used to prevent at least a portion of the identified object from being occluded based on the model associated with the identified object.

[0028] Such functionality can be performed in any suitable manner. As an example, an exemplary object masking system may perform operations such as identifying identified objects within an image depicted by an image that can be used as a model for the system. For example, in the surgical example described above, the image may be an endoscopic image depicting a surgical site, and the identified object may be an instrument or other known object whose model (e.g., a 3D CAD model, etc.) is available. Thus, the object masking system may access a model of the identified object in response to its recognition and may associate the model with the identified object. This association may include registering the model to the identified object as depicted in the image in any way. In some examples, the object masking system may segment the image (e.g., based on the association between the model and the identified object, or as part of it) to distinguish pixels of the image depicting the identified object from pixels of the image not depicting the identified object. For example, based on models of surgical instruments in the above examples, and based on data tracked or accessed (such as the position and orientation of surgical instruments in some examples and / or camera calibration parameters of endoscopes), object masking systems can accurately and precisely identify which pixels in a particular endoscopic image are part of a surgical instrument, and which pixels are part of something else besides the surgical instrument.

[0029] After associating the model with the identified object in this way (e.g., and, in some examples, segmenting the image), the object masking system may perform another operation in which rendering data is generated. In some examples, the rendering data may include image data (e.g., image data representing an enhanced version of the original image, etc.), and in the same or other examples, it may include masking data representing the segmented image. In either case, the rendering data may be generated for use by the rendering system (augmented reality rendering system, etc.) that provides the rendering data. The rendering system may render an enhanced version (i.e., an augmented reality image) of the original image based on the original image (e.g., an endoscopic image depicting a surgical site in the example above). In the augmented reality image, synthetic elements (e.g., augmented reality overlays, such as views of the underlying anatomical model configured to facilitate the surgical procedure in the example above, or other such information) may be added to the original image in such a way that the synthetic elements prevent at least a portion of the identified object from being occluded based on the model associated with the identified object.

[0030] For example, composite elements can be applied only to pixels of an image that do not depict the identified object. More specifically, in the surgical instrument example, the depiction of the surgical instrument can be filtered so that it is not covered by composite elements, so that the composite elements appear to be behind the identified object in the enhanced version of the original image. In this way, the surgical instrument can be clearly seen in the surgical augmented reality example, and the anatomical model (or other composite elements or augmented reality overlays) can be depicted behind the instrument (e.g., appearing to be projected directly onto the tissue behind the instrument).

[0031] The systems and methods described herein for masking identified objects during the application of synthetic elements to an original image offer a variety of benefits and advantages, and are associated with each other. For example, by depicting synthetic elements (e.g., augmented reality overlays) in front of certain images or objects in an image and behind other images or objects, the systems and methods described herein facilitate the generation of augmented reality images that appear more natural and integrated, more useful and more attractive, and less distracting and / or disorienting compared to conventional augmented reality images in which synthetic elements are overlaid in front of all images and other objects.

[0032] Conventional techniques used to create depth maps to aid object segmentation rely solely on conventional depth detection techniques, such as stereo depth detection, single-view structure from motion (“SfM”), simultaneous localization and mapping (“SLAM”), and others of this kind. However, these techniques inevitably present various challenges and limitations that leave room for improvement when applied to the problem of masking certain objects rather than others, in order to apply synthetic elements to the original image when generating augmented reality images.

[0033] As an example of such challenges, attempts to detect object depth data in real time (e.g., during the rendering time of presenting augmented reality images to the user) have had limited success due to real-time processing and latency constraints that make it difficult or impossible for a given system with limited processing resources to capture scene depth data completely and accurately in real time. Because of these limitations, and as will be described in more detail below, the depth data captured in typical examples may result in a relatively sparse depth map, and the edges of objects to appear in front of the augmented reality overlay may be poorly defined or exhibit unwanted artifacts that can be distracting and degrade the effect.

[0034] Advantageously, the systems and methods described herein can generate depth maps for certain objects based on detailed models (e.g., 2D models, 3D models, etc.) that are already available to the system (i.e., have been generated rather than need to be generated at runtime). In this way, the systems and methods described herein gain access to highly dense and accurate depth maps, which, when associated with an image, allow for accurate and precise segmentation of the image to distinguish the depiction of the identified object from other depicted images. This significantly improves the appearance of realistic objects appearing in front of augmented reality overlays.

[0035] Various embodiments will now be described in more detail with reference to the accompanying drawings. The disclosed systems and methods may provide one or more of the benefits described above and / or various additional and / or alternative benefits that will become apparent herein.

[0036] Figure 1 An exemplary image 100 depicting images comprising various objects is shown. Specifically, for example, object 102 is shown as a hollow square object, while objects 104 (i.e., objects 104-1 to 104-4) are shown as having other basic geometric shapes (e.g., rectangular shapes, circular shapes, etc.). As shown, each of shapes 102 and 104 is shaded by different styles of intersecting shading lines representing different colors, textures, and / or depths (relative positions) that each object may have. Additionally, each of objects 102 and 104 is depicted in front of a background 106 that is unshaded (i.e., white) in image 100.

[0037] Although for the sake of clarity... Figure 1 Simple geometric objects are depicted in the accompanying drawings and other figures, but it should be understood that objects and other images depicted in images (such as image 100) may include any type of object and / or image that may be used in a particular embodiment. For example, objects 102 and / or 104 may represent any of the types of objects mentioned in the examples above (e.g., a person, a car, a house, etc.) or any other type of object that may be depicted in a particular image. While the principles described herein can be applied to a wide variety of use cases, one particular use case, which will be described in more detail below, is the surgical use case, in which the objects depicted in the images are each associated with a surgical procedure. For example, as will be described, objects and images in such an example may include anatomical objects and images of the interior of the body on which a surgical procedure is performed, surgical instruments and / or tools for performing the surgical procedure, etc.

[0038] Figure 2The following illustrates an exemplary enhanced version of the original image 100, which will be referred to as Augmented Reality Image 200. Within Augmented Reality Image 200, a compositing element 202 (also referred to as Augmented Reality Overlay 202) is applied to image 100. As shown, compositing element 202 is an elliptical object masked by pure black. (As mentioned above...) Figure 1 As mentioned, although basic geometry is used Figure 2 For illustrative purposes only, but it should be understood that composite element 202 may represent any suitable type of enhancement as may be used in a particular implementation. For example, composite element 202 may represent a hypothetical creature presented as hidden in the real world for a user to find in an augmented reality game, virtual furniture presented in a location in a home to help a user imagine how the furniture would look and be located in a room, a data graphic configured to inform the user of one of the real objects 102 or 104 depicted in image 100, or another suitable type of enhancement as may be appropriate in another augmented reality implementation. As another example, composite element 202 may be implemented as an anatomical structure (e.g., a preoperative 3D scan model of a subsurface anatomy, etc.) or another surgically related object in an example of augmented reality-enhanced surgical procedures already mentioned.

[0039] For example, in a typical implementation of augmented reality, synthetic element 202 in Figure 2 The image is shown as being superimposed on all objects 102 and 104 in the scene that are near the composition element 202. That is, the composition element 202 is shown as "occluding" each of the other objects and background, "on top of" each of the other objects and background, or "in front of" each of the other objects and background, because the composition element 202 occludes these objects and images rather than being occluded by them (in this case, it can be said that the composition element 202 is "occluded", "behind", "below", or "behind" other objects, etc.).

[0040] Although Figure 2As not shown, but to be understood, in certain scenarios and use cases, it may be desirable for the composite element 202 to be depicted to be occluded by one or more of the objects depicted in image 100. For example, for the purposes of the following description, it will be assumed that object 102 is desired to be depicted in front of composite element 202 (i.e., for composite element 202 to be depicted behind object 102), even though composite element 202 is still depicted in front of object 104 and background 106. To achieve this effect of composite element 202 being occluded (or at least partially occluded) by object 102, a depth map may be generated for image 100 that will allow object 102 to be distinguished from other objects and images in the image, such that pixels representing object 102 can be masked so as not to be covered by composite element 202 when applied to image 100 to form augmented reality image 200.

[0041] To illustrate, Figure 3 Exemplary aspects of how depth data can be detected and used to mask object 102 during the application of synthetic element 202 to the original image 100 to form an augmented reality image, according to the principles described herein, are illustrated. Specifically, Figure 3 Representation 300-1 shows the actual depth data of the objects and images of image 100, representation 300-2 shows how a relatively sparse depth map of the objects and images of image 100 can be generated based on conventional depth detection techniques, and representation 300-3 shows an implementation of augmented reality image 200, wherein the composite element 202 is depicted as being located behind object 102 (albeit in a non-ideal way, due to the relatively sparse depth map generated by the limitations of conventional depth detection techniques). Each of representations 300 (i.e., representations 300-1 to 300-3) will now be described in more detail.

[0042] Representation 300-1 depicts the actual depth data of each of objects 102 and 104 in image 100, as well as the background 106. Specifically, for illustrative purposes, a simple notation using a single-digit number (e.g., 0, 1, 2, etc.) is employed to distinguish regions of image 100 depicting images located at different depths relative to the image capturing device (e.g., camera, endoscope, etc.) capturing image 100. For example, in this notation, depth "0" (i.e., the depth of background 106) would be understood as any farthest depth shown as a vantage point relative to the image capturing device, while depth "9" (i.e., the depth of object 102) would be understood as the closest depth shown as a vantage point relative to the image capturing device. Other depths shown in 300-1, such as depth "1" (i.e., the depth of objects 104-1 and 104-3), depth "2" (i.e., the depth of object 104-2), and depth "3" (i.e., the depth of object 104-4), will be understood as being closer to the vantage point than depth "0" of background 106, but farther from the vantage point than depth "9" of object 102.

[0043] If the image depicted in image 100 can be analyzed using one or more conventional depth detection techniques, and there are no limitations on the time or resources available for performing these techniques, the system can generate a complete and accurate depth map of the captured image. However, unfortunately, significant limitations exist in many real-world scenarios, making the generation of such an ideal depth map challenging or impossible. For example, when a user looks around at the world, augmented reality often must be generated and presented to the user in real time, meaning that very strict time constraints may be associated with any processing or analysis of the world being performed, including analysis of the depth of objects in an image captured by an image capture device. Therefore, when considering the practical time constraints associated with a given augmented reality application, generating a dense depth map of an image using conventional depth detection techniques may be impossible or infeasible.

[0044] Representation 300-2 illustrates the consequences of such practical limitations. Specifically, unlike what might be ideal, given sufficient time and resources, to detect depth at every point (e.g., every pixel) depicted by image 100, practical limitations constrain real-world systems to detect depth at significantly fewer points in image 100. As shown, for example, a system might only have the time and resources to detect depth at each point in image 100 (divided by black “x”s in Representation 300-2) using conventional depth detection techniques, but might not have enough time and resources to determine the depth at other points in image 100 (e.g., before needing to continue processing subsequent frames in a video image). Thus, depth maps can be generated in real time, but these maps may be relatively sparse and may not capture well the complexities of object edges and other parts of the image where depth values ​​change abruptly. For example, as shown in Representation 300-2, the depth of object 102 can be captured at several points, but details about the precise location of the edges of object 102 and the fact that object 102 is hollow may not be discernible from the relatively sparse depth data that can be detected.

[0045] Representation 300-3 illustrates the result of using a relatively sparse depth map, as shown in Representation 300-2, to distinguish object 102 and attempt to mask the object so that it appears to occlude composite element 202 (i.e., so that it appears to be in front of composite element 202). As shown in Representation 300-3, a portion of composite element 202 has been masked in a manner that roughly approximates the position of object 102 so as not to cover object 102 when composite element 202 is applied to image 100. However, due to the relatively sparse depth map, the edges of the masked portion are not well aligned with the edges of object 102 and the hollow portion is not properly considered. Therefore, object 102 appears to interact with composite element 202 to some extent, but this effect may be unconvincing (e.g., and may be distracting, disorienting, etc.) to a viewer expecting to see composite element 202 depicted behind object 102. Although for illustrative purposes, Figure 3 While the sparsity of depth data and its impact may be slightly exaggerated, the principle to be understood is that if augmented reality overlays (such as composite element 202) are to be convincingly depicted as being occluded by objects (such as object 102), a denser depth map may be required than the depth map actually generated in real time using conventional depth detection techniques.

[0046] to this end, Figure 4An exemplary object masking system 400 (“System 400”) for masking identified objects during the application of synthetic elements to an original image, according to the principles described herein, is illustrated. In some examples, System 400 may be implemented by an augmented reality device or a general-purpose computing device (e.g., a mobile device such as a smartphone or tablet) for various purposes. In some examples of augmented reality-enhanced surgical procedures described herein, System 400 may be included in a computer-assisted surgical system (such as those described below in conjunction with…) Figure 9 One or more components of the exemplary computer-assisted surgical system described herein, implemented therein or connected to the computer-assisted surgical system (such as those described below) Figure 9 The exemplary computer-assisted surgical system described herein may be comprised of one or more components. For example, in such an example, system 400 may be implemented by one or more components of a computer-assisted surgical system, such as a manipulation system, a user control system, or an auxiliary system. In other examples, system 400 may be implemented by a standalone computing system, such as a standalone computing system communicatively coupled to the computer-assisted surgical system or implementing another non-surgical application or use case.

[0047] like Figure 4 As shown, system 400 may include, but is not limited to, storage facility 402 and processing facility 404 selectively and communicatively coupled to each other. Facilities 402 and 404 may each include or be implemented by one or more physical computing devices, which include hardware and / or software components such as processors, memory, memory drives, communication interfaces, instructions stored in memory for execution by the processor, etc. Although facilities 402 and 404 are... Figure 4 Facilities 402 and 404 are shown as separate facilities, but facilities 402 and 404 may be combined into fewer facilities (such as being combined into a single facility), or divided into more facilities as may be used in a particular implementation. In some examples, each of facilities 402 and 404 may be distributed among multiple devices and / or multiple locations, as may be used in a particular implementation.

[0048] Storage facility 402 may maintain (e.g., store) executable data used by processing facility 404 to perform any of the functions described herein. For example, storage facility 402 may store instructions 406 that can be executed by processing facility 404 to perform one or more of the operations described herein. Instructions 406 may be implemented by any suitable application, software, code, and / or other instance of executable data. Storage facility 402 may also maintain any data received, generated, managed, used, and / or transmitted by processing facility 404.

[0049] Processing facility 404 may be configured to perform (e.g., execute instructions 406 stored in storage facility 402 to perform) various operations associated with masking identified objects during the application of compositing elements to an original image (i.e., masking identified objects during the application of augmented reality overlay). Such operations may include, for example, accessing a model of the identified object depicted in the original image of the scene, associating the model with the identified object, and generating rendering data for use by a rendering system to render an enhanced version of the original image (e.g., an augmented reality image, wherein the model associated with the identified object is used to prevent compositing elements added to the original image from occluding at least a portion of the identified object).

[0050] Such operations can be performed in any suitable manner. For example, processing facility 404 may be configured to recognize (e.g., within an image depicted by an image such as image 100) a recognized object whose model is available to system 400, and to access the model in response to the recognition. The recognized object may be an object to be presented in front of an augmented reality overlay, such as object 102 in the example above (assuming the model is available for object 102).

[0051] When accessing the model, processing facility 404 can associate the model with the identified object by accessing information indicating how the identified object is depicted within an image or in any other manner that may be used in a particular implementation. In one example, as will be described in more detail below, if spatial data is available for the identified object and / or the image capture device capturing the image, such spatial data may be used by processing facility 404 to register the position and / or orientation of the identified object relative to the capture device. Spatial data may include data supporting kinematics-based tracking, computer vision-based tracking, electromagnetic marker tracking, and / or other methods described herein or that may be used in a particular implementation. Camera calibration parameters (e.g., inherent parameters) associated with the capture device may also be accessed for processing facility 404 to use for registering the model based on the position and / or orientation of the identified object and based on the 3D model, or otherwise to associate the model with the identified object as depicted in the image.

[0052] Based on the association between the model and the identified object as depicted in the image (i.e., based on any or all of the accessed information, such as the model of the identified object, camera calibration parameters, spatial data for determining position and / or orientation, etc.), the processing device 404 can segment the image to distinguish pixels of the image depicting the identified object from pixels of the image not depicting the identified object. For example, referring to the above... Figure 1 For example, processing facility 404 can use this information to distinguish pixels representing object 102 from pixels representing object 104 and / or background 106.

[0053] Based on the association between the model and the identified object, processing facility 104 can generate rendering data for a rendering system (e.g., an augmented reality rendering system) to render an enhanced version of the original image (e.g., similar to augmented reality image 200). In some examples, the rendering data may include or be implemented as image data configured to be rendered and rendered by the rendering system. In other examples, the rendering data may include data from which such a renderable image can be constructed. For example, the rendering data may include masking data corresponding to the image data of the original image (e.g., masking data representing the segmentation of the image), as will be described and illustrated in more detail below. In still other examples, the rendering data may be another suitable type of rendering data configured for use by the rendering system to render an enhanced version of the original image. In any case, the rendering data may allow the rendering system to render the enhanced version such that compositing elements added to the original image prevent occlusion (i.e., at least partial occlusion) of at least a portion of the identified object. As will be described in more detail below, the rendering can be generated to have such characteristics based on the accessed and associated model of the identified object.

[0054] Based on the presentation data, a presentation system (e.g., implemented by system 400 or a system communicatively coupled to system 400) can present augmented reality images (such as an implementation of augmented reality image 200), which are based on the original image and wherein augmented reality overlays (such as composite element 202) are applied only to pixels of the original image that do not depict the identified object. In this way, in the augmented reality image, the augmented reality overlay can be depicted as being behind the identified object. For example, referring to the example above, augmented reality overlay 202 can be applied only to pixels of image 100 that do not depict object 102, such that augmented reality overlay 202 is depicted as being behind object 102 in the resulting augmented reality image.

[0055] As described above, due to the nature of augmented reality and various augmented reality use cases, various implementations of system 400 can be configured to mask identified objects in real time during the application of augmented reality overlays, such as by performing the above or other operations when presenting augmented reality images to a user. As used herein, an operation can be performed "in real time" when it is performed immediately without excessive delay. In some examples, real-time data processing operations can be performed relative to highly dynamic and time-sensitive data (e.g., data that becomes irrelevant after a very short time, such as image data captured by an image capture device by a user moving and reorienting to capture an image sequence representing an image at a location of the image capture device). Therefore, real-time operations will be understood as those operations designed to mask identified objects during the application of composite elements based on relevant and up-to-date data, even though it will also be understood that real-time operations are not performed instantaneously.

[0056] This document describes the above operations and other suitable operations that may be performed by processing facility 404 in more detail. In the following description, any reference to functions performed by system 400 shall be understood as being performed by processing facility 404 based on instructions 406 stored in storage facility 402.

[0057] Figure 5 An exemplary configuration 500 is shown in which system 400 can operate to mask identified objects during the application of compositing elements to an original image. Specifically, as will be described in more detail below, configuration 500 depicts a scene 502 captured by image capture device 504 to generate image data 506 provided to or otherwise accessed by system 400. In some embodiments, by using one or more models 508 with some additional depth data 510 and / or spatial data 512, system 400 generates a set of presentation data 514 provided to or otherwise accessed by presentation system 516. Presentation system 516 presents an enhanced version of image data 506 (e.g., an augmented reality image based on image data 506) to user 520 via monitor 518. Reference will now be made to... Figure 5 and reference Figure 6-8 Describe each of the 500 components in the configuration.

[0058] Scene 502 can be implemented as any type of real-world scene (as opposed to purely virtual), workplace, location, area, or other type of scene captured by an image capture device (such as image capture device 504) (e.g., photography, video recording, etc.). In some examples, scene 502 can be a real-world scene, large or small, existing indoors or outdoors, and including any type of objects and / or scenery (e.g., people, cars, houses, furniture, etc.) described herein. In other examples, as will be referenced below. Figure 9 and Figure 10 In more specific detail, scenario 502 may be associated with a specific real-world scenario, such as a surgical site within a body on which a surgical procedure is being performed (e.g., the body of a living patient, a corpse, a training device, an animal, etc.).

[0059] Image capture device 504 can be implemented as any suitable device for capturing images at scene 502. For example, if scene 502 is a relatively large-scale real-world scene, such as an outdoor scene, home, workplace, etc., image capture device 504 can be implemented as a camera (e.g., a still camera, video camera, etc.). Such a camera can be a single-field-of-view camera that captures images of scene 502 from a single vantage point, or a stereo camera that captures images of scene 502 from a stereo vantage point (shown using dashed lines extending from the corresponding elements of image capture device 504 to the corners of scene 502), as indicated by the dual right (“R”) and left (“L”) elements of image capture device 504. In yet another example, image capture device 504 may have additional elements configured to allow image capture device 504 to capture wider-angle images such as panoramic images (e.g., 360° images, spherical images, etc.).

[0060] The following will be combined Figure 9 and Figure 10 As described in more detail, some implementations of system 400 can be implemented in the context of a computer-aided surgical procedure. In such implementations, image capture device 504 may be implemented as an endoscope (e.g., a single-field or stereoscopic endoscope) or configured to capture images at a surgical site included in scenario 502 or another suitable medical imaging modality.

[0061] Image data 506 is shown as communication between image capture device 504 and system 400. For example, image data 506 may represent an image captured by image capture device 504 (e.g., the original image such as image 100), instructions for capturing such an image (e.g., commands for capturing an image, synchronization information, etc.), or any other image-related information transmitted between system 400 and image capture device 504. Depending on the use case or application, the image represented by image data 506 can be of various types and may include various types of images that can be used in a particular implementation.

[0062] For example, in some embodiments, image data 506 may represent still images, such as photographs captured by image capture device 504. In some examples, such images may comprise a collection of different images of substantially the same portion of scene 502 captured at substantially the same time. For example, a stereoscopic image may comprise two or more similar images captured simultaneously from different vantage points, such that depth information can be derived from the differences between the images. In other examples, still images may comprise a collection of different images captured (e.g., at the same or different times) of superimposed portions of scene 502, so that they can be combined to form a panoramic image (e.g., a 360° image, a spherical image, etc.). In these or other embodiments, image data 506 may represent a video image consisting of a sequence of image frames (i.e., images of the same scene 502 captured sequentially by a camera in a continuous timeframe). Each image frame in such a video image may depict an object at scene 502 (e.g., including identified objects) as the object moves relative to the remainder of the image depicted in the image.

[0063] As mentioned above, in contrast to Figure 4 As described in detail, system 400 can be configured to perform various operations to associate models with identified objects and generate rendering data representing the masking of the identified objects based on the models associated with them. For example, system 400 may receive or otherwise access image data 506 from image capture device 504 and analyze the image represented in the image data to identify the identified objects whose models (e.g., 3D models) are available in model 508. Model 508 may represent a repository of models (e.g., 2D models, 3D models, etc.) included within system 400 (e.g., stored in storage facility 402) or stored in a repository (e.g., a database or other such data storage device) communicatively coupled to and accessible by system 400.

[0064] Each of the models 508 can be any type of representation of an object generated at runtime, prior to the moment when system 400 is performing operations to generate a depth map and analyze captured images of scene 502 to generate presentation data 514. For example, model 508 can represent a detailed CAD model of certain objects. Such CAD models can be used for the design of various types of objects and are available when the objects are purchased and used. For example, surgical instruments can be associated with a highly detailed and accurate CAD model that is available for use by the person and system using the surgical instruments. In other examples, model 508 can represent a detailed, high-density scan (e.g., a 3D scan) of an object that has been performed previously (e.g., prior to runtime during a period of time when the limitations of real-time processing described herein are relaxed). For example, a 3D scanner can be used to generate high-density 3D models of various types of objects that are intended to be part of a particular augmented reality experience and are expected to be in the foreground (i.e., in front of the augmented reality overlay) in the augmented reality experience.

[0065] Based on one or more of image data 506 and model 508, system 400 can associate one or more models with one or more identified objects depicted in image data 506, and use this association to perform further analysis of the original image of scene 502. For example, once a model is associated with an identified object depicted in the original image, system 400 can accurately segment the original image to distinguish pixels depicting one or more identified objects from pixels not depicting one or more identified objects. This segmentation of the image represented by image data 506 can be performed in any suitable manner, including by using semantic scene segmentation techniques, where each pixel in the image is assigned to correspond to a specific object or set of images (e.g., one of the identified objects such as object 102, another object such as object 104, another part of the image such as background 106, etc.).

[0066] For certain images, and at least to some extent, image segmentation can be performed solely based on data representing the color and / or shadow of each point on the surface of an object or scene captured within image data 506 (i.e., what is captured when image capturing device 504 captures light reflected or originating at such points). This type of data will be referred to herein as color data, although it will be understood that in some examples, such color data may be implemented from grayscale image data, infrared image data, or other types of image data not explicitly associated with visible colors. While color data can be used to perform image segmentation, color data alone may not provide a sufficient basis for performing accurate, detailed, and real-time segmentation of certain images. In such examples, depth data may be used instead of color data or in addition to color data to segment the image accurately and efficiently. While color data represents the appearance (e.g., color, texture, etc.) of surface points of an object at a location, depth data represents the position of surface points relative to a particular location (such as a vantage point associated with the image capturing device) (i.e., the depth of the surface point, how far each surface point is from the vantage point, etc.).

[0067] Therefore, system 400 can generate a depth map of an image depicted by an image captured by image capture device 504 and represented by image data 506. Such a depth map can be generated based on one or more models 508 that have been associated with (e.g., registered or otherwise bound to or corresponded to) one or more identified objects in a scene, and in some examples, additional depth data 510, spatial data 512, and / or other data that may be used in a particular implementation. The depth map may include first depth data for depicting identified objects within the image and second depth data for the remainder of the image. For example, if the image is image 100 and object 102 is the identified object, system 400 can use model 508, additional depth data 510, and / or spatial data 512 to generate a depth map that includes detailed depth data of object 102 (e.g., in...). Figure 3 (depth data near level "9" in the example) and additional depth data for object 104 and / or background 106 (e.g., in...) Figure 3 (Depth data near levels “0”-“3” in the example).

[0068] The first depth data (e.g., depth data of object 102) may be denser than the second depth data and may be based on one or more of the aforementioned models 508 (e.g., a 3D model of object 102) and on spatial data 512 already used for registration or otherwise associating the model with object 102 (e.g., camera calibration data for image capture device 504, kinematic data indicating the position and / or orientation of object 102 relative to image capture device 504, etc.). The second depth data (e.g., stereo depth detection performed by stereo image capture device 504, time-of-flight depth detection performed by a time-of-flight scanner built into or otherwise associated with image capture device 504, SLAM technology, single-field-of-view (SfM) technology, etc.) can be generated or accessed using conventional real-time technologies and techniques based on the additional depth data 510.

[0069] Once a depth map is generated, system 400 can use the depth map to segment the original image to distinguish pixels of the original image depicting object 102 (the identified object) from pixels of the original image not depicting object 102. For example, this segmentation can be performed by identifying pixels of the image depicting object 102 based on first depth data and pixels of the image not depicting the identified object based on second depth data. Based on the segmentation of the image performed in the manner described herein, or based on other operations not involving explicit segmentation as already described, system 400 can generate rendering data 514, and in some examples such as those shown in configuration 500, the generated rendering data can be provided to rendering system 516 for presenting an enhanced version of the original image (e.g., an augmented reality image) to user 520.

[0070] As already mentioned, in some examples, system 400 may generate rendering data 514 as image data representing an enhanced version of the original image, and this rendering data may be rendered immediately by rendering system 516. However, in other examples, system 400 may generate rendering data 514 in a different form, configured to otherwise facilitate rendering system 516 in rendering the enhanced version of the original image. For example, in some implementations, rendering data 514 may include original image data (e.g., image data 506) and generated masking data and / or other data (e.g., metadata, data representing compositing elements that will be used to enhance the original image) to guide rendering system 516 itself in constructing and rendering the enhanced version of the original image.

[0071] Figure 6An exemplary representation 600 of such masking data that may be included in rendering data 514 is shown. Specifically, representation 600 includes black and white pixels associated with each of the original pixels of image 100. As shown, the white pixels in representation 600 correspond to pixels in image 100 that have been determined to correspond to object 102 (i.e., the identified object in this example) based on the association of model 508 with the identified object and the resulting dense segmentation of the object from the rest of the image. In contrast, the black pixels in representation 600 correspond to pixels in image 100 that have been determined to correspond to objects or images other than object 102 (e.g., object 104, background 106, etc.) based on the same segmentation. While black and white pixels are used to depict masking data 514 in representation 600, it should be understood that masking data 514 may take any suitable form, such as that which pixels of the image correspond to the identified object and which do not, which can be used to indicate to the rendering system which pixels of the image correspond to the identified object and which do not. For example, in some examples, the black and white color may be switched, other colors may be used, or another data structure indicating whether each pixel depicts an identified object may be employed.

[0072] System 400 can generate a mask, such as the one shown by representation 600, which can be applied to composite elements, such as composite element 202 (e.g., augmented reality overlay). When such a mask is applied, if a pixel of composite element 202 is to be covered by object 102 (i.e., to make it appear to be behind object 102), that pixel can be subtracted from or removed from composite element 202.

[0073] To illustrate, Figure 7 The composite element 702 is shown. Composite element 702 will be understood as described above in the following manner. Figure 2 The version of the composite element 202 superimposed on image 100 at the same location shown. However, as shown, in the figure, the composite element 202 is superimposed on image 100. Figure 6 The exemplary masking data shown in representation 600 has been applied to composite element 702 after composite element 202. Therefore, pixels that will be depicted as composite element 202 behind object 102 have been removed or masked from composite element 702.

[0074] To illustrate how the segmentation performed by system 400 based on model 508 can improve the masking of the identified object 102 during the application of augmented reality overlay, the above description can be used as an example. Figure 3 and Figure 8 Compare them.

[0075] Similar to Figure 3 , Figure 8Exemplary aspects of how depth data can be detected and used to mask object 102 when synthetic elements are applied to an original image to form an enhanced version of the original image are illustrated. However, with Figure 3 on the contrary, Figure 8 An example is shown in which depth data from model 508 of the identified object is used in conjunction with regular depth data (e.g., additional depth data 510) to improve the segmentation and masking of the identified object 102 when synthetic elements are applied.

[0076] Specifically, Figure 8 The representation 800-1 in the figure shows, for example, Figure 3 Representation 300-1 shows the same actual depth data for the objects and images of image 100; representation 800-2 shows how first and second depth data of different densities are combined in a single depth map of the objects and images of image 100; and representation 800-3 shows an implementation of augmented reality image 200 in which composite element 702 is depicted in such a way that it actually appears to be behind object 102. Each of representations 800 (i.e., representations 800-1 to 800-3) will now be described in more detail.

[0077] Representation 800-1 depicts the actual depth data of each of objects 102 and 104 in image 100, as well as the background 106. Representation 800-1 is the same as representation 300-1, and as described above, uses a simple notation with a single digit (e.g., 0, 1, 2, etc.) to distinguish areas of image 100 depicting images located at different depths relative to the image capture device capturing image 100. (See above regarding...) Figure 3 As described, generating highly detailed or dense depth maps in real time to capture all the nuances of the actual depth data shown in 800-1 may be impossible or infeasible.

[0078] Similar to representation 300-2 above, representation 800-2 illustrates some practical limitations imposed by real-time depth detection using conventional depth detection techniques. As in representation 300-2, instead of capturing depth for image 100 at every point (e.g., every pixel), depth is detected only at each point of image 100 divided by black "x"s in representation 800-2. While the density of the depth map in representation 800-2 is the same as that in representation 300-2 for both object 104 and background 106, Figure 8 The density of the depth map of the identified object 102 may differ significantly in representation 800-2 compared to representation 300-2. Specifically, the black "x"s are shown as being so densely packed on object 102 that they are even visible in [the image / image]. Figure 8The objects are indistinguishable from each other (making object 102 appear almost as solid black grids). This is because the depth data of the identified object 102 is not based on (or at least not specifically based on) real-time depth detection techniques performed on the rest of the image in image 100. Instead, as described above, the depth data of the identified object 102 (“first depth data”) is generated based on model 508 of object 102, which represents object 102 in a large amount of detail and has been registered or otherwise associated with object 102 in image 100, so that the details of object 102 do not need to be scanned and determined in real time. Because the depth data of object 102 is so dense, representation 800-2 shows that each edge of object 102 can be well defined by system 400 to generate rendering data that depicts the accurate application of masking to the appropriate portion of composite element 202, or at least includes masking data that enables rendering system 516 to accurately apply masking data to composite element 202 (see Figure 7 Very accurate masking data (see) Figure 6 ).

[0079] The results and some of their benefits are illustrated in Representation 800-3. As shown, a portion of the augmented reality overlay 202 has been masked to form composite element 702, which very accurately and precisely represents object 102 so as not to cover object 102. Therefore, Representation 800-3 provides a more convincing representation of the identified object 102 in front of the augmented reality overlay compared to Representation 300-3. As shown in Representation 800-3, composite element 702 is well aligned so as to convincingly appear to be behind (i.e., occluded by) object 102, while still in front of object 104 and background 106 (i.e., still used to occlude them).

[0080] Due to dynamic changes in the image captured at scene 502 (e.g., due to movement of objects 102 and 104 relative to background 106, etc.), it may be desirable to track the identified object 102 frame-by-frame in the video image so that the presentation data can be continuously updated to provide the appearance of object 102 in front of the composite element 202, even when object 102 and / or the composite element are in motion. To this end, system 400 may be configured to continuously identify the identified object 102 within the image depicted by the video image by: initially identifying the identified object 102 in the first image frame of the image frame sequence, and tracking the identified object 102 frame-by-frame (e.g., based on the initial identification) as the identified object 102 moves relative to the remainder of the image in subsequent image frames of the image frame sequence. For example, the identified object 102 may be identified in the first image frame of the video image represented by image data 506 using object recognition techniques such as computer vision and / or relying on color data, depth data, previous identification of object 102 (e.g., machine-learned). Once object 102 has been identified, system 400 can avoid having to perform object identification technology again for each frame by tracking object 102 frame by frame as the object moves gradually in scene 502.

[0081] Return to Figure 5 Spatial data 512 may also be received by or otherwise accessed by system 400 to assist in initially identifying object 102 in a first image frame, tracking object 102 frame-by-frame in subsequent image frames, associating one of models 508 with object 102, segmenting the image based on this association between model 508 and object 102 (as described above), and / or for any other purpose as may be used in a particular implementation. More specifically, spatial data 512 may include any of various types of data used to determine the spatial characteristics of the identified object (particularly with respect to image capture device 504 and the images captured by image capture device 504). For example, spatial data 512 may include data associated with computer vision and / or object recognition techniques (e.g., techniques that utilize machine-learned data and are trained using data obtained from previously captured and analyzed images). Thus, while model 508 may be configured to define the geometric details of the identified object 102, spatial data 512 may be generated or accessed to correlate how model 508 relates to object 102 in the image. For example, spatial data 512 may include any of a variety of types of data representing the spatial pose of the identified object 102 (i.e., information about the precise location and manner in which the identified object 102 is positioned relative to scene 502 and / or image capture device 504 at any given moment).

[0082] In some implementations, spatial data 512 may include kinematic data tracked by a computer-assisted medical system configured to move a robotic arm to perform robot-assisted surgery as described in some of the examples herein. In such examples, precise kinematic data can be used for each robotic arm and any surgical instruments or other objects held by such robotic arms to allow a user (e.g., a surgeon, etc.) precise control over the robotic arms. Thus, by accessing the kinematic data included in spatial data 512, system 400 can identify identified objects (e.g., including initially identified objects, later-tracked objects, etc.) at least in part based on the kinematic data, and can precisely determine how the identified objects are positioned, oriented, etc., to associate model 508 with the identified objects.

[0083] In the same or other embodiments, spatial data 512 may be configured to support methods other than kinematic methods for identifying objects and / or determining the position and orientation of objects. For example, some embodiments may rely on computer vision techniques as described above, and spatial data 512 may include data configured to support computer vision techniques (e.g., training datasets for machine learning, etc.). As another example, some embodiments may involve an identified object in which an electromagnetic tracker is embedded, and the position and orientation of the identified object are tracked by monitoring the movement of the electromagnetic tracker through an electromagnetic field. In this example, spatial data 512 may include data associated with the position, orientation, and movement of the electromagnetic field and / or the electromagnetic tracker within the field.

[0084] Along with data providing the position and orientation of objects (e.g., including identified objects) at scene 502, spatial data 512 may also include camera calibration data of image capture device 504. For example, spatial data 512 may include data representing intrinsic or extrinsic parameters of image capture device 504, including data representing the focal length of image capture device 504, lens distortion parameters of image capture device 504, principal point of image capture device 504, etc. Once the position and orientation of the identified objects have been determined, such data can help to accurately generate dense portions of the depth map used to generate rendering data (e.g., distinguishing identified objects, segmenting images, generating masking data, etc.). This is because camera calibration parameters allow system 400 to accurately associate the model with the identified objects by precisely determining how to depict the identified objects in the image captured by image capture device 504 for a given position and orientation of the identified objects relative to image capture device 504.

[0085] The rendering system 516 may receive rendering data 514 from the system 400 and render (e.g., render) the rendering data, or construct a renderable image based on the properties of the rendering data 514 that may be suitable for the provided rendering data 514. For example, if the rendering data 514 includes, for example, elements such as those derived from... Figure 6 If the masking data represented by 600 is used, then the rendering system 516 can apply the mask represented by the masking data to the compositing elements to be integrated with the original image to form an enhanced version of the original image (i.e., an augmented reality image) that will be presented to the user 520 via the monitor 518. For example, as shown in the figure. Figure 8 As shown in 800-3, the presentation system 516 can present an augmented reality image depicting a composite element 702 located behind the identified object 102. For this purpose, the presentation system 516 can be implemented by any suitable presentation system configured to present an augmented reality experience or other such experience to a user, including but not limited to augmented reality media player devices (e.g., dedicated head-mounted augmented reality devices), standard mobile devices (such as smartphones) that can be held at arm's length or mounted on the head, surgical consoles or auxiliary consoles of computer-assisted medical systems, such as those described in more detail below, or any other suitable presentation system. In some examples, the presentation system 516 may be incorporated into system 400 (i.e., built into the system, integrated with the system, etc.), while in other examples, the presentation system 516 may be separate from but communicatively coupled to system 400.

[0086] Monitor 518 can be any suitable type of presentation screen or other monitor (or multiple monitors) configured to present augmented reality images to user 520. In some examples, monitor 518 may be implemented as a device screen such as a computer monitor, television, smartphone, or tablet. In other examples, monitor 518 may be implemented as a pair of small displays configured to present images to each eye of user 520 (e.g., a head-mounted augmented reality device, a surgeon's console presenting stereoscopic images, etc.). Thus, user 520 can represent anyone experiencing the content (e.g., augmented reality content) presented by presentation system 516 based on data received from system 400. For example, user 520 could be someone playing an augmented reality game or using another type of extended reality application, a surgeon or surgical team member assisting in a surgical procedure, or any other suitable person experiencing the content presented by presentation system 516.

[0087] The rendering system 516 may render an enhanced version of the original image based on the rendering data 514 in any manner suitable for preventing occlusion (complete or partial occlusion) of identified objects depicted in the original image by means of enhancement content (composite elements added to the original image). In some examples, for instance, the rendering system 516 may render an enhanced image based on the rendering data 514, comprising an enhanced portion of the original image enhanced by only a portion of the composite element (or only a portion of other enhancement content). The displayed portion of the composite element may be any part of the composite element, such as a portion consisting of consecutive pixels or an aggregated portion consisting of non-consecutive pixels (e.g., consecutive pixels that together form a separate group of pixels). By rendering only a portion of the composite element in the enhanced image, the rendering system 516 omits different portions of the composite element from the enhanced image. For example, instead of rendering pixels associated with the omitted portion of the composite element, the rendering system 516 may render pixels associated with the identified object to prevent the identified object from being occluded by the omitted portion of the composite element. Pixels to be rendered or not rendered in the enhanced image may be identified by the rendering system 516 based on the rendering data 514 in any suitable manner, including by performing any of the masking operations described herein.

[0088] Throughout the above description, various types of use cases have been described, all of which can be well served by System 400 and the systems and methods described herein for masking identified objects during the application of synthetic elements to the original image. As already mentioned, a specific example relating to a computer-assisted surgical procedure will now be described in more detail. For reasons that will become apparent, System 400 and its principles described herein are particularly well suited to this example of augmented reality-assisted surgery.

[0089] As used herein, a surgical procedure may include any medical procedure, including any diagnostic, medical, or therapeutic procedure in which manual and / or instrumental techniques are used on the body of a patient or other subject to investigate or treat a physical condition. A surgical procedure may refer to any stage of a medical procedure, such as the preoperative stage, the surgical (i.e., intraoperative) stage, and the postoperative stage.

[0090] In such applications of the systems and methods described herein, scenario 502 will be understood as a surgical site that includes any volumetric space associated with a surgical procedure. For example, a surgical site may include one or more parts of the body of a patient or other subject undergoing surgery within the space associated with the surgical procedure. In some examples, the surgical site may be entirely located within the body and may include a space within the body adjacent to the location of a surgical procedure that is planned, being performed, or has been performed. For example, for a minimally invasive surgical procedure performed on tissue inside a patient, the surgical site may include surface tissue, anatomical structures beneath the surface tissue, and space surrounding tissue in which, for example, surgical instruments for performing the surgical procedure are located. In other examples, the surgical site may be at least partially located outside the patient. For example, for an open surgical procedure performed on a patient, a portion of the surgical site (e.g., the tissue on which it is being operated) may be inside the patient, while another portion of the surgical site (e.g., the space surrounding tissue in which one or more surgical instruments may be located) may be outside the patient.

[0091] Figure 9 An exemplary computer-assisted surgical system 900 (“surgical system 900”) is illustrated. As already mentioned, system 400 may be implemented by or within surgical system 900, or may be decoupled from but communicatively coupled to surgical system 900. For example, system 400 may receive input from and provide output to surgical system 900, and / or may access images of surgical sites, information about surgical sites, and / or information about surgical system 900 from surgical system 900. System 400 may use the accessed images and / or information to perform any of the processes described herein to generate a composite image of the surgical site and provide data representing the composite image to surgical system 900 for display.

[0092] As shown in the figure, the surgical system 900 may include a control system 902, a user control system 904 (also referred to herein as a surgeon's console), and an auxiliary system 906 (also referred to herein as an auxiliary console) that are communicatively coupled to each other. The surgical system 900 can be used by a surgical team to perform computer-aided surgical procedures on a patient 908. As shown in the figure, the surgical team may include a surgeon 910-1, an assistant 910-2, a nurse 910-3, and an anesthesiologist 910-4, all of whom may be collectively referred to as "surgical team members 910". Additional or alternative surgical team members may be present during a surgical session, such as surgical team members that may be used in a particular implementation.

[0093] Although Figure 9A minimally invasive surgical procedure in progress is illustrated, but it should be understood that the surgical system 900 can be similarly used to perform open surgical procedures or other types of surgical procedures that can similarly benefit from the accuracy and convenience of the surgical system 900. Furthermore, it will be understood that the surgical stages in which the surgical system 900 can be employed may include not only those such as Figure 9 The surgical procedure shown may include the surgical phases, and may also include the preoperative phases, postoperative phases, and / or other suitable phases of the surgical procedure.

[0094] like Figure 9 As shown, the manipulation system 902 may include multiple surgical instruments (e.g., surgical instruments that can be identified by system 400 as having an identified object with a corresponding model 508, as described above) coupled to multiple manipulator arms 912 (e.g., manipulator arms 912-1 to 912-4). Each surgical instrument may be implemented by any suitable therapeutic instrument (e.g., a tool with tissue interaction capabilities), imaging device (e.g., an endoscope, ultrasound tool, etc.), diagnostic instrument, or analogues that can be used for computer-aided surgical procedures on patient 908 (e.g., by being at least partially inserted into and manipulated to perform computer-aided surgical procedures on patient 908). In some examples, one or more of the surgical instruments may include force sensing and / or other sensing capabilities. In some examples, the surgical instrument may be implemented by an ultrasound module, or such an ultrasound module may be connected to or coupled to one of the other surgical instruments described above. Although the manipulation system 902 is depicted and described herein as comprising four manipulator arms 912, it will be appreciated that the manipulation system 902 may comprise only a single manipulator arm 912 or any other number of manipulator arms, as may be used in a particular implementation.

[0095] The manipulator arm 912 and / or surgical instruments attached to the manipulator arm 912 may include one or more displacement transducers, orientation sensors, and / or position sensors for generating raw (i.e., uncorrected) kinematic information. For example, such kinematic information may be represented by kinematic data included within the aforementioned spatial data 512. As already mentioned, system 400 and / or surgical system 900 may be configured to use kinematic information to track surgical instruments (e.g., determine the position of surgical instruments) and / or control surgical instruments (and anything held or attached to the instruments by the instruments, such as needles, ultrasound modules, retracted tissue blocks, etc.).

[0096] User control system 904 may be configured to assist surgeon 910-1 in controlling manipulator arm 912 and surgical instruments attached to manipulator arm 912. For example, surgeon 910-1 may interact with user control system 904 to remotely move or manipulate manipulator arm 912 and surgical instruments. To this end, user control system 904 may provide surgeon 910-1 with images of surgical sites (e.g., scene 502) associated with patient 908, captured by an image capture device (e.g., image capture device 504). In some examples, user control system 904 may include a stereoscopic viewer with two displays on which stereoscopic images of the surgical sites associated with patient 908 and generated by a stereoscopic imaging system can be viewed by surgeon 910-1. As described above, in some examples, augmented reality images generated by system 400 or presentation system 516 may be displayed by user control system 904. In such cases, surgeon 910-1 may use the images displayed by user control system 904 to perform one or more procedures, wherein one or more surgical instruments are attached to manipulator arm 912.

[0097] To facilitate control of surgical instruments, the user control system 904 may include a set of master controls. These master controls can be manipulated by the surgeon 910-1 to control the movement of surgical instruments (e.g., by utilizing robotics and / or teleoperation technology). The master controls can be configured to detect a wide variety of hand, wrist, and finger movements performed by the surgeon 910-1. In this way, the surgeon 910-1 can intuitively perform procedures using one or more surgical instruments.

[0098] The auxiliary system 906 may include one or more computing devices configured to perform primary processing operations of the surgical system 900. In such a configuration, the one or more computing devices included in the auxiliary system 906 may control and / or coordinate operations performed by various other components of the surgical system 900, such as the manipulation system 902 and the user control system 904. For example, the computing device included in the user control system 904 may transmit instructions to the manipulation system 902 via the one or more computing devices included in the auxiliary system 906. As another example, the auxiliary system 906 may (e.g., from the manipulation system 902) receive and process image data representing images captured by an image capture device, such as image capture device 504.

[0099] In some examples, the assistive system 906 can be configured to present visual content to a surgical team member 910 who may not have access to the images provided to the surgeon 910-1 at the user control system 904. For this purpose, the assistive system 906 can be implemented by including a display monitor 914. Figure 5The monitor 518 is configured to display one or more user interfaces and / or augmented reality images of surgical sites, information associated with patient 908 and / or surgical procedures, and / or any other visual content that may be used in a particular implementation. For example, the display monitor 914 may display augmented reality images of surgical sites, including real-time video captures and enhancements such as text and / or graphical content (e.g., preoperatively generated anatomical models, contextual information, etc.) displayed concurrently with the images. In some embodiments, the display monitor 914 is implemented as a touchscreen display, which a surgical team member 910 may interact with (e.g., via touch gestures) to provide user input to the surgical system 900.

[0100] The operating system 902, the user control system 904, and the auxiliary system 906 can be communicatively coupled to each other in any suitable manner. For example, Figure 9 As shown, the operating system 902, user control system 904, and auxiliary system 906 can be communicatively coupled via control line 916, which can represent any wired or wireless communication link as may be used in a particular implementation. Therefore, the operating system 902, user control system 904, and auxiliary system 906 may each include one or more wired or wireless communication interfaces, such as one or more local area network interfaces, Wi-Fi network interfaces, cellular interfaces, etc.

[0101] In order to apply the principles described in this article to... Figure 9 The surgical context described above, which has been described and illustrated in a relatively general manner, can be specifically applied to a surgical context. For example, in the example of augmented reality-enhanced surgical procedures, the image depicted by the image could be an image of a surgical site where a surgical procedure is performed using computer-assisted surgical instruments, the identified object could be the computer-assisted surgical instruments, and the augmented reality overlay depicted as being located behind the computer-assisted surgical instruments could be an anatomical model generated using a preoperative imaging modality prior to the surgical procedure.

[0102] To illustrate, Figure 10 An exemplary aspect of masking identified objects during the application of synthetic elements to an original image is illustrated in a specific scenario involving a surgical procedure performed using a surgical system 900. Specifically, as shown, image 1000 depicts a surgical scene including tissues and other anatomical structures manipulated or otherwise surgically operated on by surgical instruments 1002. Image 1000 represents the original image to which no enhancement has been applied (e.g., similar to image 100 described more generally above).

[0103] Figure 10Two images 1004 (i.e., images 1004-1 and 1004-2) are also included to demonstrate the results and benefits of employing the system and method described herein. Each image 1004 may represent a processed image of the augmented reality image 200 described above, incorporating a composite element 1006 (e.g., augmented reality overlay) similar to the composite element 202 applied to image 100 using augmented reality or other such extended reality techniques.

[0104] In augmented reality image 1004-1, composite element 1006 is applied to image 1000 in a conventional manner (e.g., typical augmented reality overlay techniques). This conventional manner does not take into account any objects at the surgical site, but rather overlays composite element 1006 in front of (i.e., occluded by) all objects and other scenery depicted in image 1000. This type of overlay application may be suitable for certain use cases, but it should be noted that in this example of augmented reality enhancement for a surgical procedure, it may be undesirable, or at least non-ideal. This is partly due to the nature of composite element 1006 and what the system is designed to do by including augmented reality overlays.

[0105] Composite element 1006 can be any suitable image, depiction, or representation of information that can be used in a particular implementation to assist surgical team members in performing surgical procedures. For example, in some examples (such as the example shown), composite element 1006 can be implemented as a model or other representation of a subsurface anatomical structure that is of interest to the surgical team but is not visible during the surgical procedure. As an example, composite element 1006 can represent a vascular system located just beneath the visible surface of the tissue, which has been imaged by a modality other than endoscopic image 1000 (e.g., ultrasound imaging of the vascular system during surgery, magnetic resonance imaging (“MRI”) or computed tomography (“CT”) imaging of the vascular system preoperatively). Such a vascular system may be invisible to the surgeon while the surgeon is controlling surgical instruments 1002, but may be of interest because the precise location of certain vascular systems can influence the decisions made by the surgeon.

[0106] In other examples, composite element 1006 may represent textual or graphic information that is expected to be projected directly onto surface tissue, a clean-up rendering of the tissue itself (e.g., a representation of tissue if stagnant blood, fat, smoke, or other obstructions were not present), or other such enhancements. In all these examples, it is not necessarily expected that composite element 1006 obstructs the depiction of surgical instrument 1002. For example, as shown in Figure 1004-1, where composite element 1006 obscures surgical instrument 1002, the view of instrument 1002 obstructed in this way by composite element 1006 may be disorienting, distracting, inconvenient, aesthetically unappealing, or otherwise undesirable.

[0107] Therefore, in image 1004-2, composite element 1006 is applied to image 1000 according to the methods and techniques described herein so as to apply composite element 1006 in a manner that does not obscure or cover surgical instrument 1002 (i.e., in a manner that appears to be located behind surgical instrument 1002). For the reasons described above, this type of overlay application can be less disorienting, less distracting, more convenient, more aesthetically pleasing, etc., compared to the presentation of image 1004-1. Furthermore, since highly dense depth information can be generated for surgical instrument 1002 based on, for example, a CAD model of surgical instrument 1002, applying composite element 1006 to image 1000 in image 1004-2 will accurately align composite element 1006 and surgical instrument 1002 to provide, for example, Figure 8 The illustrated implementation features accurate and attractive augmented reality images (and with) Figure 3 The misaligned and less accurate implementations shown are presented in contrast.

[0108] Although surgical instrument 1002 is Figure 10The identified objects used as examples are, but it should be understood that in certain examples, any suitable object depicted in an image of a surgical site (such as image 1000) may appear in front of a composite element or augmented reality overlay. If the identified object is a computer-aided surgical instrument (such as surgical instrument 1002) used to perform a surgical procedure, the model accessed by system 400 may be implemented as a 3D CAD model of the computer-aided surgical instrument. However, in other examples, the identified object may be held by the computer-aided surgical instrument (rather than the instrument itself) or may be located elsewhere within the image. As an example, the identified object may be a needle and / or thread held by surgical instrument 1002 (not explicitly shown) for suturing a surgical procedure. As another example, the identified object may be an ultrasound module held or otherwise attached to surgical instrument 1002. In such examples where the identified object is held by a computer-aided surgical instrument used to perform a surgical procedure, the model accessed by system 400 may have been generated by a 3D scan of the identified object (e.g., a 3D scan performed preoperatively or intraoperatively if a CAD model is unavailable).

[0109] Figure 11 An exemplary method 1100 for masking identified objects during the application of synthetic elements to the original image is shown. Although Figure 11 Exemplary operation according to one embodiment is shown, but other embodiments may omit, add, reorder, combine and / or modify it. Figure 11 Any of the operations shown. Figure 11 One or more of the operations shown may be performed by an object masking system such as system 400, any component included therein, and / or any implementation thereof.

[0110] In operation 1102, the object masking system can access models of the identified objects depicted in the original image of the scene. Operation 1102 can be performed in any of the manner described herein.

[0111] In operation 1104, the object masking system can associate the model accessed in operation 1102 with the identified object depicted in the original image. Operation 1104 can be performed in any of the manner described herein.

[0112] In operation 1106, the object masking system can generate rendering data for use by the rendering system to render an enhanced version of the original image. In some examples, composite elements will be added to the original image for use in the enhanced version. Thus, operation 1106 can be performed in such a way that the composite elements prevent at least a portion of the identified object from being occluded based on a model associated with the identified object in operation 1104. In this way, the composite elements can be depicted in the enhanced version of the original image so that it appears as if the composite elements are located behind the identified object. Operation 1106 can be performed in any of the ways described herein.

[0113] In some examples, a non-transitory computer-readable medium for storing computer-readable instructions may be provided, based on the principles described herein. When executed by a processor of a computing device, the instructions may direct the processor and / or the computing device to perform one or more operations, including one or more operations described herein. Such instructions may be stored and / or transmitted using any of a variety of known computer-readable media.

[0114] As used herein, non-transitory computer-readable media may include any non-transitory storage medium that contributes to providing data (e.g., instructions) that can be read and / or executed by a computing device (e.g., by a processor of the computing device). For example, non-transitory computer-readable media may include, but is not limited to, any combination of non-volatile storage media and / or volatile storage media. Exemplary non-volatile storage media include, but are not limited to, read-only memory, flash memory, solid-state drives, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), ferroelectric random access memory (“RAM”), and optical discs (e.g., optical discs, digital video discs, Blu-ray discs, etc.). Exemplary volatile storage media include, but are not limited to, RAM (e.g., dynamic RAM).

[0115] Figure 12 An exemplary computing device 1200 is shown, which may be specifically configured to perform one or more of the processes described herein. Any of the systems, units, computing devices and / or other components described herein may be implemented by computing device 1200.

[0116] like Figure 12 As shown, computing device 1200 may include a communication interface 1202, a processor 1204, a storage device 1206, and an input / output (“I / O”) module 1208 that are communicatively connected to each other via communication infrastructure 1210. Although in Figure 12 An exemplary computing device 1200 is shown, but... Figure 12 The components shown are not intended to be limiting. Additional or alternative components may be used in other embodiments. A more detailed description will now follow. Figure 12 The components of the computing device 1200 shown.

[0117] Communication interface 1202 can be configured to communicate with one or more computing devices. Examples of communication interface 1202 include, but are not limited to, wired network interfaces (such as network interface cards), wireless network interfaces (such as wireless network interface cards), modems, audio / video connections, and any other suitable interfaces.

[0118] Processor 1204 generally refers to any type or form of processing unit capable of processing data and / or interpreting, executing one or more of the instructions, procedures and / or operations described herein, and / or directing their execution. Processor 1204 may perform operations by executing computer-executable instructions 1212 (e.g., application programs, software, code and / or other executable data instances) stored in storage device 1206.

[0119] Storage device 1206 may include one or more data storage media, devices, or configurations, and may take any type, form, and combination of data storage media and / or devices. For example, storage device 1206 may include, but is not limited to, any combination of non-volatile media and / or volatile media described herein. Electronic data, including the data described herein, may be stored temporarily and / or permanently in storage device 1206. For example, data representing computer-executable instructions 1212 configured to direct processor 1204 to perform any of the operations described herein may be stored within storage device 1206. In some examples, data may be arranged in one or more databases residing within storage device 1206.

[0120] I / O module 1208 may include one or more I / O modules configured to receive user input and provide user output. I / O module 1208 may include any hardware, firmware, software, or a combination thereof that supports input and output capabilities. For example, I / O module 1208 may include hardware and / or software for capturing user input, including but not limited to a keyboard or keypad, a touchscreen component (e.g., a touchscreen display), a receiver (e.g., an RF or infrared receiver), a motion sensor, and / or one or more input buttons.

[0121] I / O module 1208 may include one or more devices for presenting output to a user, including but not limited to a graphics engine, a display (e.g., a screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In some embodiments, I / O module 1208 is configured to provide graphical data to the display for presentation to a user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be used in a particular implementation.

[0122] In some examples, any of the facilities described herein may be implemented by or within one or more components of computing device 1200. For example, one or more applications 1212 residing within storage device 1206 may be configured to direct the implementation of processor 1204 to perform one or more operations or functions associated with processing facility 404 of system 400. Similarly, storage facility 402 of system 400 may be implemented by or within an implementation of storage device 1206.

[0123] Various exemplary embodiments have been described in the foregoing description with reference to the accompanying drawings. However, it will be apparent that various modifications and changes can be made thereto, and additional embodiments can be implemented without departing from the scope of the invention as set forth in the appended claims. For example, certain features of one embodiment described herein may be combined with or substituted for features of another embodiment described herein. Therefore, this specification and the accompanying drawings are to be considered illustrative rather than restrictive.

Claims

1. A system comprising: The memory stores instructions; as well as A processor, communicatively coupled to the memory and configured to execute the instructions, to: Access the model of the identified object depicted in the original image of the scene; Associate the model with the identified object; as well as Rendering data is generated for use by the rendering system to render an enhanced version of the original image, wherein the model associated with the identified object is used to prevent synthetic elements added to the original image from occluding at least a portion of the identified object.

2. The system of claim 1, wherein the association between the model and the identified object includes: Generate a depth map of an image depicted by the original image, the depth map including first depth data for depicting the identified object within the image and second depth data for the remainder of the image, the first depth data being based on the model of the identified object and being denser than the second depth data; and The original image is segmented to distinguish pixels of the original image depicting the identified object from pixels of the original image not depicting the identified object in the following manner. Based on the first depth data, identify the pixels of the original image depicting the identified object; as well as Based on the second depth data, the pixels of the original image that do not depict the identified object are identified.

3. The system according to claim 1, wherein: The original image is a video image composed of a sequence of image frames, each image frame depicting the identified object as it moves relative to other images depicted by the original image; and The processor is further configured to execute the instructions to identify the identified object within the image depicted by the video image in the following manner: The identified object is initially identified in the first image frame of the image frame sequence, and Based on the initial identification, the identified object is tracked frame by frame as it moves relative to the other images in subsequent image frames of the image frame sequence.

4. The system of claim 1, wherein the processor is further configured to execute the instructions to: Access kinematic data representing the pose of the identified object; and The identified object is identified in the image depicted by the original image based on the kinematic data.

5. The system according to claim 1, wherein: The scene depicted by the original image includes a surgical site where a surgical procedure is performed using computer-aided surgical instruments for performing surgical procedures. The identified object is the computer-assisted surgical instrument; and The synthetic element added to the original image for use in the enhanced version of the original image is an anatomical model generated using a preoperative imaging modality prior to the surgical procedure.

6. The system of claim 1, wherein the identified object is a computer-aided surgical instrument for performing a surgical procedure, and the model is a three-dimensional ("3D") computer-aided design ("CAD") model of the computer-aided surgical instrument.

7. The system of claim 1, wherein the identified object is held by a computer-aided surgical instrument for performing a surgical procedure, and the model is generated by a 3D scan of the identified object.

8. The system of claim 1, wherein the processor is further configured to execute the instructions to provide the generated rendering data to the rendering system for rendering the enhanced version of the original image.

9. A method comprising: The object masking system accesses models of the identified objects depicted in the original image of the scene; The object masking system associates the model with the identified object; as well as The object masking system generates rendering data for use by the rendering system to render an enhanced version of the original image, wherein the model associated with the identified object is used to prevent synthetic elements added to the original image from occluding at least a portion of the identified object.

10. The method of claim 9, wherein the association between the model and the identified object comprises: Generate a depth map of an image depicted by the original image, the depth map including first depth data for depicting the identified object within the image and second depth data for the remainder of the image, the first depth data being based on the model of the identified object and being denser than the second depth data; and The original image is segmented to distinguish pixels of the original image depicting the identified object from pixels of the original image not depicting the identified object in the following manner. Based on the first depth data, identify the pixels of the original image depicting the identified object; as well as Based on the second depth data, the pixels of the original image that do not depict the identified object are identified.