Data processing apparatus, medical observation apparatus and method
By analyzing and processing the input image data using data processing equipment, generating and combining stereoscopic images, the bandwidth and memory limitations of stereoscopic visualization are solved, achieving low-cost 3D visualization effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, the application of stereoscopic visualization in the medical and scientific fields is limited by bandwidth and memory capacity, and requires specialized stereoscopic imaging devices, making it impossible to effectively achieve three-dimensional visualization of complex multidimensional data.
The input image data is analyzed by data processing equipment to determine different categories and generate multiple stereo images. Different parallaxes are assigned and they are combined into a composite stereo image to achieve stereo visualization without relying on stereo imaging compatible devices.
It achieves stereoscopic visualization under low bandwidth and low memory conditions, accurately represents the three-dimensional structure of complex multidimensional data, and enhances the three-dimensional depth perception effect.
Smart Images

Figure CN121844559A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a data processing device for a medical viewing apparatus, such as an endoscope, a microscope or any other type of medical imaging device. The present invention further relates to a medical viewing apparatus comprising such a data processing device. Further, the present invention relates to a computer implemented method as well as to a computer readable medium and a computer program product. BACKGROUND
[0002] In the field of medical and biomedical viewing, there is a need to capture and view three-dimensional scenes containing various types of objects. For example, in such a scene, a large number of blood vessels, tissues, tumors, organs, living beings, materials and / or substances can be spread throughout three-dimensional space and partially or completely overlap each other.
[0003] In view of this, the visualization of such a scene can be a demanding process, as complex multi-dimensional data reflecting the various compositions and positions of the objects needs to be accurately represented. To enable human vision to accurately perceive and understand the visualization of such a scene, the compositions can be conveyed by color and the positions are typically conveyed by stereoscopic visualization.
[0004] The stereoscopic visualization method presents the viewer with a left-right pair of two-dimensional images with a certain perspective disparity. The left image is presented to the left eye and the right image is presented to the right eye. When viewed, the human brain perceives the images as a single three-dimensional view, giving the viewer the perception of three-dimensional depth. It is through the perspective disparity between the images seen by the viewer's left and right eyes (so-called binocular parallax) and the viewer's accommodation that the three-dimensional view is accomplished.
[0005] The separate presentation of the images - one for the left eye and one for the right eye - is typically merged by using specialized glasses and / or displays. There are two categories of three-dimensional viewing technology: active and passive. Active viewing utilizes electronic devices, such as specialized glasses, which interact with specialized displays. Passive viewing, for example, utilizes specialized displays to filter the constant stream of binocular input to the appropriate eye, with the display directing the images into the viewer's binoculars.
[0006] Typically, the two-dimensional image pair required for stereoscopic visualization needs to be captured using specialized equipment compatible with stereoscopic imaging. Therefore, so far, stereoscopic visualization has not been an option if no stereoscopic imaging equipment was available when the scene was captured.
[0007] Furthermore, bandwidth and / or memory capacity limitations can prohibit the use of stereoscopic visualization for certain applications. After all, the two-dimensional image pair occupies twice the bandwidth during transmission and twice the memory capacity for storage compared to a single image of the same resolution.
[0008] Therefore, there is a need to improve the applicability of stereoscopic visualization in the field of medical and other scientific disciplines.
[0009] It is therefore an object of the present application to provide means that generally facilitate the visualization of complex multidimensional data, in particular means that enable stereoscopic visualization of such data. SUMMARY
[0010] This object is achieved by a data processing device for a medical observation device, such as an endoscope or a microscope, the data processing device being configured to acquire input image data, the input image data representing a scene obtained by the medical observation device, analyze the input image data to determine different categories in the scene, generate a plurality of stereoscopic images from the input image data, each of the stereoscopic images representing a determined different category in the scene, assign a different parallax to each of the plurality of stereoscopic images based on the determined categories to produce a plurality of processed stereoscopic images, and combine the plurality of processed stereoscopic images to generate a combined stereoscopic image.
[0011] As will be described in further detail below, the categories determined by the data processing device serve to group the input image data based on the content of the input image data and based on commonalities of the content. In other words, parts of the content that have something in common with other parts of the content are determined to belong to the same category. For example, all parts of the content that show the same object, such as a blood vessel, a tissue, a tumor, an organ, a living being, a material, a substance, can belong to the same category.
[0012] Once the categories are determined, the categories are mapped onto a plurality of stereoscopic images, each of which then represents a determined category. The categories are objectively mapped onto the stereoscopic images, i.e. there is a one-to-one assignment between the categories and the stereoscopic images. A category is assigned to one stereoscopic image only, and likewise, each stereoscopic image belongs to one category only.
[0013] Here, a stereoscopic image consists of two individual images, each of which represents a different viewing channel, i.e. a different viewing direction. For example, each of the stereoscopic images comprises a digital image pair, wherein the left digital image of the digital image pair represents an image to be presented to the left eye of a viewer and the right digital image of the digital image pair represents an image to be presented to the right eye of the viewer. A processed stereoscopic image can refer to a stereoscopic image that is processed according to the parallax assigned to it.
[0014] To achieve the above-mentioned objects of the present application, the data processing device according to the present application is advantageous for the following reasons:
[0015] Primarily, the data processing device does not necessarily need the input image data to come from a device that is compatible with stereoscopic imaging. Thus, the amount of bandwidth and capacity used by the input image data is relatively low. For example, the input image data can be an RGB image, a color image with at least two color layers, a monochrome image, a hyperspectral imaging data cube, a multispectral imaging data cube, etc. Of course, the input image data can already be a pair of digital images representing a stereoscopic image, i.e. an input stereoscopic image. However, this is not a necessary condition for the data processing device to be able to generate a combined stereoscopic image.
[0016] Thus, the data processing device according to the present application enables stereoscopic visualization, thereby solving the above-mentioned objects.
[0017] The above-mentioned objects are further achieved by a medical viewing apparatus comprising the above-mentioned data processing device, an optical instrument configured to capture input image data and to provide the input image data to the data processing device, a user interface configured to receive user input and / or user selection input, and a display device configured to receive the combined stereoscopic image.
[0018] As the medical viewing apparatus comprises the data processing device according to the present application, it benefits from the functionalities and advantages of the data processing device described above. Thus, the medical viewing apparatus also achieves the objects of the present application. Further, the medical viewing apparatus is ready for use in the field of imaging for medical and biomedical viewing.
[0019] Further, the above-mentioned objects are solved by a computer-implemented method for processing input image data from a medical viewing apparatus, the method comprising the steps of: acquiring input image data representing a scene imaged by the medical viewing apparatus, analyzing the input image data to determine different classes in the imaged scene, generating a plurality of stereoscopic images from the input image data, each of the stereoscopic images representing a determined different class in the imaged scene, assigning a different disparity to each of the plurality of stereoscopic images based on the determined classes, resulting in a plurality of processed stereoscopic images, and combining the plurality of processed stereoscopic images with each other, resulting in a combined stereoscopic image.
[0020] The computer-implemented method achieves the above-mentioned objects, as it produces stereoscopic images from input image data, which themselves are not necessarily stereoscopic.
[0021] Finally, the above-mentioned objects are also solved by a computer-readable medium and a computer program product, each of which comprises instructions which, when executed by a computer, cause the computer to perform the method. In particular, the computer-readable medium and the computer program product allow the method of the present application to be implemented on a general-purpose computer, such as a personal computer (PC).
[0022] The data processing apparatus, medical observation device, and method can be further improved by adding one or more features described below. Each of these features can be added to the method and / or data processing apparatus independently of the others. In particular, a person skilled in the art, possessing knowledge of the inventive data processing apparatus, can configure the inventive method to operate the inventive data processing apparatus. Furthermore, each feature has its own advantageous technical effect, as will be explained below.
[0023] According to one possible implementation, generating multiple stereo images representing different categories may include the following steps: generating identical pairs of digital images, each digital image having a non-zero pixel value representing a category determined in the scene across all pixels. Pixels in the digital images with a value of zero do not represent a category represented by pixels with non-zero values in the same digital image. The identical pairs of digital images can be used as starting points for generating left and right digital images. This will be described in more detail below.
[0024] Preferably, the data processing device is further configured to apply the assigned disparity to multiple stereoscopic images to obtain multiple processed stereoscopic images. In other words, the data processing device can be configured to generate processed stereoscopic images by applying disparity assigned to the stereoscopic images. Therefore, the computer-implemented method may further include the step of generating multiple processed stereoscopic images by applying the assigned disparity to the stereoscopic images. In the context of this disclosure, assigning disparity can be used as a synonym for applying disparity to stereoscopic images.
[0025] For example, a disparity transformation, or disparity shift, can be applied to each processed stereo image based on the assigned disparity. The resulting processed stereo image differs from the (original) stereo image in that the individual digital images (comprising each stereo image) now have disparities based on the assigned disparities. Specifically, the data processing device can be configured to apply the corresponding disparities assigned to the stereo images in a plurality of stereo images by offsetting (i.e., moving) pixels in the left digital image of the stereo image in the opposite direction to represent the amount of assigned disparity, and by offsetting (i.e. moving) pixels in the right digital image of the stereo image in the opposite direction to represent the amount of assigned disparity.
[0026] When offsetting, pixels can move along a horizontal axis, which represents the horizontal distance direction of the viewer's eye. Furthermore, offsetting pixels refers to the process of moving the position of a pixel within a raster image where the rest remains unchanged. Typically, if a stereoscopic image appears closer to the viewer, pixels in the left and right digital images are offset from each other. Conversely, offsetting pixels in the left and right digital images further apart will make the stereoscopic image appear farther away from the viewer.
[0027] Preferably, the data processing device can be configured to assign parallax to multiple stereoscopic images based on user selection input. That is, the user can select which parallax to assign to which stereoscopic image via an input device (e.g., the user interface of a medical observation device). Additionally or optionally, the data processing device can be configured to automatically assign parallax to multiple stereoscopic images.
[0028] To automatically determine the disparity to be assigned to a particular stereo image, the following heuristic can be used: the greater the disparity assigned to a stereo image among multiple stereo images, the smaller the sum of the non-zero pixels in the stereo image representing the category assigned to that stereo image.
[0029] Therefore, categories represented by stereo images with a larger number of pixels will be assigned to smaller parallaxes, while categories represented by stereo images with a smaller number of pixels will be assigned to larger parallaxes.
[0030] Another rule of thumb is that the larger the range of a segment in a stereo image, the smaller the parallax assigned to that stereo image. Here, the segment represents a region / pixel (specifically, non-zero pixels) in the stereo image that illustrates the category of the stereo image's assignment. In other words, a segment in a stereo image defines the entirety or group of pixels within the stereo image. Furthermore, the segment can provide positional and / or shape information about the location of objects in the stereo image.
[0031] Specifically, the segment can be a binary mask, where non-zero pixels (i.e., pixels with non-zero values) are distributed according to the segment's position and shape, and zero-value pixels are treated as transparent pixels that do not belong to the segment. When superimposed on a digital image, the binary mask creates separation between pixels in the digital image that fall within the area / region included by the binary mask and pixels in the digital image that are excluded by the binary mask.
[0032] Optionally, a single image of each stereo image can be a binary mask in which non-zero pixels represent content, such as a portion of a defined category assigned to the stereo image, and zero-value pixels in a single image are treated as transparent pixels, which do not accumulate into a visible image contribution when the single image is superimposed on another single image from another stereo image.
[0033] According to another embodiment, when generating a combined stereoscopic image, multiple processed stereoscopic images can be superimposed on each other. Specifically, the multiple processed stereoscopic images can be superimposed along a central axis perpendicular to the image plane of each of the processed stereoscopic images, and this central axis passes through the center of each of the processed stereoscopic images. Further, the multiple processed stereoscopic images can be superimposed in an order determined by the disparity assigned to each processed stereoscopic image. For example, the processed stereoscopic image with the highest assigned disparity can be placed on top, followed by the remaining processed stereoscopic images arranged in descending order of disparity.
[0034] Placing the processed stereoscopic image with the highest assigned parallax at the top enhances the perceived 3D depth by the viewer. As briefly mentioned above, (binocular) parallax refers to the difference in the position of an object as seen by the left and right eyes due to binocular horizontal separation (parallax). In stereoscopic vision, the brain uses binocular parallax to extract depth information from two-dimensional retinal images. In other words, parallax represents the depth or distance of the content in a stereoscopic image that a viewer will perceive when viewing a composite stereoscopic image. The greater the parallax of an object, the closer its position appears to the viewer; therefore, the highest assigned parallax has a top position in the overlay image.
[0035] Preferably, the data processing device is configured to output the combined stereoscopic image to the display device of the medical observation apparatus. Therefore, the computer-implemented method may further include the step of outputting the combined stereoscopic image to the display device.
[0036] To further enhance the perceived 3D depth by the viewer, the data processing device can be configured to modify the aspect ratio of multiple stereoscopic images based on the parallax assigned to the stereoscopic image. Here, the aspect ratio represents the angular or lateral range of the processed stereoscopic image. Specifically, the aspect ratio represents the maximum angular diameter that the stereoscopic image will cover on the display device after multiple processed stereoscopic images are superimposed to generate a combined stereoscopic image.
[0037] For example, a data processing device can be configured to magnify one of multiple processed stereo images based on the parallax assigned to the processed stereo image. Here, magnification refers to scaling the stereo image by a scaling factor. The scaling factor can be greater than or less than 1. When magnifying a stereo image, the aspect ratio remains unchanged when the size of the stereo image is modified accordingly. That is, the scaled stereo image has the same aspect ratio as the stereo image before scaling was applied.
[0038] Preferably, the data processing device is configured to modify the scale before generating the processed stereoscopic image from the stereoscopic image. In other words, the data processing device can be configured to apply parallax transformation after the scale of the stereoscopic image has been modified. Therefore, the modification of the scale will not produce unwanted additional offsets after the parallax transformation.
[0039] Preferably, the smaller the size ratio / scaling factor, the smaller the parallax assigned to the processed stereoscopic image. For the viewer, this amplifies the impression that objects displayed in a processed stereoscopic image with smaller parallax are farther away than objects displayed in a processed stereoscopic image with larger parallax.
[0040] Especially for applications utilizing active viewing systems, the data processing device can be configured to receive user input indicating the user's viewing position relative to the display device. Based on the received user input, the data processing device can be configured to adjust the allocated parallax. Adjustments to the allocated parallax can be made when the received user input indicates a change in the user's viewing angle and / or viewing distance relative to the display device.
[0041] For example, user input could be facial detection data representing the user's eyes / gaze. The data processing device can be configured to derive from the facial detection data a value representing the user's viewing angle relative to the display device and / or another value representing the viewing distance. Further, the data processing device can be configured to determine changes in viewing angle and / or viewing distance when the value or other value exceeds a threshold. This threshold can be derived from previously received user input or facial detection data.
[0042] Additionally, the data processing device can be configured to determine changes in viewing angle and / or viewing distance when a value or other value exceeds a predetermined range. The predetermined range can be a neighborhood derived from previously received user input, a value, or another value. Conversely, the neighborhood can be defined based on previously received user input, a previous viewing angle, or a previous viewing distance.
[0043] The aforementioned adjustment of the assigned parallax may include determining updated parallax to be assigned to multiple stereoscopic images. The updated parallax reflects changes in viewing angle and / or viewing distance and can be used to update the combined stereoscopic image. In other words, the parallax assigned to multiple stereoscopic images is updated for each of the stereoscopic images to reflect changes in the user's viewing angle and / or viewing distance. That is, the data processing device is configured to dynamically compensate the user for changes in their viewing angle and / or viewing distance.
[0044] To activate or deactivate the adjustment of the parallax of the assigned image, the data processing device can be configured to receive user selection input. Therefore, the data processing device can be configured to activate or deactivate the updating of the combined stereoscopic image based on the received user selection input.
[0045] According to another possible implementation, the data processing device can be configured to apply a post-processing transformation to at least one subgroup of the multiple processed stereoscopic images before combining them. This post-processing transformation can be at least one of transparency adjustment, brightness adjustment, depth of focus adjustment, perspective scaling, and surface thickening effects.
[0046] The transparency adjustment increases the transparency level of the stereoscopic image when overlaying it. Increased transparency allows you to see through the stereoscopic image and observe the stereoscopic image below.
[0047] Brightness adjustment increases the brightness level of the stereoscopic image when overlaying the stereoscopic image. Increased brightness allows certain stereoscopic elements to be highlighted while others are softened.
[0048] Depth-of-focus adjustment increases the blur level of the stereo image when it is overlaid. For example, a blur filter, such as a Gaussian filter with a kernel size proportional to the desired blur level of the processed stereo image, can be applied to the stereo image. In particular, the same filter is applied to both the left and right images of the stereo image.
[0049] Perspective scaling involves reducing or enlarging the size of a stereoscopic image before it is overlaid, while maintaining the aspect ratio of the stereoscopic image. Therefore, smaller objects appear farther away than larger objects.
[0050] The thickening surface effect can include cloning pixel values from the midpoint of the parallax adjustment to the endpoint. The midpoint is located between the start and end points of the parallax adjustment.
[0051] Preferably, the same post-processing transform is applied to each individual digital image in the digital image pair constituting the stereoscopic image or the processed stereoscopic image. The user selection input mentioned above can also be used to activate or deactivate the application of the post-processing transform. Furthermore, the user selection input can be used to select the post-processing transform to be applied.
[0052] Specifically, the post-processing transformation can be applied to the processed stereoscopic image based on the assigned disparity. Further, the data processing device can be configured to designate a stereoscopic image as a background from among multiple processed stereoscopic images, wherein the disparity assigned to the background is the global minimum or maximum value among the disparities assigned to the multiple stereoscopic images. The background can serve as a reference for the viewer.
[0053] Optionally, the data processing device can be configured to increase or decrease the weight of the post-processing transformation when applying it to the processed stereo image, based on the distance between the background and the stereo image. Here, the difference in disparity between two stereo images among multiple stereo images can define the distance between the two stereo images.
[0054] According to one possible implementation, the disparity applied to each of the multiple stereo images can be different from all other assigned disparities. That is, each of the multiple stereo images can be assigned a unique disparity.
[0055] Alternatively, stereo images representing different categories of the same whole can have the same assigned disparity. For example, two or more stereo images can be designated as the background, so they share the same disparity.
[0056] The optical instruments in a medical observation device can be a stereo digital camera, but it can also be a digital camera, a multispectral camera, a hyperspectral camera, or a time-of-flight camera. More specifically, the digital camera can be a digital RGB camera and / or a digital reflectance camera. The multispectral camera can include a digital fluorescence camera, and optionally include additional digital fluorescence cameras. The response spectra of the digital fluorescence camera can preferably be non-overlapping and particularly complementary to each other.
[0057] Regardless of the camera type, the input image data provided by the optical instruments can contain a raster image composed of pixels.
[0058] For example, the input image data could contain an RGB image representing a scene and consisting of three equally sized raster images. Each raster image layer represents the color intensity of the scene in one of the corresponding color channels: red, green, and blue.
[0059] Similarly, the input image data may also comprise a multispectral or hyperspectral imaging data cube having layers of raster images of equal size in one of the respective multispectral or hyperspectral channels. For example, the input image data may comprise a digital fluorescence input image representing the fluorescence of one or more fluorophores present in the scene. In particular, the digital fluorescence input image may comprise a fluorescence signal representing the fluorescence of fluorophores artificially added to the scene (e.g., the fluorophore could be pPIX (protoporphyrin IX) obtained by applying 5-ALA (5-aminolevulinic acid) to expose a tumor). Further, the digital fluorescence input image may also comprise an autofluorescence signal representing the fluorescence of naturally occurring fluorophores in the scene (e.g., human tissue), and a reflected signal representing light reflected from objects in the scene (e.g., fluorophore excitation light).
[0060] Similarly, the input image data can include 3D images acquired by a time-of-flight camera. The difference between 3D images and RGB images is that 3D images also include a range image. The range image provides depth information for each pixel. Therefore, a 3D image comprises one layer containing depth information and at least another layer containing pixel intensities representing light signals received from the scene.
[0061] For applications where the input image data contains cubes of multispectral or hyperspectral imaging data, the data processing device can be configured to decompose the mixed pixels in the input image data into groups of endmembers and fractions through spectral unmixing, where each of the multiple stereo images represents a different endmember in the group of endmembers obtained from spectral unmixing. The categories mentioned above can be one of spectral bands, narrow spectral bands, and wide spectral bands in the multispectral or hyperspectral data.
[0062] Spectral unmixing is an analytical method derived from satellite imaging. The result of spectral unmixing is a dataset containing a set of endmembers or constituent spectra and corresponding fractions or abundances, indicating the proportion of each endmember present in the analyzed pixel. In other words, the dataset reveals what substances are present at the location represented by the analyzed pixel, and how abundant those substances are relative to other present substances.
[0063] In other words, each blended pixel in the input image data can be a mixture of several spectral materials in the scene. The purpose of spectral unmixing is to identify the constituent spectra from the blended pixels and calculate the proportion of each constituent spectrum in the blended pixels to quantitatively decompose or "unmix" them.
[0064] The mixed pixels mentioned above are mixtures that appear to be larger than a single different substance within a single pixel, and they exist for one of two reasons. First, if the spatial resolution of the sensor is low enough that different materials can share a single pixel, the resulting spectral measurement will be a combination of individual spectra. Second, mixed pixels can be obtained when different materials combine to form a homogeneous mixture. This situation can occur independently of the sensor's spatial resolution.
[0065] Endmembers typically correspond to familiar macroscopic or microscopic objects, such as blood vessels, skin tissue, cancerous tissue, etc. To identify these endmembers, spectral demixing can be performed using a known reference spectrum. Therefore, as a result of spectral demixing, the compositional spectrum corresponding to the macroscopic or microscopic objects in the scene is obtained. Further, scores are assigned to the compositional spectrum and associated with pixels. Each score for a pixel indicates the proportion of the compositional spectrum that contributes to the data value of that pixel.
[0066] When spectral unmixing is applied to fluorescence images, it allows for the differentiation of multiple fluorophore features originating from the same source location (or, in the case of digital images, the same pixel) within a scene. Therefore, the data processing device can be configured to decompose a digital fluorescence input image into fluorescence, autofluorescence, and reflection signals. Specifically, the data processing device can be configured to identify each of the fluorescence, autofluorescence, and reflection signals as a separate category. Furthermore, the data processing device can be configured to assign a different disparity to each of the fluorescence, autofluorescence, and reflection signals.
[0067] Additionally or optionally, the data processing device can be configured to decompose the input image data into objects through image segmentation and assign the objects to stereo images in a plurality of stereo images. Here, each object has a category associated with it. For example, the object class can be an associated category. Similarly, the category can be a semantic label of an identified and / or located object in the scene. Therefore, each object also has a location associated with it in the input image data. In particular, pixels in the input image data can be associated with objects.
[0068] Preferably, image segmentation can be semantic image segmentation. Further, the data processing device can be configured to generate multiple stereo images from the results of semantic image segmentation of the input image data. That is, the data processing device can be configured to generate a new stereo image for each object instance identified in the scene. Optionally, the data processing device can be configured to merge each object instance of the same category into a single stereo image.
[0069] Therefore, in a computer-implemented method, the step of analyzing the input image data may further include spectral unmixing of the input image data and / or semantic image segmentation of the input image data. Further, the computer-implemented method may include the following steps: obtaining labels defining the categories of objects identified in the scene as a result of semantic image segmentation, and further obtaining segmentations (including pixel locations) associated with the labels of the locations of the objects in the identified scene.
[0070] For applications where the input image data is a single two-dimensional input image, the data processing device can be configured to generate image layers, for example, through spectral unmixing and / or semantic image segmentation. All image layers have the same size, and each image layer represents a different category / object identified in the scene. Further, each image layer can represent a binary mask with pixels having zero values; pixels with zero values do not represent content in the scene associated with the category / object of the layer. In this case, the data processing device can be configured to generate multiple stereo images from the image layers, preferably by producing pairs of identical image layers by copying each of the image layers, where one image layer in the pair represents the left image and the other image layer in the pair represents the right image.
[0071] In applications where the input image data is a stereo image that already includes a pair of digital input images, the digital input image pair can be used as the left and right images. In this case, the data processing device can be configured to generate multiple left-channel image layers from the left image and multiple right-channel image layers from the right image, each image layer representing a different category / object identified in the scene. In other words, if a stereo image already exists in the input image data, spectral unmixing and / or semantic image segmentation can be applied to the left and right images separately.
[0072] When the input image data is a three-dimensional image, the data processing device can be configured to calculate the projection of the three-dimensional image onto a plane. Therefore, the data processing device acquires and processes two-dimensional input images in a manner similar to that described above for a single input image.
[0073] Alternatively, when the input image data is a three-dimensional image, the data processing device can be configured to calculate two projections onto a plane, from which a stereo image can be extracted. The stereo image, in turn, comprises a pair of digital input images. The data processing device is configured to process the extracted stereo image in a manner similar to that described above when the input image data is a stereo image.
[0074] According to another possible implementation, the data processing device can be configured to receive another user input indicating a disparity value. Further, the data processing device can be configured to update the combined stereoscopic image based on the disparity value, wherein when generating the combined stereoscopic image, the superposition of processed stereoscopic images with disparities exceeding the disparity value is omitted. In this way, the user can exclude certain stereoscopic images at the top to view the stereoscopic image below.
[0075] The input image data may also include multiple input images that collectively represent a series of images. The data processing device can be configured to acquire and analyze these multiple input images sequentially and individually, or to acquire and analyze them in parallel and in combination.
[0076] As used in this document, the term “and / or” includes any and all combinations of one or more related listed items and may be abbreviated as “ / ”. Attached Figure Description
[0077] The invention will now be described by way of example using sample embodiments shown in the accompanying drawings. In the drawings, the same reference numerals are used for features that correspond to each other with respect to at least functional and / or design aspects.
[0078] The combinations of features shown in the appended embodiments are for illustrative purposes only and may be modified. For example, features of embodiments that have technical effects but are not essential may be omitted for a particular application. Similarly, if the technical effect associated with a feature is required for a particular application, that feature, which is not shown as a part of the embodiments, may be added.
[0079] Figure 1 : A schematic diagram showing a flowchart of a computer-implemented method according to an exemplary embodiment of the present invention;
[0080] Figure 2 It shows Figure 1 A detailed schematic diagram;
[0081] Figure 3 A schematic diagram of a medical observation device according to an exemplary embodiment of the present invention is shown;
[0082] Figure 4 A schematic diagram of a data processing apparatus according to an exemplary embodiment of the present invention is shown.
[0083] Figure 5 A schematic diagram of a data processing apparatus according to another exemplary embodiment of the present invention is shown;
[0084] Figure 6 A schematic diagram of a data processing apparatus according to another exemplary embodiment of the present invention is shown;
[0085] Figure 7 This diagram illustrates how parallax can be used to enhance depth perception.
[0086] Figure 8 This illustrates an illustration of enhancing depth perception through perspective scaling.
[0087] Figure 9 This diagram illustrates how to enhance depth perception by thickening the surface.
[0088] Figure 10 This diagram illustrates how depth perception is enhanced through depth-of-focus adjustment.
[0089] Figure 11 This illustrates a diagram of enhancing depth perception through brightness adjustment.
[0090] Figure 12 : This illustrates an illustration of enhancing depth perception through transparency adjustments; and
[0091] Figure 13 : This shows a schematic diagram of a system including a microscope. Detailed Implementation
[0092] First, refer to Figure 1 and Figure 2 The method 100 implemented by the computer is explained. Subsequently, refer to... Figures 3 to 13 The structure and function of the data processing device 300 and the medical observation device 304, such as the microscope 310 or endoscope, are explained.
[0093] The computer-implemented method 100 is used to process input image data 115 from the medical observation device 304 to achieve stereoscopic visualization of the input image data 115. For example... Figure 1 As can be seen, method 100 includes a step 110 of acquiring input image data 115, which represents scene 200 imaged by medical observation device 304. Here, input image data 115 may contain a raster image 116 composed of pixels.
[0094] For example, input image data 115 may comprise a multispectral or hyperspectral imaging data cube 117 having layers of raster images of equal size in one of the respective multispectral or hyperspectral channels. In particular, input image data 115 may comprise a digital fluorescence input image 118 representing the fluorescence of one or more fluorophores 306 present in scene 200.
[0095] Optionally, the input image data 115 may contain an RGB image 119, which represents the scene 200 and consists of three equally sized raster images. Each raster image layer represents the color intensity of the scene in one of the corresponding color channels: red, green, and blue.
[0096] Further, method 100 includes step 120 of analyzing input image data 115 to determine different categories 125 in the imaged scene 200. As part of this analysis, the input image data 115 can be decomposed into objects 225 by image segmentation 126. Here, each object 225 has an associated category 125. For example, object class 127 could be an associated category 125.
[0097] Preferably, image segmentation 126 can be semantic image segmentation 126. Therefore, category 125 can be a semantic label 129 for an object 225 identified and / or located in scene 200. Thus, each object 225 also has a location associated with it in the input image data 115. Specifically, pixels in the input image data 115 can be associated with an object 225.
[0098] Object 225 can be macroscopic or microscopic. For example, in medical applications, object 225 could be blood vessels, skin tissue, cancerous tissue, etc., found in scenario 200. For ease of understanding, object 225 is... Figure 2 The cross symbol 202, smiley face symbol 204, and flashing light symbol 206 are used to represent it. For example... Figure 3 As can be seen, objects 225 can also overlap each other. Here, objects 225 are circles 332, triangles 334, and striped patterns 336, each with a different color.
[0099] For applications where the input image data 115 contains a multispectral or hyperspectral imaging data cube 117, the aforementioned analysis of the input image data 115 can be performed via spectral demixing 128. If the objects 225 overlap, such as... Figure 3 This is particularly useful in situations like this. Therefore, category 125 can be a spectral band from multispectral or hyperspectral data. Figure 3 In the example, each spectral band belongs to one of the following: a circle 332, a triangle 334, or a stripe pattern 336.
[0100] When the input image data 115 contains a multispectral or hyperspectral imaging data cube 117, image segmentation 126, particularly semantic image segmentation 126, can also be utilized. In this case, the spectral layers / channels of the multispectral or hyperspectral imaging data cube 117 can be segmented individually or in combination with each other.
[0101] like Figure 1 As can be further seen, method 100 also includes a step 130 of generating a plurality of stereo images 135 from input image data 115. Each of the stereo images 135 represents a different category 125 determined in the imaging scene 200. If image segmentation 126 is applied, objects 225 can be assigned to each stereo image 135. If spectral demixing 128 is used, spectral bands can be assigned to each stereo image 135. That is, for each object 225 / spectral band determined in scene 200, a new stereo image 135 can be generated.
[0102] For applications where the input image data 115 is a two-dimensional input image 324, image layers 326 can be generated, for example, through spectral unmixing 128 and / or semantic image segmentation 126. All image layers 326 have the same size and each image layer 326 represents a different category 125 / object 225 determined in scene 200.
[0103] Furthermore, step 130, which generates multiple stereo images 135 representing different categories 125, may include step 136, which generates identical pairs of digital images (left digital image L and right digital image R), wherein digital images L and R have non-zero pixel values in all pixels 312, and all pixels 312 represent the category 125 determined in scene 200. Pixels 314 in digital images L and R with zero pixel values do not represent one of the categories 125 represented by non-zero value pixels 312 in the same digital images L and R (see [link to documentation]). Figure 4 In other words, digital images L and R can be binary masks 316, where non-zero pixels 312 represent content, such as a portion of a defined category 125 assigned to stereoscopic image 135, and zero-value pixels 314 in digital images L and R are treated as transparent pixels 318, which do not accumulate into the visible image contribution when digital images L and R are superimposed on other images in other stereoscopic images. Generating the same digital image pair L and R can be accomplished by copying each of the image layers 326.
[0104] The digital image L on the left is presented to the left eye of viewer 400, and the digital image R on the right is presented to the right eye of viewer 400. When viewed, the human brain should perceive images L and R as a single three-dimensional view 404, thereby enabling viewer 400 to perceive three-dimensional depth. The separate presentation of images L and R—one for the left eye and one for the right eye—can be combined by using specialized glasses 406 and / or a specialized display 408.
[0105] To achieve this stereoscopic visualization, images L and R need to have a certain perspective difference. After all, it is through this perspective difference between images L and R seen by the left and right eyes of the viewer 400, the so-called binocular parallax, and the viewer's adjustment through focusing, that the three-dimensional view 404 is completed.
[0106] To this end, method 100 includes step 140 of assigning different disparities A, B, C to each of a plurality of stereo images 135 based on a determined category 125, resulting in a plurality of processed stereo images 145. To automatically determine the disparities A, B, C to be assigned to the stereo images 135, the following heuristic rule can be used: the larger the disparities A, B, C assigned to the stereo images 135, the smaller the sum of the non-zero pixels 312 in the stereo images 135 representing the category 125 assigned to the stereo images 135. Therefore, category 125 represented by stereo images 135 with a large number of non-zero pixels 312 will be assigned smaller disparities A, B, C, while category 125 represented by stereo images 135 with fewer non-zero pixels 312 will be assigned larger disparities A, B, C.
[0107] exist Figure 4 In the example shown, circle 332 has the fewest non-zero pixels 312, while stripe pattern 336 has the most non-zero pixels 312. Therefore, circle 332 is assigned a larger parallax A, and stripe pattern 336 is assigned a smaller parallax C. Figure 4 Triangle 334 in the triangle is assigned a disparity B between A and C due to the median of its non-zero pixel 312.
[0108] Optionally, parallaxes A, B, and C can be assigned to multiple stereoscopic images 135 based on user selection input 195. That is, user 402, particularly viewer 400, can select which stereoscopic image 135 to assign parallaxes A, B, and C to via input device 328. Input device 328 can be connected to the user interface 330 of a medical observation device 304 configured to receive user selection input 195.
[0109] Step 140, which assigns disparities A, B, and C, preferably includes step 190, which applies the assigned disparities A, B, and C to the stereoscopic image 135. Specifically, the processed stereoscopic image 145 can be obtained by applying disparities A, B, and C to the stereoscopic image 135.
[0110] For example, a disparity transformation 410, or disparity offset 412, can be applied to each stereo image 135 according to the assigned disparities A, B, and C. The resulting processed stereo image 145 differs from the original stereo image 135 in that the individual digital images L and R (which constitute each stereo image 135, 145) now have disparities according to the assigned disparities A, B, and C.
[0111] The parallax offset 412 can be applied in the following way: the pixels in the left digital image L of the stereoscopic image 135 are offset to represent the amount of assigned parallax A, B, C, and the pixels in the right digital image R of the stereoscopic image 135 are offset in the opposite direction to represent the amount of assigned parallax A, B, C (see...).Figure 4 Offset pixels refer to the process of moving the position of a pixel within a raster image 116 where the rest remains unchanged.
[0112] When offsetting, pixels can move along a horizontal axis 414, which represents the horizontal distance direction of the viewer's eye. Typically, if the stereoscopic image 135 appears closer to the viewer 400, the pixels of the left digital image L and the right digital image R are offset towards each other. Conversely, offsetting the pixels of the left digital image L and the right digital image R away from each other will make the processed stereoscopic image 145 appear farther away from the viewer 400. This is in... Figure 4 The right side is described, showing a three-dimensional view 404 as perceived by viewer 400.
[0113] To present the image to viewer 400, method 100 may include step 150 of combining multiple processed stereoscopic images 145 to form a combined stereoscopic image 155. Specifically, in generating the combined stereoscopic image 155, the multiple processed stereoscopic images 145 may be superimposed on each other. Preferably, the multiple processed stereoscopic images 145 are superimposed along a central axis perpendicular to the image plane of each of the processed stereoscopic images 145. Furthermore, the central axis passes through the center of each of the processed stereoscopic images 145. Further, the multiple processed stereoscopic images 145 may be superimposed in an order determined by the parallaxes A, B, C assigned to them.
[0114] For example, the processed stereoscopic image 145 with the highest assigned parallax A can be placed at the top, followed by the remaining processed stereoscopic images 145 arranged in descending order of parallax B, C. Figure 4 In the example, circle 332 with the highest parallax A overlays triangle 334 and stripe pattern 336. Triangle 334 overlays stripe pattern 336 because its parallax B is greater than the parallax C of stripe pattern 336.
[0115] The combined stereoscopic image 155 can be used as the output of the specialized display 408 mentioned above. Depending on the type of display, the viewer 400 may have to wear the specialized glasses 406 mentioned above to properly view the combined stereoscopic image 155. The specialized glasses 406 and the display 408 can be part of the display device 320 of the medical observation device 304.
[0116] The medical observation device 304 may also include an optical instrument 322 configured to capture input image data 115. The optical instrument 322 may be a stereo digital camera, but it may also be a digital camera, a multispectral camera, a hyperspectral camera, or a time-of-flight camera. More specifically, the digital camera may be a digital RGB camera and / or a digital reflectance camera. The multispectral camera may include a digital fluorescence camera, and optionally include additional digital fluorescence cameras.
[0117] The data processing device 300 may also be part of the medical observation device 304. In particular, the data processing device 300 may be integrated into the microscope 310 as an embedded processor 302 or as part of such an embedded processor 302.
[0118] When the display device 320 utilizes the so-called active viewing system 500, the viewer 400 can obtain an immersive experience. In this case, the data processing device 300 can be configured to receive user input 175 via the input device 328 and the user interface 330. The user input 175 may indicate the viewer's viewing position relative to the display device 320. Based on the received user input 175, the data processing device 300 can be configured to adjust the assigned parallaxes A, B, and C. When the received user input 175 indicates a change in the viewer's angle and / or viewing distance relative to the display device 320, the assigned parallaxes A, B, and C can be adjusted 180. Figure 5 and Figure 6 Arrow 502 in the diagram indicates this change.
[0119] For example, user input 175 could be face detection data 504 received from a face-tracking camera 506 that analyzes the viewer's eyes / gaze. The data processing device 300 and / or the face-tracking camera 506 can be configured to derive another value 510 representing the viewing distance of the viewer 400 relative to the display device 320 and / or a value 508 representing the viewing angle from the face detection data 504. Further, the data processing device 300 and / or the face-tracking camera 506 can be configured to determine changes in viewing angle and / or viewing distance when value 508 or other value 510 exceeds a threshold. This threshold can be derived from previously received user input 175 or face detection data 504.
[0120] The adjustment 180 of the assigned disparities A, B, C may include determining updated disparities A', B', C' to be assigned to the multiple stereoscopic images 135. The updated disparities A', B', C' reflect changes in viewing angle and / or viewing distance and can be used to update the combined stereoscopic image 155. In other words, the disparities A, B, C assigned to the multiple stereoscopic images 135 are updated for each of the stereoscopic images 135 to reflect changes in viewing angle and / or viewing distance. That is, the data processing device 300 is configured to dynamically compensate for changes in viewing angle and / or viewing distance.
[0121] Optionally, the assigned disparities A, A', B, B', C, C' can differ between the left digital image L and the right digital image R. For example, if the face-tracking camera 506 registers a viewer's movement to the right, the disparities A, A', B, B', C, C' assigned to the right digital image R can be greater than those assigned to the left digital image L. Additionally, the larger the original disparities A, B, C, the larger the updated disparities A', B', C' can be. This is in... Figure 5 As shown in the image, and by tilting the head to the side, viewers are allowed to "peek" at the content behind the top layer.
[0122] Additionally, method 100 may include step 170 of receiving another user input 196 indicating a disparity value. The combined stereoscopic image 155 can then be updated based on the disparity value, wherein, when generating the combined stereoscopic image 155, the overlay of the processed stereoscopic image 145, which has already been assigned a disparity 230 exceeding the disparity value, is omitted (see [link to documentation]). Figure 6 ).
[0123] To enhance the perceived 3D depth by the viewer 400, the data processing device 300 can be configured to modify the scale of the processed stereoscopic image 145 based on corresponding assigned parallaxes A, B, and C. Here, the scale represents the angular or lateral range of the processed stereoscopic image 145. Specifically, the scale represents the maximum angular diameter that the stereoscopic image 155 will cover on the display device 320 after multiple processed stereoscopic images 145 are superimposed to generate a combined stereoscopic image 155.
[0124] For example, the data processing device 300 can be configured to magnify the processed stereoscopic image 145 based on the assigned parallaxes A, B, and C. Here, magnification refers to scaling the processed stereoscopic image 145 using a scaling factor. The scaling factor can be greater than or less than 1. When the stereoscopic image is magnified, the aspect ratio remains unchanged when the size of the stereoscopic image is modified accordingly. That is, the scaled stereoscopic image has the same aspect ratio as before the scaling was applied.
[0125] Preferably, the smaller the size ratio / scaling factor, the smaller the parallaxes A, B, and C assigned to the processed stereoscopic image 145. For the viewer, this amplifies the impression that an object 225 displayed in the processed stereoscopic image 145 with smaller parallax is farther away than an object 225 displayed in the processed stereoscopic image with larger parallax. Figure 6 In the example, striped pattern 336 has the smallest scaling factor for its application because its parallax C is also the smallest. Triangle 334, with the second smallest parallax B, receives the second smallest scaling factor. Circle 332, with the largest parallax A, can remain unchanged or even be scaled up by a scaling factor greater than 1.
[0126] Optionally, the size scale / scaling factor can also be updated based on face detection data 504, specifically a value 510 representing the viewing distance of viewer 400 relative to display device 320. This is in Figure 6 As shown in the image.
[0127] Figures 8 to 12 An embodiment is shown in which a data processing device is configured to apply a post-processing transformation 160 to at least one subgroup of a plurality of processed stereoscopic images 145 before combining them. This post-processing transformation 160 may be perspective scaling 800 (see...). Figure 8 ), thickening surface effect 900 (see Figure 9 ), depth of focus adjusted to 1000 (see Figure 10 Brightness adjustment 1100 (see) Figure 11 ) and transparency adjustment 1200 (see Figure 12 ).
[0128] Perspective scaling 800 involves reducing or increasing the size / size of the stereoscopic image 135 before overlaying the stereoscopic image 145, while maintaining the aspect ratio of the stereoscopic image 135. Therefore, smaller objects appear to be farther away than larger objects.
[0129] The thickened surface effect 900 may include cloning pixel values from the middle position 902 of the parallax adjustment 700 to the end position 904. The middle position 902 is located between the starting position 906 and the end position 904 of the parallax adjustment 700.
[0130] The depth-of-focus adjustment of 1000 increases the blur level of the stereoscopic image 135 in the overlay stereoscopic image 145. For example, a blur filter, such as a Gaussian filter with a kernel size proportional to the desired blur level of the processed stereoscopic image 145, can be applied to the stereoscopic image 135. The stronger the applied blur effect, the farther away the object appears.
[0131] Brightness adjustment 1100 increases the brightness level of the stereoscopic image 135 at the overlay stereoscopic image 145. The increased brightness allows certain stereoscopic elements in the foreground to be highlighted, while other stereoscopic elements in the background to be faded.
[0132] The transparency adjustment of 1200 increases the transparency level of the stereoscopic image 135 when the stereoscopic image 145 is overlaid. The increased transparency allows the stereoscopic image to be seen through and the stereoscopic image below to be observed.
[0133] Preferably, the same post-processing transform 160 is applied to each individual digital image 324 in L and R that constitute the stereoscopic image 135 or the processed stereoscopic image 145. For example, the same blur filter is applied to the left digital image L and the right digital image R.
[0134] Specifically, the post-processing transformation 160 can be applied to the processed stereoscopic image 145 based on the assigned disparity. Further, the data processing device 300 can be configured to designate a stereoscopic image 135 as a background 908 from among a plurality of processed stereoscopic images 135, wherein the disparity D assigned to the background 908 is the global minimum or maximum value among the disparities A, B, C, and D assigned to the plurality of stereoscopic images 135. The background 908 can serve as a reference for the viewer 400.
[0135] Optionally, the data processing device 300 can be configured to increase or decrease the weight of the post-processing transformation 160 when applying it to the stereoscopic image 135, based on the distance between the background 908 and the stereoscopic image 135. Here, the difference between the disparities A, B, C, and D assigned to two of the plurality of stereoscopic images 135 can define the distance between the two stereoscopic images 135.
[0136] like Figure 9 As can be seen, the parallax applied to each of the multiple stereoscopic images 135 can be different from the parallax assigned to all other images. That is, each stereoscopic image 135 can be assigned a unique parallax. Therefore, the circle 332, the triangle 334, the striped pattern 336, and the background 908 all appear to be at different depths in the three-dimensional view 404.
[0137] Optionally, stereo images representing different categories of the same whole can have the same assigned disparity. For example, two or more stereo images can be designated as background 908, thus sharing the same disparity. Figure 7 In the background 908 and stripe pattern 336, there is the same parallax, so they appear to be at the same depth in the 3D view 404.
[0138] Although some aspects are already described in the context of the apparatus, it is clear that these aspects also represent descriptions of the corresponding methods, where a block or device corresponds to a method step or feature of a method step. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks, items, or features of the corresponding apparatus.
[0139] Some implementations involve microscopes, which include information about Figures 1 to 12 One or more of the systems described herein. Optionally, the microscope may be about Figures 1 to 12 It is part of or connected to one or more of the systems described in the text. Figure 13 A schematic diagram of a system 1300 configured to perform the methods described herein is shown. System 1300 includes a microscope 310 and a computer system 1320. The microscope 310 is configured to capture images and is connected to the computer system 1320. The computer system 1320 is configured to perform at least a portion of the methods described herein. The computer system 1320 may be configured to execute machine learning algorithms. The computer system 1320 and the microscope 310 may be separate entities, but may also be integrated together in a common housing. The computer system 1320 may be part of the central processing system of the microscope 310 and / or the computer system 1320 may be part of a sub-component of the microscope 310, such as a sensor, actuator, camera, or illumination unit of the microscope 310.
[0140] Computer system 1320 may be a local computer device (e.g., a personal computer, laptop computer, tablet computer, or mobile phone) having one or more processors and one or more storage devices, or it may be a distributed computer system (e.g., a cloud computing system having one or more processors and one or more storage devices distributed in different locations, such as at local clients and / or one or more remote server farms and / or data centers). Computer system 1320 may include any circuitry or combination of circuitry. In one embodiment, computer system 1320 may include one or more processors, which may be of any type. As used herein, a processor may mean any type of computing circuitry, such as, but not limited to, a microprocessor, microcontroller, complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, graphics processor, digital signal processor (DSP), multi-core processor, field-programmable gate array (FPGA) (e.g., a microscope or microscope component (e.g., a camera)), or any other type of processor or processing circuitry. Other types of circuitry that may be included in computer system 1320 may be custom circuitry, application-specific integrated circuits (ASICs), etc., such as one or more circuits (e.g., communication circuitry) used in wireless devices such as mobile phones, tablet computers, laptop computers, two-way radios, and similar electronic systems. Computer system 1320 may include one or more storage devices, which may include one or more memory elements suitable for a particular application, such as main memory in the form of random access memory (RAM), one or more hard disk drives, and / or one or more drives for disposing of removable media such as CDs, flash memory cards, digital video discs (DVDs), etc. Computer system 1320 may also include a display device, one or more speakers and a keyboard and / or controller, the controller including a mouse, trackball, touchscreen, voice recognition device, or any other device that allows system users to input and receive information from computer system 1320.
[0141] Some or all of the method steps can be performed by (or using) hardware devices (e.g., processors, microprocessors, programmable computers, or electronic circuits). In some embodiments, such devices can perform one or more of the most important method steps.
[0142] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be carried out using a non-transitory storage medium (such as a digital storage medium, e.g., floppy disk, DVD, Blu-ray disc, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having electronically readable control signals stored thereon, which cooperate with (or are capable of cooperating with) a programmable computer system to cause the corresponding methods to be executed. Therefore, the digital storage medium can be computer-readable.
[0143] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0144] Generally, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0145] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0146] In other words, therefore, an embodiment of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.
[0147] Therefore, another embodiment of the invention is a storage medium (or data carrier, or computer-readable medium) comprising a computer program stored thereon, which, when executed by a processor, is used to perform one of the methods described herein. Data carriers, digital storage media, or recording media are generally tangible and / or non-transitory. Another embodiment of the invention is an apparatus as described herein, comprising a processor and a storage medium.
[0148] Therefore, another embodiment of the invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection (e.g., via the Internet).
[0149] Another embodiment includes a processing component, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.
[0150] Another embodiment includes a computer on which a computer program is installed for performing one of the methods described herein.
[0151] Another embodiment of the invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0152] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device. Figure Labels
[0153] 100 methods
[0154] 110 steps
[0155] 115 Input image data
[0156] 116 raster images
[0157] 117 Data Cube
[0158] 118 digital fluorescence input images
[0159] 119 RGB image
[0160] 120 steps
[0161] 125 categories
[0162] 126 Image Segmentation
[0163] 127 Object Class
[0164] 128 Spectral Demixing
[0165] 129 Semantic Tags
[0166] 130 steps
[0167] 135 Stereoscopic Images
[0168] 136 steps
[0169] 140 steps
[0170] 145 Processed stereoscopic image
[0171] 150 steps
[0172] 155 Combined Stereo Images
[0173] 160 Post-processing Transformation
[0174] 170 steps
[0175] 175 User Input
[0176] 180 adjustment
[0177] 190 steps
[0178] 195 User selection input
[0179] 196 User Input
[0180] 200 scenes
[0181] 202 Cross Symbol
[0182] 204 smiley face symbols
[0183] 206 flashing symbols
[0184] 225 objects
[0185] 230 parallax
[0186] 300 Data Processing Equipment
[0187] 302 Embedded Processor
[0188] 304 Medical Observation Device
[0189] 306 fluorophore
[0190] 310 Microscope
[0191] 312 non-zero pixels
[0192] 314 Zero Pixels
[0193] 316 binary mask
[0194] 318 transparent pixels
[0195] 320 display device
[0196] 322 Optical Instruments
[0197] 324 Two-dimensional input images
[0198] 326 Image Layers
[0199] 328 Input Devices
[0200] 330 User Interface
[0201] 332 Circular
[0202] 334 Triangle
[0203] 336 striped pattern
[0204] 400 viewers
[0205] 402 users
[0206] 404 3D View
[0207] 406 Specialized Glasses
[0208] 408 Dedicated Monitor
[0209] 410 Parallax Transformation
[0210] 412 Parallax Shift
[0211] 414 Horizontal axis
[0212] 500 Active Viewing System
[0213] 502 arrow
[0214] 504 Facial Detection Data
[0215] 506 Face Tracking Camera
[0216] 508 value
[0217] 510 value
[0218] 700 parallax adjustment
[0219] 800 perspective zoom
[0220] 900 Thickened Surface Effect
[0221] 902 Midpoint
[0222] 904 Finish Line
[0223] 906 Starting position
[0224] 908 Background
[0225] 1000 Depth of Focus Adjustment
[0226] 1100 brightness adjustment
[0227] 1200 Transparency Adjustment
[0228] 1300 System
[0229] 1320 Computer System
[0230] Parallax of A, B, and C
[0231] Updated parallax of A', B', and C'
[0232] D is assigned to the parallax of the background.
[0233] L Left side digital image
[0234] R right-hand digital image
Claims
1. A data processing device (300) for a medical observation apparatus, such as an endoscope or microscope, said data processing device being configured to: - Acquire (110) input image data (115), the input image data representing a scene obtained by the medical observation device; -Analyze (120) the input image data to determine different categories (125) in the scene; - Generate (130) a plurality of stereo images (135) from the input image data (115), each of the stereo images representing a different category (125) determined in the scene; - Based on the determined category, assign (140) different disparities to each of the plurality of stereo images to generate a plurality of processed stereo images (145). and - Combine (150) the plurality of processed stereo images (145) to generate a combined stereo image (155).
2. The data processing apparatus (300) according to claim 1 is further configured to modify (160) the size ratio of the stereoscopic images in the plurality of stereoscopic images (135) based on the parallax assigned to the stereoscopic image.
3. The data processing apparatus (300) according to claim 1 or 2 is further configured to: - Receive (170) user input (175), the user input (175) representing the user's viewing position relative to the display device, and - Adjust (180) the assigned parallax based on the received user input (175).
4. The data processing apparatus (300) according to any one of claims 1 to 3 is further configured to: - Prior to the combination (150), at least one subgroup of the plurality of processed stereo images (145) is subjected to a post-processing transformation (160).
5. The data processing apparatus (300) according to claim 4, wherein the post-processing transformation (160) is at least one of the following: transparency adjustment, brightness adjustment, depth of focus adjustment, perspective scaling, and surface thickening effect.
6. The data processing apparatus (300) according to claim 4 or 5. The post-processing transformation (160) is applied to the processed stereoscopic image (145) based on the assigned disparity, and the data processing device is configured to: - Designate a stereo image as a background from the plurality of processed stereo images (145), wherein the disparity assigned to the background is the global minimum or maximum value of the disparities (230) assigned to the plurality of stereo images (135). The difference between the disparity (230) assigned to two stereo images in the plurality of stereo images (135) defines the distance between the two stereo images; -Based on the distance between the background and the stereo image, when the post-processing transformation is applied to the processed stereo image, the closer the stereo image is to the background, the greater or less the weight of the post-processing transformation (160) is increased or decreased.
7. The data processing apparatus (300) according to any one of claims 1 to 6 is further configured to decompose (130) mixed pixels in the input image data (115) into groups of endmembers and fractions by spectral unmixing, wherein each of the plurality of stereo images (135) represents a different endmember in the group of endmembers obtained from the spectral unmixing.
8. The data processing apparatus (300) according to any one of claims 1 to 7 is further configured to decompose (130) the input image data (115) into objects by image segmentation and to assign the objects to stereo images in the plurality of stereo images (135).
9. The data processing apparatus (300) according to claim 8, wherein the image segmentation is semantic image segmentation.
10. The data processing apparatus (300) according to any one of claims 3 to 9, further configured to: - Receive (170) user select input (195); and - When the data processing device further relies on at least claim 3, it deactivates the adjustment (180) of the combined stereoscopic image (155) based on the received user selection input (195); and / or - When the data processing device further depends on at least claim 4, the application of the post-processing transformation (160) is deactivated based on the received user selection input (195).
11. The data processing apparatus (300) according to any one of claims 1 to 10, further configured to: - Receive (170) another user input (196) indicating the disparity value; and - Update (180) the combined stereo image (155) based on the disparity value (196), wherein when generating the combined stereo image (155), the superposition (150) of the processed stereo image (145) that has been assigned disparities (A, B, C) exceeding the disparity value is omitted.
12. A medical observation device (304), comprising a data processing device according to any one of claims 1 to 11, further comprising: Optical instruments, which are configured as follows: - Capture the input image data (115); and - Provide the input image data (115) to the data processing device; The user interface is configured to receive user input (175, 196) and / or user selection input (195). and A display device configured to receive the combined stereoscopic image (155).
13. A computer-implemented method (100) for processing input image data from a medical observation device, such as a microscope or endoscope, the method comprising the steps of: - Acquire (110) the input image data (115) representing the scene imaged by the medical observation device. - Analyze (120) the input image data (115) to determine different categories (125) in the scene of the imaging; - Generate (130) a plurality of stereo images (135) from the input image data (115), each of the stereo images representing a different category (125) determined in the scene of the imaging. - Based on the determined category, a different disparity (AC) is assigned (140) to each of the plurality of stereo images (135), resulting in a plurality of processed stereo images (145); and - The multiple processed stereo images are combined with each other (150) to result in a combined stereo image (155).
14. The computer-implemented method (100) according to claim 13, wherein analyzing the input image data further comprises: - Perform semantic image segmentation on the input image data (115); and / or - Spectral demixing of the input image data (115).
15. A computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method of claim 13 or 14.
16. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of claim 13 or 14.